11 releases

✓ Uses Rust 2018 edition

new 0.3.1 Feb 17, 2020
0.3.0 Feb 14, 2020
0.2.1 Feb 12, 2020
0.1.6 Feb 7, 2020

#88 in Text processing

Download history 101/week @ 2020-01-31 75/week @ 2020-02-07

64 downloads per month
Used in lindera-cli

MIT license

18KB
360 lines

Lindera

License: MIT Join the chat at https://gitter.im/lindera-morphology/lindera

A Japanese morphological analysis library in Rust. This project fork from fulmicoton's kuromoji-rs.

Lindera aims to build a library which is easy to install and provides concise APIs for various Rust applications.

Build

The following products are required to build:

  • Rust >= 1.39.0
  • make >= 3.81
% make build

Usage

Basic example

This example covers the basic usage of Lindera.

It will:

  • Create a tokenizer in normal mode
  • Tokenize the input text
  • Output the tokens
use lindera::tokenizer::Tokenizer;

fn main() -> std::io::Result<()> {
    // create tokenizer
    let mut tokenizer = Tokenizer::default_normal();

    // tokenize the text
    let tokens = tokenizer.tokenize("関西国際空港限定トートバッグ");

    // output the tokens
    for token in tokens {
        println!("{}", token.text);
    }

    Ok(())
}

The above example can be run as follows:

% cargo run --example basic_example

You can see the result as follows:

関西国際空港
限定
トートバッグ

API reference

The API reference is available. Please see following URL:

Dependencies

~17MB
~126K SLoC