Link Search Menu Expand Document Documentation Menu

Kuromoji stemmer token filter

The kuromoji_stemmer token filter normalizes katakana words by removing a trailing long vowel mark (ー) from words that meet a minimum length threshold. Many foreign loanwords in Japanese are written in katakana with a trailing ー that denotes a lengthened final vowel (for example, コンピューター, プリンター). In practice, Japanese speakers often drop the trailing ー in informal or technical writing, resulting in two surface forms for the same word. This filter collapses those variants to a single canonical form.

The filter makes conversions such as the following:

  • コンピューター (computer) becomes コンピュータ.
  • プリンター (printer) becomes プリンタ.

Installation

The kuromoji_stemmer token filter requires the analysis-kuromoji plugin. For installation instructions, see Kuromoji analyzer.

Parameters

The following table lists the parameters for the kuromoji_stemmer token filter.

Parameter Data type Description
minimum_length Integer The minimum number of characters a token must have for the trailing long vowel mark to be removed. Tokens shorter than this threshold are passed through unchanged. Default is 4.

The default minimum length of 4 prevents short words such as カー (car, two characters) from being incorrectly stemmed to カ.

Example

The following example creates an index with a custom analyzer that uses kuromoji_stemmer:

PUT /kuromoji-stemmer-index
{
  "settings": {
    "analysis": {
      "filter": {
        "katakana_stemmer": {
          "type": "kuromoji_stemmer",
          "minimum_length": 4
        }
      },
      "analyzer": {
        "stemmer_analyzer": {
          "type": "custom",
          "tokenizer": "kuromoji_tokenizer",
          "filter": ["katakana_stemmer"]
        }
      }
    }
  }
}

Test the analyzer with katakana loanwords (meaning “I use a computer and printer”):

POST /kuromoji-stemmer-index/_analyze
{
  "analyzer": "stemmer_analyzer",
  "text": "コンピューターとプリンターを使う"
}

The response shows the trailing long vowel marks removed from both words:

{
  "tokens": [
    {
      "token": "コンピュータ",
      "start_offset": 0,
      "end_offset": 7,
      "type": "word",
      "position": 0
    },
    {
      "token": "と",
      "start_offset": 7,
      "end_offset": 8,
      "type": "word",
      "position": 1
    },
    {
      "token": "プリンタ",
      "start_offset": 8,
      "end_offset": 13,
      "type": "word",
      "position": 2
    },
    {
      "token": "を",
      "start_offset": 13,
      "end_offset": 14,
      "type": "word",
      "position": 3
    },
    {
      "token": "使う",
      "start_offset": 14,
      "end_offset": 16,
      "type": "word",
      "position": 4
    }
  ]
}
350 characters left

Have a question? .

Want to contribute? or .