Chunx.Chunker.Semantic (chunx v0.2.0)

Copy Markdown View Source

Splits text where adjacent sentence embeddings become less similar.

Summary

Functions

Splits text using sentence embeddings.

Types

chunk_opts()

@type chunk_opts() :: [
  chunk_size: pos_integer(),
  threshold: float() | :auto,
  min_sentences: pos_integer(),
  min_chunk_size: pos_integer(),
  threshold_step: float(),
  separator: String.t(),
  delimiters: [String.t(), ...],
  min_chars_per_sentence: non_neg_integer(),
  similarity_window: non_neg_integer()
]

Functions

chunk(text, tokenizer, embedding_fun, opts \\ [])

@spec chunk(
  binary(),
  Chunx.Tokenizer.t(),
  Chunx.Chunker.embedding_fun(),
  chunk_opts()
) :: {:ok, [Chunx.SentenceChunk.t()]} | {:error, term()}

Splits text using sentence embeddings.

Options

  • :chunk_size - Target maximum content-token count (default: 512). :min_sentences takes precedence over this target.
  • :threshold - Split when average adjacent similarity is at or below this value, or use :auto to select a value from the input (default: :auto).
  • :min_sentences - Minimum sentence count considered when placing boundaries (default: 1).
  • :min_chunk_size - Target minimum content-token count used while choosing an automatic threshold (default: 2).
  • :threshold_step - Search precision for an automatic threshold (default: 0.01).
  • :delimiters - Strings that end sentences (default: [".", "!", "?", "\n"]).
  • :min_chars_per_sentence - Minimum trimmed character count used when joining short sentence fragments (default: 12).
  • :similarity_window - Number of neighboring sentences included on each side of each embedding input (default: 1).
  • :separator - Compatibility option. It must be non-empty and otherwise has no effect.