Splits text at token offsets, with optional overlap.
Summary
Functions
Splits text into overlapping chunks using the given tokenizer.
Types
@type chunk_opts() :: [ chunk_size: pos_integer(), chunk_overlap: non_neg_integer() | float() ]
Functions
@spec chunk(binary(), Chunx.Tokenizer.t(), chunk_opts()) :: {:ok, [Chunx.Chunk.t()]} | {:error, term()}
Splits text into overlapping chunks using the given tokenizer.
Options
:chunk_size- Target maximum number of content tokens per chunk (default: 512). A byte-indivisible grapheme may contain more tokens and is kept intact.:chunk_overlap- Overlap as a token count or a fraction in the range[0.0, 1.0)(default:0.25).
Examples
iex> {:ok, tokenizer} = Tokenizers.Tokenizer.from_pretrained("distilbert/distilbert-base-uncased")
iex> Chunx.Chunker.Token.chunk("Some text to split", tokenizer, chunk_size: 3, chunk_overlap: 1)
{
:ok,
[
%Chunx.Chunk{end_byte: 12, start_byte: 0, text: "Some text to", token_count: 3},
%Chunx.Chunk{end_byte: 18, start_byte: 10, text: "to split", token_count: 2}
]
}