PaddingSide

Side on which padding tokens are inserted.

Union cases

NameDescription
Left
Right

Instance members

NameDescription
this.IsLeft
this.IsRight

TruncationSide

Side removed when a token sequence exceeds the target length.

Union cases

NameDescription
Left
Right

Instance members

NameDescription
this.IsLeft
this.IsRight

CollationLength

Target sequence length for padding and truncation.

Union cases

NameDescription
BatchMax maxLengthPad to the longest sequence in the batch, optionally truncating first.
Fixed intPad and truncate every sequence to this length.

Instance members

NameDescription
this.IsBatchMax
this.IsFixed

CollationOptions

Text collation policy.

Record fields

NameDescription
Length
PadTokenId
PaddingSide
TruncationSide

CollationOptions module

Constructors for text collation policies.

FunctionDescription
create padTokenId lengthCreate a right-padded policy that truncates tokens from the right.

create

create padTokenId length

Create a right-padded policy that truncates tokens from the right.

Parameters

  • padTokenId : int64
  • length : CollationLength

Returns CollationOptions

EncodedBatch

A model-ready batch produced from text.

Record fields

NameDescription
AttentionMaskBoolean attention mask with shape [batch, sequence].
InputIdsToken IDs with shape [batch, sequence].
LengthsOriginal token count for each batch item before truncation.

Collation

Text-to-tensor collation.

FunctionDescription
batch tokenizer texts options deviceEncode texts into a batch with mask and retained lengths.
toTensor tokenizer text options deviceEncode one text into a token tensor. Fixed length pads to that length;
batch-max pads only to the (possibly truncated) sequence itself.

batch

batch tokenizer texts options device

Encode texts into a batch with mask and retained lengths.

Parameters

  • tokenizer : Tokenizer
  • texts : string list
  • options : CollationOptions
  • device : Device

Returns EncodedBatch


toTensor

toTensor tokenizer text options device

Encode one text into a token tensor. Fixed length pads to that length; batch-max pads only to the (possibly truncated) sequence itself.

Parameters

  • tokenizer : Tokenizer
  • text : string
  • options : CollationOptions
  • device : Device

Returns Tensor