CausalLmInput

Tensor input shared by cached causal language models.

Record fields

NameDescription
AttentionMaskOptional 0/1 padding mask with shape [batch, total sequence].
CacheOptional model-specific key/value cache.
InputIdsToken IDs with shape [batch, sequence].
PositionIdsOptional absolute positions with shape [sequence] or [batch, sequence].

CausalLmOutput

Tensor output shared by cached causal language models.

Record fields

NameDescription
CacheThe cache supplied in the input, after the successful forward pass.
LogitsVocabulary logits with shape [batch, sequence, vocabulary].

CausalLm contract

Typed operations and metadata required for causal language-model generation.

Record fields

NameDescription
CacheLengthReturn the number of tokens currently stored by a cache.
ContextLengthMaximum total number of prompt and generated tokens.
CreateCacheAllocate an empty request-local cache for a batch and token capacity.
DeviceDevice on which token inputs are created.
DisposeCacheRelease a request-local cache and its tensors.
EosTokenIdsToken IDs that terminate generation after being selected.
ForwardRun the model for a prefill or single-token decode input.

CausalLm

Tensor-level prefill and decode operations for causal language models.

FunctionDescription
decode input modelDecode exactly one new token using a populated cache.
prefill input modelPopulate an empty cache from one or more input tokens.

decode

decode input model

Decode exactly one new token using a populated cache.

Parameters

  • input : CausalLmInput<'Cache>
  • model : CausalLm<'Cache>

Returns CausalLmOutput<'Cache>


prefill

prefill input model

Populate an empty cache from one or more input tokens.

Parameters

  • input : CausalLmInput<'Cache>
  • model : CausalLm<'Cache>

Returns CausalLmOutput<'Cache>