GenerationSampling

Token-selection strategy used during generation.

Union cases

NameDescription
GreedySelect the highest-logit token.
Temperature temperatureSample from logits divided by a positive temperature.

Instance members

NameDescription
this.IsGreedy
this.IsTemperature

GenerationOptions

Options for one causal language-model generation session.

Record fields

NameDescription
CancellationTokenCancellation signal checked before each model invocation.
MaxNewTokensMaximum number of tokens to generate after the prompt.
SamplingToken-selection strategy applied to each next-token distribution.

GenerationOptions module

Constructors and validation for generation options.

FunctionDescription
greedy maxNewTokensCreate greedy generation options without cancellation.
temperature temperature maxNewTokensCreate temperature-sampling options without cancellation.

greedy

greedy maxNewTokens

Create greedy generation options without cancellation.

Parameters

  • maxNewTokens : int

Returns GenerationOptions


temperature

temperature temperature maxNewTokens

Create temperature-sampling options without cancellation.

Parameters

  • temperature : float
  • maxNewTokens : int

Returns GenerationOptions

GenerationSession

A single-request generation session that owns its token sequence and key/value cache.

Instance members

NameDescription
this.GenerateGenerate until EOS or the configured maximum token count is reached.
this.GeneratedTokenIdsTokens generated so far, including a terminating EOS token when one was selected.
this.IsFinishedWhether EOS or the configured maximum token count has been reached.
this.PromptTokenIdsTokens supplied as the prompt.
this.StepGenerate at most one token. Returns None after the session has finished.
this.TokenIdsPrompt and generated tokens accumulated by this session.
this.TokensGenerate tokens until EOS or the configured maximum token count is reached.

Generation

Session-based causal language-model generation.

FunctionDescription
createSession options promptTokenIds modelCreate a request-local generation session. The caller owns the returned session.
generate options promptTokenIds modelGenerate token IDs and dispose the request-local cache before returning.
tokens sessionEnumerate generated token IDs until the session finishes.

createSession

createSession options promptTokenIds model

Create a request-local generation session. The caller owns the returned session.

Parameters

  • options : GenerationOptions
  • promptTokenIds : int64 list
  • model : CausalLm<'a>

Returns GenerationSession<'a>


generate

generate options promptTokenIds model

Generate token IDs and dispose the request-local cache before returning.

Parameters

  • options : GenerationOptions
  • promptTokenIds : int64 list
  • model : CausalLm<'a>

Returns int64 list


tokens

tokens session

Enumerate generated token IDs until the session finishes.

Parameters

  • session : GenerationSession<'Cache>

Returns int64 seq