SmolLm2Attention

Projection layers in one SmolLM2 grouped-query attention block.

Record fields

NameDescription
Key
Output
Query
Value

SmolLm2Mlp

Projection layers in one SmolLM2 SwiGLU feed-forward block.

Record fields

NameDescription
Down
Gate
Up

SmolLm2Block

One SmolLM2 transformer block.

Record fields

NameDescription
Attention
InputNorm
Mlp
PostAttentionNorm

SmolLm2

Construction, state, cache, and forward operations for SmolLM2.

FunctionDescription
asCausalLm modelBind a SmolLM2 model to the common typed causal language-model interface.
create config dtype deviceCreate a SmolLM2 model from a validated configuration.
createCache batchSize capacity modelAllocate a reusable fixed-capacity key/value cache for a model.
descriptorState descriptor whose canonical names match Hugging Face SmolLM2 weight names.
dispose modelDispose tensors owned by a SmolLM2 model.
forward input modelRun SmolLM2. When a cache is supplied, only new tokens should be passed after prefill.
loadConfig pathLoad and validate a Hugging Face config.json file.
loadFromDirectory directory deviceLoad config and a strict single-file or sharded SafeTensors state from a local directory.
state modelCreate a validated named state view using Hugging Face weight names.

asCausalLm

asCausalLm model

Bind a SmolLM2 model to the common typed causal language-model interface.

Parameters

  • model : SmolLm2

Returns CausalLm<SmolLm2Cache>


create

create config dtype device

Create a SmolLM2 model from a validated configuration.

Parameters

  • config : SmolLm2Config
  • dtype : ScalarType
  • device : Device

Returns SmolLm2


createCache

createCache batchSize capacity model

Allocate a reusable fixed-capacity key/value cache for a model.

Parameters

  • batchSize : int64
  • capacity : int64
  • model : SmolLm2

Returns SmolLm2Cache


descriptor

descriptor

State descriptor whose canonical names match Hugging Face SmolLM2 weight names.

Returns ModelDescriptor<SmolLm2>


dispose

dispose model

Dispose tensors owned by a SmolLM2 model.

Parameters

  • model : SmolLm2

forward

forward input model

Run SmolLM2. When a cache is supplied, only new tokens should be passed after prefill.

Parameters

  • input : SmolLm2Input
  • model : SmolLm2

Returns SmolLm2Output


loadConfig

loadConfig path

Load and validate a Hugging Face config.json file.

Parameters

  • path : string

Returns SmolLm2Config


loadFromDirectory

loadFromDirectory directory device

Load config and a strict single-file or sharded SafeTensors state from a local directory.

Parameters

  • directory : string
  • device : Device

Returns SmolLm2 * LoadReport


state

state model

Create a validated named state view using Hugging Face weight names.

Parameters

  • model : SmolLm2

Returns ModelState