Octrees as an
Explicit 3D Language

OctLLM

† Corresponding author

Abstract

Existing 3D large language models (LLMs) compromise on two fronts: they compress shapes into latent codebook indices or coordinate text, which removes spatial structure from what the model observes, and they acquire the 3D modality by fine-tuning the backbone, which overwrites its general language ability. We present OctLLM, which addresses both limitations. Geometry enters as an explicit 3D sequence of octree occupancy tokens. However, full octree sequences grow rapidly with depth; OctLLM therefore randomly empties penultimate-level nodes and omits descendants while preserving shape, yielding a shorter coordinate- and depth-anchored Sparse Octree (S-Octree) for position-aware mask-modeling generation and 3D understanding. On the other front, existing methods introduce a new modality with full fine-tuning or LoRA, but full fine-tuning is costly, LoRA limits 3D capacity, and both modify the language pathway. OctLLM instead adds 3D capacity in parameters separate from the pretrained ones: mesh tokens are routed through independent trainable branches in a subset of blocks while text and image tokens retain the frozen vision-language pathway, and the two streams interact through shared self-attention. It trains far fewer parameters than full fine-tuning, yet sets a new state of the art among unified multimodal LLMs, lowering image-to-3D FID by 17.4% and raising render-grounded captioning by 28.7 points over ShapeLLM-Omni, while matching the backbone on general language benchmarks.

Explore in 3D.

From coarse to complete.

OctLLM builds a 6-level octree layer by layer, completes it with a 3D U-Net, and decodes the final mesh with a flow-based decoder.

One 3D language.

Generation and understanding through S-Octrees.

Sparse octrees

Occupancy tokens anchored to 3D coordinates and depth.

Shared attention

Trainable 3D branches alongside a frozen text-image pathway.

Mesh reconstruction

A 3D U-Net completes occupancy; a pretrained flow-based decoder produces the textured mesh.

From shape to language.

Excerpts from OctLLM’s descriptions.