pip install mlbricks-kit[ EFFICIENT INTELLIGENCE · BUILT IN BRICKS ]
Build more
intelligence.
Waste less
compute.
MLBricks researches and engineers reusable neural building blocks for models that are lighter, faster, composable and practical — from cloud GPUs to edge devices.
Architecture research · kernels · deployment
[ INSTALL / GET STARTED ]
pip install mlbricks-Studio[ WHO WE ARE ]
Our vision is to make
capable intelligence
dramatically more efficient.
AI should not need ever-growing compute, memory and infrastructure just to become more useful. MLBricks is building a modular intelligence stack where sequence processing, feed-forward computation, residual flow, precision and memory can each be redesigned as efficient components.
Instead of treating the neural network as one monolith, we treat it as a system of bricks — measurable, replaceable and optimizable independently.
RETHINK THE BLOCK
BENCHMARK THE CLAIM
SHIP THE KERNEL
DEPLOY ANYWHERE
[ MLBRICKS / MLB STUDIO ]
Build AI models visually.
Learn what every brick does.
MLBricks Studio turns model architecture into an interactive learning workspace. Students and young learners can create real models by arranging understandable components on a canvas instead of starting with a wall of code.
Start with a template or build from scratch. Add inputs, embeddings, ESA, feed-forward blocks, residual paths and outputs, inspect how every brick is connected, then change the architecture and immediately see how the model evolves.
VISUAL MODEL BUILDING•REAL MLBRICKS COMPONENTS•BUILD → TRAIN → GENERATE

Learners can move from individual components to a complete model they can inspect, modify, build, train and use for generation.
BUILD VISUALLY
Drag understandable components into a connected model graph and learn architecture by constructing it yourself.
INSPECT EVERY BRICK
Select a component to explore its role, settings, ports and connections so model internals become familiar over time.
BUILD → TRAIN → GENERATE
Move through the full workflow in one place and repeat experiments until the relationship between architecture and results becomes intuitive.
LEARN THE COMPONENTSMLB Studio makes the MLBricks library tangible — now explore the bricks behind the canvas.
Explore ESA, Bolt, SOUP & more↓[ ESA / THROUGHPUT SNAPSHOT ]
ESA is our proof that a core AI
building block can be rethought.
Entangled State Attention explores state-based sequence processing as an alternative path to standard token-to-token attention. The objective is simple: retain useful sequence capability while reducing practical overhead and opening new kernel-level optimization paths.
tokens / sec
ESA training · 64K contextMB
ESA reported peak · 64K contextbest ESA perplexity
16K contexttokens / sec
ESA average over 512, 1K, 4K, and 8KSelected internal / research benchmark results. Hardware, precision, sequence length and implementation materially affect performance.
Measure architecture at the point
where math becomes hardware.
We profile quality and throughput together, then optimize the implementation rather than relying on theory alone.
[ FEATURED PRODUCTS ]
Three core ideas.
Three different paths to efficiency.
ESA, Bolt and SOUP form the featured MLBricks sequence stack: state-based recurrence, optimized causal attention, and a state-memory architecture with built-in observation and fusion.
Entangled State Attention
A recurrent, state-based sequence mixer built around a compact evolving state. ESA supports full-sequence training and recurrent generation while keeping a consistentauto / native / pytorchexecution contract.
Bolt Attention
An optimized causal-attention brick for language models, exposed throughBoltandBoltAttention. Bolt keeps a PyTorch reference path alongside supported native acceleration and the MLBricks automatic execution planner.
SOUP Architecture
A state-and-memory architecture with built-in Observer State Memory and SOUP Fusion. It supports configurable mixers and FFNs, plus prepared recurrent generation with packed inference work for compatible paths.
[ MLBRICKS PRODUCT LIBRARY ]
The current stack,
from sequence to deployment.
The product library exposes each current MLBricks component independently — including StateAwareFFN, VirtualStateAwareFFN and MicroVirtualFFN as separate feed-forward bricks.
ESA
Recurrent state-based sequence processing for efficient full-sequence training and compact-state generation.
SEQUENCE · VIEW PRODUCT ↗Bolt
An optimized causal-attention brick with compact latent processing, cached decode support, and automatic selection between supported execution paths.
ATTENTION · VIEW PRODUCT ↗SOUP
A state-and-memory architecture with Observer State Memory, SOUP Fusion, configurable mixers and FFNs, and recurrent generation support.
ARCHITECTURE · MEMORY · VIEW PRODUCT ↗StateAwareFFN
A mixer-conditioned recurrent feature network that carries feature state through physical depth and reacts explicitly to mixer change.
COMPUTE · STATE-AWARE · VIEW PRODUCT ↗VirtualStateAwareFFN
A shared condition-aware state refiner with pass-specific identity, creating virtual computation depth without duplicating full physical FFN blocks.
COMPUTE · VIRTUAL DEPTH · VIEW PRODUCT ↗MicroVirtualFFN
Pass-specific narrow gated refinements that add virtual depth with a fraction of the parameter footprint of another full-width FFN.
COMPUTE · MICRO DEPTH · VIEW PRODUCT ↗ResController
An adaptive RMS-based residual controller that bounds update energy and residual-stream growth so deeper blocks retain room to influence the representation.
FLOW · RESIDUAL · VIEW PRODUCT ↗ElasticBit
Adaptive 4–32-bit matrix storage that measures the smallest precision meeting an error target, then executes through the next supported compute bucket.
PRECISION · STORAGE · VIEW PRODUCT ↗VESA
Visual computing powered by ESA across Serpentine, ViT, CNN, Diffusion and autoregressive visual engines.
VISION · ESA · VIEW PRODUCT ↗VisualBolt
Visual computing powered by Bolt across the same five MLBricks visual engines.
VISION · BOLT · VIEW PRODUCT ↗[ FUTURE VISION ]
From model components
to intelligence everywhere.
MLBricks is building toward a future where capable AI can run closer to where data, people and machines actually are.
Robotics
Persistent state, efficient perception and local reasoning for machines that continuously interact with the physical world.
ROBOTSMobile AI
Private, responsive intelligence that runs directly on phones instead of depending on a datacenter round trip.
MOBILEIoT & Edge
Compact intelligence for sensors, embedded systems and connected devices operating under strict power and memory limits.
IOTAutonomous Systems
Low-latency architectures for agents that must sense, update state and act continuously.
AUTONOMYAI PCs & OS
Always-available local assistants and intelligent system software that can reason without sending every task to the cloud.
LOCAL AIEfficient Cloud Intelligence
The same principle at datacenter scale: make every GPU cycle, memory transfer and model parameter do more useful work.
CLOUD[ BUILD THE NEXT BRICK ]
Intelligence should become
more capable — not
simply more expensive.
Research collaborations, commercial licensing, engineering partnerships and deployment work.
Start a conversation↗