Embedded AI
Model sizing: fitting AI into the device
A model that scores well on a laptop is of no use if it does not fit the device. We size models against the memory, compute and energy budget of the target before training starts.
· by BCF Embedded engineering team
What sets the size of a model
Flash
Weights and the inference code must fit next to the application and an OTA slot.
RAM
Activations and buffers peak during inference; the peak, not the average, decides.
Compute
Operations per inference set the latency on a given core or accelerator.
Energy
Energy per inference times inferences per day decides battery life.
Typical model classes
Indicative ranges after INT8 quantisation. The exact footprint depends on the architecture and runtime.
Tiny
- Model size
- Up to ~50k parameters
- Memory footprint
- Tens of KB flash, < 64 KB RAM
- Typical use
- Anomaly detection, simple classifiers, wake triggers
Small
- Model size
- ~50k – 1M parameters
- Memory footprint
- 100 KB – 1 MB flash, 64–512 KB RAM
- Typical use
- Keyword spotting, gesture and activity recognition
Medium
- Model size
- ~1M – 20M parameters
- Memory footprint
- Several MB, external RAM
- Typical use
- Image classification, low-resolution object detection
Large
- Model size
- Above ~20M parameters
- Memory footprint
- Hundreds of MB and more
- Typical use
- Multi-camera video, small language and vision-language models
How we fit a model to the device
- 01
Budget
Memory, latency and energy limits taken from the hardware and the product requirements.
- 02
Baseline
The simplest model that solves the task, often signal features plus a small classifier.
- 03
Compress
INT8 quantisation, pruning and knowledge distillation where they pay off.
- 04
Measure on target
Latency, peak RAM and accuracy profiled on the device, not only in simulation.
Smaller is often better
On a microcontroller, well-chosen signal features (spectra, statistics) with a compact network often match a large end-to-end model at a fraction of the memory. We compare both before committing to an architecture, and leave headroom for model updates delivered over the air.
FAQ
Discuss your project
Describe your device and goal. An engineer replies within one business day.