Embedded AI
Hardware targeting: matching the model to the silicon
The same model can run ten times faster or slower depending on the core, the accelerator and the toolchain. We pick the target with the model in mind and optimise for it.
· by BCF Embedded engineering team
Target classes
Microcontrollers
Cortex-M4/M7/M33 and ESP32-S3 with optimised kernels such as CMSIS-NN. Lowest power.
MCUs with an NPU
Cortex-M55 with Ethos-U, STM32N6 and similar: much faster inference at a microcontroller power budget.
DSP and vector extensions
Arm Helium and audio DSPs for signal processing and audio models.
Application processors
Embedded Linux SoCs, often with an NPU, for vision and larger models.
Edge GPUs
NVIDIA Jetson class for multi-camera video and the largest edge models.
FPGA and SoC FPGA
Custom datapaths with deterministic latency, written in VHDL / Verilog or HLS.
How we target hardware
- 01
Requirements
Accuracy, latency, power, unit cost and long-term availability.
- 02
Benchmark
Candidate models run on evaluation boards of the shortlisted platforms.
- 03
Toolchain
Conversion and compilation with the runtime and vendor tools for that target.
- 04
Integration
Memory placement, DMA, RTOS scheduling and low-power modes in the firmware.
Toolchains and platforms
Runtimes
- TensorFlow Lite Micro / LiteRT
- ONNX Runtime
- Edge Impulse
Compilers and libraries
- ST Edge AI (STM32Cube.AI)
- Arm Vela
- CMSIS-NN / CMSIS-DSP
- Apache TVM
- TensorRT
Platforms
- STM32
- Nordic nRF
- ESP32
- Embedded Linux
Choose the board after the benchmark
The cheapest moment to change the processor is before the PCB exists. We run a feasibility benchmark first and, when a custom board is needed, design it with our hardware team, including power, thermal and memory for future model versions.
FAQ
Discuss your project
Describe your device and goal. An engineer replies within one business day.