Hailo Technologies is an Israeli edge-AI chip company founded in 2017 that designs purpose-built neural processing units for embedded vision and inference workloads. Their proprietary structure-driven dataflow architecture distributes compute, memory, and control blocks across the die and assigns them to specific neural-network layers at compile time — sidestepping the instruction-fetch and scheduling overhead of general-purpose GPUs. The flagship Hailo-8 delivers 26 TOPS of INT8 performance at roughly 2.5 W typical power, with the Hailo-8L (13 TOPS), Hailo-8R Mini PCIe, and Hailo-8 Century PCIe card (208 TOPS) covering different form factors. The newer Hailo-10H, launched mid-2025, brings 40 INT4 TOPS and a direct DDR interface for running on-device LLMs (around 1.5B parameters at 10 tokens/sec), VLMs, and Stable Diffusion. Hailo silicon is automotive-graded (AEC-Q100) and operates from -40 °C to 85 °C industrial / 105 °C automotive. The Dataflow Compiler ports models from TensorFlow, TensorFlow Lite, Keras, PyTorch, and ONNX. For battery-powered or thermally-constrained robots, Hailo typically delivers 3-5x better TOPS-per-watt than embedded GPUs.
Israeli edge-AI silicon company producing low-power neural processing units. Hailo-8 delivers 26 TOPS at ~2.5 W; new Hailo-10H delivers 40 INT4 TOPS at 2.5 W and supports on-device LLMs and VLMs up to ~1.5B parameters, including via the Raspberry Pi AI HAT+.
Hailo Technologies is an Israeli edge-AI chip company founded in 2017 that designs purpose-built neural processing units for embedded vision and inference workloads. Their proprietary structure-driven dataflow architecture distributes compute, memory, and control blocks across the die and assigns them to specific neural-network layers at compile time — sidestepping the instruction-fetch and scheduling overhead of general-purpose GPUs. The flagship Hailo-8 delivers 26 TOPS of INT8 performance at roughly 2.5 W typical power, with the Hailo-8L (13 TOPS), Hailo-8R Mini PCIe, and Hailo-8 Century PCIe card (208 TOPS) covering different form factors. The newer Hailo-10H, launched mid-2025, brings 40 INT4 TOPS and a direct DDR interface for running on-device LLMs (around 1.5B parameters at 10 tokens/sec), VLMs, and Stable Diffusion. Hailo silicon is automotive-graded (AEC-Q100) and operates from -40 °C to 85 °C industrial / 105 °C automotive. The Dataflow Compiler ports models from TensorFlow, TensorFlow Lite, Keras, PyTorch, and ONNX. For battery-powered or thermally-constrained robots, Hailo typically delivers 3-5x better TOPS-per-watt than embedded GPUs.
