Loading the catalog…
INTEL GAUDI 3 AI ACCELERATOR (Intel Gaudi)
Blurred values are subscriber-only. Sign in to start a free trial — 2 full pages a day — or subscribe for unlimited access.
| Memory Bandwidth Tb S | 3.7 |
|---|---|
| Process Node | TSMC 5nm process |
| Host Interface | PCIe Gen5 X16 |
| Tdp W | 900 |
| Memory Type | HBM2E |
| Gpu Memory Gb | 128 |
| Fp8 Tflops Dense | 1800 |
| Fp32 Tflops | 229 |
| Fp16 Bf16 Tflops Dense | 1678 |
| Product Name | |
| Tf32 Tflops Dense |
| SKU | INTEL GAUDI 3 AI ACCELERATOR |
|---|
| Document title | |
|---|---|
| Document version/date | |
| Architecture generation | |
| Compute dies |
| MME engines |
|---|
| TPC engines |
|---|
| RDMA NIC ports |
|---|
| HBM chips |
|---|
| Unified HBM capacity |
|---|
| FP8 and BF16 compute |
|---|
| HBM2e memory capacity |
|---|
| BF16 MME TFLOPS |
|---|
| FP8 MME TFLOPS |
|---|
| BF16 Vector TFLOPS |
|---|
| MME Units |
|---|
| TPC Units |
|---|
| On-die SRAM Capacity |
|---|
| On-die SRAM Bandwidth (read/write) |
|---|
| Networking (bidirectional) |
|---|
| Host Interface Peak BW |
|---|
| Media Decoders |
|---|
| Key functions: Compute Engines - MMEs |
|---|
| Key functions: Compute Engines - TPCs |
|---|
| Media Engines - Media Decoder Engines (DECs) |
|---|
| Media Engines - Rotator Engines (ROT) |
|---|
| Memory - L2 Cache |
|---|
| Memory - HBM2e |
|---|
| Networking - Host port |
|---|
| Networking - Network ports |
|---|
| Physical partitioning - DCOREs |
|---|
| Per-DCORE MMEs |
|---|
| Per-DCORE TPCs |
|---|
| Per-DCORE L2 Cache |
|---|
| PCIe total bandwidth |
|---|
| Peak TFLOP/Sec - MME (Matrix) FP8 |
|---|
| Peak TFLOP/Sec - MME (Matrix) BF16 |
|---|
| Peak TFLOP/Sec - MME (Matrix) FP16 (signed) |
|---|
| Peak TFLOP/Sec - MME (Matrix) TF32 |
|---|
| Peak TFLOP/Sec - MME (Matrix) FP32 |
|---|
| Peak TFLOP/Sec - TPC (Vector) FP8 |
|---|
| Peak TFLOP/Sec - TPC (Vector) BF16 |
|---|
| Peak TFLOP/Sec - TPC (Vector) FP16 |
|---|
| Peak TFLOP/Sec - TPC (Vector) FP32 |
|---|
| MME count |
|---|
| MME parallel operations |
|---|
| MACs per MME |
|---|
| Peak throughput per MME chip |
|---|
| MACs per cycle per MME |
|---|
| MME large unit input data per cycle |
|---|
| MME array organization |
|---|
| L2 cache (pipelining context) |
|---|
| MME supported datatypes |
|---|
| MME accumulator |
|---|
| TPC generation |
|---|
| TPC width |
|---|
| TPC supported floating datatypes |
|---|
| TPC supported integer datatypes |
|---|
| Media decoding units |
|---|
| Video format - HEVC |
|---|
| Video format - H.264/SVC/MVC |
|---|
| Video format - VP9 |
|---|
| Image format - JPEG |
|---|
| Image format - Progressive JPEG |
|---|
| Post processing - max down-scale output |
|---|
| Post processing - up-scaling |
|---|
| Post processing channels per decoder block |
|---|
| Decoder performance - HEVC |
|---|
| Decoder performance - VP9 |
|---|
| Decoder performance - H.264 |
|---|
| Decoder performance - Jpeg 420 |
|---|
| Decoder input stream format |
|---|
| Decoder output stream format |
|---|
| Rotator engine transformations |
|---|
| On-die SRAM size |
|---|
| On-die SRAM L2 split |
|---|
| HBM2e frequency |
|---|
| HBM peak bandwidth |
|---|
| HBM2e device capacity |
|---|
| HBM total capacity |
|---|
| PCIe |
|---|
| PCIe Peak BW |
|---|
| On-die-SRAM |
|---|
| On-die-SRAM BW |
|---|
| Cache L2 throughput |
|---|
| Cache L3 throughput |
|---|
| Cache capacity & set-associativity |
|---|
| HBM instances |
|---|
| HBM total capacity (memory subsystem) |
|---|
| HBM total bandwidth (memory subsystem) |
|---|
| NIC ports |
|---|
| Aggregation Engines |
|---|
| NIC aggregated bandwidth per direction |
|---|
| NIC RDMA protocol |
|---|
| Collective offload min buffer size |
|---|
| Congestion control schemes |
|---|
| In-network reduction operations |
|---|
| In-network reduction datatypes |
|---|
| Example cluster scale |
|---|
| Sub-cluster leaf switches |
|---|
| Cluster spine switches |
|---|
| Product line - generations |
|---|
| Copyright/legal code |
|---|