AMD Instinct
AMD's data-center accelerator line, paired with the open ROCm software stack. The most credible alternative to CUDA for large-scale training.
Visit AMD Instinct →Training and inference hardware is the most expensive decision in an AI stack and the hardest to reverse. This is a plain directory of who makes the silicon, who rents it by the hour, and who packages it into machines — grouped so you can compare like with like.
26 listings · 5 categories · updated September 2026Grouped by where the hardware actually sits: in a data centre, at the edge, or in someone else's cloud that you rent.
AMD's data-center accelerator line, paired with the open ROCm software stack. The most credible alternative to CUDA for large-scale training.
Visit AMD Instinct →Builds wafer-scale processors — a single enormous chip instead of many linked ones — to avoid the cost of moving data between devices.
Visit Cerebras →Google's tensor processing units, available only as cloud capacity. Tightly integrated with JAX and TensorFlow.
Visit Google Cloud TPU →Intelligence Processing Units built around large on-chip memory and fine-grained parallelism.
Visit Graphcore →Deterministic inference hardware built for very low latency token generation, offered both as silicon and as a hosted API.
Visit Groq →Intel's deep-learning accelerator family, positioned on price-performance for training and inference rather than peak specification.
Visit Intel Gaudi →The default AI training platform. Its advantage is as much CUDA and the surrounding software ecosystem as the silicon itself.
Visit NVIDIA →Reconfigurable dataflow architecture sold as a complete system rather than as loose accelerator cards.
Visit SambaNova →RISC-V based AI processors with an open software stack, sold as cards, systems, and licensable IP.
Visit Tenstorrent →In-memory computing accelerators for edge vision workloads, sold as modules and evaluation systems.
Visit Axelera AI →Graph-native edge processors and a low-code toolchain for building vision pipelines on them.
Visit Blaize →Energy-efficient edge inference processors and a compiler stack that targets them from standard frameworks.
Visit EdgeCortix →Korean designer of inference accelerators targeting efficient data-center and enterprise deployment.
Visit FuriosaAI →Small Edge TPU modules and boards for on-device inference, aimed at prototyping and low-power products.
Visit Google Coral →Edge AI processors for cameras, industrial equipment and vehicles, aimed at high inference throughput inside a small power budget.
Visit Hailo →Embedded modules that bring the CUDA ecosystem to robots, drones and industrial devices.
Visit NVIDIA Jetson →On-device AI across phones, PCs and automotive, with a focus on performance per watt in battery-powered products.
Visit Qualcomm AI →AI inference chip company focused on energy efficiency for data-center and enterprise inference.
Visit Rebellions →Specialised GPU cloud built for AI and rendering workloads rather than general-purpose computing.
Visit CoreWeave →GPU cloud and on-premise deep-learning workstations and servers from the same vendor.
Visit Lambda →On-demand and spot GPU containers aimed at developers who want to rent capacity by the minute.
Visit RunPod →GPU clusters and hosted open-model inference, with training and fine-tuning on the same infrastructure.
Visit Together AI →Marketplace that matches GPU renters with owners of idle capacity, usually at lower cost and more variable reliability.
Visit Vast.ai →Validated AI server and storage designs for enterprises that want a supported, pre-integrated stack.
Visit Dell AI Infrastructure →Builds the GPU server chassis and rack systems that most accelerator deployments physically live in.
Visit Supermicro →Photonic interconnect and computing, attacking the data-movement bottleneck rather than the arithmetic.
Visit Lightmatter →No matches. Try a different search.
Short, direct answers.
A GPU is a general parallel processor that happens to be excellent at the matrix maths neural networks need. A dedicated AI accelerator is purpose-built for that maths and usually trades flexibility for throughput or efficiency. GPUs win on ecosystem and software maturity; accelerators often win on performance per watt for a narrower set of models.
Rent while your workload is still changing shape, because the wrong purchase locks up capital in silicon that depreciates fast. Buy when utilisation is high and steady — sustained usage above roughly half of capacity is the point where owning usually beats renting, though the exact crossover depends on your power and hosting costs.
Running a trained model on a device near where the data is created — a camera, a vehicle, a factory sensor — instead of sending the data to a data centre. It lowers latency, cuts bandwidth cost, and keeps data local, which is often the deciding factor for privacy or regulation.
Large models spend much of their time moving weights rather than multiplying them. When a chip cannot feed its own compute units fast enough, headline throughput numbers are unreachable in practice. For large-model inference, memory capacity and bandwidth usually constrain real performance before the arithmetic units do.
No. Listings are editorial and unpaid. Advertising on the site is labelled as advertising and has no bearing on what gets listed or where it appears.