QuantaFONS
Talk to an architect

C.1 · Artificial Intelligence & Machine Learning

AI Training Frameworks

PyTorch and TensorFlow, tuned for the interconnect

  • PyTorch
  • TensorFlow
  • NCCL
  • RCCL
  • ZeRO-3
  • FP8
  • BF16
GPU accelerators lined up in a rack-mount server

What we build

Extensions and optimizations for PyTorch and TensorFlow. Includes custom NCCL/RCCL communicators for multi-GPU all-reduce, ZeRO-3 partitioning (model states, gradients, and optimizers), mixed-precision trainers (FP8/BF16), asynchronous checkpointing, and elastic training that dynamically adds/removes nodes mid-run.

Capabilities

  • Extensions and optimizations for PyTorch and TensorFlow
  • Custom NCCL and RCCL communicators for multi-GPU all-reduce
  • ZeRO-3 partitioning of model states, gradients, and optimizer states
  • Mixed-precision trainers in FP8 and BF16
  • Asynchronous checkpointing and elastic training that adds or removes nodes mid-run

Related services

How it connects

Where it sits in the stack.

This system, and the two it hands off to. None of them can be optimized alone.

01You are here

AI Training Frameworks

Extensions and optimizations for PyTorch and TensorFlow.

02

AI Inference Optimization

A serving engine built on TensorRT and ONNX Runtime with custom optimization passes.

AI & Machine Learning · see service
03

Interconnect Standards

Protocol stack implementations and bridge architectures for UCIe (die-to-die chiplet links), CXL 3.x (memory pooling and coherency), PCIe 6.0/7.0 (peripheral I/O), and 400G/800G Ethernet fabrics.

Semiconductor & Silicon · see service

Bring us the whole stack.

Tell us where latency is costing you, from the die to the data center to the control room. An architect replies with a first read of the problem, not a sales deck.