Distributed Training & Inference: From CPUs and GPUs to a Cluster
An illustrated guide to distributed training: CPUs vs GPUs, CUDA, data, tensor and pipeline parallelism, FSDP vs DeepSpeed ZeRO, Ray, and inference.
Sep 29, 202617 min read

Search for a command to run...
Articles tagged with #distributed-systems
An illustrated guide to distributed training: CPUs vs GPUs, CUDA, data, tensor and pipeline parallelism, FSDP vs DeepSpeed ZeRO, Ray, and inference.

How async GRPO decouples rollout generation, CPU gym execution, and policy updates to remove 80% GPU idle time in environment-heavy agent RL.
