Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
Escape the Sequential Training Trap: 16x Higher Throughput for LLM Experimentation
Learn how RapidFire AI parallelizes LLM fine‑tuning by chunking data, sharing memory, dynamically managing configs, and orchestrating multi‑GPU FSDP automatically.
I’ll be presenting RapidFire AI, a new open-source framework that transforms LLM fine-tuning and post-training from sequential one-config-at-a-time training into hyperparallelized experimentation with dynamic real-time experiment control and automatic multi-GPU orchestration.
The core innovation is an adaptive execution engine that allows for multiple configs to be compared on even a single GPU by automatically chunking the data into subsets and cycling configs across them via a new shared memory subsystem. It enables “Interactive Control Operations” - dynamic modification of running experiments. Stop underperforming configurations, clone promising ones, and warm-start variants from parent checkpoints. The RapidFire AI scheduler intelligently manages multi-GPU orchestration to optimize GPU utilization and uses FSDP automatically for sharding large models across GPUs. The framework supports multiple popular LLM customization workflows from Hugging Face TRL, including SFT, DPO, and GRPO.
Adaptive engine enables hyperparallel LLM fine-tuning with real-time IC Ops control.
Compose Email
Loading recent emails...