Hacker News Logo

Offline

dayweek

Show HN: Minimal LLM Post-Training Experiments on an 8GB GPU (SFT, DPO, GRPO)

15 points|github.com|
popopanda|8hrs