PinnedRitvik Rastogi·18h agoStart Here: A Guide to Papers ExplainedA map to the Papers Explained series, organized by architecture, topic, and lab here’s how to find the one you need.”
Ritvik Rastogi·Jun 26Papers Explained 585: VibeThinker-3BVibeThinker-3B is a compact dense model with 3B parameters, developed to investigate how far verifiable reasoning can be pushed within a…
Ritvik Rastogi·Jun 25Papers Explained 584: VibeThinker-1.5BVibeThinker-1.5B, a 1.5B-parameter dense model developed using an innovative post-training methodology centered on the “Spectrum-to-Signal…
Ritvik Rastogi·Jun 24Papers Explained 583: Mellum2Mellum 2 is an open-weight 12B-parameter Mixture-of-Experts (MoE) language model with 2.5B active parameters per token, specialized in…
Ritvik Rastogi·Jun 23Papers Explained 582: MellumMellum is a family of open-weight, 4B-parameter code completion models designed by JetBrains for interactive use in IDEs, specifically…
Ritvik Rastogi·Jun 22Papers Explained 581: Rubric Guided Self DistillationRubric-Guided Self-Distillation (RGSD) is a verifier-free training method in which the base policy, conditioned on the rubric, serves as…
Ritvik Rastogi·Jun 19Papers Explained 580: Nemotron 3 UltraNemotron 3 Ultra is a 550B total and 55B active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. It is pretrained on 20…
Ritvik Rastogi·Jun 18Papers Explained 579: Policy-Aware Rubric Reward (POW3R)Rubric-based reinforcement learning for language models often aggregates multi-criteria rewards using static human-assigned weights, but…
Ritvik Rastogi·Jun 17Papers Explained 578: Reward Hacking in Rubric-Based RLThe study analyzes reward hacking in rubric-based RL, where policies are optimized against a training verifier but evaluated by stronger…
Ritvik Rastogi·Jun 16Papers Explained 577: MAI-Thinking-1MAI-Thinking-1 is a large reasoning model with 35B active/1T parameters, trained from scratch, on 30T tokens of exclusively clean…