<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Ibrahim Ahmed</title><description>personal website and blog</description><link>https://ibrahimahmed.ca/</link><item><title>Fast RL using off-policy sampling</title><link>https://ibrahimahmed.ca/posts/fast-rl/</link><guid isPermaLink="true">https://ibrahimahmed.ca/posts/fast-rl/</guid><description>First open-source implementation of Soft Policy Optimization, an off-policy RL algorithm that works with LMs. This makes many RL experiments faster and cheaper.</description></item><item><title>Real Work</title><link>https://ibrahimahmed.ca/posts/real-work/</link><guid isPermaLink="true">https://ibrahimahmed.ca/posts/real-work/</guid><description>We are developing the first open-source LLM RL environment framework for real work.</description></item><item><title>How to Vibe Code Effectively</title><link>https://ibrahimahmed.ca/posts/vibe-coding/</link><guid isPermaLink="true">https://ibrahimahmed.ca/posts/vibe-coding/</guid><description>My intuitive and counterintuitive learnings</description></item><item><title>Proposal: Self-Refined RL (SRRL)</title><link>https://ibrahimahmed.ca/posts/self-refined-rl/</link><guid isPermaLink="true">https://ibrahimahmed.ca/posts/self-refined-rl/</guid><description>Policy gradient RL algorithms like GRPO have been used to improve LLMs&apos; performance on verifiable tasks like math and coding problems.</description></item><item><title>How scaling pretraining affects RL sample efficiency</title><link>https://ibrahimahmed.ca/posts/warm-start-rl/</link><guid isPermaLink="true">https://ibrahimahmed.ca/posts/warm-start-rl/</guid><description>Insights from a small-scale transformers experiment.</description></item><item><title>Bugs in LLM Benchmark Grading</title><link>https://ibrahimahmed.ca/posts/eval-qa/</link><guid isPermaLink="true">https://ibrahimahmed.ca/posts/eval-qa/</guid><description>I used Claude to audit the grading code of 8 major LLM benchmarks and found issues throughout all of them.</description></item><item><title>The path from Fable to superintelligence</title><link>https://ibrahimahmed.ca/posts/fable-to-superintelligence/</link><guid isPermaLink="true">https://ibrahimahmed.ca/posts/fable-to-superintelligence/</guid><description>How real-world feedback loops could turn capable AI agents into recursively improving systems.</description></item></channel></rss>