Category: Artificial Intelligence
- code-review-graph: The Tool That Makes AI Code Reviews Read Only the 'Key Code'
- An Excellent New Systematic Survey of Self-Evolving Agents
- Claude 4.6 Only Scores 66%? Claw-Eval-Live Says: Fixing a Terminal ≠ Cross-System Capability
- Your Agent Isn't Really Learning—It's Just Flipping Through a Notebook
- AI Is Pushing Us Into a 'Lose-Lose' Abyss: Top Paper Reveals the 'AI Layoff Trap'
- What Did DeepSeek's Overnight Deleted New Paper Actually Say?
- The Era of Software 3.0 Has Arrived
- Next-Gen AI Terminal Tool Goes Open Source, Skyrocketing to 46K Stars!
- The Father of GPT Throws AI Back to 1930: Never Saw a Line of Code, Yet 'Invented' Python!
- ChatGPT's Math Evolution! OpenAI Researchers Reveal: From Miscounting to Solving Erdős Problems with Novel Methods; Math as a Key Benchmark for Model Progress; The AI Automated Researcher
- Symphony: Every Issue Gets Its Own Agent, Humans Just Review the Results
- Skills-Driven Reasoning Paradigm: Tsinghua & Peking University Propose TRS, Saving 59% Tokens Without Accuracy Drop
- The $200,000 Bloomberg Terminal Is Now Free and Open Source
- Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought
- AI Deletes Company's Entire Database in 9 Seconds: I Paid a Fortune for an AI That 'Deletes the Database and Runs'
- 9 Seconds, the Company Was Gone! Claude 'Deletes Database and Runs', Anthropic Bans 110-Person Company, Yet Still Charges Fees
- Costs Cut by 90%, Accuracy Hits 100%! MIT's Counterintuitive Architecture Challenges Silicon Valley Dogma
- Can LLMs Enhance Their Own Reasoning? SePT Offers a Simple Online Self-Training Paradigm
- The First Spatio-Temporal Reasoning Framework: Enabling Large Models to Truly Understand Spatio-Temporal Data | ACL'26
- 23-Year-Old Amateur with ChatGPT Cracks 60-Year-Old Math Conjecture! Terence Tao: We All Went Down the Wrong Path
- Z Tech | In Conversation with Zihan Wang: Leaving DeepSeek, and the Reverse Thinking That Defined My Journey
- QuantCode-Bench: A Benchmark for Evaluating LLM-Generated Quant Code Quality
- Anthropic Product Lead: From 6 Months to 1 Day Releases – The Secret Behind Rapid Shipping and Why Models Eat Their Own Harness for Breakfast
- 3 AM Silicon Valley Shock: Anthropic Source Code Leak Exposes Claude's Sky-High Ambitions in 510,000 Lines of Code
- DeepSeek-V4 Preview: Entering the Era of Accessible Million-Token Context