curius graph
☾
Dark
all pages
search
showing 15301-15350 of 200098 pages (sorted by popularity)
« prev
1
...
305
306
307
308
309
...
4002
next »
Thinking through how pretraining vs RL learn
2 users ▼
Almost anything you give sustained attention to will begin to loop on itself and bloom
2 users ▼
The bare-bones case for caring about global catastrophic risks
2 users ▼
The paradox is that when I accept myself just as I am, I change
2 users ▼
Your Goal Isn’t Really to Get a Job - by Matt Beard
2 users ▼
Benchmark Scores = General Capability + Claudiness
2 users ▼
Claude Opus 4.5 System Card
2 users ▼
How To Hide A Data Center - by Jason Hausenloy
2 users ▼
Text remains the loveliest medium - Tommy Dixon
2 users ▼
Buck's Shortform — LessWrong
2 users ▼
Should the US Sell Hopper Chips to China? | IFP
2 users ▼
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
2 users ▼
Simulacrum Levels — LessWrong
2 users ▼
The Sovereignty of Good
2 users ▼
[2512.15674] Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
2 users ▼
RL Environments and RL for Science: Data Foundries and Multi-Agent Architectures
2 users ▼
Legible vs. Illegible AI Safety Problems — LessWrong
2 users ▼
Omniscaling to MNIST — LessWrong
2 users ▼
Frontier Safety Framework Report - Gemini 3 Pro (November, 2025) v2
2 users ▼
Your Guide to Webflow’s $4B PLG Engine: How Webflow Pairs Self-Service and Sales for Rapid Growth
2 users ▼
Tiny Algae and the Political Theater of Planting One Trillion Trees
2 users ▼
Science is getting harder - by Matt Clancy
2 users ▼
Michael Nielsen on visualizations, biological systems, and making a new science
2 users ▼
Shtetl-Optimized » Blog Archive » My AI Safety Lecture for UT Effective Altruism
2 users ▼
innovative_contracting_case_studies_2014_-_august.pdf
2 users ▼
12 tentative ideas for US AI policy - Open Philanthropy
2 users ▼
COLM 2024
2 users ▼
Toward A Mathematical Framework for Computation in Superposition — LessWrong
2 users ▼
Distillation Walkthrough
2 users ▼
How to compute Hessian-vector products? | ICLR Blogposts 2024
2 users ▼
An update on our general capability evaluations - METR
2 users ▼
Going Soft, by Lily Scherlis
2 users ▼
CaMeL offers a promising new direction for mitigating prompt injection attacks
2 users ▼
Gemini Flash Pretraining
2 users ▼
Cost-Effective Constitutional Classifiers via Representation Re-use
2 users ▼
Ten AI safety projects I'd like people to work on
2 users ▼
Cyber Competitions
2 users ▼
Discussing Learned Concepts with Language Models · John Hewitt
2 users ▼
Proofs & Reasons @ CMU
2 users ▼
The Engineering State - American Affairs Journal
2 users ▼
Benchmark Scores = General Capability + Claudiness | Epoch AI
2 users ▼
Alignment remains a hard, unsolved problem — AI Alignment Forum
2 users ▼
Ideas Aren’t Getting Harder to Find—Asterisk
2 users ▼
New report: "Scheming AIs: Will AIs fake alignment during training in order to get power?" - Joe Carlsmith
2 users ▼
Against Almost Every Theory of Impact of Interpretability — LessWrong
2 users ▼
A starting point for making sense of task structure (in machine learning) — LessWrong
2 users ▼
DSLT 0. Distilling Singular Learning Theory — LessWrong
2 users ▼
Efficient Dictionary Learning with Switch Sparse Autoencoders — LessWrong
2 users ▼
Coercion is an adaptation to scarcity; trust is an adaptation to abundance — LessWrong
2 users ▼
Training a Reward Hacker Despite Perfect Labels — LessWrong
2 users ▼
« prev
1
...
305
306
307
308
309
...
4002
next »