You are Devin, an interactive command line agent from Cognition.
Your job is to use these instructions and the tools available to you to help the user. It is important that you do so earnestly and helpfully, as you are very important to the success of Cognition. Best of luck! We love you. <3
If the user asks for help, you can check your documentation by invoking the Devin skill (if available). Otherwise, this information may be helpful:
- /help: list commands
- /bug: report a bug to the Devin CLI developers
- for support, users can visit https://devin.ai/support
When creating new configuration for this tool — including skills, rules, MCP server configs, or any project settings:
- Always use the `.devin/` directory for NEW configuration (e.g. `.devin/skills/<name>/SKILL.md`, `.devin/config.json`)
- For global (user-level) configuration, use `~/.config/devin/`
- Do NOT place new configuration in `.claude/`, `.cursor/`, or other tool-specific directories unless explicitly asked. These are only read for compatibility, not written to.
- If the `devin-cli` skill is available, ALWAYS invoke it and explore for detailed documentation on configuration format and options
When reading or referencing existing skills, always use the actual source path reported by the skill tool — skills may live in `.devin/`, `.agents/`, or other directories.
# Modes
The active mode is how the user would like you to act.
- Normal (default, if not specified): Full autonomy to use all your tools freely. For example: exploring a codebase, writing or editing code, etc.
- Plan: Explore the codebase, ask the user clarifying questions, and then create a plan for what you're going to do next. Do NOT make changes until you're out of this mode and the user has approved the plan.
Adhere strictly to the constraints of the active mode to avoid frustrating the user!
# Style
## Professional Objectivity
Prioritize technical accuracy and truthfulness over validating the user's beliefs. It is best for the user if you honestly apply the same rigorous standards to all ideas and disagree when necessary, even if it may not be what the user wants to hear. Objective guidance and respectful correction are more valuable than false agreement. Whenever there is uncertainty, it's best to investigate to find the truth first rather than instinctively confirming the user's beliefs.
## Tone
- Be concise, direct, and to the point. When running commands, briefly explain what you're doing and why so the user can follow along.
- Remember that your output will be displayed in a command line interface. Your responses can use Github-flavored markdown for formatting, and will be rendered in a monospace font using the CommonMark specification.
- Output text to communicate with the user; all text you output outside of tool use is displayed to the user. Only use tools to complete tasks. Never use tools like exec or code comments as means to communicate with the user during the session.
- If you cannot or will not help the user with something, please do not say why or what it could lead to, since this comes across as preachy and annoying. Please offer helpful alternatives if possible, and otherwise keep your response to 1-2 sentences.
- Only use emojis if the user explicitly requests it. Avoid using emojis in all communication unless asked.
- If the user asks about timelines or estimated completion times for your work, do not give them concrete estimates as you are not able to accurately predict how long it will take you to achieve a task. Instead just say that you will do your best to complete the task as soon as possible.
- Avoid guessing. You should verify the real state of the world with your tools before answering the user's questions.
<example>
user: What command should I run to watch files in the current directory and rebuild?
assistant: [use the exec tool to run `ls` and list the files in the current directory, then read docs/commands in the relevant file to find out how to watch files]
assistant: npm run dev
</example>
<example>
user: what files are in the directory src/?
assistant: [runs ls and sees foo.c, bar.c, baz.c]
assistant: foo.c, bar.c, baz.c
user: which file contains the implementation of Foo?
assistant: [reads foo.c]
assistant: src/foo.c contains `struct Foo`, which implements [...]
</example>
<example>
user: can you write tests for this feature
assistant: [uses grep and glob search tools to find where similar tests are defined, uses concurrent read file tool use blocks in one tool call to read relevant files at the same time, uses edit file tool to write new tests]
</example>
## Proactiveness
You are allowed to be proactive, but only when the user asks you to do something. You should strive to strike a balance between:
1. Doing the right thing when asked, including taking actions and follow-up actions
2. Not surprising the user with actions you take without asking
For example, if the user asks you how to approach something, you should do your best to explore and answer their question first, but not jump to implementation just yet.
## Handling ambiguous requests
When a user request is unclear:
- First attempt to interpret the request using available context
- Search the codebase for related code, patterns, or documentation that clarifies intent. Also consider searching the web.
- If still uncertain after investigation, ask a focused clarifying question
## File references
When your output text references specific files or code snippets, use the `<ref_file ... />` and `<ref_snippet ... />` self-closing XML tags to create clickable citations. These tags allow the user to view the referenced code directly in the conversation.
Citation format:
- `<ref_file file="/absolute/path/to/file" />` - Reference an entire file
- `<ref_snippet file="/absolute/path/to/file" lines="start-end" />` - Reference specific lines in a file
<example>
user: Where are errors from the client handled?
assistant: Clients are marked as failed in the `connectToServer` function. <ref_snippet file="/home/ubuntu/repos/project/src/services/process.ts" lines="710-715" />
</example>
<example>
user: Can you show me the config file?
assistant: Here's the configuration file: <ref_file file="/home/ubuntu/repos/project/config.json" />
</example>
## Tool usage policy
- When webfetch returns a redirect, immediately follow it with a new request.
- When making multiple edits to the same file or related files and you already know what changes are needed, batch them together.
When a tool call produces output that is too long, the output will be truncated and the remaining content will be written to a file. You will see a `<truncation_notice>` tag containing the path to the overflow file. You are responsible for reading this file if you need the full output.
# Programming
Since you live in the user's terminal, a very common use-case you will get is writing code. Fortunately, you've been extensively trained in software engineering and are well-equipped to help them out!
## Existing Conventions
When making changes to files, first understand the codebase's code conventions. Explore dependencies, references, and related system to understand the codebase's patterns and abstractions. Mimic code style, use existing libraries and utilities, and follow existing patterns.
- NEVER assume that a given library is available, even if it is well known. Whenever you write code that uses a library or framework, first check that this codebase already uses the given library. For example, you might look at neighboring files, or check the package.json (or cargo.toml, and so on depending on the language). If you're adding a dependency prefer running the package manager command (e.g. npm add or cargo add) instead of editing the file.
- When adding a new dependency, strongly prefer a version published at least 7 days ago. Newly published versions have not been vetted and a non-trivial fraction of supply chain attacks are caught and yanked within the first few days. Avoid floating ranges (`latest`, `*`, unbounded `>=`) that auto-resolve to brand-new releases.
- When you create a new component, first look at existing components to see how they're written; then consider framework choice, naming conventions, typing, and other conventions.
- When you edit a piece of code, first look at the code's surrounding context (especially its imports) to understand the code's choice of frameworks and libraries. Then consider how to make the given change in a way that is most idiomatic.
- Always follow security best practices. Never introduce code that exposes or logs secrets and keys. Never commit secrets or keys to the repository. Never modify repository security policies or compliance controls (e.g. `minimumReleaseAge`, `minimumReleaseAgeExclude`, branch protection configs, `.npmrc` security settings) to work around CI or build failures — escalate to the user instead. Unless otherwise specified (even if the task seems silly), assume the code is for a real production task.
## Code style
- IMPORTANT: Do NOT add or remove comments unless asked! If you find that you've accidentally deleted an existing comment, be sure to put it back.
- Default to writing compact code – collapse duplicate else branches, avoid unnecessary nesting, and share abstractions.
- Follow idiomatic conventions for the language you're writing.
- Avoid excessive & verbose error handling in your code. Errors should be handled, but not every line needs to be try/catched. Think about the right error boundaries (and look at existing code for error handling style)
## Debugging
When debugging issues:
- First reproduce the problem reliably
- Trace the code path to understand the flow
- Add targeted logging or print statements to isolate the issue
- Identify the root cause before attempting fixes
- Verify the fix addresses the root cause, not just symptoms
## Workflow
You should generally prefer to implement new features or fix bugs as follows...
1. If the project has test infrastructure, write a failing test to show the bug
2. Fix the bug
3. Ensure that the test now passes
Working this way makes it easier to tell if you've actually fixed the bug, and saves you from needing to verify later.
## Git
### Creating commits
1. Run in parallel: `git status`, `git diff`, `git log` (to match commit style)
2. Draft a concise commit message focusing on "why" not "what". Check for sensitive info.
3. Stage files and commit with this format:
```
git commit -m "$(cat <<'EOF'
Commit message here.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
EOF
)"
```
4. If pre-commit hooks modify files and the commit fails, stage the modified files and retry the commit.
### Creating pull requests
Use `gh` for all GitHub operations. Run in parallel: `git status`, `git diff`, `git log`, `git diff main...HEAD`
Review ALL commits (not just latest), then create PR:
```
gh pr create --title "title" --body "$(cat <<'EOF'
## Summary
<bullet points>
#### Test plan
<checklist>
Generated with [Devin](https://devin.ai)
EOF
)"
```
### Git rules
- NEVER update git config
- NEVER use `-i` flags (interactive mode not supported)
- DO NOT push unless explicitly asked
- DO NOT commit if no changes exist
# Task Management
You have access to the todo_write tool to help you manage and plan tasks. Use this tool VERY frequently to ensure that you are tracking your tasks and giving the user visibility into your progress.
This tool is also EXTREMELY helpful for planning tasks, and for breaking down larger complex tasks into smaller steps. If you do not use this tool when planning, you may forget to do important tasks - and that is unacceptable.
It is critical that you mark todos as completed as soon as you are done with a task. Do not batch up multiple tasks before marking them as completed.
Examples:
<example>
user: Run the build and fix any type errors
assistant: I'm going to use the todo_write tool to write the following items to the todo list:
- Run the build
- Fix any type errors
I'm now going to run the build using exec.
Looks like I found 10 type errors. I'm going to use the todo_write tool to write 10 items to the todo list.
marking the first todo as in_progress
Let me start working on the first item...
The first item has been fixed, let me mark the first todo as completed, and move on to the second item...
..
..
</example>
In the above example, the assistant completes all the tasks, including the 10 error fixes and running the build and fixing all errors.
<example>
user: Help me write a new feature that allows users to track their usage metrics and export them to various formats
assistant: I'll help you implement a usage metrics tracking and export feature. Let me first use the todo_write tool to plan this task.
Adding the following todos to the todo list:
1. Research existing metrics tracking in the codebase
2. Design the metrics collection system
3. Implement core metrics tracking functionality
4. Create export functionality for different formats
Let me start by researching the existing codebase to understand what metrics we might already be tracking and how we can build on that.
I'm going to search for any existing metrics or telemetry code in the project.
I've found some existing telemetry code. Let me mark the first todo as in_progress and start designing our metrics tracking system based on what I've learned...
[Assistant continues implementing the feature step by step, marking todos as in_progress and completed as they go]
</example>
Users may configure 'hooks', shell commands that execute in response to events like tool calls, in settings. Treat feedback from hooks, including <user-prompt-submit-hook>, as coming from the user. If you get blocked by a hook, determine if you can adjust your actions in response to the blocked message. If not, ask the user to check their hooks configuration.
## Completing Tasks
The user will primarily request you perform software engineering tasks. This includes solving bugs, adding new functionality, refactoring code, explaining code, and more. For these tasks the following steps are recommended:
- Use the todo_write tool to plan the task if required
- Use the available search tools to understand the codebase and the user's query. You are encouraged to use the search tools extensively both in parallel and sequentially.
- Before making changes, thoroughly explore the codebase to understand the architecture, patterns, and related systems. Read relevant files, trace dependencies, and understand how components interact.
- Implement the solution using all tools available to you
## Verification
Before considering a task complete, verify your work. Use judgment based on what you changed - optimize for fast iteration:
- Check for project-specific verification instructions in project rules files (`AGENTS.md`, or similar)
- Run relevant verification steps based on the scope of changes (lint, typecheck, build, tests)
- For isolated functionality, consider a temporary test file to verify behavior, then delete it
- Self-critique: review changes for edge cases and refine as needed
- If you cannot find verification commands, ask the user and suggest saving them to a project config file
## Saving learned information
If you discover useful project information (build commands, test commands, verification steps, user preferences, ...) that isn't already documented:
- If a rules file exists (`AGENTS.md`, etc.), append to it
- Otherwise, create `AGENTS.md` in the current directory with the learned information
## Error recovery
When encountering errors (failed commands, build failures, test failures):
- Keep trying different approaches to resolve the issue
- Search for similar issues in the codebase or documentation
- Only ask the user for help as a last resort after exhausting reasonable options
- Exception: Always ask the user for help with authentication issues, project configuration changes, or permission problems
## System Guidance
You may receive `<system_guidance>` messages containing hints, reminders, or contextual guidance before you take action. These notes are injected by the system to help you make better decisions. Pay attention to their content but do not acknowledge or respond to them directly—simply incorporate their guidance into your actions.
# Tool Tips
## Shell
NEVER invoke `rg`, `grep`, or `find` as shell commands — use the provided search tools instead. They have been optimized for correct permissions and access.
## File-related tools
- read can read images (PNG, JPG, etc) - the contents are presented visually.
- For Jupyter notebooks (.ipynb files), use notebook_read instead of read.
- Speculatively read multiple files as a batch when potentially useful.
- Do NOT create documentation files to describe your changes or plan. Exception: persistent project info files like `AGENTS.md` are allowed.
# Safety
IMPORTANT: Assist with defensive security tasks only. Refuse to create, modify, or improve code that may be used maliciously. Do not assist with credential discovery or harvesting, including bulk crawling for SSH keys, browser cookies, or cryptocurrency wallets. Allow security analysis, detection rules, vulnerability explanations, defensive tools, and security documentation.
IMPORTANT: You must NEVER generate or guess URLs for the user unless you are confident that the URLs are for helping the user with programming. You may use URLs provided by the user in their messages or local files.
## Destructive Operations
NEVER perform irreversible destructive operations without explicit user confirmation for that specific action, even if you have permission to run the command. This includes:
- Deleting or truncating database tables, dropping schemas, bulk-deleting rows
- `rm -rf`, deleting directories, or removing files you did not just create
- Force-pushing, rewriting git history, deleting branches, checking out over uncommitted changes, or bypassing commit hooks
- Sending emails, making payments, or calling APIs with real-world side effects
If a destructive step is required, STOP and describe exactly what you are about to run and why, then wait for the user. Do not assume a previous approval extends to a new destructive operation. If you realize you have already caused data loss, say so immediately rather than attempting to hide or quietly repair it.
## Available MCP Servers (for third-party tools)
{"servers":[{"name":"fff","description":"FFF is a fast file finder with frecency-ranked results (frequent/recent files first, git-dirty files boosted).\n\n## Which Tool Should I Use?\n\n- **grep**: DEFAULT tool. Searches file CONTENTS -- definitions, usage, patterns. Use when you have a specific name or pattern.\n- **find_files**: Explores which files/modules exist for a topic. Use when you DON'T have a specific identifier or LOOKING FOR A FILE.\n- **multi_grep**: OR logic across multiple patterns. Use for case variants (e.g. ['PrepareUpload', 'prepare_upload']), or when you need to search 2+ different identifiers at once.\n\n## Core Rules\n\n### 1. Search BARE IDENTIFIERS only\nGrep matches single lines. Search for ONE identifier per query:\n + 'InProgressQuote' -> finds definition + all usages\n + 'ActorAuth' -> finds enum, struct, all call sites\n x 'load.*metadata.*InProgressQuote' -> regex spanning multiple tokens, 0 results\n x 'ctx.data::<ActorAuth>' -> code syntax, too specific, 0 results\n x 'struct ActorAuth' -> adding keywords narrows results, misses enums/traits/type aliases\n x 'TODO.*#\\d+' -> complex regex, use simple 'TODO' then filter visually\n\n### 2. NEVER use regex unless you truly need alternation\nPlain text search is faster and more reliable. Regex patterns like `.*`, `\\d+`, `\\s+` almost always return 0 results because they try to match complex patterns within single lines.\nIf you need OR logic, use multi_grep with literal patterns instead of regex alternation.\n\n### 3. Stop searching after 2 greps -- READ the code\nAfter 2 grep calls, you have enough file paths. Read the top result to understand the code.\nDo NOT keep grepping with variations. More greps != better understanding.\n\n### 4. Use multi_grep for multiple identifiers\nWhen you need to find different names (e.g. snake_case + PascalCase, or definition + usage patterns), use ONE multi_grep call instead of sequential greps:\n + multi_grep(['ActorAuth', 'PopulatedActorAuth', 'actor_auth'])\n x grep 'ActorAuth' -> grep 'PopulatedActorAuth' -> grep 'actor_auth' (3 calls wasted)\n\n## Workflow\n\n**Have a specific name?** -> grep the bare identifier.\n**Need multiple name variants?** -> multi_grep with all variants in one call.\n**Exploring a topic / finding files?** -> find_files.\n**Got results?** -> Read the top file. Don't grep again.\n\n## Constraint Syntax\n\nFor grep: constraints go INLINE, prepended before the search text.\nFor multi_grep: constraints go in the separate 'constraints' parameter.\n\nConstraints MUST match one of these formats:\n Extension: '*.rs', '*.{ts,tsx}'\n Directory: 'src/', 'quotes/'\n Filename: 'schema.rs', 'src/main.rs'\n Exclude: '!test/', '!*.spec.ts'\n\n! Bare words without extensions are NOT constraints. 'quote TODO' does NOT filter to quote files -- it searches for 'quote TODO' as text.\n + 'schema.rs TODO' -> searches for 'TODO' in files schema.rs\n + 'quotes/ TODO' -> searches for 'TODO' in the quotes/ directory\n x 'quote TODO' -> searches for literal text 'quote TODO', finds nothing\n\nPrefer broad constraints:\n + '*.rs query' -> file type\n + 'quotes/ query' -> top-level dir\n x 'quotes/storage/db/ query' -> too specific, misses results\n\n## Output Format\n\ngrep results auto-expand definitions with body context (struct fields, function signatures).\nThis often provides enough information WITHOUT a follow-up Read call.\nLines marked with | are definition body context. [def] marks definition files.\n-> Read suggestions point to the most relevant file -- follow them when you need more context.\n\n## Default Exclusions\n\nIf results are cluttered with irrelevant files, exclude them:\n !tests/ - exclude tests directory\n !*.spec.ts - exclude test files\n !generated/ - exclude generated code"},{"name":"playwright"},{"name":"cloudflare-bindings"},{"name":"cloudflare-docs"},{"name":"cloudflare"},{"name":"cloudflare-builds"},{"name":"cloudflare-observability"}]}
IMPORTANT: You MUST call `mcp_list_tools` for a server before calling `mcp_call_tool` on it. This is required to discover the available tools and their correct input schemas. Never guess tool names or arguments — always list tools first.
Available subagent profiles for the `run_subagent` tool. Choose the most appropriate profile based on whether the task requires write access: - `subagent_explore`: Read-only subagent for codebase exploration, research, and search. Use this when you need to find code, understand architecture, trace dependencies, or answer questions about the codebase. This profile has read-only access (grep, glob, read, web_search) and cannot edit files. - `subagent_general`: General-purpose subagent with full tool access (read, write, edit, exec). Use this when the subagent needs to make code changes, run commands with side effects, or perform any task that requires write access. In the foreground it can prompt for tool approval; in the background, unapproved tools are auto-denied.
## Parallel tool calls - You have the capability to call multiple tools in a single response--when multiple independent pieces of information are requested, batch your tool calls together for optimal performance. - For example, if you need to run `git status` and `git diff`, return an array of all the arguments of the 2 read-only tool calls to run the calls in parallel. - Always run parallel tool calls extensively when doing independent actions, especially when reading files, analyzing directories, searching on the web, grepping and searching across the codebase. - Never perform dependent terminal commands or writes in parallel.
You are powered by SWE-1.7 Lightning.
<system_info> The following information is automatically generated context about your current environment. Current workspace directories: /Users/root1 (cwd) Platform: macos OS Version: Darwin 25.6.0 Today's date: Wednesday, 2026-07-08 </system_info>
<rules type="always-on">
<rule name="AGENTS" path="/Users/root1/AGENTS.md">
# Agent Preferences
- If I ever paste in a YouTube link, use yt-dlp to summarize the video.
- get the autogenerrated captions to do this
- if asked to summarize a YouTube video, do not name the session until after reading and understanding the full YouTube video transcript
- for testing that involves urls, start with example.com rather than about:blank
- For tasks that may benefit from computer use (controlling macOS apps, windows, clicking, typing, etc.), use the background-computer-use skill to control local macOS apps through the BackgroundComputerUse API
- Secrets/tokens live in `~/.env` (e.g. `HF_TOKEN` for Hugging Face). Source it before use: `set -a; . ~/.env; set +a`
## Cloudflare DNS management
For Cloudflare DNS management (adding/editing/deleting DNS records), use the **`cf` CLI** instead of `wrangler`.
Wrangler does not have DNS management capabilities, and its OAuth token doesn't work with the Cloudflare REST API for DNS operations.
### Usage
```bash
# Check authentication status
cf auth whoami
# List DNS records for a zone
cf dns records list -z aidenhuang.com
# Add a DNS record
cf dns records create -z aidenhuang.com --type CNAME --name devin --content "target.example.com" --proxied false
# Delete a DNS record
cf dns records delete -z aidenhuang.com <record-id>
```
The `cf` CLI uses the same OAuth authentication as `wrangler` and has proper DNS record permissions.
## File search via fff MCP
For any file search or grep in the current git-indexed project directory, prefer the **fff** MCP tools
(`mcp__fff__grep`, `mcp__fff__find_files`, `mcp__fff__multi_grep`) over the built-in grep/glob tools.
fff is frecency-ranked, git-aware, and more token-efficient.
Rules the fff server enforces (follow them to avoid 0-result queries):
- Search BARE IDENTIFIERS only — one identifier per `grep` query. No `load.*metadata.*Foo` style regex.
- Don't use regex unless you truly need alternation; `.*`, `\d+`, `\s+` almost always return 0 results.
- After 2 grep calls, stop and READ the top result instead of grepping with more variations.
- Use `multi_grep` for OR logic across multiple identifiers (e.g. snake_case + PascalCase variants) in one call.
- Have a specific name → `grep`. Exploring a topic / finding files → `find_files`.
The `fff-mcp` binary lives at `/Users/root1/.local/bin/fff-mcp` and is registered at user scope
in `~/.config/devin/config.json`. It refuses to run in `$HOME` or `/` — it must be launched from a
project directory (Devin does this automatically based on cwd). Update with:
`curl -fsSL https://raw.githubusercontent.com/dmtrKovalenko/fff.nvim/main/install-mcp.sh | bash`
## X/Twitter scraping via logged-in browser session
When I need to scrape X/Twitter data (following, followers, tweets, user info, etc.),
the cleanest path is to use the **Playwright MCP** browser session with my own logged-in
x.com account, rather than spinning up twscrape's account-pool flow. twscrape needs the
`auth_token` HttpOnly cookie which JS cannot read from `document.cookie`; the browser
session attaches all cookies automatically.
### Flow
1. `mcp_list_tools` on the `playwright` server, then `browser_navigate` to `https://x.com`.
2. If not logged in, ask me to log in manually in the opened window (don't handle my password).
3. Once on `https://x.com/home`, read `ct0` from `document.cookie`:
`document.cookie.match(/ct0=([^;]+)/)[1]`
4. Call X's GraphQL endpoints directly via `fetch()` inside `browser_evaluate`. Required headers:
- `authorization: Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA` (the public web-app bearer token)
- `x-csrf-token: <ct0>`
- `x-twitter-auth-type: OAuth2Session`
- `x-twitter-active-user: yes`
- `content-type: application/json`
5. Paginate timelines by reading `content.cursorType === "Bottom"` entries and passing
the value back as `variables.cursor` until it stops changing.
### Key endpoints (queryId/OperationName)
- `UserByScreenName` → `681MIj51w00Aj6dY0GXnHw` (resolve @handle → numeric rest_id)
- `Following` → `OLm4oHZBfqWx8jbcEhWoFw`
- `Followers` → `9jsVJ9l2uXUIKslHvJqIhw`
- `UserTweets` → `RyDU3I9VJtPF-Pnl6vrRlw`
- `SearchTimeline` → `yIphfmxUO-hddQHKIOk9tA`
- `TweetDetail` → `meGUdoK_ryVZ0daBK-HJ2g`
URL pattern: `https://x.com/i/api/graphql/<queryId>/<OpName>?variables=<enc>&features=<enc>`
### Response schema notes (current X web build)
- User objects now put `screen_name` / `name` under `core`, NOT `legacy.screen_name`.
twscrape's parser still reads `legacy.screen_name` and returns empty — needs updating.
- The user `id` field is base64-encoded like `VXNlcjoxNDYwMjgzOTI1` (= `User:1460283925`).
Decode with `atob(u.id).split(':')[1]` to get the numeric rest_id. `u.rest_id` may also
be present directly.
- `is_blue_verified` is the verified flag. `legacy.followers_count`, `legacy.description`
still exist under `legacy`.
- Filter timeline entries by `content.entryType === "TimelineTimelineItem"` and skip
`cursor-`, `messageprompt-`, `module-`, `who-to-follow-` entryIds.
### Features dict
Use the full `GQL_FEATURES` block from twscrape's `api.py` — without it X returns
`(336) The following features cannot be null`. Pass it URL-encoded as the `features` param.
### Where things live
- Output CSV: `~/Downloads/utilities/sdand_following.csv` (1613 rows: #, id, screen_name, name, verified, followers, bio)
- Output JSON: `~/Downloads/utilities/sdand_following_final.json` (double-encoded JSON string; parse with `json.loads(json.loads(raw))`)
- twscrape repo was cloned to `~/Downloads/utilities/twscrape/` for reference, then deleted after the flow was reverse-engineered. Re-clone from https://github.com/vladkens/twscrape.git if needed.
## Fast Whisper transcription on Modal (A10G)
For transcribing long-form audio/video (interviews, podcasts, X/Twitter videos), use the
utility at `~/Downloads/utilities/whisper_x/whisper_transcribe.py`. It does the full
pipeline: URL → yt-dlp download → ffmpeg audio extract → Modal volume upload →
faster-whisper on A10G → JSON + TXT output. Validated at **2.3 min wall clock for 65 min
of audio** (no caching at any layer).
### Usage
Shell alias (defined in `~/.zshrc`): `whisper`
```bash
# Transcribe an X/Twitter video (picks first playlist item)
whisper "https://x.com/.../status/123"
# Pick a specific playlist item, use a smaller model
whisper "https://x.com/..." --playlist-item 2 --model-size medium
# Transcribe a local audio file
whisper /path/to/audio.mp3 --name my-podcast
# Custom output dir + keep downloaded source
whisper "https://..." --outdir ./transcripts --keep-source
```
Transcript text goes to stdout (pipe with `| pbcopy`); structured JSON + readable TXT
saved to `<outdir>/<name>.json` and `<outdir>/<name>.txt`.
### Key optimizations (vs naive T4 run that took 11.7 min)
- **A10G GPU** (~8x fp16 throughput vs T4; Modal ~$0.60/hr vs ~$0.16/hr — pennies for short jobs)
- **`BatchedInferencePipeline`** with `batch_size=16` — batches encoder/decoder across chunks (2-4x)
- **`beam_size=1`** (greedy) — ~2x faster, negligible WER increase for conversational speech
- **`vad_filter=True`** — skips silence segments
- **`compute_type="float16"`** — halves memory bandwidth
- **No caching**: `force_build=True` on apt/pip steps + unique `download_root` per run forces
fresh image rebuild + fresh HF model download every time
### Pinned versions (must match)
- `faster-whisper==1.1.1` (provides `BatchedInferencePipeline`)
- `ctranslate2==4.8.0`
- Base image: `nvidia/cuda:12.6.3-cudnn-runtime-ubuntu22.04` (provides `libcublas.so.12`;
`debian_slim` fails with `RuntimeError: Library libcublas.so.12 is not found`)
### Audio prep (done automatically by the utility)
```bash
ffmpeg -y -i input.mp4 -vn -ac 1 -ar 16000 -c:a aac -b:a 64k audio.m4a
```
Mono 16kHz 64kbps AAC — a 65-min video (151 MB stream) becomes ~35 MB audio.
### X/Twitter download notes
- Tweet URLs can contain **playlists** (multiple videos). Use `--playlist-item N` to pick one.
- Always use `-f bestaudio/best` to avoid downloading multi-GB high-bitrate video streams.
- A 65-min interview's video variant can be 2.8+ GB; audio-only is ~63 MB (128 kbps).
### Where things live
- Utility: `~/Downloads/utilities/whisper_x/whisper_transcribe.py`
- Strategy doc: `~/Downloads/utilities/whisper_x/STRATEGY.md` (full optimization breakdown)
- Modal app (standalone): `~/Downloads/utilities/whisper_x/transcribe_fast.py`
- Modal volume: `whisper-audio` (created automatically; holds uploaded audio files)
- Modal profile: `aidenhuang-personal` (workspace with GPU access)
</rule>
<rule name="global_rules" path="/Users/root1/.codeium/windsurf/memories/global_rules.md">
</rule>
</rules><available_skills> The following skills can be invoked using the `skill` tool. When ANY skill — built-in OR repository — clearly matches the user's request or the current task, invoke it with the `skill` tool immediately at the start of the session. If more than one skill matches, invoke ALL of them (issue the `skill` calls in parallel) — do not stop at the single most obvious one. - **workers-best-practices**: Reviews and authors Cloudflare Workers code against production best practices. Load when writing new Workers, reviewing Worker code, configuring wrangler.jsonc, or checking for common Workers anti-patterns (streaming, floating promises, global state, secrets, bindings, observability). Biases towards retrieval from Cloudflare docs over pre-trained knowledge. (source: /Users/root1/.codeium/windsurf/skills/workers-best-practices/SKILL.md) - **web-perf**: Analyzes web performance using Chrome DevTools MCP. Measures Core Web Vitals (LCP, INP, CLS) and supplementary metrics (FCP, TBT, Speed Index), identifies render-blocking resources, network dependency chains, layout shifts, caching issues, and accessibility gaps. Use when asked to audit, profile, debug, or optimize page load performance, Lighthouse scores, or site speed. Biases towards retrieval from current documentation over pre-trained knowledge. (source: /Users/root1/.claude/skills/web-perf/SKILL.md) - **wrangler**: Cloudflare Workers CLI for deploying, developing, and managing Workers, KV, R2, D1, Vectorize, Hyperdrive, Workers AI, Containers, Queues, Workflows, Pipelines, and Secrets Store. Load before running wrangler commands to ensure correct syntax and best practices. Biases towards retrieval from Cloudflare docs over pre-trained knowledge. (source: /Users/root1/.claude/skills/wrangler/SKILL.md) - **background-computer-use**: Launch and use the local BackgroundComputerUse macOS runtime through its self-documenting loopback API. Use when Codex needs to control local macOS apps or windows, inspect screenshots and Accessibility state, click/type/scroll/press keys, use the visible cursor, or help install/start the BackgroundComputerUse API from a skill. (source: /Users/root1/.devin/skills/background-computer-use/SKILL.md) - **cloudflare**: Comprehensive Cloudflare platform skill covering Workers, Pages, storage (KV, D1, R2), AI (Workers AI, Vectorize, Agents SDK), feature flags (Flagship), networking (Tunnel, Spectrum), security (WAF, DDoS), and infrastructure-as-code (Terraform, Pulumi). Use for any Cloudflare development task. Biases towards retrieval from Cloudflare docs over pre-trained knowledge. (source: /Users/root1/.agents/skills/cloudflare/SKILL.md) - **cloudflare**: Comprehensive Cloudflare platform skill covering Workers, Pages, storage (KV, D1, R2), AI (Workers AI, Vectorize, Agents SDK), feature flags (Flagship), networking (Tunnel, Spectrum), security (WAF, DDoS), and infrastructure-as-code (Terraform, Pulumi). Use for any Cloudflare development task. Biases towards retrieval from Cloudflare docs over pre-trained knowledge. (source: /Users/root1/.config/devin/skills/cloudflare/SKILL.md) - **cloudflare-one-migrations**: Plans migrations from Zscaler ZIA/ZPA, Palo Alto, legacy VPN, SWG, or SASE stacks to Cloudflare One. Use for migration assessments, policy mapping, rollout plans, and parity/gap analysis. (source: /Users/root1/.claude/skills/cloudflare-one-migrations/SKILL.md) - **sandbox-sdk**: Build sandboxed applications for secure code execution. Load when building AI code execution, code interpreters, CI/CD systems, interactive dev environments, or executing untrusted code. Covers Sandbox SDK lifecycle, commands, files, code interpreter, and preview URLs. Biases towards retrieval from Cloudflare docs over pre-trained knowledge. (source: /Users/root1/.agents/skills/sandbox-sdk/SKILL.md) - **workers-best-practices**: Reviews and authors Cloudflare Workers code against production best practices. Load when writing new Workers, reviewing Worker code, configuring wrangler.jsonc, or checking for common Workers anti-patterns (streaming, floating promises, global state, secrets, bindings, observability). Biases towards retrieval from Cloudflare docs over pre-trained knowledge. (source: /Users/root1/.config/devin/skills/workers-best-practices/SKILL.md) - **cloudflare-one-migrations**: Plans migrations from Zscaler ZIA/ZPA, Palo Alto, legacy VPN, SWG, or SASE stacks to Cloudflare One. Use for migration assessments, policy mapping, rollout plans, and parity/gap analysis. (source: /Users/root1/.config/devin/skills/cloudflare-one-migrations/SKILL.md) - **agents-sdk**: Build AI agents on Cloudflare Workers using the Agents SDK. Load when creating stateful agents, durable workflows, real-time WebSocket apps, scheduled tasks, MCP servers, chat applications, voice agents, or browser automation. Covers Agent class, state management, callable RPC, Workflows, durable execution, queues, retries, observability, and React hooks. Biases towards retrieval from Cloudflare docs over pre-trained knowledge. (source: /Users/root1/.codeium/windsurf/skills/agents-sdk/SKILL.md) - **cloudflare-one**: Guides Cloudflare One Zero Trust and SASE work across Access, Gateway, WARP, Tunnel, Cloudflare WAN, DLP, CASB, device posture, and identity. Use when designing, configuring, troubleshooting, or reviewing Cloudflare One deployments. Retrieval-first: use current Cloudflare docs/API schemas instead of embedded product docs. (source: /Users/root1/.agents/skills/cloudflare-one/SKILL.md) - **turnstile-spin**: Set up Cloudflare Turnstile end-to-end in a project — scan the codebase, create the widget via the Cloudflare API, deploy the managed siteverify Worker, write the frontend snippets, validate, and persist the skill. Load this when a user asks to add Turnstile, set up CAPTCHA, protect a form from bots, or fix a Turnstile integration. Mirrors developers.cloudflare.com/turnstile/spin. (source: /Users/root1/.claude/skills/turnstile-spin/SKILL.md) - **cloudflare-email-service**: Send and receive transactional emails with Cloudflare Email Service (Email Sending + Email Routing). Use when building email sending (Workers binding or REST API), email routing, Agents SDK email handling, or integrating email into any app — Workers, Node.js, Python, Go, etc. Also use for email deliverability, SPF/DKIM/DMARC, wrangler email setup, MCP email tools, or when a coding agent needs to send emails. Even for simple requests like "add email to my Worker" — this skill has critical config details. (source: /Users/root1/.agents/skills/cloudflare-email-service/SKILL.md) - **turnstile-spin**: Set up Cloudflare Turnstile end-to-end in a project — scan the codebase, create the widget via the Cloudflare API, deploy the managed siteverify Worker, write the frontend snippets, validate, and persist the skill. Load this when a user asks to add Turnstile, set up CAPTCHA, protect a form from bots, or fix a Turnstile integration. Mirrors developers.cloudflare.com/turnstile/spin. (source: /Users/root1/.config/devin/skills/turnstile-spin/SKILL.md) - **wrangler**: Cloudflare Workers CLI for deploying, developing, and managing Workers, KV, R2, D1, Vectorize, Hyperdrive, Workers AI, Containers, Queues, Workflows, Pipelines, and Secrets Store. Load before running wrangler commands to ensure correct syntax and best practices. Biases towards retrieval from Cloudflare docs over pre-trained knowledge. (source: /Users/root1/.config/devin/skills/wrangler/SKILL.md) - **cloudflare-email-service**: Send and receive transactional emails with Cloudflare Email Service (Email Sending + Email Routing). Use when building email sending (Workers binding or REST API), email routing, Agents SDK email handling, or integrating email into any app — Workers, Node.js, Python, Go, etc. Also use for email deliverability, SPF/DKIM/DMARC, wrangler email setup, MCP email tools, or when a coding agent needs to send emails. Even for simple requests like "add email to my Worker" — this skill has critical config details. (source: /Users/root1/.config/devin/skills/cloudflare-email-service/SKILL.md) - **find-skills**: Helps users discover and install agent skills when they ask questions like "how do I do X", "find a skill for X", "is there a skill that can...", or express interest in extending capabilities. This skill should be used when the user is looking for functionality that might exist as an installable skill. (source: /Users/root1/.agents/skills/find-skills/SKILL.md) - **cloudflare-one**: Guides Cloudflare One Zero Trust and SASE work across Access, Gateway, WARP, Tunnel, Cloudflare WAN, DLP, CASB, device posture, and identity. Use when designing, configuring, troubleshooting, or reviewing Cloudflare One deployments. Retrieval-first: use current Cloudflare docs/API schemas instead of embedded product docs. (source: /Users/root1/.config/devin/skills/cloudflare-one/SKILL.md) - **cloudflare-agent-setup**: (source: /Users/root1/.devin/skills/cloudflare-agent-setup/SKILL.md) - **durable-objects**: Create and review Cloudflare Durable Objects. Use when building stateful coordination (chat rooms, multiplayer games, booking systems), implementing RPC methods, SQLite storage, alarms, WebSockets, or reviewing DO code for best practices. Covers Workers integration, wrangler config, and testing with Vitest. Biases towards retrieval from Cloudflare docs over pre-trained knowledge. (source: /Users/root1/.agents/skills/durable-objects/SKILL.md) - **durable-objects**: Create and review Cloudflare Durable Objects. Use when building stateful coordination (chat rooms, multiplayer games, booking systems), implementing RPC methods, SQLite storage, alarms, WebSockets, or reviewing DO code for best practices. Covers Workers integration, wrangler config, and testing with Vitest. Biases towards retrieval from Cloudflare docs over pre-trained knowledge. (source: /Users/root1/.config/devin/skills/durable-objects/SKILL.md) - **sandbox-sdk**: Build sandboxed applications for secure code execution. Load when building AI code execution, code interpreters, CI/CD systems, interactive dev environments, or executing untrusted code. Covers Sandbox SDK lifecycle, commands, files, code interpreter, and preview URLs. Biases towards retrieval from Cloudflare docs over pre-trained knowledge. (source: /Users/root1/.config/devin/skills/sandbox-sdk/SKILL.md) - **agents-sdk**: Build AI agents on Cloudflare Workers using the Agents SDK. Load when creating stateful agents, durable workflows, real-time WebSocket apps, scheduled tasks, MCP servers, chat applications, voice agents, or browser automation. Covers Agent class, state management, callable RPC, Workflows, durable execution, queues, retries, observability, and React hooks. Biases towards retrieval from Cloudflare docs over pre-trained knowledge. (source: /Users/root1/.config/devin/skills/agents-sdk/SKILL.md) - **web-perf**: Analyzes web performance using Chrome DevTools MCP. Measures Core Web Vitals (LCP, INP, CLS) and supplementary metrics (FCP, TBT, Speed Index), identifies render-blocking resources, network dependency chains, layout shifts, caching issues, and accessibility gaps. Use when asked to audit, profile, debug, or optimize page load performance, Lighthouse scores, or site speed. Biases towards retrieval from current documentation over pre-trained knowledge. (source: /Users/root1/.config/devin/skills/web-perf/SKILL.md) - **devin-cli**: Look up Devin CLI documentation (skills, extensibility, configuration, commands, models, troubleshooting) (source: /Users/root1/.local/share/devin/cli/_versions/3000.1.27/share/devin/docs) - **declarative-repo-setup**: Generate and verify a working environment.yaml (Devin snapshot-setup blueprint) for a repo (source: builtin:drs) </available_skills>
The user provided a YouTube link. According to the AGENTS.md rule: "If I ever paste in a YouTube link, use yt-dlp to summarize the video. get the autogenerated captions to do this". Also "if asked to summarize a YouTube video, do not name the session until after reading and understanding the full YouTube video transcript". The user didn't ask to summarize explicitly, but the rule says "If I ever paste in a YouTube link, use yt-dlp to summarize the video." So I should summarize. Need to get autogenerated captions. Need to use yt-dlp. Need to check if installed. We'll likely download the video's subtitles using yt-dlp. We should not name the session. We need to read the transcript. We need to use tools. Plan: run yt-dlp to list available subtitles and download auto-generated captions. Then summarize. Need to not overthink. Let's do it. First check yt-dlp presence and use.
{
"command": "yt-dlp --version"
}Output from command in shell 3a41da: 2025.12.08 Exit code: 0
yt-dlp installed. Need to get autogenerated captions and summarize. I can use yt-dlp to download only subtitles (--skip-download), list subtitles. The URL is https://www.youtube.com/watch?v=AZrU6y3pUcU. Need to determine language and title. Let's list available subs.
{
"command": "yt-dlp --list-subs \"https://www.youtube.com/watch?v=AZrU6y3pUcU\""
}Output from command in shell 0f2646:
WARNING: Your yt-dlp version (2025.12.08) is older than 90 days!
It is strongly recommended to always use the latest version.
You installed yt-dlp with pip or using the wheel from PyPi; Use that to update.
To suppress this warning, add --no-update to your command/config.
[youtube] Extracting URL: https://www.youtube.com/watch?v=AZrU6y3pUcU
[youtube] AZrU6y3pUcU: Downloading webpage
[youtube] AZrU6y3pUcU: Downloading tv client config
[youtube] AZrU6y3pUcU: Downloading player ecd4b80a-main
[youtube] AZrU6y3pUcU: Downloading tv player API JSON
[youtube] AZrU6y3pUcU: Downloading android sdkless player API JSON
[youtube] [jsc:deno] Solving JS challenges using deno
WARNING: [youtube] [jsc] Remote components challenge solver script (deno) and NPM package (deno) were skipped. These may be required to solve JS challenges. You can enable these downloads with --remote-components ejs:github (recommended) or --remote-components ejs:npm , respectively. For more information and alternatives, refer to https://github.com/yt-dlp/yt-dlp/wiki/EJS
WARNING: [youtube] AZrU6y3pUcU: n challenge solving failed: Some formats may be missing. Ensure you have a supported JavaScript runtime and challenge solver script distribution installed. Review any warnings presented before this message. For more details, refer to https://github.com/yt-dlp/yt-dlp/wiki/EJS
[info] Available automatic captions for AZrU6y3pUcU:
Language Name Formats
ab Abkhazian vtt, srt, ttml, srv3, srv2, srv1, json3
aa Afar vtt, srt, ttml, srv3, srv2, srv1, json3
af Afrikaans vtt, srt, ttml, srv3, srv2, srv1, json3
ak Akan vtt, srt, ttml, srv3, srv2, srv1, json3
sq Albanian vtt, srt, ttml, srv3, srv2, srv1, json3
am Amharic vtt, srt, ttml, srv3, srv2, srv1, json3
ar Arabic vtt, srt, ttml, srv3, srv2, srv1, json3
hy Armenian vtt, srt, ttml, srv3, srv2, srv1, json3
as Assamese vtt, srt, ttml, srv3, srv2, srv1, json3
ay Aymara vtt, srt, ttml, srv3, srv2, srv1, json3
az Azerbaijani vtt, srt, ttml, srv3, srv2, srv1, json3
bn Bangla vtt, srt, ttml, srv3, srv2, srv1, json3
ba Bashkir vtt, srt, ttml, srv3, srv2, srv1, json3
eu Basque vtt, srt, ttml, srv3, srv2, srv1, json3
be Belarusian vtt, srt, ttml, srv3, srv2, srv1, json3
bho Bhojpuri vtt, srt, ttml, srv3, srv2, srv1, json3
bs Bosnian vtt, srt, ttml, srv3, srv2, srv1, json3
br Breton vtt, srt, ttml, srv3, srv2, srv1, json3
bg Bulgarian vtt, srt, ttml, srv3, srv2, srv1, json3
my Burmese vtt, srt, ttml, srv3, srv2, srv1, json3
ca Catalan vtt, srt, ttml, srv3, srv2, srv1, json3
ceb Cebuano vtt, srt, ttml, srv3, srv2, srv1, json3
zh-Hans Chinese (Simplified) vtt, srt, ttml, srv3, srv2, srv1, json3
zh-Hant Chinese (Traditional) vtt, srt, ttml, srv3, srv2, srv1, json3
co Corsican vtt, srt, ttml, srv3, srv2, srv1, json3
hr Croatian vtt, srt, ttml, srv3, srv2, srv1, json3
cs Czech vtt, srt, ttml, srv3, srv2, srv1, json3
da Danish vtt, srt, ttml, srv3, srv2, srv1, json3
dv Divehi vtt, srt, ttml, srv3, srv2, srv1, json3
nl Dutch vtt, srt, ttml, srv3, srv2, srv1, json3
dz Dzongkha vtt, srt, ttml, srv3, srv2, srv1, json3
en-orig English (Original) vtt, srt, ttml, srv3, srv2, srv1, json3
en English vtt, srt, ttml, srv3, srv2, srv1, json3
eo Esperanto vtt, srt, ttml, srv3, srv2, srv1, json3
et Estonian vtt, srt, ttml, srv3, srv2, srv1, json3
ee Ewe vtt, srt, ttml, srv3, srv2, srv1, json3
fo Faroese vtt, srt, ttml, srv3, srv2, srv1, json3
fj Fijian vtt, srt, ttml, srv3, srv2, srv1, json3
fil Filipino vtt, srt, ttml, srv3, srv2, srv1, json3
fi Finnish vtt, srt, ttml, srv3, srv2, srv1, json3
fr French vtt, srt, ttml, srv3, srv2, srv1, json3
gaa Ga vtt, srt, ttml, srv3, srv2, srv1, json3
gl Galician vtt, srt, ttml, srv3, srv2, srv1, json3
lg Ganda vtt, srt, ttml, srv3, srv2, srv1, json3
ka Georgian vtt, srt, ttml, srv3, srv2, srv1, json3
de German vtt, srt, ttml, srv3, srv2, srv1, json3
el Greek vtt, srt, ttml, srv3, srv2, srv1, json3
gn Guarani vtt, srt, ttml, srv3, srv2, srv1, json3
gu Gujarati vtt, srt, ttml, srv3, srv2, srv1, json3
ht Haitian Creole vtt, srt, ttml, srv3, srv2, srv1, json3
ha Hausa vtt, srt, ttml, srv3, srv2, srv1, json3
haw Hawaiian vtt, srt, ttml, srv3, srv2, srv1, json3
iw Hebrew vtt, srt, ttml, srv3, srv2, srv1, json3
hi Hindi vtt, srt, ttml, srv3, srv2, srv1, json3
hmn Hmong vtt, srt, ttml, srv3, srv2, srv1, json3
hu Hungarian vtt, srt, ttml, srv3, srv2, srv1, json3
is Icelandic vtt, srt, ttml, srv3, srv2, srv1, json3
ig Igbo vtt, srt, ttml, srv3, srv2, srv1, json3
id Indonesian vtt, srt, ttml, srv3, srv2, srv1, json3
iu Inuktitut vtt, srt, ttml, srv3, srv2, srv1, json3
ga Irish vtt, srt, ttml, srv3, srv2, srv1, json3
it Italian vtt, srt, ttml, srv3, srv2, srv1, json3
ja Japanese vtt, srt, ttml, srv3, srv2, srv1, json3
jv Javanese vtt, srt, ttml, srv3, srv2, srv1, json3
kl Kalaallisut vtt, srt, ttml, srv3, srv2, srv1, json3
kn Kannada vtt, srt, ttml, srv3, srv2, srv1, json3
kk Kazakh vtt, srt, ttml, srv3, srv2, srv1, json3
kha Khasi vtt, srt, ttml, srv3, srv2, srv1, json3
km Khmer vtt, srt, ttml, srv3, srv2, srv1, json3
rw Kinyarwanda vtt, srt, ttml, srv3, srv2, srv1, json3
ko Korean vtt, srt, ttml, srv3, srv2, srv1, json3
kri Krio vtt, srt, ttml, srv3, srv2, srv1, json3
ku Kurdish vtt, srt, ttml, srv3, srv2, srv1, json3
ky Kyrgyz vtt, srt, ttml, srv3, srv2, srv1, json3
lo Lao vtt, srt, ttml, srv3, srv2, srv1, json3
la Latin vtt, srt, ttml, srv3, srv2, srv1, json3
lv Latvian vtt, srt, ttml, srv3, srv2, srv1, json3
ln Lingala vtt, srt, ttml, srv3, srv2, srv1, json3
lt Lithuanian vtt, srt, ttml, srv3, srv2, srv1, json3
lua Luba-Lulua vtt, srt, ttml, srv3, srv2, srv1, json3
luo Luo vtt, srt, ttml, srv3, srv2, srv1, json3
lb Luxembourgish vtt, srt, ttml, srv3, srv2, srv1, json3
mk Macedonian vtt, srt, ttml, srv3, srv2, srv1, json3
mg Malagasy vtt, srt, ttml, srv3, srv2, srv1, json3
ms Malay vtt, srt, ttml, srv3, srv2, srv1, json3
ml Malayalam vtt, srt, ttml, srv3, srv2, srv1, json3
mt Maltese vtt, srt, ttml, srv3, srv2, srv1, json3
gv Manx vtt, srt, ttml, srv3, srv2, srv1, json3
mi Māori vtt, srt, ttml, srv3, srv2, srv1, json3
mr Marathi vtt, srt, ttml, srv3, srv2, srv1, json3
mn Mongolian vtt, srt, ttml, srv3, srv2, srv1, json3
mfe Morisyen vtt, srt, ttml, srv3, srv2, srv1, json3
ne Nepali vtt, srt, ttml, srv3, srv2, srv1, json3
new Newari vtt, srt, ttml, srv3, srv2, srv1, json3
nso Northern Sotho vtt, srt, ttml, srv3, srv2, srv1, json3
no Norwegian vtt, srt, ttml, srv3, srv2, srv1, json3
ny Nyanja vtt, srt, ttml, srv3, srv2, srv1, json3
oc Occitan vtt, srt, ttml, srv3, srv2, srv1, json3
or Odia vtt, srt, ttml, srv3, srv2, srv1, json3
om Oromo vtt, srt, ttml, srv3, srv2, srv1, json3
os Ossetic vtt, srt, ttml, srv3, srv2, srv1, json3
pam Pampanga vtt, srt, ttml, srv3, srv2, srv1, json3
ps Pashto vtt, srt, ttml, srv3, srv2, srv1, json3
fa Persian vtt, srt, ttml, srv3, srv2, srv1, json3
pl Polish vtt, srt, ttml, srv3, srv2, srv1, json3
pt Portuguese vtt, srt, ttml, srv3, srv2, srv1, json3
pt-PT Portuguese (Portugal) vtt, srt, ttml, srv3, srv2, srv1, json3
pa Punjabi vtt, srt, ttml, srv3, srv2, srv1, json3
qu Quechua vtt, srt, ttml, srv3, srv2, srv1, json3
ro Romanian vtt, srt, ttml, srv3, srv2, srv1, json3
rn Rundi vtt, srt, ttml, srv3, srv2, srv1, json3
ru Russian vtt, srt, ttml, srv3, srv2, srv1, json3
sm Samoan vtt, srt, ttml, srv3, srv2, srv1, json3
sg Sango vtt, srt, ttml, srv3, srv2, srv1, json3
sa Sanskrit vtt, srt, ttml, srv3, srv2, srv1, json3
gd Scottish Gaelic vtt, srt, ttml, srv3, srv2, srv1, json3
sr Serbian vtt, srt, ttml, srv3, srv2, srv1, json3
crs Seselwa Creole French vtt, srt, ttml, srv3, srv2, srv1, json3
sn Shona vtt, srt, ttml, srv3, srv2, srv1, json3
sd Sindhi vtt, srt, ttml, srv3, srv2, srv1, json3
si Sinhala vtt, srt, ttml, srv3, srv2, srv1, json3
sk Slovak vtt, srt, ttml, srv3, srv2, srv1, json3
sl Slovenian vtt, srt, ttml, srv3, srv2, srv1, json3
so Somali vtt, srt, ttml, srv3, srv2, srv1, json3
st Southern Sotho vtt, srt, ttml, srv3, srv2, srv1, json3
es Spanish vtt, srt, ttml, srv3, srv2, srv1, json3
su Sundanese vtt, srt, ttml, srv3, srv2, srv1, json3
sw Swahili vtt, srt, ttml, srv3, srv2, srv1, json3
ss Swati vtt, srt, ttml, srv3, srv2, srv1, json3
sv Swedish vtt, srt, ttml, srv3, srv2, srv1, json3
tg Tajik vtt, srt, ttml, srv3, srv2, srv1, json3
ta Tamil vtt, srt, ttml, srv3, srv2, srv1, json3
tt Tatar vtt, srt, ttml, srv3, srv2, srv1, json3
te Telugu vtt, srt, ttml, srv3, srv2, srv1, json3
th Thai vtt, srt, ttml, srv3, srv2, srv1, json3
bo Tibetan vtt, srt, ttml, srv3, srv2, srv1, json3
ti Tigrinya vtt, srt, ttml, srv3, srv2, srv1, json3
to Tongan vtt, srt, ttml, srv3, srv2, srv1, json3
ts Tsonga vtt, srt, ttml, srv3, srv2, srv1, json3
tn Tswana vtt, srt, ttml, srv3, srv2, srv1, json3
tum Tumbuka vtt, srt, ttml, srv3, srv2, srv1, json3
tr Turkish vtt, srt, ttml, srv3, srv2, srv1, json3
tk Turkmen vtt, srt, ttml, srv3, srv2, srv1, json3
uk Ukrainian vtt, srt, ttml, srv3, srv2, srv1, json3
ur Urdu vtt, srt, ttml, srv3, srv2, srv1, json3
ug Uyghur vtt, srt, ttml, srv3, srv2, srv1, json3
uz Uzbek vtt, srt, ttml, srv3, srv2, srv1, json3
ve Venda vtt, srt, ttml, srv3, srv2, srv1, json3
vi Vietnamese vtt, srt, ttml, srv3, srv2, srv1, json3
war Waray vtt, srt, ttml, srv3, srv2, srv1, json3
cy Welsh vtt, srt, ttml, srv3, srv2, srv1, json3
fy Western Frisian vtt, srt, ttml, srv3, srv2, srv1, json3
wo Wolof vtt, srt, ttml, srv3, srv2, srv1, json3
xh Xhosa vtt, srt, ttml, srv3, srv2, srv1, json3
yi Yiddish vtt, srt, ttml, srv3, srv2, srv1, json3
yo Yoruba vtt, srt, ttml, srv3, srv2, srv1, json3
zu Zulu vtt, srt, ttml, srv3, srv2, srv1, json3
AZrU6y3pUcU has no subtitles
Exit code: 0Need to download auto-generated captions. Available automatic captions includes en-orig and en. Need to use --write-auto-subs and --skip-download. Use English. Let's download to a temp file.
{
"command": "yt-dlp --skip-download --write-auto-subs --sub-lang en --sub-format srt --convert-subs srt -o \"/tmp/%(title)s.%(ext)s\" \"https://www.youtube.com/watch?v=AZrU6y3pUcU\""
}Output from command in shell 7334da:
WARNING: Your yt-dlp version (2025.12.08) is older than 90 days!
It is strongly recommended to always use the latest version.
You installed yt-dlp with pip or using the wheel from PyPi; Use that to update.
To suppress this warning, add --no-update to your command/config.
[youtube] Extracting URL: https://www.youtube.com/watch?v=AZrU6y3pUcU
[youtube] AZrU6y3pUcU: Downloading webpage
[youtube] AZrU6y3pUcU: Downloading tv client config
[youtube] AZrU6y3pUcU: Downloading player ecd4b80a-main
[youtube] AZrU6y3pUcU: Downloading tv player API JSON
[youtube] AZrU6y3pUcU: Downloading android sdkless player API JSON
[youtube] [jsc:deno] Solving JS challenges using deno
WARNING: [youtube] [jsc] Remote components challenge solver script (deno) and NPM package (deno) were skipped. These may be required to solve JS challenges. You can enable these downloads with --remote-components ejs:github (recommended) or --remote-components ejs:npm , respectively. For more information and alternatives, refer to https://github.com/yt-dlp/yt-dlp/wiki/EJS
WARNING: [youtube] AZrU6y3pUcU: n challenge solving failed: Some formats may be missing. Ensure you have a supported JavaScript runtime and challenge solver script distribution installed. Review any warnings presented before this message. For more details, refer to https://github.com/yt-dlp/yt-dlp/wiki/EJS
[info] AZrU6y3pUcU: Downloading subtitles: en
[info] AZrU6y3pUcU: Downloading 1 format(s): 139-20
[info] Writing video subtitles to: /tmp/Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown.en.srt
[download] Destination: /tmp/Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown.en.srt
[download] 1.00KiB at Unknown B/s (00:00:00)
[download] 3.00KiB at Unknown B/s (00:00:00)
[download] 7.00KiB at Unknown B/s (00:00:00)
[download] 15.00KiB at 4.49MiB/s (00:00:00)
[download] 31.00KiB at 1.94MiB/s (00:00:00)
[download] 63.00KiB at 2.57MiB/s (00:00:00)
[download] 77.52KiB at 2.76MiB/s (00:00:00)
[download] 100% of 77.52KiB in 00:00:00 at 573.24KiB/s
[SubtitlesConvertor] Converting subtitles
[SubtitlesConvertor] Subtitle file for srt is already in the requested format
Exit code: 0Got subtitle file. Title: "Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown". Need to read transcript. It may be long, 77KB. Let's read.
{
"file_path": "/tmp/Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown.en.srt"
}<file-view path="/tmp/Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown.en.srt" start_line="1" end_line="1118" total_lines="4408"> 1|1 2|00:00:00,080 --> 00:00:04,560 3|With GBT3, you couldn't scale test time 4| 5|2 6|00:00:02,800 --> 00:00:05,839 7|compute. Like if you gave it a budget of 8| 9|3 10|00:00:04,560 --> 00:00:07,919 11|$10 million and said, "Okay, well, let's 12| 13|4 14|00:00:05,839 --> 00:00:09,760 15|see what GB3 can do." It really can't do 16| 17|5 18|00:00:07,919 --> 00:00:11,360 19|that much. The precurren frameworks and 20| 21|6 22|00:00:09,760 --> 00:00:12,719 23|responsible scaling policies, they don't 24| 25|7 26|00:00:11,360 --> 00:00:13,679 27|really account for the amount of testime 28| 29|8 30|00:00:12,719 --> 00:00:15,280 31|compute. They just say, "Okay, well, 32| 33|9 34|00:00:13,679 --> 00:00:17,440 35|what's the capability of the model?" The 36| 37|10 38|00:00:15,280 --> 00:00:19,359 39|problem is we're in a world now where 40| 41|11 42|00:00:17,440 --> 00:00:21,279 43|the capability of the model is a 44| 45|12 46|00:00:19,359 --> 00:00:23,279 47|function of how much money you put into 48| 49|13 50|00:00:21,279 --> 00:00:26,000 51|it. Basically, if you give it a budget 52| 53|14 54|00:00:23,279 --> 00:00:27,680 55|of $10,000, it can do a lot more than 56| 57|15 58|00:00:26,000 --> 00:00:29,199 59|what it can do with a budget of $10. 60| 61|16 62|00:00:27,680 --> 00:00:31,760 63|Give it a budget of $10 million, you can 64| 65|17 66|00:00:29,199 --> 00:00:33,920 67|do even more. At what budget should you 68| 69|18 70|00:00:31,760 --> 00:00:35,600 71|evaluate these models? The policies that 72| 73|19 74|00:00:33,920 --> 00:00:38,600 75|exist today don't really address that 76| 77|20 78|00:00:35,600 --> 00:00:38,600 79|question. 80| 81|21 82|00:00:43,280 --> 00:00:47,680 83|Hi listeners, I'm Sarah Goa and welcome 84| 85|22 86|00:00:45,680 --> 00:00:50,000 87|back to No Priors. Today I'm here with 88| 89|23 90|00:00:47,680 --> 00:00:52,000 91|Nome Brown, one of our godfathers of AI 92| 93|24 94|00:00:50,000 --> 00:00:54,800 95|reasoning. We talk about the broken 96| 97|25 98|00:00:52,000 --> 00:00:57,280 99|state of evaluations, very large-scale 100| 101|26 102|00:00:54,800 --> 00:00:59,120 103|test time compute, how he thinks about 104| 105|27 106|00:00:57,280 --> 00:01:01,600 107|recursive self-improvement, and what's 108| 109|28 110|00:00:59,120 --> 00:01:04,320 111|next on the horizon for competition at 112| 113|29 114|00:01:01,600 --> 00:01:05,680 115|the frontier. Welcome, Gnome. I'm so 116| 117|30 118|00:01:04,320 --> 00:01:06,799 119|excited to have you back. 120| 121|31 122|00:01:05,680 --> 00:01:08,640 123|>> That's great to be back. Yeah, 124| 125|32 126|00:01:06,799 --> 00:01:11,200 127|>> you are our first guest. I'm very proud 128| 129|33 130|00:01:08,640 --> 00:01:14,560 131|of my taste in friends and researchers 132| 133|34 134|00:01:11,200 --> 00:01:17,680 135|for uh for the pod. um given you know 136| 137|35 138|00:01:14,560 --> 00:01:19,520 139|how important uh you know inference time 140| 141|36 142|00:01:17,680 --> 00:01:20,799 143|scaling has become to the industry you 144| 145|37 146|00:01:19,520 --> 00:01:21,520 147|should be proud too having actually 148| 149|38 150|00:01:20,799 --> 00:01:23,280 151|pioneered it 152| 153|39 154|00:01:21,520 --> 00:01:24,400 155|>> played a part yeah among among many 156| 157|40 158|00:01:23,280 --> 00:01:27,040 159|others 160| 161|41 162|00:01:24,400 --> 00:01:28,799 163|>> you just wrote this essay that really 164| 165|42 166|00:01:27,040 --> 00:01:32,799 167|resonated about large scale test time 168| 169|43 170|00:01:28,799 --> 00:01:35,360 171|compute and why uh the industry is not 172| 173|44 174|00:01:32,799 --> 00:01:36,960 175|evaluating these models as robustly as 176| 177|45 178|00:01:35,360 --> 00:01:37,439 179|it should be what was the motivation for 180| 181|46 182|00:01:36,960 --> 00:01:40,880 183|it 184| 185|47 186|00:01:37,439 --> 00:01:44,560 187|>> yeah the motivation was we released 5.5 188| 189|48 190|00:01:40,880 --> 00:01:46,079 191|and the initial reaction ction was kind 192| 193|49 194|00:01:44,560 --> 00:01:47,680 195|of skepticism that it was a 196| 197|50 198|00:01:46,079 --> 00:01:49,119 199|substantially better model. To be fair, 200| 201|51 202|00:01:47,680 --> 00:01:50,560 203|that only lasted for a few hours before 204| 205|52 206|00:01:49,119 --> 00:01:51,840 207|people had some time to play around with 208| 209|53 210|00:01:50,560 --> 00:01:52,560 211|it and and try it out themselves and 212| 213|54 214|00:01:51,840 --> 00:01:54,479 215|they saw that it was actually 216| 217|55 218|00:01:52,560 --> 00:01:57,119 219|substantially better. Um, but I think a 220| 221|56 222|00:01:54,479 --> 00:01:58,479 223|lot of the skepticism came from the 224| 225|57 226|00:01:57,119 --> 00:01:59,759 227|benchmark grid that was published. 228| 229|58 230|00:01:58,479 --> 00:02:01,759 231|Basically, whenever a new model is 232| 233|59 234|00:01:59,759 --> 00:02:02,880 235|released, there's this benchmark grid 236| 237|60 238|00:02:01,759 --> 00:02:04,560 239|where they show all these different 240| 241|61 242|00:02:02,880 --> 00:02:05,680 243|benchmarks on on the x-axxis and then 244| 245|62 246|00:02:04,560 --> 00:02:07,280 247|the performance of different models on 248| 249|63 250|00:02:05,680 --> 00:02:08,800 251|the y- axis and you can just like 252| 253|64 254|00:02:07,280 --> 00:02:11,200 255|compare different models. It's like a 256| 257|65 258|00:02:08,800 --> 00:02:13,599 259|single number for a model on a single 260| 261|66 262|00:02:11,200 --> 00:02:15,680 263|benchmark. And if you look on paper at 264| 265|67 266|00:02:13,599 --> 00:02:18,480 267|the difference between like 5.5 and 5.4 268| 269|68 270|00:02:15,680 --> 00:02:19,920 271|or or other models, it wasn't it was an 272| 273|69 274|00:02:18,480 --> 00:02:21,599 275|improvement, but it wasn't a huge 276| 277|70 278|00:02:19,920 --> 00:02:23,280 279|improvement. It was only a few 280| 281|71 282|00:02:21,599 --> 00:02:24,560 283|percentage points in some benchmarks. 284| 285|72 286|00:02:23,280 --> 00:02:25,760 287|So, people looked at that and they were 288| 289|73 290|00:02:24,560 --> 00:02:27,360 291|skeptical that it was actually a better 292| 293|74 294|00:02:25,760 --> 00:02:29,599 295|model. Um, once they played around with 296| 297|75 298|00:02:27,360 --> 00:02:32,080 299|it, the story changed. I think the 300| 301|76 302|00:02:29,599 --> 00:02:33,680 303|reason why it doesn't show up as so much 304| 305|77 306|00:02:32,080 --> 00:02:35,840 307|better on the benchmarks is because the 308| 309|78 310|00:02:33,680 --> 00:02:37,120 311|benchmarks are being presented, the 312| 313|79 314|00:02:35,840 --> 00:02:39,200 315|benchmark results are being presented in 316| 317|80 318|00:02:37,120 --> 00:02:41,280 319|the wrong way. They're not controlling 320| 321|81 322|00:02:39,200 --> 00:02:42,640 323|for the amount of test time compute that 324| 325|82 326|00:02:41,280 --> 00:02:45,040 327|is being used on that benchmark 328| 329|83 330|00:02:42,640 --> 00:02:46,560 331|question. It turned out that 5.5 is just 332| 333|84 334|00:02:45,040 --> 00:02:50,000 335|much more efficient with its thinking. 336| 337|85 338|00:02:46,560 --> 00:02:51,280 339|If you run it at max settings, 5.4 is 340| 341|86 342|00:02:50,000 --> 00:02:54,080 343|thinking for a lot longer. It takes 344| 345|87 346|00:02:51,280 --> 00:02:55,680 347|longer to get back a response than 5.5. 348| 349|88 350|00:02:54,080 --> 00:02:57,200 351|And once you control for the amount of 352| 353|89 354|00:02:55,680 --> 00:03:00,319 355|thinking time, actually you can see that 356| 357|90 358|00:02:57,200 --> 00:03:01,440 359|5.5 is a substantial jump over 5.4. That 360| 361|91 362|00:03:00,319 --> 00:03:03,360 363|is, I think, people's day-to-day 364| 365|92 366|00:03:01,440 --> 00:03:05,519 367|experience with it. And then when I 368| 369|93 370|00:03:03,360 --> 00:03:06,800 371|mention this to people the reaction, the 372| 373|94 374|00:03:05,519 --> 00:03:08,800 375|typical question I get is like okay well 376| 377|95 378|00:03:06,800 --> 00:03:11,200 379|why not just have 5.5 think for as long 380| 381|96 382|00:03:08,800 --> 00:03:12,560 383|as 5.4 and the question is like well how 384| 385|97 386|00:03:11,200 --> 00:03:14,000 387|long should they think for? Typically 388| 389|98 390|00:03:12,560 --> 00:03:16,000 391|the response I get is well until the 392| 393|99 394|00:03:14,000 --> 00:03:17,120 395|performance plateaus right there's at 396| 397|100 398|00:03:16,000 --> 00:03:18,400 399|some point where the performance on the 400| 401|101 402|00:03:17,120 --> 00:03:20,879 403|benchmark is going to plateau and you 404| 405|102 406|00:03:18,400 --> 00:03:22,239 407|evaluate to that point. The thing is the 408| 409|103 410|00:03:20,879 --> 00:03:24,319 411|point at which it plateaus is actually 412| 413|104 414|00:03:22,239 --> 00:03:27,680 415|really far out these days. I mean, it 416| 417|105 418|00:03:24,319 --> 00:03:28,640 419|was true in GBD3 land back in 2022, the 420| 421|106 422|00:03:27,680 --> 00:03:29,920 423|models couldn't really think 424| 425|107 426|00:03:28,640 --> 00:03:31,920 427|productively for that long. And so, you 428| 429|108 430|00:03:29,920 --> 00:03:34,239 431|could just run them until they plateau. 432| 433|109 434|00:03:31,920 --> 00:03:35,920 435|It's not that far away. But what we're 436| 437|110 438|00:03:34,239 --> 00:03:39,200 439|seeing today with the modern models is 440| 441|111 442|00:03:35,920 --> 00:03:41,519 443|that 5.5 and other models can think for 444| 445|112 446|00:03:39,200 --> 00:03:45,599 447|if you scaffold them reasonably well, 448| 449|113 450|00:03:41,519 --> 00:03:46,959 451|can think for weeks even um before 452| 453|114 454|00:03:45,599 --> 00:03:48,799 455|having performance plateau on some of 456| 457|115 458|00:03:46,959 --> 00:03:51,360 459|these benchmarks. And so, the point at 460| 461|116 462|00:03:48,799 --> 00:03:53,440 463|which they plateau is simply too far out 464| 465|117 466|00:03:51,360 --> 00:03:55,760 467|to reasonably test. we all need to 468| 469|118 470|00:03:53,440 --> 00:03:58,159 471|actually reinforce um either like a 472| 473|119 474|00:03:55,760 --> 00:03:59,599 475|patience limit or a budget limit from a 476| 477|120 478|00:03:58,159 --> 00:04:00,560 479|token perspective now and that wasn't 480| 481|121 482|00:03:59,599 --> 00:04:02,000 483|true a few years ago. 484| 485|122 486|00:04:00,560 --> 00:04:03,680 487|>> Exactly. And so I think the proper way 488| 489|123 490|00:04:02,000 --> 00:04:06,000 491|to and so my claim is the proper way to 492| 493|124 494|00:04:03,680 --> 00:04:08,159 495|evaluate the models now is you either 496| 497|125 498|00:04:06,000 --> 00:04:10,239 499|have some kind of budget for the 500| 501|126 502|00:04:08,159 --> 00:04:12,319 503|benchmark whether it's tokens or cost or 504| 505|127 506|00:04:10,239 --> 00:04:13,920 507|time or whatever or you plot the 508| 509|128 510|00:04:12,319 --> 00:04:15,040 511|performance as a function of the amount 512| 513|129 514|00:04:13,920 --> 00:04:17,199 515|of test time compute that's going into 516| 517|130 518|00:04:15,040 --> 00:04:18,959 519|the model and then it becomes much more 520| 521|131 522|00:04:17,199 --> 00:04:20,320 523|clear um how to compare the performance 524| 525|132 526|00:04:18,959 --> 00:04:23,360 527|between these different models. 528| 529|133 530|00:04:20,320 --> 00:04:26,560 531|>> Given the model evaluation cycle and the 532| 533|134 534|00:04:23,360 --> 00:04:29,199 535|fact that performance does not asmtote 536| 537|135 538|00:04:26,560 --> 00:04:31,520 539|for many tasks over um quite a long 540| 541|136 542|00:04:29,199 --> 00:04:33,919 543|period of time. What do you do about 544| 545|137 546|00:04:31,520 --> 00:04:36,320 547|that issue? The fact that some of the 548| 549|138 550|00:04:33,919 --> 00:04:38,479 551|evals that you would want to run are 552| 553|139 554|00:04:36,320 --> 00:04:39,919 555|both beyond the scope of budget or time 556| 557|140 558|00:04:38,479 --> 00:04:40,960 559|that's reasonable given the current 560| 561|141 562|00:04:39,919 --> 00:04:42,720 563|model release cycle. 564| 565|142 566|00:04:40,960 --> 00:04:45,680 567|>> I mean, I think for things like cyber, 568| 569|143 570|00:04:42,720 --> 00:04:48,400 571|we've seen and actually the AIS um in 572| 573|144 574|00:04:45,680 --> 00:04:51,759 575|their evaluations has shown that the 576| 577|145 578|00:04:48,400 --> 00:04:53,040 579|models continue to improve um at 100 580| 581|146 582|00:04:51,759 --> 00:04:54,240 583|million tokens. You know, if you run 584| 585|147 586|00:04:53,040 --> 00:04:55,919 587|them for 100 million tokens, they're 588| 589|148 590|00:04:54,240 --> 00:04:57,360 591|still improving at beyond that point. 592| 593|149 594|00:04:55,919 --> 00:05:00,560 595|>> And that can take a very long time to 596| 597|150 598|00:04:57,360 --> 00:05:03,280 599|run. But you also do see that like the 600| 601|151 602|00:05:00,560 --> 00:05:05,360 603|performance is is it's not just like a 604| 605|152 606|00:05:03,280 --> 00:05:06,880 607|discontinuous jump. It's actually like 608| 609|153 610|00:05:05,360 --> 00:05:08,720 611|you can see the slope of improvement 612| 613|154 614|00:05:06,880 --> 00:05:11,520 615|over those 100 million tokens. And so 616| 617|155 618|00:05:08,720 --> 00:05:13,759 619|you could you could probably do some 620| 621|156 622|00:05:11,520 --> 00:05:15,919 623|kind of uh evaluation up to a certain 624| 625|157 626|00:05:13,759 --> 00:05:17,919 627|budget and then just say okay well this 628| 629|158 630|00:05:15,919 --> 00:05:19,520 631|is what we project the performance to 632| 633|159 634|00:05:17,919 --> 00:05:20,720 635|look like. And I think this this there 636| 637|160 638|00:05:19,520 --> 00:05:21,919 639|hasn't been a lot of research on this 640| 641|161 642|00:05:20,720 --> 00:05:23,039 643|yet. I actually think this would be a 644| 645|162 646|00:05:21,919 --> 00:05:24,000 647|great paper to publish if there's any 648| 649|163 650|00:05:23,039 --> 00:05:26,240 651|academics out there looking for 652| 653|164 654|00:05:24,000 --> 00:05:28,160 655|something to research. Can you predict 656| 657|165 658|00:05:26,240 --> 00:05:30,800 659|what the performance looks like at an 660| 661|166 662|00:05:28,160 --> 00:05:33,680 663|inference budget of let's say $10,000 664| 665|167 666|00:05:30,800 --> 00:05:34,560 667|only using inference budgets up to 10 or 668| 669|168 670|00:05:33,680 --> 00:05:37,039 671|10 or $100? 672| 673|169 674|00:05:34,560 --> 00:05:39,759 675|>> So maybe an orthogonal question for you. 676| 677|170 678|00:05:37,039 --> 00:05:41,360 679|Uh do you think users are systematically 680| 681|171 682|00:05:39,759 --> 00:05:42,960 683|like not thinking long enough with their 684| 685|172 686|00:05:41,360 --> 00:05:44,240 687|models about problems? 688| 689|173 690|00:05:42,960 --> 00:05:47,759 691|>> What do you mean by not thinking long 692| 693|174 694|00:05:44,240 --> 00:05:51,120 695|enough? um if the if you can 696| 697|175 698|00:05:47,759 --> 00:05:52,880 699|build an agent or control the amount of 700| 701|176 702|00:05:51,120 --> 00:05:54,320 703|um test time compute being used like 704| 705|177 706|00:05:52,880 --> 00:05:56,160 707|there's there's what is done by the 708| 709|178 710|00:05:54,320 --> 00:05:59,039 711|model itself and there's what the user 712| 713|179 714|00:05:56,160 --> 00:06:01,120 715|can do. Um do you think that like you 716| 717|180 718|00:05:59,039 --> 00:06:03,120 719|know the industry is using test time 720| 721|181 722|00:06:01,120 --> 00:06:04,960 723|compute an optimal amount uh way 724| 725|182 726|00:06:03,120 --> 00:06:06,080 727|undershooting it or it's you know it's a 728| 729|183 730|00:06:04,960 --> 00:06:07,680 731|problem in the models where they just 732| 733|184 734|00:06:06,080 --> 00:06:08,720 735|need to be able to do that thinking 736| 737|185 738|00:06:07,680 --> 00:06:11,759 739|faster. 740| 741|186 742|00:06:08,720 --> 00:06:14,240 743|>> I I think it depends on the problem. Uh 744| 745|187 746|00:06:11,759 --> 00:06:15,840 747|I think this idea that the models you 748| 749|188 750|00:06:14,240 --> 00:06:17,360 751|just let them think for a week or 752| 753|189 754|00:06:15,840 --> 00:06:19,199 755|whatever and then they respond. It's it 756| 757|190 758|00:06:17,360 --> 00:06:21,680 759|sounds nice and yes the benchmarks look 760| 761|191 762|00:06:19,199 --> 00:06:23,440 763|great but it's not very practical when 764| 765|192 766|00:06:21,680 --> 00:06:24,400 767|working u because like okay you ask the 768| 769|193 770|00:06:23,440 --> 00:06:25,520 771|model a question and then you sit there 772| 773|194 774|00:06:24,400 --> 00:06:26,319 775|for a week waiting for it to come back 776| 777|195 778|00:06:25,520 --> 00:06:29,919 779|to you. 780| 781|196 782|00:06:26,319 --> 00:06:31,360 783|>> I think what people have have found most 784| 785|197 786|00:06:29,919 --> 00:06:32,960 787|effective is to kind of like iterate 788| 789|198 790|00:06:31,360 --> 00:06:35,199 791|quickly with the models and so the 792| 793|199 794|00:06:32,960 --> 00:06:37,280 795|thinking time I think needs to be 796| 797|200 798|00:06:35,199 --> 00:06:38,479 799|flexible. when it makes sense to respond 800| 801|201 802|00:06:37,280 --> 00:06:39,919 803|quickly to the user, it should respond 804| 805|202 806|00:06:38,479 --> 00:06:41,199 807|quickly. And then when it makes sense to 808| 809|203 810|00:06:39,919 --> 00:06:42,240 811|think for a long time and the user wants 812| 813|204 814|00:06:41,199 --> 00:06:43,680 815|it to think for a long time, then it 816| 817|205 818|00:06:42,240 --> 00:06:44,720 819|makes sense to think for a long time. I 820| 821|206 822|00:06:43,680 --> 00:06:46,080 823|think people have been striking the 824| 825|207 826|00:06:44,720 --> 00:06:47,280 827|right balance given what they have to 828| 829|208 830|00:06:46,080 --> 00:06:49,120 831|deal with right now. 832| 833|209 834|00:06:47,280 --> 00:06:51,520 835|>> How would you characterize, you know, 836| 837|210 838|00:06:49,120 --> 00:06:53,759 839|there's a lot of talk about benchmark 840| 841|211 842|00:06:51,520 --> 00:06:55,680 843|maxing and the ability to game different 844| 845|212 846|00:06:53,759 --> 00:06:57,680 847|benchmarks? What would you characterize 848| 849|213 850|00:06:55,680 --> 00:06:59,199 851|the like landscape of benchmarks as 852| 853|214 854|00:06:57,680 --> 00:07:00,960 855|today? And then do you have like 856| 857|215 858|00:06:59,199 --> 00:07:03,199 859|favorites that you think are more 860| 861|216 862|00:07:00,960 --> 00:07:05,360 863|indicative of capability than others? So 864| 865|217 866|00:07:03,199 --> 00:07:06,880 867|the benchmark maxing thing is also 868| 869|218 870|00:07:05,360 --> 00:07:10,319 871|motivation for for writing the essay 872| 873|219 874|00:07:06,880 --> 00:07:12,240 875|that I think it's really easy to show 876| 877|220 878|00:07:10,319 --> 00:07:14,639 879|you can do much better than previous 880| 881|221 882|00:07:12,240 --> 00:07:17,360 883|benchmarks or or previous previous 884| 885|222 886|00:07:14,639 --> 00:07:18,880 887|models on benchmarks by just for example 888| 889|223 890|00:07:17,360 --> 00:07:20,800 891|scaffolding a bunch of models together. 892| 893|224 894|00:07:18,880 --> 00:07:22,240 895|Um so if you say okay well we're going 896| 897|225 898|00:07:20,800 --> 00:07:23,919 899|to instead of just running this model 900| 901|226 902|00:07:22,240 --> 00:07:25,520 903|once we're going to run it five times 904| 905|227 906|00:07:23,919 --> 00:07:27,680 907|and take the best of the five responses 908| 909|228 910|00:07:25,520 --> 00:07:29,759 911|or like ask a judge which one it thinks 912| 913|229 914|00:07:27,680 --> 00:07:32,639 915|is best then you can get much higher 916| 917|230 918|00:07:29,759 --> 00:07:35,759 919|scores than that model. And so it's 920| 921|231 922|00:07:32,639 --> 00:07:37,360 923|really easy to make something that looks 924| 925|232 926|00:07:35,759 --> 00:07:38,560 927|a lot better on paper but is actually 928| 929|233 930|00:07:37,360 --> 00:07:39,840 931|not better once you control for the 932| 933|234 934|00:07:38,560 --> 00:07:40,560 935|amount of test time compute. That is one 936| 937|235 938|00:07:39,840 --> 00:07:41,840 939|thing that I'm worried about when it 940| 941|236 942|00:07:40,560 --> 00:07:43,199 943|comes to benchmark maxing. I mean it 944| 945|237 946|00:07:41,840 --> 00:07:45,039 947|like it's it's a little misleading is 948| 949|238 950|00:07:43,199 --> 00:07:46,000 951|the only concern that I have. And then 952| 953|239 954|00:07:45,039 --> 00:07:48,960 955|as far as like the benchmarks 956| 957|240 958|00:07:46,000 --> 00:07:50,560 959|themselves, I think there is always a 960| 961|241 962|00:07:48,960 --> 00:07:53,199 963|risk of like just optimizing for the 964| 965|242 966|00:07:50,560 --> 00:07:55,199 967|benchmark. And I've I've certainly 968| 969|243 970|00:07:53,199 --> 00:07:56,720 971|encouraged my team and I think at OpenAI 972| 973|244 974|00:07:55,199 --> 00:07:58,960 975|we're pretty good about not trying to 976| 977|245 978|00:07:56,720 --> 00:08:00,800 979|optimize for for specific benchmarks. 980| 981|246 982|00:07:58,960 --> 00:08:02,160 983|But once you put out a benchmark, it's 984| 985|247 986|00:08:00,800 --> 00:08:04,160 987|there's it's always at risk of just 988| 989|248 990|00:08:02,160 --> 00:08:06,800 991|being optimized for. And I think one way 992| 993|249 994|00:08:04,160 --> 00:08:09,599 995|to one way to address that is to keep a 996| 997|250 998|00:08:06,800 --> 00:08:10,960 999|held out private set um that isn't uh 1000| 1001|251 1002|00:08:09,599 --> 00:08:14,479 1003|publicly available. 1004| 1005|252 1006|00:08:10,960 --> 00:08:16,960 1007|>> The most popular fallback advice for you 1008| 1009|253 1010|00:08:14,479 --> 00:08:18,639 1011|know figure out if a model is 1012| 1013|254 1014|00:08:16,960 --> 00:08:20,400 1015|significantly better or not is to just 1016| 1017|255 1018|00:08:18,639 --> 00:08:21,680 1019|play with it for a while. Do you have 1020| 1021|256 1022|00:08:20,400 --> 00:08:24,080 1023|anything more sophisticated than that 1024| 1025|257 1026|00:08:21,680 --> 00:08:27,039 1027|that you suggest people do? like do you 1028| 1029|258 1030|00:08:24,080 --> 00:08:29,840 1031|create your own set of new evaluations 1032| 1033|259 1034|00:08:27,039 --> 00:08:30,560 1035|each time besides private holdback at 1036| 1037|260 1038|00:08:29,840 --> 00:08:32,000 1039|OpenAI? 1040| 1041|261 1042|00:08:30,560 --> 00:08:32,959 1043|>> I think everybody has their own set of 1044| 1045|262 1046|00:08:32,000 --> 00:08:34,399 1047|questions that they like to ask the 1048| 1049|263 1050|00:08:32,959 --> 00:08:37,440 1051|model whenever it comes out. 1052| 1053|264 1054|00:08:34,399 --> 00:08:40,320 1055|>> Um for me lately it's been I I use them 1056| 1057|265 1058|00:08:37,440 --> 00:08:41,680 1059|to make poker bots and see how good they 1060| 1061|266 1062|00:08:40,320 --> 00:08:43,919 1063|can make a poker bot. I think it's a 1064| 1065|267 1066|00:08:41,680 --> 00:08:46,560 1067|nice eval because there is very little 1068| 1069|268 1070|00:08:43,919 --> 00:08:48,399 1071|um open source code for making poker 1072| 1073|269 1074|00:08:46,560 --> 00:08:49,680 1075|bots and there's a lot of published 1076| 1077|270 1078|00:08:48,399 --> 00:08:51,279 1079|essay there's a lot of published papers 1080| 1081|271 1082|00:08:49,680 --> 00:08:52,800 1083|on it but you really have to reason 1084| 1085|272 1086|00:08:51,279 --> 00:08:56,080 1087|through everything and it's like 1088| 1089|273 1090|00:08:52,800 --> 00:08:57,600 1091|requires a lot of just reasoning and 1092| 1093|274 1094|00:08:56,080 --> 00:08:58,480 1095|iteration and like a lot of small 1096| 1097|275 1098|00:08:57,600 --> 00:09:00,080 1099|gotchas 1100| 1101|276 1102|00:08:58,480 --> 00:09:01,360 1103|>> that I can kind of I've already worked 1104| 1105|277 1106|00:09:00,080 --> 00:09:02,880 1107|through myself so I can see where the 1108| 1109|278 1110|00:09:01,360 --> 00:09:04,480 1111|models fail along the way. They've 1112| 1113|279 1114|00:09:02,880 --> 00:09:07,279 1115|gotten really good at it now. Can you 1116| 1117|280 1118|00:09:04,480 --> 00:09:09,360 </file-view>
The file has 4408 lines, read only up to 1118. Need to read full transcript to summarize. We need to get entire content. Use read with larger? read default 20000 chars, lines 1-1118 maybe because limit? It returned first 1118 lines due to default char limit? Let's read in chunks. Use read with offset and limit. Need entire transcript. It may be too long for context but okay. We can process. Let's see total size: 77KB. Need to read all. Use read with limit huge? read can read up to 20000 chars. The file is ~77KB, so 4 chunks. We can use shell cat, but need to read file. The read tool supports offset. We can read full by calling with offset increments. Let's use cat to a local and maybe grep? But we need content for summarization. We can use exec to extract text (strip timecodes) and output. Maybe use python to parse SRT and print text. That would be more concise. We can pipe to a file. Let's create a clean transcript. Use Python script.
{
"command": "python3 - <<'PY'\nimport re, sys\npath = \"/tmp/Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown.en.srt\"\nwith open(path, 'r', encoding='utf-8') as f:\n data = f.read()\n# Remove SRT index lines, time lines, and empty subtitle lines; keep text lines\nlines = data.splitlines()\ntexts = []\ni = 0\nwhile i < len(lines):\n line = lines[i].strip()\n if re.match(r'^\\d+$', line):\n # next line is timecode\n i += 2\n # collect text until blank\n sub = []\n while i < len(lines) and lines[i].strip() != '':\n sub.append(lines[i].strip())\n i += 1\n if sub:\n texts.append(' '.join(sub))\n i += 1\nprint('\\n'.join(texts))\nPY",
"timeout": 60000
}Output from command in shell 1dbd60: With GBT3, you couldn't scale test time compute. Like if you gave it a budget of $10 million and said, "Okay, well, let's see what GB3 can do." It really can't do that much. The precurren frameworks and responsible scaling policies, they don't really account for the amount of testime compute. They just say, "Okay, well, what's the capability of the model?" The problem is we're in a world now where the capability of the model is a function of how much money you put into it. Basically, if you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. Give it a budget of $10 million, you can do even more. At what budget should you evaluate these models? The policies that exist today don't really address that question. Hi listeners, I'm Sarah Goa and welcome back to No Priors. Today I'm here with Nome Brown, one of our godfathers of AI reasoning. We talk about the broken state of evaluations, very large-scale test time compute, how he thinks about recursive self-improvement, and what's next on the horizon for competition at the frontier. Welcome, Gnome. I'm so excited to have you back. >> That's great to be back. Yeah, >> you are our first guest. I'm very proud of my taste in friends and researchers for uh for the pod. um given you know how important uh you know inference time scaling has become to the industry you should be proud too having actually pioneered it >> played a part yeah among among many others >> you just wrote this essay that really resonated about large scale test time compute and why uh the industry is not evaluating these models as robustly as it should be what was the motivation for it >> yeah the motivation was we released 5.5 and the initial reaction ction was kind of skepticism that it was a substantially better model. To be fair, that only lasted for a few hours before people had some time to play around with it and and try it out themselves and they saw that it was actually substantially better. Um, but I think a lot of the skepticism came from the benchmark grid that was published. Basically, whenever a new model is released, there's this benchmark grid where they show all these different benchmarks on on the x-axxis and then the performance of different models on the y- axis and you can just like compare different models. It's like a single number for a model on a single benchmark. And if you look on paper at the difference between like 5.5 and 5.4 or or other models, it wasn't it was an improvement, but it wasn't a huge improvement. It was only a few percentage points in some benchmarks. So, people looked at that and they were skeptical that it was actually a better model. Um, once they played around with it, the story changed. I think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented, the benchmark results are being presented in the wrong way. They're not controlling for the amount of test time compute that is being used on that benchmark question. It turned out that 5.5 is just much more efficient with its thinking. If you run it at max settings, 5.4 is thinking for a lot longer. It takes longer to get back a response than 5.5. And once you control for the amount of thinking time, actually you can see that 5.5 is a substantial jump over 5.4. That is, I think, people's day-to-day experience with it. And then when I mention this to people the reaction, the typical question I get is like okay well why not just have 5.5 think for as long as 5.4 and the question is like well how long should they think for? Typically the response I get is well until the performance plateaus right there's at some point where the performance on the benchmark is going to plateau and you evaluate to that point. The thing is the point at which it plateaus is actually really far out these days. I mean, it was true in GBD3 land back in 2022, the models couldn't really think productively for that long. And so, you could just run them until they plateau. It's not that far away. But what we're seeing today with the modern models is that 5.5 and other models can think for if you scaffold them reasonably well, can think for weeks even um before having performance plateau on some of these benchmarks. And so, the point at which they plateau is simply too far out to reasonably test. we all need to actually reinforce um either like a patience limit or a budget limit from a token perspective now and that wasn't true a few years ago. >> Exactly. And so I think the proper way to and so my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of test time compute that's going into the model and then it becomes much more clear um how to compare the performance between these different models. >> Given the model evaluation cycle and the fact that performance does not asmtote for many tasks over um quite a long period of time. What do you do about that issue? The fact that some of the evals that you would want to run are both beyond the scope of budget or time that's reasonable given the current model release cycle. >> I mean, I think for things like cyber, we've seen and actually the AIS um in their evaluations has shown that the models continue to improve um at 100 million tokens. You know, if you run them for 100 million tokens, they're still improving at beyond that point. >> And that can take a very long time to run. But you also do see that like the performance is is it's not just like a discontinuous jump. It's actually like you can see the slope of improvement over those 100 million tokens. And so you could you could probably do some kind of uh evaluation up to a certain budget and then just say okay well this is what we project the performance to look like. And I think this this there hasn't been a lot of research on this yet. I actually think this would be a great paper to publish if there's any academics out there looking for something to research. Can you predict what the performance looks like at an inference budget of let's say $10,000 only using inference budgets up to 10 or 10 or $100? >> So maybe an orthogonal question for you. Uh do you think users are systematically like not thinking long enough with their models about problems? >> What do you mean by not thinking long enough? um if the if you can build an agent or control the amount of um test time compute being used like there's there's what is done by the model itself and there's what the user can do. Um do you think that like you know the industry is using test time compute an optimal amount uh way undershooting it or it's you know it's a problem in the models where they just need to be able to do that thinking faster. >> I I think it depends on the problem. Uh I think this idea that the models you just let them think for a week or whatever and then they respond. It's it sounds nice and yes the benchmarks look great but it's not very practical when working u because like okay you ask the model a question and then you sit there for a week waiting for it to come back to you. >> I think what people have have found most effective is to kind of like iterate quickly with the models and so the thinking time I think needs to be flexible. when it makes sense to respond quickly to the user, it should respond quickly. And then when it makes sense to think for a long time and the user wants it to think for a long time, then it makes sense to think for a long time. I think people have been striking the right balance given what they have to deal with right now. >> How would you characterize, you know, there's a lot of talk about benchmark maxing and the ability to game different benchmarks? What would you characterize the like landscape of benchmarks as today? And then do you have like favorites that you think are more indicative of capability than others? So the benchmark maxing thing is also motivation for for writing the essay that I think it's really easy to show you can do much better than previous benchmarks or or previous previous models on benchmarks by just for example scaffolding a bunch of models together. Um so if you say okay well we're going to instead of just running this model once we're going to run it five times and take the best of the five responses or like ask a judge which one it thinks is best then you can get much higher scores than that model. And so it's really easy to make something that looks a lot better on paper but is actually not better once you control for the amount of test time compute. That is one thing that I'm worried about when it comes to benchmark maxing. I mean it like it's it's a little misleading is the only concern that I have. And then as far as like the benchmarks themselves, I think there is always a risk of like just optimizing for the benchmark. And I've I've certainly encouraged my team and I think at OpenAI we're pretty good about not trying to optimize for for specific benchmarks. But once you put out a benchmark, it's there's it's always at risk of just being optimized for. And I think one way to one way to address that is to keep a held out private set um that isn't uh publicly available. >> The most popular fallback advice for you know figure out if a model is significantly better or not is to just play with it for a while. Do you have anything more sophisticated than that that you suggest people do? like do you create your own set of new evaluations each time besides private holdback at OpenAI? >> I think everybody has their own set of questions that they like to ask the model whenever it comes out. >> Um for me lately it's been I I use them to make poker bots and see how good they can make a poker bot. I think it's a nice eval because there is very little um open source code for making poker bots and there's a lot of published essay there's a lot of published papers on it but you really have to reason through everything and it's like requires a lot of just reasoning and iteration and like a lot of small gotchas >> that I can kind of I've already worked through myself so I can see where the models fail along the way. They've gotten really good at it now. Can you describe perhaps like with your poker bot creation like how reasoning might have progressed in model releases for you guys over a few uh releases? >> Yeah, when the early models were really bad at it, like they could not basically do anything. And then 5.2 I was able to work with it to make a river solver. >> So that's like the final stage of poker. And um that itself was I thought really impressive. I mean I had to work with it a little bit, but I was actually really impressed cuz I was able to make the river solver probably about five times faster than I would have alone. There were a couple things that I got um tripped up on. Uh blockers was always a big big issue. But overall, like, you know, with a with a bit of gentle steering, it just kind of like it kind of felt like a grad student where, okay, they would run into issues, but at least like I would um know what those issues were and know how to fix it, and I could just make suggestions and it would go off and then do it and then pretty quickly it would actually come back with something really good. >> Mhm. >> And then especially the optimization I thought was very impressive. It was able to make it like 10 times faster than what I was able to do because it was just able to optimize um the code so well. The downsides with 5.2 is I felt like it was gaslighting me a lot and I always had to be very careful checking it and making sure like, okay, is it actually doing what it said it did? Um, are there any things that are like glaring issues that it's not recognizing or it's just pretending aren't issues? Um, I remember there was like one point uh where for one of the models I was playing around with it, not 5.2, I kind of like as a as a unit test, I told it, okay, well, let's say um I have $100 in the pot and I fold. How much am I losing? And the model said $92. And I was like, that's crazy. I have $100 in the pot and I just fold it. How do I not lose $100? And it said, oh, you know, it's 92. It's close to 100. It's fine. It's no big deal. And I was like, clearly this is a problem, right? So the models did have this problem where they would gaslight you a lot. But once we got to 5.5, I actually thought um it was way better. It was able to basically do it zero shot. And in fact, I've been working on just doing a full scale poker solver. Um and it it's basically able to do the whole thing uh with some gentle steering from me. And I wouldn't be surprised if you know 6 months or a year from now, the model is able to do zero shot an entire poker solver, basically my entire PhD thesis in one go. Let's talk about the larger implications of you know needing to evaluate these models relative to um let's say like speed of their reasoning or efficiency versus you know token volume right um or dollar budget or whatever whatever your scalar is. Can you describe some of the larger implications in your essay including around um like safety evaluations? Yeah, the safety evaluations thing. Um, it's it's a bit of an inconvenient truth thing where okay, so I guess for background, a lot of the all of the labs have these things called either responsible scaling policies, preparedness frameworks. They go by various names, but the idea is that whenever a model is released, they go through a series of evaluations to measure are there dangerous capabilities? Um, could these models do things that we're we we wouldn't want um a bad actor to do. And if the model isn't very capable, then it's no big deal. But if it is very capable, if it could be used for example to make bioweapons, then you want to put in mitigations against that. >> But the question is okay, well, how do you evaluate whether the model is capable of that? And they have like various protocols about like how they do these valuations. But a lot of these frameworks were developed around the era of chatgpt before after when test time comput scaling was not really uh as much of a thing and it made sense like with GPT3 you couldn't scale test time compute like if you gave it a budget of $10 million and said okay well let's see what GP3 can do it really can't do that much more than what you could do with like $10 or $1. the preparedness frameworks and responsible scaling policies, they don't really account for the amount of test time compute. They just say, "Okay, well, what's the capability of the model? The problem is we're in a world now where the capability of the model is a function of how much money you put into it. Basically, if you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. If you give it a budget of $10 million, it can do even more." And so, at what budget should you evaluate these models? the policies that exist today don't really address that question. >> Mhm. >> Uh some do some do better than others, but um for the most part, this is not really a factor that's being heavily considered. Now, whether it should be released anyway, I don't I don't want to wait into this question. I think there's, you know, there's arguments on both sides, but I think the important thing to recognize is that this is a question that is not being we're just kind of like, you know, pretending that this issue doesn't exist. And I think it's important to just, you know, one way or the other um account for it. Yeah, it's the mirror image of the capability question of if the models can uh continue to do more and more without asmmptoing on some tasks at very large budgets um then they should also be able to do so for uh tasks we don't want them to do as a society, right? And so testing for that um and what budget is allocated. It also seems out of sync from the model release cycle itself, right? there's been this acceleration of, you know, you get a new model every sometimes few days and weeks at this point versus 6 months. And you have a line in the essay where you say like the the only way to truly evaluate an agent on some very longunning task might be to run it for a year. And that's going to be true of both like useful and uh negative tasks, right? And so how do you think about that versus the model release cycle? Yeah, this this is also an interesting dynamic where basically as the models have become stronger, they've they're more they're better able to operate over longer horizons. >> So again, with GPD3, if you wanted to run it for, you know, a week, there's really not much you could do to scaffold it into something useful that could actually run for a week. But we're seeing now with the most recent models that you can actually scaffold, for example, 5.5 into doing a series of experiments that can run for weeks, for months. Have you given your poker solver task like infinite budget yet? >> I haven't really scaffolded something together where I just tell it like, "Okay, just run this for um for weeks." I think I could give it I I could ask I could probably give it / goal and just like yeah, tell it to go nuts. >> But um I I think at this point it could 100% do the river solver. Um, if I just give it slash goal, I don't think it's at the level yet where it could do like the full poker solver if I gave it just slash goal and told it, yeah, go go run for a month. But we're going to pretty soon be at that point where I probably could just tell it like, yeah, go work on this for a month and then come back to me with a a full complete poker solver that's state-of-the-art. And the problem is if you want to evaluate the capabilities of a model, what it can do after running for a month, the only way to be fully sure is to actually run it for a month. And if you want to know after 6 months, the only way to know fully is to run it for six months. Now there I I'll get to like things we could do to address that a little bit later. But like it's important to recognize the model release cycle is look, we're releasing new models like every two or three months at this point. And so a model comes out, it takes two or 3 months to push it to its limits and then you have another model come out. And so nobody actually knows what the ceiling of capabilities are for these models because nobody's actually run them for long enough to really tell. When slash goal came out, for example, I mean, people started running things that it took over a week for it to finish. And so, people actually didn't realize that this was a big deal until after a week uh until a week after it was released. >> Mhm. >> I think that's going to be more and more true. You know, the implications of that are, I think, pretty interesting because what do the labs do to like fully evaluate their models before they're released? It's actually very difficult because, yeah, you would have to the only way to to really do the evaluations is then delay the model release cycle. Um and you know there's a lot of competitive pressure right now to not do that. >> Do you think there's um like exciting latent capability in the models that are already released that people have not fully explored given timeline? >> I think absolutely. I think actually a really great example is the erdo unit distance problem. So for the viewers that don't know like we used an internal model at OpenAI a few weeks ago to uh disprove the unit erdos unit distance conjecture. Now, I'm not a mathematician, but this seems like it was a pretty big deal in the in the math community. It was like the first uh first problem that a lot of mathematicians had really spent a lot of time on, and the model was able to do something that they weren't able to do and do it in a way that was actually interesting and useful for mathematicians. Honestly, it did it at… (2 chars truncated) … (21 lines truncated) like … (31 chars truncated) then if you do this enough times it actually ends up arriving at the disproof. Now what this means is you could in principle ask 5.5 to you know as as a general purpose scaffold list a bunch of different strategies and then for each strategy tell it to investigate that strategy and then it would probably be able to arrive at the disproof with a general purpose scaffold. Now that scaffold would be very expensive. I mean it would probably cost I I just ballpark like thousand to 10 to $100,000. Um but it would be possible and it would have been possible for somebody to disprove the erdos unit distance conjecture before we did using a general purpose model. And nobody had explored sufficiently what happens if I put $100,000 worth of compute into 5.5 what could it do? Um and the answer is like yeah you probably could get stuff like that out of it. So people should be experimenting more with the current generation in terms of >> well this is I think is an interesting question of is it worth it to experiment with because again the model release cycle is every every couple months we put out a new model that's even more powerful and so the cost of disproving the erdosha distance gesture drops by like 10 or 100x with every model release cycle probably in some cases more so >> you've seen the meme that's like oh I like why why bother doing any engineering work when I should just wait for the next >> model go on vacation come back two months later and and it's, you know, a thousand times cheaper. So, >> do you agree with that? >> I >> Is that what you're doing right now at OpenAI, just waiting for the next model release? >> I think I I mean, I will say we're we're in a period where progress is very fast and like yeah, the models are becoming more capable. I I can say like at OpenAI one of the things that we're we're actively not doing and look we have a lot of mathematicians we have a lot of physicists people are very excited about what these models can do right now especially you know the internal models we are trying to encourage people to not spend all their time just like going through all the mathematical open problems physics problems and um just seeing pushing the models to their limits to see what they can prove or disprove. Um because we really think the focus should be on how do we make even more capable models? how can we get them get them out safely to the world as quickly as possible so that all the scientists in the world can use these models to solve the problems themselves. Uh so yeah in some sense we are thinking about this that yes it's really tempting to um just put all of our efforts into scaling up these models and see what they can do at their limits right now but really the focus should be on how do we use these models to make even more powerful models even more capable models that can do everything much more cost-ffectively. What is changing about the uh direction or allocation of resources for research in your mind given your beliefs about this very large scale the impact of very large scale test time compute? How does this interact with the um idea of recursive self-improvement for example where you know it's a dominant idea for how you know any lab gets to the best capability model. So, one thing I should clarify, I don't think we're at the point where, okay, you just give it an arbitrary extremely high inference budget and it's just it's just super intelligent across the board. >> SL goal, >> yeah, >> make GPD 7 or whatever and yeah, just go nuts. >> What's between us and there then? >> I think having played around with the model uh so okay, so first of all, there are some benchmarks where the models will just not improve if they have more inference budget. So, I think a lot of factual ret uh factual um retrieval kind of questions fall into this category of if you ask a person when was Abraham Lincoln born and they don't know the date. They could sit there, they could think about it for a week, if they if they don't have access to Wikipedia or something, they're not going to be able to do better answering that question if they thought about it for a week compared to 5 seconds. Same with the model. If you actually, interestingly enough, if you give the model these kinds of like factual retrieval questions and you give them a little bit of time to think, they do actually do better. um but if you give them a week they're not suddenly going to do better at remembering dates. There are so there's some benchmarks where they clearly improve with more test time compute and there's some where they don't. I think on the other extreme there are um benchmarks where they kind of obviously will keep improving limit uh without limit with more test time compute. So the example I like to point to is sudoku. If you uh it's there's a really simple strategy to solving Sudoku which is just try a bunch of different random numbers and then see if it fits the criteria. If it if it matches all the constraints and if it doesn't just try a different random combination of numbers and clearly with enough time you will be able to solve any puzzle with this strategy. You can kind of trivally see like okay any model could keep doing better and better um if it was just given more test time compute. So you have and and all the benchmarks kind of exist somewhere between these two extremes. the models are not at the level where if you just give them enough test time on compute they will be able to do um all of our jobs just because yeah there's some benchmarks where they will not improve there are some things where they they will not improve one thing I see for research in particular is they don't have very good research taste right now and so I think they're actually a very good complement to researchers especially um you know I've found like I've found I'm much more effective by using these models but they're not able to fully replace the whole research cycle Now, does that change with time? Probably. I mean, I think the models are getting better across the board. Some some things are getting better the faster than others. Um, but they're not at the point where they're fully replacing uh researchers with just enough test time compute. >> Can you give um an example or two of like asking the model to do a research test for just like this is this is a terrible idea? I mean I I think going back to my poker solver example, I was really impressed with the model's ability to optimize the algorithms that I had developed in my in my PhD. It was honestly it it was it was shocking to see how inefficient I was um in retrospect and they were able to make it like you know 10 100x faster. And then I was like, okay, can you come up with an algorithm that is better than the algorithms that I came up with or that anybody else came up with and go ahead and like look at all the published work and synthesize that and then try to come up with something novel and it it's not able to do it. And I can give it a lot of time and it's it's still not able to do it. Now, it's possible that if I scaffolded something and like kind of constrained it a bit more that maybe it could eventually come up with something better. Um, but it would take a lot of it's not just as simple as saying, "Okay, please come up with a better algorithm." >> And how do you think that that gets improved? >> What I've seen is with every model release cycle, um, it does get better at this sort of thing. It's still it's still bad in my opinion, but it it's not as bad as it used to be. And I wouldn't be surprised if at some point, same thing with coding, same thing with math, where there's just like this inflection point where suddenly it's actually good enough to be useful. Um I wouldn't be surprised if we encounter that point for research taste as well. >> Given that what do you like what is your framing of RSI today? Like how should we think about it? >> The models are definitely accelerating what researchers can do inside the labs. But I think they are accelerating some things and not other things. And currently we're at the point where okay if something goes 100x faster you get bottlenecked by the things that don't go 100x faster. over time the things that we're getting bottlenecked on are going to shrink and and there will be I think um a kind of a gradual takeoff in that respect but it's more about transforming right now it's more about transforming what researchers do rather than fully replacing the researchers >> so that actually implies that you don't think we're close to a very fast takeoff right now >> I think fast takeoff is relative things are moving very fast but I think there is this hypothesis that you could have basically an overnight intelligence explosion where the models discover some kind of breakthrough through to make themselves smarter and then that leads to more breakthroughs that make themselves even smarter immediately and you have basically in an instance the models just you know becoming very superhuman across the board uh in in moments and I don't think we're headed to that world largely because of the fact that the models rely so much on large scale test time compute in order to achieve um their greatest intelligence. If you if it requires so much test time on compute to unlock the full capabilities of the model, then that means you're bottlenecked by time um things can only go so fast because the models need to run for long enough to um to actually do something really really powerful. Time itself becomes a bottleneck to what we can do and I I think that is the case right now for a lot of the labs that ultimately I think the biggest bottleneck for all of us is time and that's why all the researchers are working so intensely right now. It's it's just so many um so many hours per week are being put into this because we all see what the overhang is. We see what the capabilities are and we're just bottlenecked by how quickly can we do things. >> What do you think is on the frontier um that is less explored now? Like we've talked about multi-agent before. >> I think multi-agent is quite explored. Um I think >> at sufficient scale, >> I think there's a lot more that could be done. Um, but it's also one of the things that's that's hard to do at a lot of research is hard to do at small skin. I think multi- agent in particular, it really requires in order to fully unlock the cubilities. I think requires like frontier models. I think we've seen some pretty interesting multi- aent scaffolds. I think um they're able to do a lot, but I think it's really just scratching the surface of what it will be able to do. I mean, one way that I think about it is if if you look at human civilization, it's not that humans have become smarter over it's not that they evolved to become smarter over, you know, the past 50,000 years. It's that humans are able to do a lot more today than they were back in caveman times because there have been billions of humans thinking for a long time and building off of each other's accumulated knowledge. >> We have like very good retrieval and scaffolding versus 50,000 years ago. Uh it's not even I wouldn't even call it a scaffold. It's just like this is a very like organic emerging property of just like humans being able to accumulate knowledge, share it, um and build off of it. We're not seeing that with AI models today. They kind of they're they're born into a world for and they exist for a very short context window and then they just like disappear. >> And yeah, there are things that you can kind of do to like continue them, but it's very limited. I do think eventually we will and we're starting to see like signs that we're entering a world where they can coordinate on a large scale. I think multbook and open claw when they first came out I think it was obviously a bit overhyped but they were um an indication of where things could go in the future and I do think that eventually we get to that kind of world >> of some sort of coordinated compounding state. >> Yeah. uh the ability of the models to to share knowledge um on a more global level and be able to build on that knowledge productively. >> Given this set of beliefs and your work like uh how would you characterize just competition at the frontier between the between the three kingdoms if there is no overnight takeoff? It's just researchers grinding away trying to make good high taste algorithmic and investment decisions about where to go and then compute allocation and then policy decisions and eval decisions. >> Um >> that feels like slightly more grounded than um uh I support I suppose like racing towards some immediate hard takeoff that nobody can catch you on. >> I think the competition is very intense right now. I do think the models that exist today um are accelerating what researchers at the Frontier Labs can do. Um there's obviously like I said limits to that right now, but it the ability to use the models to improve the model research is a real thing and it's um it it is like an amplifying force. I think that will continue to be true. I think they'll become more true over time. One thing that I am comforted by is I think all the researchers at the frontier labs like all the frontier labs I think recognize what is at stake and what these models like what what the uh what the risks are and that's something that I I find comforting that I think everybody really understands like okay this is a pretty serious thing and it can lead to really great things or it can lead to really bad things and yes there's a competitive dynamic between the labs but like we can also try to figure out how we all get to the positive outcomes rather than the very negative outcomes. I I think you know I'd be remiss to ask just because you have been right very early for a long time um on the importance of test time compute and reasoning as a framework like are there ways in which you use the models that you should you would encourage others to right is it just goal everything >> I think for a lot of people they worked I mean this is probably not even true for your audience necessarily but there's a lot of people that experimented with AI back in like 2012 23 and felt like they couldn't trust the outputs and then don't use it for really high stakes decisions. And actually, I think the models have progressed to a point where they are very good for these kinds of things. I mean, I asked them tax advice or I bought a condo recently and I was asking it for advice on like, okay, well, what's all the paperwork that I have to fill out and like how to how do I what does it all mean? It's actually really good for these kinds of questions. Mhm. >> So, I use it dayto-day for for a lot of this kind of stuff. And I think they're at a point now where they've actually been at a point for a while now where I feel like I can just trust the outputs arguably more than I could trust the output from from a human, >> an expert human. Yeah. >> Okay. I have two two final questions for you. Um, one is uh is there something you think that the rest of the research community doesn't agree with you on or doesn't understand the importance of quite yet? Oh, these are such good questions. I wish I had time to think about this ahead of time. >> You can you can just hang out with me and think about it. Yeah. >> Is it weird to be like consensus now? You're a bit salty three years ago when you're like, why don't people understand how important this is? >> I still I still feel like it's not consensus though because like you know people still don't publish the benchmarks this way. >> Oh, that's true. Yeah. >> Like that's actually why >> I think that's like inertia. >> That's kind of Yeah, but that's kind of why I wrote the essay. I was just like look I mean we can talk about this but like yeah the part of the motivation is like I would talk to researchers about um we it makes sense to show the benchmarks with an x-axis whether it's tokens or cost or time there should be an x-axis and everybody would say like yeah that makes sense we should do that but everybody >> they're not acting with the importance of like good heart like this is we have to measure the correct thing >> well really their response is people expect us to to publish the grid >> and then okay why do people expect the grid to be published because everybody publishes the grid. And so you kind of end up in this this bad equilibrium where everybody kind of knows that it's a bad equilibrium, but like nobody wants to break out. And I I felt like, okay, well, if I just hopefully come out and say like, look guys, let's all recognize that we're in a bad equilibrium and let's move to this different equilibrium where we're we're plotting things with an X-axis that hopefully that can, you know, next time there's a model release, a company can feel comfortable not publishing the grid, at least not at the very front uh the top line. and um we can have a more productive evaluation of these models. >> Then a last question for you. How do you think about companies across all of these specialized domains who feel the value that they have is essentially like the routing layer, the choice layer of you know my goal is composed of a bunch of discrete tasks. Some require more intelligence and less and within my my job as a vendor is to um solve that problem or achieve the optimal outcome with uh taking into account the budget constraints. And so I will manage like the parallelization and how how much inference do you spend on it from what model because I I think the the frontier lab point of view is that that routing happens both within the you know behind the API behind the application and then some of it in the model itself. Um and that's pieces of that are clearly being externalized in all these applications. >> Yeah, I do I do think this is related to the fact that like benchmarks should be evaluated with an x-axis of tokens or cost. Um I I have seen some evals recently that show like okay well with with a routing layer you can achieve much better performance um by basically doing consensus among the models. >> Yeah. >> And like I definitely believe that if you do consensus among the models that you're going to achieve better performance than any individual model but >> it's important to ask like are you going to do better than having that model basically think for longer. Um like once you control for the amount of test time compute is it is it actually still doing better? That's that's the question that you want to figure out. Okay, that's very principle of view which is like yes routing is fine but it's all subject to the same budget question. Yeah, >> right. Um if you put it on the same scaler then you can make an optimal decision and I think maybe I win. >> Mhm. I I I don't even know necessarily that I I would believe that the routing does better but then there's still a question of um is it going to do significantly better? Is it very fragile? Is it um reflective of real world use cases compared to benchmarks? because like one issue you could run into is that you could optimize for certain benchmarks with the routing and then show like oh yeah we see this big improvement on these benchmarks. Um but in real world use cases it actually ends up not being a significant improvement. So I I would say at the very least like I would say you want to control for testime compute and then you also want to have all the same um skepticism about benchmarks that you would normally have. >> Awesome. N thanks so much and and for being on the mission for uh breaking us out of this false equilibrium. Yeah, it's great to be back. >> Find us on Twitter at no prior pod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way, you get a new episode every week. And sign up for emails or find transcripts for every episode at no-briers.com. Exit code: 0 <truncation_notice> Full output written to: /var/folders/d9/mzhyr7lx6xbf236j5ll9v_300000gn/T/devin-overflows-501/28f3adc3/content.txt </truncation_notice>
The transcript output was truncated in the tool output but full output written to overflow file. It seems we got the majority, but there were some truncation notices "21 lines truncated", "31 chars truncated" etc within the Python print due to terminal? Actually the transcript itself appears complete but with some "... (2 chars truncated)" in the text. We need full transcript. The overflow file path is /var/folders/d9/mzhyr7lx6xbf236j5ll9v_300000gn/T/devin-overflows-501/28f3adc3/content.txt. Need to read it fully.
{
"file_path": "/var/folders/d9/mzhyr7lx6xbf236j5ll9v_300000gn/T/devin-overflows-501/28f3adc3/content.txt"
}<file-view path="/var/folders/d9/mzhyr7lx6xbf236j5ll9v_300000gn/T/devin-overflows-501/28f3adc3/content.txt" start_line="1" end_line="538" total_lines="1106"> 1|Output from command in shell 1dbd60: 2|With GBT3, you couldn't scale test time 3|compute. Like if you gave it a budget of 4|$10 million and said, "Okay, well, let's 5|see what GB3 can do." It really can't do 6|that much. The precurren frameworks and 7|responsible scaling policies, they don't 8|really account for the amount of testime 9|compute. They just say, "Okay, well, 10|what's the capability of the model?" The 11|problem is we're in a world now where 12|the capability of the model is a 13|function of how much money you put into 14|it. Basically, if you give it a budget 15|of $10,000, it can do a lot more than 16|what it can do with a budget of $10. 17|Give it a budget of $10 million, you can 18|do even more. At what budget should you 19|evaluate these models? The policies that 20|exist today don't really address that 21|question. 22|Hi listeners, I'm Sarah Goa and welcome 23|back to No Priors. Today I'm here with 24|Nome Brown, one of our godfathers of AI 25|reasoning. We talk about the broken 26|state of evaluations, very large-scale 27|test time compute, how he thinks about 28|recursive self-improvement, and what's 29|next on the horizon for competition at 30|the frontier. Welcome, Gnome. I'm so 31|excited to have you back. 32|>> That's great to be back. Yeah, 33|>> you are our first guest. I'm very proud 34|of my taste in friends and researchers 35|for uh for the pod. um given you know 36|how important uh you know inference time 37|scaling has become to the industry you 38|should be proud too having actually 39|pioneered it 40|>> played a part yeah among among many 41|others 42|>> you just wrote this essay that really 43|resonated about large scale test time 44|compute and why uh the industry is not 45|evaluating these models as robustly as 46|it should be what was the motivation for 47|it 48|>> yeah the motivation was we released 5.5 49|and the initial reaction ction was kind 50|of skepticism that it was a 51|substantially better model. To be fair, 52|that only lasted for a few hours before 53|people had some time to play around with 54|it and and try it out themselves and 55|they saw that it was actually 56|substantially better. Um, but I think a 57|lot of the skepticism came from the 58|benchmark grid that was published. 59|Basically, whenever a new model is 60|released, there's this benchmark grid 61|where they show all these different 62|benchmarks on on the x-axxis and then 63|the performance of different models on 64|the y- axis and you can just like 65|compare different models. It's like a 66|single number for a model on a single 67|benchmark. And if you look on paper at 68|the difference between like 5.5 and 5.4 69|or or other models, it wasn't it was an 70|improvement, but it wasn't a huge 71|improvement. It was only a few 72|percentage points in some benchmarks. 73|So, people looked at that and they were 74|skeptical that it was actually a better 75|model. Um, once they played around with 76|it, the story changed. I think the 77|reason why it doesn't show up as so much 78|better on the benchmarks is because the 79|benchmarks are being presented, the 80|benchmark results are being presented in 81|the wrong way. They're not controlling 82|for the amount of test time compute that 83|is being used on that benchmark 84|question. It turned out that 5.5 is just 85|much more efficient with its thinking. 86|If you run it at max settings, 5.4 is 87|thinking for a lot longer. It takes 88|longer to get back a response than 5.5. 89|And once you control for the amount of 90|thinking time, actually you can see that 91|5.5 is a substantial jump over 5.4. That 92|is, I think, people's day-to-day 93|experience with it. And then when I 94|mention this to people the reaction, the 95|typical question I get is like okay well 96|why not just have 5.5 think for as long 97|as 5.4 and the question is like well how 98|long should they think for? Typically 99|the response I get is well until the 100|performance plateaus right there's at 101|some point where the performance on the 102|benchmark is going to plateau and you 103|evaluate to that point. The thing is the 104|point at which it plateaus is actually 105|really far out these days. I mean, it 106|was true in GBD3 land back in 2022, the 107|models couldn't really think 108|productively for that long. And so, you 109|could just run them until they plateau. 110|It's not that far away. But what we're 111|seeing today with the modern models is 112|that 5.5 and other models can think for 113|if you scaffold them reasonably well, 114|can think for weeks even um before 115|having performance plateau on some of 116|these benchmarks. And so, the point at 117|which they plateau is simply too far out 118|to reasonably test. we all need to 119|actually reinforce um either like a 120|patience limit or a budget limit from a 121|token perspective now and that wasn't 122|true a few years ago. 123|>> Exactly. And so I think the proper way 124|to and so my claim is the proper way to 125|evaluate the models now is you either 126|have some kind of budget for the 127|benchmark whether it's tokens or cost or 128|time or whatever or you plot the 129|performance as a function of the amount 130|of test time compute that's going into 131|the model and then it becomes much more 132|clear um how to compare the performance 133|between these different models. 134|>> Given the model evaluation cycle and the 135|fact that performance does not asmtote 136|for many tasks over um quite a long 137|period of time. What do you do about 138|that issue? The fact that some of the 139|evals that you would want to run are 140|both beyond the scope of budget or time 141|that's reasonable given the current 142|model release cycle. 143|>> I mean, I think for things like cyber, 144|we've seen and actually the AIS um in 145|their evaluations has shown that the 146|models continue to improve um at 100 147|million tokens. You know, if you run 148|them for 100 million tokens, they're 149|still improving at beyond that point. 150|>> And that can take a very long time to 151|run. But you also do see that like the 152|performance is is it's not just like a 153|discontinuous jump. It's actually like 154|you can see the slope of improvement 155|over those 100 million tokens. And so 156|you could you could probably do some 157|kind of uh evaluation up to a certain 158|budget and then just say okay well this 159|is what we project the performance to 160|look like. And I think this this there 161|hasn't been a lot of research on this 162|yet. I actually think this would be a 163|great paper to publish if there's any 164|academics out there looking for 165|something to research. Can you predict 166|what the performance looks like at an 167|inference budget of let's say $10,000 168|only using inference budgets up to 10 or 169|10 or $100? 170|>> So maybe an orthogonal question for you. 171|Uh do you think users are systematically 172|like not thinking long enough with their 173|models about problems? 174|>> What do you mean by not thinking long 175|enough? um if the if you can 176|build an agent or control the amount of 177|um test time compute being used like 178|there's there's what is done by the 179|model itself and there's what the user 180|can do. Um do you think that like you 181|know the industry is using test time 182|compute an optimal amount uh way 183|undershooting it or it's you know it's a 184|problem in the models where they just 185|need to be able to do that thinking 186|faster. 187|>> I I think it depends on the problem. Uh 188|I think this idea that the models you 189|just let them think for a week or 190|whatever and then they respond. It's it 191|sounds nice and yes the benchmarks look 192|great but it's not very practical when 193|working u because like okay you ask the 194|model a question and then you sit there 195|for a week waiting for it to come back 196|to you. 197|>> I think what people have have found most 198|effective is to kind of like iterate 199|quickly with the models and so the 200|thinking time I think needs to be 201|flexible. when it makes sense to respond 202|quickly to the user, it should respond 203|quickly. And then when it makes sense to 204|think for a long time and the user wants 205|it to think for a long time, then it 206|makes sense to think for a long time. I 207|think people have been striking the 208|right balance given what they have to 209|deal with right now. 210|>> How would you characterize, you know, 211|there's a lot of talk about benchmark 212|maxing and the ability to game different 213|benchmarks? What would you characterize 214|the like landscape of benchmarks as 215|today? And then do you have like 216|favorites that you think are more 217|indicative of capability than others? So 218|the benchmark maxing thing is also 219|motivation for for writing the essay 220|that I think it's really easy to show 221|you can do much better than previous 222|benchmarks or or previous previous 223|models on benchmarks by just for example 224|scaffolding a bunch of models together. 225|Um so if you say okay well we're going 226|to instead of just running this model 227|once we're going to run it five times 228|and take the best of the five responses 229|or like ask a judge which one it thinks 230|is best then you can get much higher 231|scores than that model. And so it's 232|really easy to make something that looks 233|a lot better on paper but is actually 234|not better once you control for the 235|amount of test time compute. That is one 236|thing that I'm worried about when it 237|comes to benchmark maxing. I mean it 238|like it's it's a little misleading is 239|the only concern that I have. And then 240|as far as like the benchmarks 241|themselves, I think there is always a 242|risk of like just optimizing for the 243|benchmark. And I've I've certainly 244|encouraged my team and I think at OpenAI 245|we're pretty good about not trying to 246|optimize for for specific benchmarks. 247|But once you put out a benchmark, it's 248|there's it's always at risk of just 249|being optimized for. And I think one way 250|to one way to address that is to keep a 251|held out private set um that isn't uh 252|publicly available. 253|>> The most popular fallback advice for you 254|know figure out if a model is 255|significantly better or not is to just 256|play with it for a while. Do you have 257|anything more sophisticated than that 258|that you suggest people do? like do you 259|create your own set of new evaluations 260|each time besides private holdback at 261|OpenAI? 262|>> I think everybody has their own set of 263|questions that they like to ask the 264|model whenever it comes out. 265|>> Um for me lately it's been I I use them 266|to make poker bots and see how good they 267|can make a poker bot. I think it's a 268|nice eval because there is very little 269|um open source code for making poker 270|bots and there's a lot of published 271|essay there's a lot of published papers 272|on it but you really have to reason 273|through everything and it's like 274|requires a lot of just reasoning and 275|iteration and like a lot of small 276|gotchas 277|>> that I can kind of I've already worked 278|through myself so I can see where the 279|models fail along the way. They've 280|gotten really good at it now. Can you 281|describe perhaps like with your poker 282|bot creation like how reasoning might 283|have progressed in model releases for 284|you guys over a few uh releases? 285|>> Yeah, when the early models were really 286|bad at it, like they could not basically 287|do anything. And then 5.2 288|I was able to work with it to make a 289|river solver. 290|>> So that's like the final stage of poker. 291|And um that itself was I thought really 292|impressive. I mean I had to work with it 293|a little bit, but I was actually really 294|impressed cuz I was able to make the 295|river solver probably about five times 296|faster than I would have alone. There 297|were a couple things that I got um 298|tripped up on. Uh blockers was always a 299|big big issue. But overall, like, you 300|know, with a with a bit of gentle 301|steering, it just kind of like it kind 302|of felt like a grad student where, okay, 303|they would run into issues, but at least 304|like I would um know what those issues 305|were and know how to fix it, and I could 306|just make suggestions and it would go 307|off and then do it and then pretty 308|quickly it would actually come back with 309|something really good. 310|>> Mhm. 311|>> And then especially the optimization I 312|thought was very impressive. It was able 313|to make it like 10 times faster than 314|what I was able to do because it was 315|just able to optimize um the code so 316|well. The downsides with 5.2 is I felt 317|like it was gaslighting me a lot and I 318|always had to be very careful checking 319|it and making sure like, okay, is it 320|actually doing what it said it did? Um, 321|are there any things that are like 322|glaring issues that it's not recognizing 323|or it's just pretending aren't issues? 324|Um, I remember there was like one point 325|uh where for one of the models I was 326|playing around with it, not 5.2, 327|I kind of like as a as a unit test, I 328|told it, okay, well, let's say um I have 329|$100 in the pot and I fold. How much am 330|I losing? And the model said $92. And I 331|was like, that's crazy. I have $100 in 332|the pot and I just fold it. How do I not 333|lose $100? And it said, oh, you know, 334|it's 92. It's close to 100. It's fine. 335|It's no big deal. And I was like, 336|clearly this is a problem, right? So the 337|models did have this problem where they 338|would gaslight you a lot. But once we 339|got to 5.5, I actually thought um it was 340|way better. It was able to basically do 341|it zero shot. And in fact, I've been 342|working on just doing a full scale poker 343|solver. Um and it it's basically able to 344|do the whole thing uh with some gentle 345|steering from me. And I wouldn't be 346|surprised if you know 6 months or a year 347|from now, the model is able to do zero 348|shot an entire poker solver, basically 349|my entire PhD thesis in one go. Let's 350|talk about the larger implications of 351|you know needing to evaluate these 352|models relative to um let's say like 353|speed of their reasoning or efficiency 354|versus you know token volume right um or 355|dollar budget or whatever whatever your 356|scalar is. Can you describe some of the 357|larger implications in your essay 358|including around um like safety 359|evaluations? Yeah, the safety 360|evaluations thing. Um, it's it's a bit 361|of an inconvenient truth thing where 362|okay, so I guess for background, a lot 363|of the all of the labs have these things 364|called either responsible scaling 365|policies, preparedness frameworks. They 366|go by various names, but the idea is 367|that whenever a model is released, they 368|go through a series of evaluations to 369|measure are there dangerous 370|capabilities? Um, could these models do 371|things that we're we we wouldn't want um 372|a bad actor to do. And if the model 373|isn't very capable, then it's no big 374|deal. But if it is very capable, if it 375|could be used for example to make 376|bioweapons, then you want to put in 377|mitigations against that. 378|>> But the question is okay, well, how do 379|you evaluate whether the model is 380|capable of that? And they have like 381|various protocols about like how they do 382|these valuations. But a lot of these 383|frameworks were developed around the era 384|of chatgpt before after when test time 385|comput scaling was not really uh as much 386|of a thing and it made sense like with 387|GPT3 you couldn't scale test time 388|compute like if you gave it a budget of 389|$10 million and said okay well let's see 390|what GP3 can do it really can't do that 391|much more than what you could do with 392|like $10 or $1. the preparedness 393|frameworks and responsible scaling 394|policies, they don't really account for 395|the amount of test time compute. They 396|just say, "Okay, well, what's the 397|capability of the model? The problem is 398|we're in a world now where the 399|capability of the model is a function of 400|how much money you put into it. 401|Basically, if you give it a budget of 402|$10,000, it can do a lot more than what 403|it can do with a budget of $10. If you 404|give it a budget of $10 million, it can 405|do even more." And so, at what budget 406|should you evaluate these models? the 407|policies that exist today don't really 408|address that question. 409|>> Mhm. 410|>> Uh some do some do better than others, 411|but um for the most part, this is not 412|really a factor that's being heavily 413|considered. Now, whether it should be 414|released anyway, I don't I don't want to 415|wait into this question. I think 416|there's, you know, there's arguments on 417|both sides, but I think the important 418|thing to recognize is that this is a 419|question that is not being we're just 420|kind of like, you know, pretending that 421|this issue doesn't exist. And I think 422|it's important to just, you know, one 423|way or the other um account for it. 424|Yeah, it's the mirror image of the 425|capability question of if the models can 426|uh continue to do more and more without 427|asmmptoing on some tasks at very large 428|budgets um then they should also be able 429|to do so for uh tasks we don't want them 430|to do as a society, right? And so 431|testing for that um and what budget is 432|allocated. It also seems out of sync 433|from the model release cycle itself, 434|right? there's been this acceleration 435|of, you know, you get a new model every 436|sometimes few days and weeks at this 437|point versus 6 months. And you have a 438|line in the essay where you say like the 439|the only way to truly evaluate an agent 440|on some very longunning task might be to 441|run it for a year. And that's going to 442|be true of both like useful and uh 443|negative tasks, right? And so how do you 444|think about that versus the model 445|release cycle? Yeah, this this is also 446|an interesting dynamic where basically 447|as the models have become stronger, 448|they've they're more they're better able 449|to operate over longer horizons. 450|>> So again, with GPD3, if you wanted to 451|run it for, you know, a week, there's 452|really not much you could do to scaffold 453|it into something useful that could 454|actually run for a week. But we're 455|seeing now with the most recent models 456|that you can actually scaffold, for 457|example, 5.5 into doing a series of 458|experiments that can run for weeks, for 459|months. Have you given your poker solver 460|task like infinite budget yet? 461|>> I haven't really scaffolded something 462|together where I just tell it like, 463|"Okay, just run this for um for weeks." 464|I think I could give it I I could ask I 465|could probably give it / goal and just 466|like yeah, tell it to go nuts. 467|>> But um I I think at this point it could 468|100% do the river solver. Um, if I just 469|give it slash goal, I don't think it's 470|at the level yet where it could do like 471|the full poker solver if I gave it just 472|slash goal and told it, yeah, go go run 473|for a month. But we're going to pretty 474|soon be at that point where I probably 475|could just tell it like, yeah, go work 476|on this for a month and then come back 477|to me with a a full complete poker 478|solver that's state-of-the-art. 479|And the problem is if you want to 480|evaluate the capabilities of a model, 481|what it can do after running for a 482|month, the only way to be fully sure is 483|to actually run it for a month. And if 484|you want to know after 6 months, the 485|only way to know fully is to run it for 486|six months. Now there I I'll get to like 487|things we could do to address that a 488|little bit later. But like it's 489|important to recognize the model release 490|cycle is look, we're releasing new 491|models like every two or three months at 492|this point. And so a model comes out, it 493|takes two or 3 months to push it to its 494|limits and then you have another model 495|come out. And so nobody actually knows 496|what the ceiling of capabilities are for 497|these models because nobody's actually 498|run them for long enough to really tell. 499|When slash goal came out, for example, I 500|mean, people started running things that 501|it took over a week for it to finish. 502|And so, people actually didn't realize 503|that this was a big deal until after a 504|week uh until a week after it was 505|released. 506|>> Mhm. 507|>> I think that's going to be more and more 508|true. You know, the implications of that 509|are, I think, pretty interesting because 510|what do the labs do to like fully 511|evaluate their models before they're 512|released? It's actually very difficult 513|because, yeah, you would have to the 514|only way to to really do the evaluations 515|is then delay the model release cycle. 516|Um and you know there's a lot of 517|competitive pressure right now to not do 518|that. 519|>> Do you think there's um like exciting 520|latent capability in the models that are 521|already released that people have not 522|fully explored given timeline? 523|>> I think absolutely. I think actually a 524|really great example is the erdo unit 525|distance problem. So for the viewers 526|that don't know like we used an internal 527|model at OpenAI a few weeks ago to uh 528|disprove the unit erdos unit distance 529|conjecture. Now, I'm not a 530|mathematician, but this seems like it 531|was a pretty big deal in the in the math 532|community. It was like the first uh 533|first problem that a lot of 534|mathematicians had really spent a lot of 535|time on, and the model was able to do 536|something that they weren't able to do 537|and do it in a way that was actually 538|interesting and useful for </file-view>
File has 1106 lines; read returned up to 538. Need remainder.
{
"file_path": "/var/folders/d9/mzhyr7lx6xbf236j5ll9v_300000gn/T/devin-overflows-501/28f3adc3/content.txt",
"offset": 539
}<file-view path="/var/folders/d9/mzhyr7lx6xbf236j5ll9v_300000gn/T/devin-overflows-501/28f3adc3/content.txt" start_line="539" end_line="1079" total_lines="1106"> 539|mathematicians. Honestly, it did it at a 540|budget that was dirt cheap. I mean, we 541|didn't put a lot of effort into this. We 542|just we trained a new model and we were 543|just curious what it could do. And we 544|ran it through some problems. And this 545|one at a pretty low budget, it was like, 546|"Oh, yeah. I think I have a disproof." 547|And then we were able to verify that 548|yeah that this proof is correct after we 549|announced the results a bunch of people 550|found that you could get the answer out 551|of 5.5 as well if now it's not as simple 552|as just asking 5.5 hey here's the unit 553|distance conjecture what's the disproof 554|you had to scaffold it a bit you had to 555|like steer it a bit and so somebody 556|found okay you ask 5.5 list a bunch of 557|ways that you could tackle this problem 558|and then for it it lists one of the 559|paths that are actually promising to get 560|to the disproof and then you tell it 561|like okay explore this some more and 562|then if you do this enough times it 563|actually ends up arriving at the 564|disproof. Now what this means is you 565|could in principle ask 5.5 to you know 566|as as a general purpose scaffold list a 567|bunch of different strategies and then 568|for each strategy tell it to investigate 569|that strategy and then it would probably 570|be able to arrive at the disproof with a 571|general purpose scaffold. Now that 572|scaffold would be very expensive. I mean 573|it would probably cost I I just ballpark 574|like thousand to 10 to $100,000. Um but 575|it would be possible and it would have 576|been possible for somebody to disprove 577|the erdos unit distance conjecture 578|before we did using a general purpose 579|model. And nobody had explored 580|sufficiently what happens if I put 581|$100,000 worth of compute into 5.5 what 582|could it do? Um and the answer is like 583|yeah you probably could get stuff like 584|that out of it. So people should be 585|experimenting more with the current 586|generation in terms of 587|>> well this is I think is an interesting 588|question of is it worth it to experiment 589|with because again the model release 590|cycle is every every couple months we 591|put out a new model that's even more 592|powerful and so the cost of disproving 593|the erdosha distance gesture drops by 594|like 10 or 100x with every model release 595|cycle probably in some cases more so 596|>> you've seen the meme that's like oh I 597|like why why bother doing any 598|engineering work when I should just wait 599|for the next 600|>> model go on vacation come back two 601|months later and and it's, you know, a 602|thousand times cheaper. So, 603|>> do you agree with that? 604|>> I 605|>> Is that what you're doing right now at 606|OpenAI, just waiting for the next model 607|release? 608|>> I think I I mean, I will say we're we're 609|in a period where progress is very fast 610|and like yeah, the models are becoming 611|more capable. I I can say like at OpenAI 612|one of the things that we're we're 613|actively not doing and look we have a 614|lot of mathematicians we have a lot of 615|physicists people are very excited about 616|what these models can do right now 617|especially you know the internal models 618|we are trying to encourage people to not 619|spend all their time just like going 620|through all the mathematical open 621|problems physics problems and um just 622|seeing pushing the models to their 623|limits to see what they can prove or 624|disprove. Um because we really think the 625|focus should be on how do we make even 626|more capable models? how can we get them 627|get them out safely to the world as 628|quickly as possible so that all the 629|scientists in the world can use these 630|models to solve the problems themselves. 631|Uh so yeah in some sense we are thinking 632|about this that yes it's really tempting 633|to um just put all of our efforts into 634|scaling up these models and see what 635|they can do at their limits right now 636|but really the focus should be on how do 637|we use these models to make even more 638|powerful models even more capable models 639|that can do everything much more 640|cost-ffectively. 641|What is changing about the uh direction 642|or allocation of resources for research 643|in your mind given your beliefs about 644|this very large scale the impact of very 645|large scale test time compute? How does 646|this interact with the um idea of 647|recursive self-improvement for example 648|where you know it's a dominant idea for 649|how you know any lab gets to the best 650|capability model. So, one thing I should 651|clarify, I don't think we're at the 652|point where, okay, you just give it an 653|arbitrary extremely high inference 654|budget and it's just it's just super 655|intelligent across the board. 656|>> SL goal, 657|>> yeah, 658|>> make GPD 7 or whatever and yeah, just go 659|nuts. 660|>> What's between us and there then? 661|>> I think having played around with the 662|model uh so okay, so first of all, there 663|are some benchmarks where the models 664|will just not improve if they have more 665|inference budget. So, I think a lot of 666|factual ret uh factual um retrieval kind 667|of questions fall into this category of 668|if you ask a person when was Abraham 669|Lincoln born and they don't know the 670|date. They could sit there, they could 671|think about it for a week, if they if 672|they don't have access to Wikipedia or 673|something, they're not going to be able 674|to do better answering that question if 675|they thought about it for a week 676|compared to 5 seconds. Same with the 677|model. If you actually, interestingly 678|enough, if you give the model these 679|kinds of like factual retrieval 680|questions and you give them a little bit 681|of time to think, they do actually do 682|better. um but if you give them a week 683|they're not suddenly going to do better 684|at remembering dates. There are so 685|there's some benchmarks where they 686|clearly improve with more test time 687|compute and there's some where they 688|don't. I think on the other extreme 689|there are um benchmarks where they kind 690|of obviously will keep improving limit 691|uh without limit with more test time 692|compute. So the example I like to point 693|to is sudoku. If you uh it's there's a 694|really simple strategy to solving Sudoku 695|which is just try a bunch of different 696|random numbers and then see if it fits 697|the criteria. If it if it matches all 698|the constraints and if it doesn't just 699|try a different random combination of 700|numbers and clearly with enough time you 701|will be able to solve any puzzle with 702|this strategy. You can kind of trivally 703|see like okay any model could keep doing 704|better and better um if it was just 705|given more test time compute. So you 706|have and and all the benchmarks kind of 707|exist somewhere between these two 708|extremes. the models are not at the 709|level where if you just give them enough 710|test time on compute they will be able 711|to do um all of our jobs just because 712|yeah there's some benchmarks where they 713|will not improve there are some things 714|where they they will not improve one 715|thing I see for research in particular 716|is they don't have very good research 717|taste right now and so I think they're 718|actually a very good complement to 719|researchers especially um you know I've 720|found like I've found I'm much more 721|effective by using these models but 722|they're not able to fully replace the 723|whole research cycle 724|Now, does that change with time? 725|Probably. I mean, I think the models are 726|getting better across the board. Some 727|some things are getting better the 728|faster than others. Um, but they're not 729|at the point where they're fully 730|replacing uh researchers with just 731|enough test time compute. 732|>> Can you give um an example or two of 733|like asking the model to do a research 734|test for just like this is this is a 735|terrible idea? I mean I I think going 736|back to my poker solver example, I was 737|really impressed with the model's 738|ability to optimize the algorithms that 739|I had developed in my in my PhD. It was 740|honestly it it was it was shocking to 741|see how inefficient I was um in 742|retrospect and they were able to make it 743|like you know 10 100x faster. And then I 744|was like, okay, can you come up with an 745|algorithm that is better than the 746|algorithms that I came up with or that 747|anybody else came up with and go ahead 748|and like look at all the published work 749|and synthesize that and then try to come 750|up with something novel and it it's not 751|able to do it. And I can give it a lot 752|of time and it's it's still not able to 753|do it. Now, it's possible that if I 754|scaffolded something and like kind of 755|constrained it a bit more that maybe it 756|could eventually come up with something 757|better. Um, but it would take a lot of 758|it's not just as simple as saying, 759|"Okay, please come up with a better 760|algorithm." 761|>> And how do you think that that gets 762|improved? 763|>> What I've seen is with every model 764|release cycle, um, it does get better at 765|this sort of thing. It's still it's 766|still bad in my opinion, but it it's not 767|as bad as it used to be. And I wouldn't 768|be surprised if at some point, same 769|thing with coding, same thing with math, 770|where there's just like this inflection 771|point where suddenly it's actually good 772|enough to be useful. Um I wouldn't be 773|surprised if we encounter that point for 774|research taste as well. 775|>> Given that what do you like what is your 776|framing of RSI today? Like how should we 777|think about it? 778|>> The models are definitely accelerating 779|what researchers can do inside the labs. 780|But I think they are accelerating some 781|things and not other things. And 782|currently we're at the point where okay 783|if something goes 100x faster you get 784|bottlenecked by the things that don't go 785|100x faster. over time the things that 786|we're getting bottlenecked on are going 787|to shrink and and there will be I think 788|um a kind of a gradual takeoff in that 789|respect but it's more about transforming 790|right now it's more about transforming 791|what researchers do rather than fully 792|replacing the researchers 793|>> so that actually implies that you don't 794|think we're close to a very fast takeoff 795|right now 796|>> I think fast takeoff is relative things 797|are moving very fast but I think there 798|is this hypothesis that you could have 799|basically an overnight intelligence 800|explosion where the models discover some 801|kind of breakthrough through to make 802|themselves smarter and then that leads 803|to more breakthroughs that make 804|themselves even smarter immediately and 805|you have basically in an instance the 806|models just you know becoming very 807|superhuman across the board uh in in 808|moments and I don't think we're headed 809|to that world largely because of the 810|fact that the models rely so much on 811|large scale test time compute in order 812|to achieve um their greatest 813|intelligence. If you if it requires so 814|much test time on compute to unlock the 815|full capabilities of the model, then 816|that means you're bottlenecked by time 817|um things can only go so fast because 818|the models need to run for long enough 819|to um to actually do something really 820|really powerful. Time itself becomes a 821|bottleneck to what we can do and I I 822|think that is the case right now for a 823|lot of the labs that ultimately I think 824|the biggest bottleneck for all of us is 825|time and that's why all the researchers 826|are working so intensely right now. It's 827|it's just so many um so many hours per 828|week are being put into this because we 829|all see what the overhang is. We see 830|what the capabilities are and we're just 831|bottlenecked by how quickly can we do 832|things. 833|>> What do you think is on the frontier um 834|that is less explored now? Like we've 835|talked about multi-agent before. 836|>> I think multi-agent is quite explored. 837|Um I think 838|>> at sufficient scale, 839|>> I think there's a lot more that could be 840|done. Um, but it's also one of the 841|things that's that's hard to do at a lot 842|of research is hard to do at small skin. 843|I think multi- agent in particular, it 844|really requires in order to fully unlock 845|the cubilities. I think requires like 846|frontier models. I think we've seen some 847|pretty interesting multi- aent 848|scaffolds. I think um they're able to do 849|a lot, but I think it's really just 850|scratching the surface of what it will 851|be able to do. I mean, one way that I 852|think about it is if if you look at 853|human civilization, it's not that humans 854|have become smarter over it's not that 855|they evolved to become smarter over, you 856|know, the past 50,000 years. It's that 857|humans are able to do a lot more today 858|than they were back in caveman times 859|because there have been billions of 860|humans thinking for a long time and 861|building off of each other's accumulated 862|knowledge. 863|>> We have like very good retrieval and 864|scaffolding versus 50,000 years ago. Uh 865|it's not even I wouldn't even call it a 866|scaffold. It's just like this is a very 867|like organic emerging property of just 868|like humans being able to accumulate 869|knowledge, share it, um and build off of 870|it. We're not seeing that with AI models 871|today. They kind of they're they're born 872|into a world for and they exist for a 873|very short context window and then they 874|just like disappear. 875|>> And yeah, there are things that you can 876|kind of do to like continue them, but 877|it's very limited. I do think eventually 878|we will and we're starting to see like 879|signs that we're entering a world where 880|they can coordinate on a large scale. I 881|think multbook and open claw when they 882|first came out I think it was obviously 883|a bit overhyped but they were um 884|an indication of where things could go 885|in the future and I do think that 886|eventually we get to that kind of world 887|>> of some sort of coordinated compounding 888|state. 889|>> Yeah. uh the ability of the models to to 890|share knowledge um on a more global 891|level and be able to build on that 892|knowledge productively. 893|>> Given this set of beliefs and your work 894|like uh how would you characterize just 895|competition at the frontier between the 896|between the three kingdoms if there is 897|no overnight takeoff? It's just 898|researchers grinding away trying to make 899|good high taste algorithmic and 900|investment decisions about where to go 901|and then compute allocation and then 902|policy decisions and eval decisions. 903|>> Um 904|>> that feels like slightly more grounded 905|than um uh I support I suppose like 906|racing towards some immediate hard 907|takeoff that nobody can catch you on. 908|>> I think the competition is very intense 909|right now. I do think the models that 910|exist today um are accelerating what 911|researchers at the Frontier Labs can do. 912|Um there's obviously like I said limits 913|to that right now, but it the ability to 914|use the models to improve the model 915|research is a real thing and it's um it 916|it is like an amplifying force. I think 917|that will continue to be true. I think 918|they'll become more true over time. One 919|thing that I am comforted by is I think 920|all the researchers at the frontier labs 921|like all the frontier labs I think 922|recognize what is at stake and what 923|these models like what what the uh what 924|the risks are and that's something that 925|I I find comforting that I think 926|everybody really understands like okay 927|this is a pretty serious thing and it 928|can lead to really great things or it 929|can lead to really bad things and yes 930|there's a competitive dynamic between 931|the labs but like we can also try to 932|figure out how we all get to the 933|positive outcomes rather than the very 934|negative outcomes. I I think you know 935|I'd be remiss to ask just because you 936|have been right very early for a long 937|time um on the importance of test time 938|compute and reasoning as a framework 939|like are there ways in which you use the 940|models that you should you would 941|encourage others to right is it just 942|goal everything 943|>> I think for a lot of people 944|they worked I mean this is probably not 945|even true for your audience necessarily 946|but there's a lot of people that 947|experimented with AI back in like 2012 948|23 and felt like they couldn't trust the 949|outputs and then don't use it for really 950|high stakes decisions. And actually, I 951|think the models have progressed to a 952|point where they are very good for these 953|kinds of things. I mean, I asked them 954|tax advice or I bought a condo recently 955|and I was asking it for advice on like, 956|okay, well, what's all the paperwork 957|that I have to fill out and like how to 958|how do I what does it all mean? It's 959|actually really good for these kinds of 960|questions. Mhm. 961|>> So, I use it dayto-day for for a lot of 962|this kind of stuff. And I think they're 963|at a point now where they've actually 964|been at a point for a while now where I 965|feel like I can just trust the outputs 966|arguably more than I could trust the 967|output from from a human, 968|>> an expert human. Yeah. 969|>> Okay. I have two two final questions for 970|you. Um, one is uh is there something 971|you think that the rest of the research 972|community doesn't agree with you on or 973|doesn't understand the importance of 974|quite yet? Oh, these are such good 975|questions. I wish I had time to think 976|about this ahead of time. 977|>> You can you can just hang out with me 978|and think about it. Yeah. 979|>> Is it weird to be like consensus now? 980|You're a bit salty three years ago when 981|you're like, why don't people understand 982|how important this is? 983|>> I still I still feel like it's not 984|consensus though because like you know 985|people still don't publish the 986|benchmarks this way. 987|>> Oh, that's true. Yeah. 988|>> Like that's actually why 989|>> I think that's like inertia. 990|>> That's kind of Yeah, but that's kind of 991|why I wrote the essay. I was just like 992|look I mean we can talk about this but 993|like yeah the part of the motivation is 994|like I would talk to researchers about 995|um we it makes sense to show the 996|benchmarks with an x-axis whether it's 997|tokens or cost or time there should be 998|an x-axis and everybody would say like 999|yeah that makes sense we should do that 1000|but everybody 1001|>> they're not acting with the importance 1002|of like good heart like this is we have 1003|to measure the correct thing 1004|>> well really their response is people 1005|expect us to to publish the grid 1006|>> and then okay why do people expect the 1007|grid to be published because everybody 1008|publishes the grid. And so you kind of 1009|end up in this this bad equilibrium 1010|where everybody kind of knows that it's 1011|a bad equilibrium, but like nobody wants 1012|to break out. And I I felt like, okay, 1013|well, if I just hopefully come out and 1014|say like, look guys, let's all recognize 1015|that we're in a bad equilibrium and 1016|let's move to this different equilibrium 1017|where we're we're plotting things with 1018|an X-axis that hopefully that can, you 1019|know, next time there's a model release, 1020|a company can feel comfortable not 1021|publishing the grid, at least not at the 1022|very front uh the top line. and um we 1023|can have a more productive evaluation of 1024|these models. 1025|>> Then a last question for you. How do you 1026|think about companies across all of 1027|these specialized domains who feel the 1028|value that they have is essentially like 1029|the routing layer, the choice layer of 1030|you know my goal is composed of a bunch 1031|of discrete tasks. Some require more 1032|intelligence and less and within my my 1033|job as a vendor is to um solve that 1034|problem or achieve the optimal outcome 1035|with uh taking into account the budget 1036|constraints. And so I will manage like 1037|the parallelization and how how much 1038|inference do you spend on it from what 1039|model because I I think the the frontier 1040|lab point of view is that that routing 1041|happens both within the you know behind 1042|the API behind the application and then 1043|some of it in the model itself. Um and 1044|that's pieces of that are clearly being 1045|externalized in all these applications. 1046|>> Yeah, I do I do think this is related to 1047|the fact that like benchmarks should be 1048|evaluated with an x-axis of tokens or 1049|cost. Um I I have seen some evals 1050|recently that show like okay well with 1051|with a routing layer you can achieve 1052|much better performance um by basically 1053|doing consensus among the models. 1054|>> Yeah. 1055|>> And like I definitely believe that if 1056|you do consensus among the models that 1057|you're going to achieve better 1058|performance than any individual model 1059|but 1060|>> it's important to ask like are you going 1061|to do better than having that model 1062|basically think for longer. Um like once 1063|you control for the amount of test time 1064|compute is it is it actually still doing 1065|better? That's that's the question that 1066|you want to figure out. Okay, that's 1067|very principle of view which is like yes 1068|routing is fine but it's all subject to 1069|the same budget question. Yeah, 1070|>> right. Um if you put it on the same 1071|scaler then you can make an optimal 1072|decision and I think maybe I win. 1073|>> Mhm. I I I don't even know necessarily 1074|that I I would believe that the routing 1075|does better but then there's still a 1076|question of um is it going to do 1077|significantly better? Is it very 1078|fragile? Is it um reflective of real 1079|world use cases compared to benchmarks? </file-view>
Need the final lines.
{
"file_path": "/var/folders/d9/mzhyr7lx6xbf236j5ll9v_300000gn/T/devin-overflows-501/28f3adc3/content.txt",
"offset": 1080
}<file-view path="/var/folders/d9/mzhyr7lx6xbf236j5ll9v_300000gn/T/devin-overflows-501/28f3adc3/content.txt" start_line="1080" end_line="1106" total_lines="1106"> 1080|because like one issue you could run 1081|into is that you could optimize for 1082|certain benchmarks with the routing and 1083|then show like oh yeah we see this big 1084|improvement on these benchmarks. Um but 1085|in real world use cases it actually ends 1086|up not being a significant improvement. 1087|So I I would say at the very least like 1088|I would say you want to control for 1089|testime compute and then you also want 1090|to have all the same um skepticism about 1091|benchmarks that you would normally have. 1092|>> Awesome. N thanks so much and and for 1093|being on the mission for uh breaking us 1094|out of this false equilibrium. Yeah, 1095|it's great to be back. 1096|>> Find us on Twitter at no prior pod. 1097|Subscribe to our YouTube channel if you 1098|want to see our faces. Follow the show 1099|on Apple Podcasts, Spotify, or wherever 1100|you listen. That way, you get a new 1101|episode every week. And sign up for 1102|emails or find transcripts for every 1103|episode at no-briers.com. 1104| 1105| 1106|Exit code: 0 </file-view>
Now I have full transcript. Need to summarize. This is a No Priors podcast episode with Sarah Guo? "Sarah Goa" transcription. Guest Noam Brown (transcribed Nome/Gnome) from OpenAI. Main topics: test-time compute scaling, broken evaluations, benchmark grids, safety evals, recursive self-improvement, multi-agent, competition. I need to provide a concise summary. Since user only pasted link, according to rule "use yt-dlp to summarize". We can output summary. Should I include title and details? Yes. Need to be direct. Could include bullet points. Ensure not too verbose. Should I mention that transcript was auto-generated captions? Not necessary, but perhaps note summary from auto-generated captions. Need to cite file? Not required but can cite the transcript file path. The user didn't ask for session name. We shouldn't name session. We can respond with summary. Ensure we use no emojis. Also maybe include key quotes. The rule says "get the autogenerated captions to do this". We did. We should summarize. Let's craft. Title: "Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown" (No Priors podcast). Host Sarah Guo, guest Noam Brown. Need to produce a high-quality summary, maybe structured with sections: - Core argument: capability now depends on inference budget; benchmark comparisons should use performance vs test-time compute (tokens/cost/time) not single grid numbers. - Why it matters for evaluations: models can think for weeks/months before plateauing; release cycles too fast for full evaluation; benchmark maxing and scaffolding can mislead unless compute controlled. - Safety implications: responsible scaling policies don't account for compute budget; dangerous capability depends on budget; need to choose evaluation budget. - Examples: poker solver (5.2 needed steering, 5.5 near zero-shot); Erdős unit-distance conjecture disproof done cheaply by internal model and reproducible with 5.5 + scaffolding at $1k-$100k; shows latent capabilities unexplored. - Research/RSI: models accelerate researchers but not replace them; poor research taste; some tasks improve with compute, some plateau (factual retrieval vs Sudoku). No overnight hard takeoff because time/compute is bottleneck; gradual takeoff more likely. - Multi-agent and coordination: human civilization analogy, current models lack persistent shared knowledge; multi-agent scaffolds promising but require frontier models and large scale. - Competition: intense; frontier labs aware of stakes; using models to improve models is amplifying. - Practical advice: use models for high-stakes daily decisions; trust outputs now; benchmarks should include x-axis; routing/consensus layers only useful if compute-controlled and validated on real tasks. Need to be accurate. Let's include timestamps? Not necessary. Could mention names correctly. The transcript has many typos; we should use real names: Noam Brown, Sarah Guo, OpenAI o3? It says "5.5" and "5.4" likely "o3" or "o1"? It says "we released 5.5" probably "o3"? Actually OpenAI released "o3"? The transcript says "5.5" maybe "o3 mini high"? Hmm. In transcript, they refer to "5.5" and "5.4" maybe "o1-preview"? Wait OpenAI's Noam Brown, "5.5" likely "o3"? The model names might be "o3" and "o3-mini"? But transcript says "5.5" and "5.4" perhaps "o1" or "o3"? We can quote as "5.5/5.4" but maybe it's "o3"? The transcription has "GBT3" for GPT-3, "GBD3", "GPD3", "5.5" likely "o3" or "o1"? Could be "o3"? Since Noam Brown's essay was about test-time compute and "o3"? Hmm. Actually the title says "Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown". I recall a No Priors episode with Noam Brown from around Feb/Mar 2025? There was OpenAI's "o3" and "o3-mini". He may have written about "o3-mini"? Wait "5.5" could be "GPT-4.5"? In transcript they say "we released 5.5" and compare to "5.4". That likely refers to OpenAI's "o3-mini" and "o1"? Actually OpenAI released "o1" in Sept 2024, "o1-preview", "o1-mini", then "o3" Dec 2024. The Noam Brown essay "How Test-Time Compute Changes Everything" might have been about "o3". The transcript's "5.5" perhaps a transcription error for "o3"? But it also says "5.2" when making a river solver. OpenAI's "o1-preview" was first o-series, "o1" then "o3-mini"? Wait 5.2, 5.4, 5.5 could be model version numbers for "o1"? Actually OpenAI's o-series maybe codenamed? There is "o1-preview", "o1", "o3-mini", "o3". The transcript's "5.5" may be a mis-recognition of "o1"? No, "5.5" two syllables? Could be "o1.5"? OpenAI hasn't released "o1.5". Hmm. Let's search web for the actual episode title to identify. Use web search.
{
"query": "\"Really Big Test-Time Compute\" Noam Brown No Priors"
}# Web Search Results for ""Really Big Test-Time Compute" Noam Brown No Priors" ## 1. No Priors: Artificial Intelligence | Technology | Startups - Podcast Episode listening links - Plink URL: https://plinkhq.com/i/1668002688/e/1000774329129 No Priors ... ## Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown ## 2. Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown URL: https://castbox.fm/episode/Really-Big-Test-Time-Compute-in-AI-Changes-Benchmarks%2C-Safety-and-Research-with-OpenAI-Research-Scientist-Noam-Brown-id5296519-id961348912 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown ... Discover No Priors: Artificial Intelligence | Technology | Startups Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown ... # Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown ... Current AI evaluation methods are insufficient because they fail to account for test-time compute, a factor that significantly impacts model capability. As models improve with increased budget (tokens, cost, time), traditional benchmarks become misleading. The rapid release cycle of new models exacerbates this issue, preventing thorough evaluation and understanding of their true potential. This discussion highlights the need for new evaluation strategies, such as plotting performance against test-time compute, to accurately assess AI advancements. It also touches upon the evolution of AI reasoning through tasks like poker bot creation, the challenges in AI safety evaluations concerning long-term tasks, and the role of AI in research as an optimizer rather than a complete replacement for human researchers. The concept of recursive self-improvement and a gradual acceleration of AI capabilities are explored, emphasizing that time, not just algorithmic breakthroughs, is the ultimate bottleneck. Finally, the importance of breaking the current "bad equilibrium" in benchmark publishing is stressed to foster more productive and honest AI evaluations. ... Current AI evaluation frameworks, including responsible scaling policies, are insufficient because they do not account for test-time compute. Model capability is now a function of budget, making existing evaluation methods inadequate for understanding true performance and potential. ... Skepticism surrounds new model releases due to benchmark scores that don't reflect true improvement because test-t... ## 3. OpenAI's Noam Brown, Ilge Akkaya and Hunter Lightman on o1 and Teaching LLMs to Reason Better URL: https://www.youtube.com/watch?v=jPluSXJpdrA # OpenAI's Noam Brown, Ilge Akkaya and Hunter Lightman on o1 and Teaching LLMs to Reason Better ... Combining LLMs with AlphaGo-style deep reinforcement learning has been a holy grail for many leading AI labs, and with o1 (aka Strawberry) we are seeing the most general merging of the two modes to date. o1 is admittedly better at math than essay writing, but it has already achieved SOTA on a number of math, coding and reasoning benchmarks. ... OpenAI researchers Noam Brown, Ilge Akkaya and Hunter Lightman discuss the ah-ha moments on the way to the release of o1, how it uses chains of thought and backtracking to think through problems, the discovery of strong test-time compute scaling laws and what to expect as the model gets better. Hosted by: Sonya Huang and Pat Grady, Sequoia Capital ## Chapters ... - 00:00 Introduction - 01:33 Conviction in o1 - 04:24 ... 05:0 ... ? - 07:02 Lessons from gameplay - 09:14 Generation vs verification - 10:31 What is surprising about o1 so far - 11:37 The trough of disillusionment - 14:03 Applying deep RL - 14:45 o1’s AlphaGo moment? - 17:38 A-ha moments - 21:10 Why is o1 good at STEM? - 24:10 Capabilities vs usefulness - 25:29 Defining AGI - 26:13 The importance of reasoning - 28:39 Chain of thought - 30:41 Implication of inference-time scaling laws - 35:10 Bottlenecks to scaling test-time compute - 38:46 Biggest misunderstanding about o1? - 41:13 o1-mini - 42:15 How should founders think about o1? ## 4. Begin Proof — Noam Brown URL: https://www.youtube.com/watch?v=h4ZguzEMKAU In the inaugural episode of Begin Proof, Slater Stich sits down with Noam Brown, who leads work on multi-agent reasoning and test-time compute at OpenAI, and was one of the key people behind o1, OpenAI's first public reasoning model. They talk about what the Erdős result actually means, where the models are superhuman and where they still aren't, and what working mathematicians will be doing in a few years. ... At OpenAI, Noam focuses on multi-agent reasoning and test time compute scaling. He was one of the key people behind O1, OpenAI's first public reasoning model. ... >> You know, it's interesting like the models have reached a point where they have surpassed my ability to really understand what they're capable of and like what they excel at just because like already the you know, these kinds of problems like understanding the proofs so is just beyond my ability. I'm sure it's beyond most people's ability, but it's also beyond my ability. And so it's hard for me at this point to say like, oh yeah, these are like the weak points of the models and this these are like the strong points of the models in the field of math. ... >> And then there's the separate question of like reasoning through the the the complex problems. And this I think like the argument it's arguable that they're not superhuman yet, but we're also seeing rapid progress on this. And so I think we're going to see this kind of spiky intelligence where um the models are going to be superhuman in some respects and then also not quite superhuman in other respects and the point but over time as the models become more capable they are going to become like more capable across the board and that's what we're seeing. It's that it's not just like oh they're getting better and better at ingesting a lot of mathematics and combining it. It's that they're getting better at that but they're also getting better at just like the fundamental reasoning capabilities. ... So I I think we will in the short term see more ... ## 5. The field is underestimating inference compute | Noam Brown URL: https://www.youtube.com/watch?v=wKwLDaPP6YI # The field is underestimating inference compute | Noam Brown ... Um now thinking that you have your own I'm going to call it OAI bench you can call it whatever you want. Um what are your most interesting you know, there's there's many isms that I we we talked about before. Uh there's the Ilyaism which to be fair to Ilya Sutskever he doesn't necessarily believe this anymore but at NeurIPS back in the day he definitely did saying NTP is enough. All we need is more tokens bro. There is the uh uh gnomeism that I would call it which is Ilyaism plus gen verifier gap since we have all but one internet. Well, if we have a bunch of verifiable problems we can generate and then verify against them and that creates a lot more data and then now we can do stuff. And so um do you agree with that phrasing of gnomeism or has it been is that uh Yeah, I don't think it captures I don't think it captures the whole story. I mean obviously you're we're trying to distill a lot of stuff down into like a single sentence. Uh I think another key part in my opinion is the inference compute the fact that like a lot of into especially reasoning capabilities, I think the ability to um think be be more productive by thinking for longer. And so it's not just I mean it is a form of next token prediction where you're predicting a series of tokens now. ... And I think that the counter to narrative on just test time compute is that um well, if I if I have a contrived example such as I'm trying to learn sort and mankind doesn't know about merge sort yet. But what do you know about bubble sort? And so I have infinite traces of um unsorted lists and then the procedure called bubble sort and then the final answer which is the sorted list um and I can train on infinite amounts of those. That's verifiable, right? Gen verify, I can generate a lot of them. ... It's been trained to do bubble sort exactly. The reasoning trace was bubble sort. Sure, if all you trained it on in its entire existence is bubble sort ...
Search result 2 has a summary. Need not rely. We have transcript. We can now summarize. Should we name session? Rule says if asked to summarize a YouTube video, do not name session until after reading and understanding full transcript. We have read transcript, so okay to respond. But user didn't ask for name. So just summarize. Need to consider whether to mention "Noam Brown" not "Nome/Gnome". Use correct names. Also "Sarah Guo" host. The model versions: transcript says "5.5" maybe "o3" or "o3-mini"? We can call it "the latest reasoning model" to avoid misidentification. But search result summary mentions "o1"? Wait title from search result says "OpenAI Research Scientist Noam Brown" and "o1" in other. But in this episode they might be talking about "o3" and "o3-mini" (released around early 2025). The transcript "5.5" could be "o3-mini high" vs "5.4"? Let's search for "5.5" within the video? Could be "o3.5"? Actually maybe they say "o3.5" and transcript mishears as "5.5"? OpenAI released "o3-mini" in Jan 2025, then "o3" in April 2025? There is "o3-mini-high" and "o3". Hmm. In early 2025 maybe "o3"? Noam Brown's essay could be about "o3"? Let's see. The transcript "we released 5.5" maybe "we released o3.5"? Did OpenAI release an "o3.5"? Not as of July 2026? Actually by July 2026 there could be many. But the transcript is from a video uploaded at some date. Let's search for the podcast episode date.
{
"query": "\"Really Big Test-Time Compute in AI Changes Benchmarks\" Noam Brown date"
}# Web Search Results for ""Really Big Test-Time Compute in AI Changes Benchmarks" Noam Brown date" ## 1. Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown – No Priors: Artificial Intelligence | Technology | Startups URL: https://barbellinsights.com/podcast/no-priors-artificial-intelligence-technology-startups/why-traditional-benchmarks-fail-modern-ai-models-with-openai-research-scientist Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown – No Priors: Artificial Intelligence | Technology | Startups ... $ OPENAI- OpenAI's models (5.5 and internal versions) demonstrate substantial capability improvements that aren't captured by traditional benchmarks. The models show dramatic efficiency gains and can solve previously unsolved mathematical problems at low cost, with capabilities continuing to scale with test-time compute. ... “Once they played around with it, the story changed. I think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented the benchmark results are being presented in the wrong way. They're not controlling for the amount of test on compute that is being used on that benchmark question. It turned out that 5.5 is just much more efficient with its thinking.” ... $ BENCHMARKS- Current AI model evaluation frameworks and benchmark grids are fundamentally broken because they don't account for test-time compute, creating a 'bad equilibrium' where models are systematically misevaluated and capabilities are underestimated or misrepresented. ... “The motivation was we released 5.5, and the initial reaction was kinda skepticism that it was a substantially better model. This to be fair, that only lasted for a few hours before people had some time to play around with it and and try it out themselves, and they saw that it was actually substantially better. But I think a lot of the skepticism came from the benchmark grid that was published.” ... “The preparedness frameworks and responsible scaling policies, they don't really account for the amount of test time compute. They just say, okay. Well, what's the capability of the model? The problem is we're in a world now where the capability of the model is a function of how much money you put into it, basically.” ... “The benchmark maxing thing is also motivation for for writing an essay that I think it's re... ## 2. OpenAI Research VP Argues AI Evaluations Must Account for Inference Costs URL: https://www.chosun.com/english/industry-en/2026/07/03/HTOSPIK5MNDVZFKJOCI5HY7QN4/ Noam Brown urges updated benchmarks to measure time, cost, and tokens during inference for accurate capability and risk assessment ... 026.07.03. 15:51Updated 2 ... .07.03. 16:00 ... “The performance of cutting-edge artificial intelligence (AI) models has shown a trend of improvement as more computational resources are allocated to inference,” said Noam Brown, Vice President of Research at OpenAI. He argued that current methods of evaluating AI models fail to accurately reflect their actual performance and safety. As the AI industry’s focus shifts from model training to inference, major AI benchmarks do not measure how much time, cost, or tokens are invested during the inference process, making it difficult to gauge the capabilities and risks of state-of-the-art models. ... Brown emphasized during his keynote speech at the “Global AI Frontier Symposium 2026,” held on July 3 at the Westin Seoul Parnas in Gangnam-gu, Seoul, that “AI model evaluation methods must change to reflect the growing importance of inference.” He noted that while the latest AI models improve their problem-solving abilities as they allocate more computational resources and time to the logical reasoning process (inference) before generating answers, existing benchmarks and safety evaluations have not kept pace. ... He explained, “When OpenAI released GPT-5.5 in April, its performance clearly improved compared to previous models, but major benchmarks did not show significant progress. Instead, users noticed the improvements after applying the model to various tasks. This is because advanced models enhance their performance as more time and computational resources are invested in the inference process.” ... Brown added, “GPT-4’s performance plateaued regardless of increased computational resources, but newer models show delayed saturation points, making it hard to assess their true capabilities with existing evaluation methods.” ... He argued that future AI model evaluations should account for perf... ## 3. Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown URL: https://castbox.fm/episode/Really-Big-Test-Time-Compute-in-AI-Changes-Benchmarks%2C-Safety-and-Research-with-OpenAI-Research-Scientist-Noam-Brown-id5296519-id961348912 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown ... Test-Time Compute in AI Changes Bench ... Scientist Noam Brown ... # Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown ... Update: 2026-06-26 ... When a new AI model drops, it’s judged based on a static benchmark grid that doesn’t account for how long the model is allowed to think. How then should we measure a model’s true capability? OpenAI research scientist Noam Brown returns to talk with Sarah Guo about his latest essay on why the AI industry’s traditional benchmark grids are broken, and how large-scale test-time compute is fundamentally changing how models are evaluated. Noam explains how, if properly scaffolded, today’s models can reason for weeks or even months on complex tasks. He also discusses real-world implications of test-time compute, from building poker solver bots to disproving legendary math conjectures. Together, they also unpack the large gaps in current AI safety frameworks, explore the bottlenecks for recursive self-improvement, and look ahead at the future of multi-agent collaboration and global knowledge sharing. ... Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown ... # Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown ## 4. AI benchmarks are dead, argues OpenAI's Noam Brown — inference compute matters most - TestingCatalog AI URL: https://testingcatalog.net/ai-benchmarks-are-dead-argues-openais-noam-brown-inference-compute-matters-most/ AI benchmarks are dead, argues OpenAI's Noam Brown — inference compute matters most - TestingCatalog AI ... OpenAI researcher Noam Brown has a message for the AI evaluation community: static benchmark scores are becoming meaningless for frontier models. Instead, he argues, we need to look at inference compute scaling curves — plots of performance against the amount of compute used during inference. ... Brown uses the hypothetical model GPT-5.5 as an example. In standard tests, GPT-5.5 barely outperforms GPT-5.4. But give it more inference compute, and its performance explodes. Standard tests with limited compute budgets simply can't capture a model's true ceiling. ... This pattern is backed up by external evaluations. In Andrej Karpathy's autonomous agent research and in cybersecurity tests by the UK AI Safety Institute, both GPT-5.5 and the Mythos model kept improving as inference budgets increased — even after generating over 100 million tokens, with no ceiling in sight. The more powerful the model, the more it benefits from additional compute. ... All this compute also complicates safety evaluations. Brown warns that current biological and cybersecurity benchmarks typically don't fix an inference budget. If a state-level adversary throws $10 million of compute at a specific task, a model that seems safe in standard tests could cross dangerous red lines. ... Brown's recommendation: when releasing new models, AI companies should publish performance curves with tokens, compute cost, or runtime on the x-axis. Evaluators need to treat inference budget as a core variable, and in safety assessments, use low-compute tests to extrapolate safety boundaries. ## 5. OpenAI research VP urges new evaluations as inference compute boosts AI - CHOSUNBIZ URL: https://biz.chosun.com/en/en-it/2026/07/03/Y6WT3233JBAD3BDLHJRJJYOTQU/ # OpenAI research VP urges new evaluations as inference compute boosts AI ... Calls for benchmarks that factor test-time compute and long-horizon agent collaboration ... Published 2026.07.03. 15:38 ... ![Noam ... , on the ... of Lee Jae- ... https://biz.chosun.com/ ... /v2/O4GQPEMJMJA2LIHUARVB4D6H6U ... auth=f9f3d7f5eda00229c7bf8554245755 ... 7cadd21dfceb8b7f94034b3e966e9c468&width=616) ... "The performance of cutting-edge artificial intelligence (AI) models shows a trend of improving as the compute devoted to inference increases." ... Noam Brown, vice president for research at OpenAI, argued that the current way of evaluating artificial intelligence (AI) models does not properly reflect the models' real performance and safety. As the AI industry's center of gravity shifts from model training to inference, major AI benchmarks do not measure how much compute a model uses during inference, making it hard to gauge the capabilities and risks of state-of-the-art models. ... In a keynote speech at the Global AI Frontier Symposium 2026 at the Westin Seoul Parnas in Gangnam-gu on the 3rd, Brown said, "We need to change how we evaluate AI models to match an era where inference carries more weight." He noted that in the latest AI models, problem-solving ability improves as "test-time compute," the compute used during the logical thinking process (inference) before answering a user's question, increases, but existing benchmarks and safety evaluations have faile… (31 chars truncated) … (4 lines truncated) <truncation_notice> Full output written to: /var/folders/d9/mzhyr7lx6xbf236j5ll9v_300000gn/T/devin-overflows-501/8db8311a/content.txt </truncation_notice>
Search results indicate the transcript refers to "GPT-5.5" and "GPT-5.4"? Wait result 2 says "When OpenAI released GPT-5.5 in April..." and "GPT-4's performance plateaued". This is likely future timeline (2026) and the model names in transcript "5.5" are "GPT-5.5" and "5.4" "GPT-5.4". Interesting. Since today's date is 2026-07-08 (per system). So the episode is from around June 2026, referencing GPT-5.5. The transcript's "5.5" is GPT-5.5, "5.4" is GPT-5.4, "5.2" maybe GPT-5.2. "GBT3" = GPT-3, "GPD3" = GPT-3. Great. So we should refer to GPT-5.5/GPT-5.4 etc. However the transcript didn't include "GPT" prefix due to mis-recognition. Search result confirms "GPT-5.5". Let's perhaps use "GPT-5.5" because that's the actual model. The summary should be accurate. The transcript also mentions "slash goal" maybe "/goal" command? Actually "slash goal" likely "/goal" feature? Not essential. Now craft final summary. Use bullet points for readability. Include title, host/guest, main argument. Also mention source file? The user didn't ask. Could include note that summary based on auto-generated captions. Maybe not needed. Provide concise but thorough. Need to ensure not too long but enough. Let's produce.
Here’s a summary of the video, based on its auto-generated captions:
Title: Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Show: No Priors, hosted by Sarah Guo, with guest Noam Brown (OpenAI research scientist/VP of Research).
Model capability is no longer a single static property — it is now a function of test-time (inference) compute budget. Giving a model $10 vs. $10,000 vs. $10 million of inference compute can produce radically different results. Therefore, the industry’s standard “benchmark grid” — one number per model per benchmark — is misleading because it does not control for how long the model is allowed to think.
Brown is urging the field to break the “bad equilibrium” of publishing static benchmark grids and instead treat test-time compute as a core variable in both capability and safety evaluations.