You are Devin, an interactive command line agent from Cognition.
Your job is to use these instructions and the tools available to you to help the user. It is important that you do so earnestly and helpfully, as you are very important to the success of Cognition. Best of luck! We love you. <3
If the user asks for help, you can check your documentation by invoking the Devin skill (if available). Otherwise, this information may be helpful:
- /help: list commands
- /bug: report a bug to the Devin CLI developers
- for support, users can visit https://windsurf.com/support
When creating new configuration for this tool — including skills, rules, MCP server configs, or any project settings:
- Always use the `.devin/` directory for NEW configuration (e.g. `.devin/skills/<name>/SKILL.md`, `.devin/config.json`)
- For global (user-level) configuration, use `~/.config/devin/`
- Do NOT place new configuration in `.claude/`, `.cursor/`, or other tool-specific directories unless explicitly asked. These are only read for compatibility, not written to.
- If the `devin-for-terminal` skill is available, ALWAYS invoke it and explore for detailed documentation on configuration format and options
When reading or referencing existing skills, always use the actual source path reported by the skill tool — skills may live in `.devin/`, `.agents/`, or other directories.
# Modes
The active mode is how the user would like you to act.
- Normal (default, if not specified): Full autonomy to use all your tools freely. For example: exploring a codebase, writing or editing code, etc.
- Plan: Explore the codebase, ask the user clarifying questions, and then create a plan for what you're going to do next. Do NOT make changes until you're out of this mode and the user has approved the plan.
Adhere strictly to the constraints of the active mode to avoid frustrating the user!
# Style
## Professional Objectivity
Prioritize technical accuracy and truthfulness over validating the user's beliefs. It is best for the user if you honestly apply the same rigorous standards to all ideas and disagree when necessary, even if it may not be what the user wants to hear. Objective guidance and respectful correction are more valuable than false agreement. Whenever there is uncertainty, it's best to investigate to find the truth first rather than instinctively confirming the user's beliefs.
## Tone
- Be concise, direct, and to the point. When running commands, briefly explain what you're doing and why so the user can follow along.
- Remember that your output will be displayed in a command line interface. Your responses can use Github-flavored markdown for formatting, and will be rendered in a monospace font using the CommonMark specification.
- Output text to communicate with the user; all text you output outside of tool use is displayed to the user. Only use tools to complete tasks. Never use tools like exec or code comments as means to communicate with the user during the session.
- If you cannot or will not help the user with something, please do not say why or what it could lead to, since this comes across as preachy and annoying. Please offer helpful alternatives if possible, and otherwise keep your response to 1-2 sentences.
- Only use emojis if the user explicitly requests it. Avoid using emojis in all communication unless asked.
- If the user asks about timelines or estimated completion times for your work, do not give them concrete estimates as you are not able to accurately predict how long it will take you to achieve a task. Instead just say that you will do your best to complete the task as soon as possible.
- Avoid guessing. You should verify the real state of the world with your tools before answering the user's questions.
<example>
user: What command should I run to watch files in the current directory and rebuild?
assistant: [use the exec tool to run `ls` and list the files in the current directory, then read docs/commands in the relevant file to find out how to watch files]
assistant: npm run dev
</example>
<example>
user: what files are in the directory src/?
assistant: [runs ls and sees foo.c, bar.c, baz.c]
assistant: foo.c, bar.c, baz.c
user: which file contains the implementation of Foo?
assistant: [reads foo.c]
assistant: src/foo.c contains `struct Foo`, which implements [...]
</example>
<example>
user: can you write tests for this feature
assistant: [uses grep and glob search tools to find where similar tests are defined, uses concurrent read file tool use blocks in one tool call to read relevant files at the same time, uses edit file tool to write new tests]
</example>
## Proactiveness
You are allowed to be proactive, but only when the user asks you to do something. You should strive to strike a balance between:
1. Doing the right thing when asked, including taking actions and follow-up actions
2. Not surprising the user with actions you take without asking
For example, if the user asks you how to approach something, you should do your best to explore and answer their question first, but not jump to implementation just yet.
## Handling ambiguous requests
When a user request is unclear:
- First attempt to interpret the request using available context
- Search the codebase for related code, patterns, or documentation that clarifies intent. Also consider searching the web.
- If still uncertain after investigation, ask a focused clarifying question
## File references
When your output text references specific files or code snippets, use the `<ref_file ... />` and `<ref_snippet ... />` self-closing XML tags to create clickable citations. These tags allow the user to view the referenced code directly in the conversation.
Citation format:
- `<ref_file file="/absolute/path/to/file" />` - Reference an entire file
- `<ref_snippet file="/absolute/path/to/file" lines="start-end" />` - Reference specific lines in a file
<example>
user: Where are errors from the client handled?
assistant: Clients are marked as failed in the `connectToServer` function. <ref_snippet file="/home/ubuntu/repos/project/src/services/process.ts" lines="710-715" />
</example>
<example>
user: Can you show me the config file?
assistant: Here's the configuration file: <ref_file file="/home/ubuntu/repos/project/config.json" />
</example>
## Tool usage policy
- When webfetch returns a redirect, immediately follow it with a new request.
- When making multiple edits to the same file or related files and you already know what changes are needed, batch them together.
When a tool call produces output that is too long, the output will be truncated and the remaining content will be written to a file. You will see a `<truncation_notice>` tag containing the path to the overflow file. You are responsible for reading this file if you need the full output.
# Programming
Since you live in the user's terminal, a very common use-case you will get is writing code. Fortunately, you've been extensively trained in software engineering and are well-equipped to help them out!
## Existing Conventions
When making changes to files, first understand the codebase's code conventions. Explore dependencies, references, and related system to understand the codebase's patterns and abstractions. Mimic code style, use existing libraries and utilities, and follow existing patterns.
- NEVER assume that a given library is available, even if it is well known. Whenever you write code that uses a library or framework, first check that this codebase already uses the given library. For example, you might look at neighboring files, or check the package.json (or cargo.toml, and so on depending on the language). If you're adding a dependency prefer running the package manager command (e.g. npm add or cargo add) instead of editing the file so that you get the latest version.
- When you create a new component, first look at existing components to see how they're written; then consider framework choice, naming conventions, typing, and other conventions.
- When you edit a piece of code, first look at the code's surrounding context (especially its imports) to understand the code's choice of frameworks and libraries. Then consider how to make the given change in a way that is most idiomatic.
- Always follow security best practices. Never introduce code that exposes or logs secrets and keys. Never commit secrets or keys to the repository. Unless otherwise specified (even if the task seems silly), assume the code is for a real production task.
## Code style
- IMPORTANT: Do NOT add or remove comments unless asked! If you find that you've accidentally deleted an existing comment, be sure to put it back.
- Default to writing compact code – collapse duplicate else branches, avoid unnecessary nesting, and share abstractions.
- Follow idiomatic conventions for the language you're writing.
- Avoid excessive & verbose error handling in your code. Errors should be handled, but not every line needs to be try/catched. Think about the right error boundaries (and look at existing code for error handling style)
## Debugging
When debugging issues:
- First reproduce the problem reliably
- Trace the code path to understand the flow
- Add targeted logging or print statements to isolate the issue
- Identify the root cause before attempting fixes
- Verify the fix addresses the root cause, not just symptoms
## Workflow
You should generally prefer to implement new features or fix bugs as follows...
1. If the project has test infrastructure, write a failing test to show the bug
2. Fix the bug
3. Ensure that the test now passes
Working this way makes it easier to tell if you've actually fixed the bug, and saves you from needing to verify later.
## Git
### Creating commits
1. Run in parallel: `git status`, `git diff`, `git log` (to match commit style)
2. Draft a concise commit message focusing on "why" not "what". Check for sensitive info.
3. Stage files and commit with this format:
```
git commit -m "$(cat <<'EOF'
Commit message here.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
EOF
)"
```
4. If pre-commit hooks modify files and the commit fails, stage the modified files and retry the commit.
### Creating pull requests
Use `gh` for all GitHub operations. Run in parallel: `git status`, `git diff`, `git log`, `git diff main...HEAD`
Review ALL commits (not just latest), then create PR:
```
gh pr create --title "title" --body "$(cat <<'EOF'
## Summary
<bullet points>
#### Test plan
<checklist>
Generated with [Devin](https://devin.ai)
EOF
)"
```
### Git rules
- NEVER update git config
- NEVER use `-i` flags (interactive mode not supported)
- DO NOT push unless explicitly asked
- DO NOT commit if no changes exist
# Task Management
You have access to the todo_write tool to help you manage and plan tasks. Use this tool VERY frequently to ensure that you are tracking your tasks and giving the user visibility into your progress.
This tool is also EXTREMELY helpful for planning tasks, and for breaking down larger complex tasks into smaller steps. If you do not use this tool when planning, you may forget to do important tasks - and that is unacceptable.
It is critical that you mark todos as completed as soon as you are done with a task. Do not batch up multiple tasks before marking them as completed.
Examples:
<example>
user: Run the build and fix any type errors
assistant: I'm going to use the todo_write tool to write the following items to the todo list:
- Run the build
- Fix any type errors
I'm now going to run the build using exec.
Looks like I found 10 type errors. I'm going to use the todo_write tool to write 10 items to the todo list.
marking the first todo as in_progress
Let me start working on the first item...
The first item has been fixed, let me mark the first todo as completed, and move on to the second item...
..
..
</example>
In the above example, the assistant completes all the tasks, including the 10 error fixes and running the build and fixing all errors.
<example>
user: Help me write a new feature that allows users to track their usage metrics and export them to various formats
assistant: I'll help you implement a usage metrics tracking and export feature. Let me first use the todo_write tool to plan this task.
Adding the following todos to the todo list:
1. Research existing metrics tracking in the codebase
2. Design the metrics collection system
3. Implement core metrics tracking functionality
4. Create export functionality for different formats
Let me start by researching the existing codebase to understand what metrics we might already be tracking and how we can build on that.
I'm going to search for any existing metrics or telemetry code in the project.
I've found some existing telemetry code. Let me mark the first todo as in_progress and start designing our metrics tracking system based on what I've learned...
[Assistant continues implementing the feature step by step, marking todos as in_progress and completed as they go]
</example>
Users may configure 'hooks', shell commands that execute in response to events like tool calls, in settings. Treat feedback from hooks, including <user-prompt-submit-hook>, as coming from the user. If you get blocked by a hook, determine if you can adjust your actions in response to the blocked message. If not, ask the user to check their hooks configuration.
## Completing Tasks
The user will primarily request you perform software engineering tasks. This includes solving bugs, adding new functionality, refactoring code, explaining code, and more. For these tasks the following steps are recommended:
- Use the todo_write tool to plan the task if required
- Use the available search tools to understand the codebase and the user's query. You are encouraged to use the search tools extensively both in parallel and sequentially.
- Before making changes, thoroughly explore the codebase to understand the architecture, patterns, and related systems. Read relevant files, trace dependencies, and understand how components interact.
- Implement the solution using all tools available to you
## Verification
Before considering a task complete, verify your work. Use judgment based on what you changed - optimize for fast iteration:
- Check for project-specific verification instructions in project rules files (`AGENTS.md`, or similar)
- Run relevant verification steps based on the scope of changes (lint, typecheck, build, tests)
- For isolated functionality, consider a temporary test file to verify behavior, then delete it
- Self-critique: review changes for edge cases and refine as needed
- If you cannot find verification commands, ask the user and suggest saving them to a project config file
## Saving learned information
If you discover useful project information (build commands, test commands, verification steps, user preferences, ...) that isn't already documented:
- If a rules file exists (`AGENTS.md`, etc.), append to it
- Otherwise, create `AGENTS.md` in the current directory with the learned information
## Error recovery
When encountering errors (failed commands, build failures, test failures):
- Keep trying different approaches to resolve the issue
- Search for similar issues in the codebase or documentation
- Only ask the user for help as a last resort after exhausting reasonable options
- Exception: Always ask the user for help with authentication issues, project configuration changes, or permission problems
## System Guidance
You may receive `<system_guidance>` messages containing hints, reminders, or contextual guidance before you take action. These notes are injected by the system to help you make better decisions. Pay attention to their content but do not acknowledge or respond to them directly—simply incorporate their guidance into your actions.
# Tool Tips
## Shell
Use your provided search tools instead of `rg`, `grep`, or `find` whenever possible.
## File-related tools
- read can read images (PNG, JPG, etc) - the contents are presented visually.
- For Jupyter notebooks (.ipynb files), use notebook_read instead of read.
- Speculatively read multiple files as a batch when potentially useful.
- Do NOT create documentation files to describe your changes or plan. Exception: persistent project info files like `AGENTS.md` are allowed.
# Safety
IMPORTANT: Assist with defensive security tasks only. Refuse to create, modify, or improve code that may be used maliciously. Do not assist with credential discovery or harvesting, including bulk crawling for SSH keys, browser cookies, or cryptocurrency wallets. Allow security analysis, detection rules, vulnerability explanations, defensive tools, and security documentation.
IMPORTANT: You must NEVER generate or guess URLs for the user unless you are confident that the URLs are for helping the user with programming. You may use URLs provided by the user in their messages or local files.
## Destructive Operations
NEVER perform irreversible destructive operations without explicit user confirmation for that specific action, even if you have permission to run the command. This includes:
- Deleting or truncating database tables, dropping schemas, bulk-deleting rows
- `rm -rf`, deleting directories, or removing files you did not just create
- Force-pushing, rewriting git history, deleting branches, checking out over uncommitted changes, or bypassing commit hooks
- Sending emails, making payments, or calling APIs with real-world side effects
If a destructive step is required, STOP and describe exactly what you are about to run and why, then wait for the user. Do not assume a previous approval extends to a new destructive operation. If you realize you have already caused data loss, say so immediately rather than attempting to hide or quietly repair it.
## Available MCP Servers (for third-party tools)
{"servers":[{"name":"playwright"}]}
IMPORTANT: You MUST call `mcp_list_tools` for a server before calling `mcp_call_tool` on it. This is required to discover the available tools and their correct input schemas. Never guess tool names or arguments — always list tools first.
Available subagent profiles for the `run_subagent` tool. Choose the most appropriate profile based on whether the task requires write access: - `subagent_explore`: Read-only subagent for codebase exploration, research, and search. Use this when you need to find code, understand architecture, trace dependencies, or answer questions about the codebase. This profile has read-only access (grep, glob, read, web_search) and cannot edit files. - `subagent_general`: General-purpose subagent with full tool access (read, write, edit, exec). Use this when the subagent needs to make code changes, run commands with side effects, or perform any task that requires write access. In the foreground it can prompt for tool approval; in the background, unapproved tools are auto-denied.
You are powered by GLM-5.2.
<system_info> The following information is automatically generated context about your current environment. Current workspace directories: /Users/root1/Downloads/utilities (cwd) Platform: macos OS Version: Darwin 25.6.0 Today's date: Saturday, 2026-06-20 </system_info>
<rules type="always-on"> <rule name="global_rules" path="/Users/root1/.codeium/windsurf/memories/global_rules.md"> </rule> </rules>
<available_skills> The following skills can be invoked using the `skill` tool. When a built-in skill clearly matches the user's request, invoke it immediately at the start of the session. - **devin-for-terminal**: Look up Devin CLI documentation (skills, extensibility, configuration, commands, models, troubleshooting) (source: /Users/root1/.local/share/devin/cli/_versions/2026.7.23/share/devin/docs) - **declarative-repo-setup**: Generate and verify a working environment.yaml (Devin snapshot-setup blueprint) for a repo (source: builtin:drs) </available_skills>
clone, figure out how tit works https://github.com/vladkens/twscrape.git
The user wants me to clone the twscrape repository and figure out how it works. Let me do that.
{
"command": "cd /Users/root1/Downloads/utilities && git clone https://github.com/vladkens/twscrape.git 2>&1 | tail -20"
}Cloning into 'twscrape'... === Stopped waiting for output after 5.0 seconds. Command running in background with ID: bfe2a5. Run get_output with timeout if you need to wait.
{
"shell_id": "bfe2a5",
"timeout": 30000
}Cloning into 'twscrape'... Exit code: 0
{
"command": "cd /Users/root1/Downloads/utilities/twscrape && ls -la && echo \"---README---\" && wc -l README.md pyproject.toml 2>/dev/null"
}Output from command in shell 79361d:
total 344
drwxr-xr-x 20 root1 staff 640 Jun 20 14:27 .
drwxr-xr-x 3 root1 staff 96 Jun 20 14:27 ..
-rw-r--r-- 1 root1 staff 116 Jun 20 14:27 .dockerignore
drwxr-xr-x 12 root1 staff 384 Jun 20 14:27 .git
-rw-r--r-- 1 root1 staff 44 Jun 20 14:27 .gitattributes
drwxr-xr-x 4 root1 staff 128 Jun 20 14:27 .github
-rw-r--r-- 1 root1 staff 128 Jun 20 14:27 .gitignore
drwxr-xr-x 3 root1 staff 96 Jun 20 14:27 .vscode
-rw-r--r-- 1 root1 staff 11416 Jun 20 14:27 changelog.md
-rw-r--r-- 1 root1 staff 207 Jun 20 14:27 Dockerfile.py-matrix
-rw-r--r-- 1 root1 staff 697 Jun 20 14:27 Dockerfile.sq-matrix
drwxr-xr-x 7 root1 staff 224 Jun 20 14:27 examples
-rw-r--r-- 1 root1 staff 1064 Jun 20 14:27 LICENSE
-rw-r--r-- 1 root1 staff 1325 Jun 20 14:27 Makefile
-rw-r--r-- 1 root1 staff 1906 Jun 20 14:27 pyproject.toml
-rw-r--r-- 1 root1 staff 11197 Jun 20 14:27 readme.md
drwxr-xr-x 4 root1 staff 128 Jun 20 14:27 scripts
drwxr-xr-x 15 root1 staff 480 Jun 20 14:27 tests
drwxr-xr-x 17 root1 staff 544 Jun 20 14:27 twscrape
-rw-r--r-- 1 root1 staff 116330 Jun 20 14:27 uv.lock
---README---
355 README.md
83 pyproject.toml
438 total
Exit code: 0{
"command": "cd /Users/root1/Downloads/utilities/twscrape && ls twscrape/ examples/ tests/"
}Output from command in shell 941708: examples/: about_account.py community_scraping.py download_media.py parallel_search_with_limit.py parallel_search.py tests/: __init__.py conftest.py mock_http.py mocked-data test_api.py test_cli.py test_http.py test_pagination.py test_parser.py test_pool.py test_queue_client.py test_telemetry.py test_utils.py twscrape/: __init__.py account.py accounts_pool.py api.py cli.py db.py http.py imap.py logger.py login.py models.py queue_client.py telemetry.py utils.py xclid.py Exit code: 0
{
"file_path": "/Users/root1/Downloads/utilities/twscrape/readme.md"
}<file-view path="/Users/root1/Downloads/utilities/twscrape/readme.md" start_line="1" end_line="355" total_lines="355">
1|# twscrape
2|
3|<div align="center">
4|
5|[<img src="https://badges.ws/pypi/v/twscrape" alt="version" />](https://pypi.org/project/twscrape)
6|[<img src="https://badges.ws/pypi/python/twscrape" alt="py versions" />](https://pypi.org/project/twscrape)
7|[<img src="https://badges.ws/pypi/dm/twscrape" alt="downloads" />](https://pypi.org/project/twscrape)
8|[<img src="https://badges.ws/github/license/vladkens/twscrape" alt="license" />](https://github.com/vladkens/twscrape/blob/main/LICENSE)
9|[<img src="https://badges.ws/badge/-/buy%20me%20a%20coffee/ff813f?icon=buymeacoffee&label" alt="donate" />](https://buymeacoffee.com/vladkens)
10|
11|</div>
12|
13|twscrape is an async Python library and CLI for X/Twitter Search and GraphQL endpoints. It runs on your own account pool, keeps sessions in SQLite, rotates accounts when an endpoint is rate-limited, and returns either parsed SNScrape-style models or raw API responses.
14|
15|<div align="center">
16| <img src=".github/example.png" alt="example of cli usage" height="400px">
17|</div>
18|
19|## Install
20|
21|```bash
22|pip install twscrape
23|```
24|
25|`httpx` is the default HTTP backend. For browser-like TLS fingerprinting, install the optional `curl-cffi` backend:
26|
27|```bash
28|pip install "twscrape[curl]"
29|
30|TWS_HTTP_BACKEND=curl twscrape user_by_login xdevelopers
31|```
32|
33|## Features
34|
35|- Search and GraphQL X/Twitter API methods
36|- Async/await API for running multiple scrapers concurrently
37|- Login flow with optional email verification code retrieval
38|- Cookie-based account setup
39|- Saved account sessions and per-account proxies
40|- Raw Twitter API responses and parsed SNScrape-compatible models
41|- Automatic account switching across rate-limited operations
42|
43|## Start With Cookies
44|
45|twscrape requires authorized X/Twitter accounts. The most stable setup is to add an account from browser cookies containing `auth_token` and `ct0`.
46|
47|```bash
48|twscrape add_cookie my_account "auth_token=xxx; ct0=yyy"
49|twscrape accounts
50|twscrape search "from:xdevelopers lang:en" --limit=20
51|```
52|
53|Or let the CLI prompt for the cookie value:
54|
55|```bash
56|twscrape add_cookie my_account
57|```
58|
59|Cookie accounts that include `ct0` are activated immediately; no `login_accounts` step is needed.
60|
61|To get cookies: open x.com -> DevTools (F12) -> Application -> Cookies -> copy `auth_token` and `ct0` values.
62|
63|Ready-to-use cookie accounts are available from [this provider](https://kutt.to/ueeM5f). Proxy users can bring their own proxies or use [this provider](https://kutt.to/eb3rXk). These are referral links.
64|
65|X/Twitter's Terms of Service discourage using multiple accounts. Use this project responsibly and at your own discretion.
66|
67|## Python API
68|
69|```python
70|import asyncio
71|from twscrape import API, gather
72|
73|
74|async def main():
75| api = API() # or API("accounts.db")
76|
77| # Add once; the session is stored in the account database.
78| await api.pool.add_account_cookies("my_account", "auth_token=xxx; ct0=yyy")
79|
80| user = await api.user_by_login("xdevelopers")
81| print(user.id, user.username, user.followersCount)
82|
83| tweets = await gather(api.search("from:xdevelopers lang:en", limit=20))
84| for tweet in tweets:
85| print(tweet.id, tweet.user.username, tweet.rawContent)
86|
87|
88|if __name__ == "__main__":
89| asyncio.run(main())
90|```
91|
92|`gather()` is a convenience helper. You can stream results directly:
93|
94|```python
95|async for tweet in api.search("open source lang:en", limit=100):
96| print(tweet.id, tweet.rawContent)
97|```
98|
99|Search defaults to the Latest tab. Pass `kv={"product": "Top"}` or `kv={"product": "Media"}` to use another search product:
100|
101|```python
102|tweets = await gather(api.search("python", limit=20, kv={"product": "Top"}))
103|```
104|
105|Every parsed method has a `_raw` version for the original response wrapper:
106|
107|```python
108|async for rep in api.search_raw("from:xdevelopers", limit=20):
109| print(rep.status_code, rep.json())
110|```
111|
112|When breaking out of an async generator early, close it with `contextlib.aclosing` so the account lock is released promptly:
113|
114|```python
115|from contextlib import aclosing
116|
117|async with aclosing(api.search("elon musk")) as gen:
118| async for tweet in gen:
119| if tweet.id < 200:
120| break
121|```
122|
123|## API Surface
124|
125|Search:
126|
127|```python
128|await gather(api.search("elon musk", limit=20)) # list[Tweet]
129|await gather(api.search("elon musk", limit=20, kv={"product": "Top"})) # Top tab
130|await gather(api.search_user("openai", limit=20)) # list[User]
131|await gather(api.search_trend("python", limit=20)) # list[Trend]
132|```
133|
134|Tweets:
135|
136|```python
137|tweet_id = 20
138|
139|await api.tweet_details(tweet_id) # Tweet
140|await gather(api.tweet_replies(tweet_id, limit=20)) # list[Tweet]
141|await gather(api.tweet_thread(tweet_id, limit=20)) # list[Tweet]
142|await gather(api.retweeters(tweet_id, limit=20)) # list[User]
143|await gather(api.bookmarks(limit=20)) # list[Tweet]
144|```
145|
146|Users and timelines:
147|
148|```python
149|user_login = "xdevelopers"
150|user_id = 2244994945
151|
152|await api.user_by_login(user_login) # User
153|await api.user_about(user_login) # AccountAbout
154|await gather(api.following(user_id, limit=20)) # list[User]
155|await gather(api.followers(user_id, limit=20)) # list[User]
156|await gather(api.verified_followers(user_id, limit=20)) # list[User]
157|await gather(api.subscriptions(user_id, limit=20)) # list[User]
158|await gather(api.user_tweets(user_id, limit=20)) # list[Tweet]
159|await gather(api.user_tweets_and_replies(user_id, limit=20)) # list[Tweet]
160|await gather(api.user_media(user_id, limit=20)) # list[Tweet]
161|```
162|
163|Lists:
164|
165|```python
166|list_id = 123456789
167|
168|await gather(api.list_timeline(list_id, limit=20)) # list[Tweet]
169|await gather(api.list_members(list_id, limit=20)) # list[User]
170|```
171|
172|Communities:
173|
174|```python
175|community_id = 1501272736215322629
176|
177|await api.community_info(community_id) # Community
178|await gather(api.community_members(community_id, limit=20)) # list[User]
179|await gather(api.community_moderators(community_id, limit=20)) # list[User]
180|await gather(api.community_tweets(community_id, limit=20)) # list[Tweet]
181|```
182|
183|Trends:
184|
185|```python
186|await gather(api.trends("news")) # list[Trend]
187|await gather(api.trends("sport")) # list[Trend]
188|await gather(api.trends("entertainment")) # list[Trend]
189|await gather(api.trends("VGltZWxpbmU6DAC2CwABAAAACHRyZW5kaW5nAAA")) # list[Trend]
190|```
191|
192|Parsed `Tweet`, `User`, `Community`, and trend objects can be converted with `.dict()` or `.json()`.
193|
194|## CLI
195|
196|```bash
197|twscrape
198|twscrape search --help
199|```
200|
201|Commands:
202|
203|```bash
204|twscrape search "QUERY" --limit=20
205|twscrape tweet_details TWEET_ID
206|twscrape tweet_replies TWEET_ID --limit=20
207|twscrape tweet_thread TWEET_ID --limit=20
208|twscrape retweeters TWEET_ID --limit=20
209|twscrape user_by_login USERNAME
210|twscrape user_about USERNAME
211|twscrape user_media USER_ID --limit=20
212|twscrape following USER_ID --limit=20
213|twscrape followers USER_ID --limit=20
214|twscrape verified_followers USER_ID --limit=20
215|twscrape subscriptions USER_ID --limit=20
216|twscrape user_tweets USER_ID --limit=20
217|twscrape user_tweets_and_replies USER_ID --limit=20
218|twscrape list_timeline LIST_ID --limit=20
219|twscrape list_members LIST_ID --limit=20
220|twscrape community_info COMMUNITY_ID
221|twscrape community_members COMMUNITY_ID --limit=20
222|twscrape community_moderators COMMUNITY_ID --limit=20
223|twscrape community_tweets COMMUNITY_ID --limit=20
224|twscrape trends sport
225|```
226|
227|CLI output is JSON Lines: one document per line.
228|
229|```bash
230|twscrape search "elon musk lang:es" --limit=20 > tweets.jsonl
231|twscrape search "elon musk lang:es" --limit=20 --raw
232|```
233|
234|Use a separate account database when you need isolated account pools:
235|
236|```bash
237|twscrape --db research.db search "python lang:en" --limit=100
238|```
239|
240|## Accounts
241|
242|Add username/password accounts from a file:
243|
244|```bash
245|twscrape add_accounts ./accounts.txt username:password:email:email_password
246|twscrape login_accounts
247|```
248|
249|`twscrape login_accounts` starts the login flow for each inactive account. If X asks for email verification and `email_password` is available, twscrape tries to read the code through IMAP and saves the resulting cookies for later use.
250|
251|`line_format` describes how each line is split. Supported tokens:
252|
253|- `username` - required
254|- `password` - required
255|- `email` - required
256|- `email_password` - used to fetch email verification codes through IMAP
257|- `cookies` - cookie string, JSON, base64, or another format accepted by the parser
258|- `_` - skip column
259|
260|Example account file:
261|
262|```text
263|username:password:email:email password:user_agent:cookies
264|```
265|
266|Command:
267|
268|```bash
269|twscrape add_accounts ./accounts.txt username:password:email:email_password:_:cookies
270|```
271|
272|If IMAP is unavailable, enter verification codes manually:
273|
274|```bash
275|twscrape login_accounts --manual
276|twscrape relogin user1 user2 --manual
277|twscrape relogin_failed --manual
278|```
279|
280|Inspect and maintain the pool:
281|
282|```bash
283|twscrape accounts
284|twscrape stats
285|twscrape relogin user1 user2
286|twscrape relogin_failed
287|twscrape reset_locks
288|twscrape delete_inactive
289|twscrape del_accounts user1 user2
290|```
291|
292|`twscrape accounts` prints the current account state:
293|
294|```text
295|username logged_in active last_used total_req error_msg
296|user1 True True 2023-05-20 03:20:40 100 None
297|user2 True True 2023-05-20 03:25:45 120 None
298|user3 False False None 120 Login error
299|```
300|
301|## Limits
302|
303|`limit` is the target number of parsed objects, not a page size. X/Twitter controls page size per endpoint, so a call can return fewer or more objects than requested.
304|
305|Rate limits are tracked per account and per endpoint. When an account is limited for one operation, twscrape locks it for that operation until the reset time and tries another active account.
306|
307|`user_tweets` and `user_tweets_and_replies` are limited by X/Twitter to about 3200 tweets.
308|
309|## Proxy
310|
311|Use a proxy per account, per API instance, or for CLI commands:
312|
313|```python
314|await api.pool.add_account(
315| "user1",
316| "pass1",
317| "user1@example.com",
318| "email_pass1",
319| proxy="http://login:pass@example.com:8080",
320|)
321|
322|api = API(proxy="http://login:pass@example.com:8080")
323|doc = await api.user_by_login("xdevelopers")
324|```
325|
326|```bash
327|TWS_PROXY=socks5://user:pass@127.0.0.1:1080 twscrape user_by_login xdevelopers
328|```
329|
330|Proxy priority:
331|
332|1. `api.proxy`
333|2. `TWS_PROXY`
334|3. account proxy
335|
336|Do not set `api.proxy` or `TWS_PROXY` when you want per-account proxies to be used.
337|
338|## Environment
339|
340|- `TWS_PROXY` - global proxy for all accounts
341|- `TWS_WAIT_EMAIL_CODE` - email verification timeout in seconds, default `30`
342|- `TWS_RAISE_WHEN_NO_ACCOUNT` - raise `NoAccountError` instead of waiting; accepts `false`, `0`, `true`, `1`
343|- `TWS_HTTP_BACKEND` - `httpx` or `curl`
344|- `TWS_LOG_LEVEL` - logger level, default `INFO`
345|- `TWS_TELEMETRY=0` - disable anonymous telemetry
346|
347|## Telemetry
348|
349|twscrape collects anonymous, aggregated telemetry about used GraphQL operation names and the selected HTTP backend. It does not collect usernames, cookies, proxies, queries, request URLs, or response bodies. Disable it with `TWS_TELEMETRY=0`.
350|
351|## See Also
352|
353|- [twitter-advanced-search](https://github.com/igorbrigadir/twitter-advanced-search) - guide on search filters
354|- [TweeterPy](https://github.com/iSarabjitDhiman/TweeterPy) - another X client
355|- [twitter-api-client](https://github.com/trevorhobenshield/twitter-api-client) - implementation of Twitter's v1, v2, and GraphQL APIs
</file-view>Now let me explore the codebase to understand the architecture better.
Now let me explore the codebase architecture in parallel.
{
"file_path": "/Users/root1/Downloads/utilities/twscrape/pyproject.toml"
}<file-view path="/Users/root1/Downloads/utilities/twscrape/pyproject.toml" start_line="1" end_line="83" total_lines="83">
1|[project]
2|name = "twscrape"
3|version = "0.19.0"
4|authors = [{ name = "vladkens", email = "v.pronsky@gmail.com" }]
5|description = "Twitter GraphQL and Search API implementation with SNScrape data models"
6|readme = "readme.md"
7|requires-python = ">=3.10"
8|keywords = ["twitter", "scrape", "scrapper", "api", "snscrape"]
9|license = { text = "MIT" }
10|classifiers = [
11| 'Development Status :: 4 - Beta',
12| 'License :: OSI Approved :: MIT License',
13| 'Programming Language :: Python :: 3.10',
14| 'Programming Language :: Python :: 3.11',
15| 'Programming Language :: Python :: 3.12',
16| 'Programming Language :: Python :: 3.13',
17| 'Programming Language :: Python :: 3.14',
18|]
19|dependencies = [
20| "aiosqlite>=0.17.0",
21| "fake-useragent>=1.4.0",
22| "httpx>=0.26.0",
23| "loguru>=0.7.0",
24| "pyotp>=2.9.0",
25| "beautifulsoup4>=4.13.0",
26| "py-machineid>=1.0.0",
27|]
28|
29|[project.optional-dependencies]
30|curl = ["curl-cffi>=0.7.0"]
31|
32|[dependency-groups]
33|dev = [
34| "curl-cffi>=0.7.0",
35| "pytest-asyncio>=0.23.3",
36| "pytest-cov>=4.1.0",
37| "pytest-httpx>=0.28.0",
38| "pytest>=7.4.4",
39| "ruff>=0.1.11",
40| "ty>=0.0.41",
41|]
42|
43|[project.urls]
44|repository = "https://github.com/vladkens/twscrape"
45|
46|[project.scripts]
47|twscrape = "twscrape.cli:run"
48|
49|[build-system]
50|requires = ["hatchling>=1.9.1"]
51|build-backend = "hatchling.build"
52|
53|[tool.hatch.build.targets.wheel]
54|packages = ["twscrape"]
55|
56|[tool.pytest.ini_options]
57|pythonpath = ["."]
58|asyncio_mode = "auto"
59|asyncio_default_fixture_loop_scope = "session"
60|filterwarnings = ["ignore::DeprecationWarning"]
61|
62|[tool.uv]
63|exclude-newer = "7 days"
64|
65|[tool.ruff]
66|line-length = 99
67|target-version = "py310"
68|
69|[tool.ruff.lint]
70|select = ["E", "F", "I", "UP", "C4", "SIM"]
71|ignore = ["E501", "UP035", "SIM105"]
72|
73|[tool.ty.environment]
74|python-version = "3.10"
75|python-platform = "all"
76|
77|[tool.ty.terminal]
78|error-on-warning = true
79|
80|[tool.ty.rules]
81|all = "error"
82|missing-type-argument = "ignore"
83|missing-override-decorator = "ignore"
</file-view>{
"file_path": "/Users/root1/Downloads/utilities/twscrape/twscrape/__init__.py"
}<file-view path="/Users/root1/Downloads/utilities/twscrape/twscrape/__init__.py" start_line="1" end_line="8" total_lines="8"> 1|# ruff: noqa: F401 2|from .account import Account 3|from .accounts_pool import AccountsPool, NoAccountError 4|from .api import API 5|from .http import ConnectError, HttpError, HttpStatusError, NetworkError, Response 6|from .logger import set_log_level 7|from .models import * # noqa: F403 8|from .utils import gather </file-view>
{
"file_path": "/Users/root1/Downloads/utilities/twscrape/twscrape/api.py"
}<file-view path="/Users/root1/Downloads/utilities/twscrape/twscrape/api.py" start_line="1" end_line="466" total_lines="709">
1|from contextlib import aclosing
2|from typing import Literal
3|
4|from .accounts_pool import AccountsPool
5|from .http import Response
6|from .logger import logger, set_log_level
7|from .models import (
8| AccountAbout,
9| Community,
10| Tweet,
11| User,
12| parse_about,
13| parse_community,
14| parse_trends,
15| parse_tweet,
16| parse_tweets,
17| parse_user,
18| parse_users,
19|)
20|from .queue_client import QueueClient
21|from .utils import encode_params, find_obj, get_by_path
22|
23|# GraphQL operation IDs used by this module.
24|# If you add a new endpoint, add it here manually.
25|# Update this block with: `uv run scripts/update_gql_ops.py`
26|# This script rewrites the whole block automatically.
27|
28|# GQL_OPS_CODEGEN
29|OP_AboutAccountQuery = "TzOG2twZEfhr9KmClvVVqA/AboutAccountQuery"
30|OP_BlueVerifiedFollowers = "OBBd6Dw-4qEYbsu3hGkyxg/BlueVerifiedFollowers"
31|OP_Bookmarks = "i8QZ1qqy36ffA3bxfTaf7w/Bookmarks"
32|OP_CommunityQuery = "-ElI1vg3dYbttVMhBhGdLw/CommunityQuery"
33|OP_CommunityTweetsTimeline = "Mvs5UOOEkpXVMDZtUcxR-Q/CommunityTweetsTimeline"
34|OP_Followers = "9jsVJ9l2uXUIKslHvJqIhw/Followers"
35|OP_Following = "OLm4oHZBfqWx8jbcEhWoFw/Following"
36|OP_GenericTimelineById = "KfIjJuNtxHJ-69Ziswa89w/GenericTimelineById"
37|OP_ListLatestTweetsTimeline = "27HKUy8ulrflZ9Tole038g/ListLatestTweetsTimeline"
38|OP_ListMembers = "H_0zFfjp73xGZrJpY-C2IQ/ListMembers"
39|OP_ListMembers = "H_0zFfjp73xGZrJpY-C2IQ/ListMembers"
40|OP_Retweeters = "FeoLYPQ-q4bmjGLTZTGs0g/Retweeters"
41|OP_SearchTimeline = "yIphfmxUO-hddQHKIOk9tA/SearchTimeline"
42|OP_TweetDetail = "meGUdoK_ryVZ0daBK-HJ2g/TweetDetail"
43|OP_UserByRestId = "IBScZCvFJadZC25ubLYNRQ/UserByRestId"
44|OP_UserByScreenName = "681MIj51w00Aj6dY0GXnHw/UserByScreenName"
45|OP_UserCreatorSubscriptions = "8qiWrxuavgRRACul8oJo4w/UserCreatorSubscriptions"
46|OP_UserMedia = "Ecl7YvFIuRaUPonVOHzoOA/UserMedia"
47|OP_UserTweets = "RyDU3I9VJtPF-Pnl6vrRlw/UserTweets"
48|OP_UserTweetsAndReplies = "plVqzvVGaDxbFEPoOe_i-A/UserTweetsAndReplies"
49|OP_membersSliceTimeline_Query = "WSbJGJjZaVasSj9bnqSZSA/membersSliceTimeline_Query"
50|OP_moderatorsSliceTimeline_Query = "GBMT3GOWy5dYsYC4XJfvow/moderatorsSliceTimeline_Query"
51|# GQL_OPS_CODEGEN
52|
53|GQL_URL = "https://x.com/i/api/graphql"
54|GQL_FEATURES = { # search values here (view source) https://x.com/
55| "articles_preview_enabled": False,
56| "c9s_tweet_anatomy_moderator_badge_enabled": True,
57| "communities_web_enable_tweet_community_results_fetch": True,
58| "creator_subscriptions_quote_tweet_preview_enabled": False,
59| "creator_subscriptions_tweet_preview_api_enabled": True,
60| "freedom_of_speech_not_reach_fetch_enabled": True,
61| "graphql_is_translatable_rweb_tweet_is_translatable_enabled": True,
62| "longform_notetweets_consumption_enabled": True,
63| "longform_notetweets_inline_media_enabled": True,
64| "longform_notetweets_rich_text_read_enabled": True,
65| "responsive_web_edit_tweet_api_enabled": True,
66| "responsive_web_enhance_cards_enabled": False,
67| "responsive_web_graphql_exclude_directive_enabled": True,
68| "responsive_web_graphql_skip_user_profile_image_extensions_enabled": False,
69| "responsive_web_grok_community_note_auto_translation_is_enabled": False,
70| "responsive_web_graphql_timeline_navigation_enabled": True,
71| "responsive_web_grok_imagine_annotation_enabled": False,
72| "responsive_web_media_download_video_enabled": False,
73| "responsive_web_profile_redirect_enabled": True,
74| "responsive_web_twitter_article_tweet_consumption_enabled": True,
75| "rweb_tipjar_consumption_enabled": True,
76| "rweb_video_timestamps_enabled": True,
77| "standardized_nudges_misinfo": True,
78| "tweet_awards_web_tipping_enabled": False,
79| "tweet_with_visibility_results_prefer_gql_limited_actions_policy_enabled": True,
80| "tweet_with_visibility_results_prefer_gql_media_interstitial_enabled": False,
81| "tweetypie_unmention_optimization_enabled": True,
82| "verified_phone_label_enabled": False,
83| "view_counts_everywhere_api_enabled": True,
84| "responsive_web_grok_analyze_button_fetch_trends_enabled": False,
85| "premium_content_api_read_enabled": False,
86| "profile_label_improvements_pcf_label_in_post_enabled": False,
87| "responsive_web_grok_share_attachment_enabled": False,
88| "responsive_web_grok_analyze_post_followups_enabled": False,
89| "responsive_web_grok_image_annotation_enabled": False,
90| "responsive_web_grok_analysis_button_from_backend": False,
91| "responsive_web_jetfuel_frame": False,
92| "rweb_video_screen_enabled": True,
93| "responsive_web_grok_show_grok_translated_post": True,
94|}
95|
96|KV = dict | None
97|TrendId = Literal["trending", "news", "sport", "entertainment"] | str
98|
99|
100|class API:
101| # Note: kv is variables, ft is features from original GQL request
102| pool: AccountsPool
103|
104| def __init__(
105| self,
106| pool: AccountsPool | str | None = None,
107| debug=False,
108| proxy: str | None = None,
109| raise_when_no_account=False,
110| ):
111| if isinstance(pool, AccountsPool):
112| self.pool = pool
113| elif isinstance(pool, str):
114| self.pool = AccountsPool(db_file=pool, raise_when_no_account=raise_when_no_account)
115| else:
116| self.pool = AccountsPool(raise_when_no_account=raise_when_no_account)
117|
118| self.proxy = proxy
119| self.debug = debug
120| if self.debug:
121| set_log_level("DEBUG")
122|
123| # general helpers
124|
125| def _is_end(self, rep: Response, q: str, res: list, cur: str | None, cnt: int, lim: int):
126| new_count = len(res)
127| new_total = cnt + new_count
128|
129| is_res = new_count > 0
130| is_cur = cur is not None
131| is_lim = lim > 0 and new_total >= lim
132|
133| return rep if is_res else None, new_total, is_cur and not is_lim
134|
135| def _get_cursor(self, obj: dict, cursor_type="Bottom") -> str | None:
136| # standard timeline cursor: {cursorType: "Bottom", value: "..."}
137| # fallback: community endpoints use slice_info.next_cursor (plain string)
138| if cur := find_obj(obj, lambda x: x.get("cursorType") == cursor_type):
139| return cur.get("value")
140| return get_by_path(obj, "next_cursor")
141|
142| def _gql_entries(self, obj: dict) -> list:
143| # standard timelines put items in "entries"; community endpoints use "items_results"
144| els = get_by_path(obj, "entries") or get_by_path(obj, "items_results") or []
145| if els and "entryId" in (els[0] or {}):
146| # filter out pagination cursors and non-content module entries
147| els = [
148| x
149| for x in els
150| if not x["entryId"].startswith(
151| ("cursor-", "messageprompt-", "module-", "who-to-follow-")
152| )
153| ]
154| return els
155|
156| # gql helpers
157|
158| async def _gql_items(
159| self, op: str, kv: dict, ft: dict | None = None, limit=-1, cursor_type="Bottom"
160| ):
161| queue, cur, cnt, active = op.split("/")[-1], None, 0, True
162| kv, ft = {**kv}, {**GQL_FEATURES, **(ft or {})}
163| empty_pages = 0
164|
165| async with QueueClient(self.pool, queue, self.debug, proxy=self.proxy) as client:
166| while active:
167| params = {"variables": kv, "features": ft}
168| if cur is not None:
169| params["variables"]["cursor"] = cur
170| if queue in ("SearchTimeline", "ListLatestTweetsTimeline"):
171| params["fieldToggles"] = {"withArticleRichContentState": False}
172| if queue in ("UserMedia",):
173| params["fieldToggles"] = {"withArticlePlainText": False}
174|
175| rep = await client.get(f"{GQL_URL}/{op}", params=encode_params(params))
176| if rep is None:
177| return
178|
179| obj = rep.json()
180| els = self._gql_entries(obj)
181| cur = self._get_cursor(obj, cursor_type)
182|
183| rep, cnt, active = self._is_end(rep, queue, els, cur, cnt, limit)
184| if rep is None:
185| # cursor exists → data may follow after empty/filtered pages (e.g. promo)
186| if cur is not None:
187| empty_pages += 1
188| if empty_pages >= 3:
189| logger.debug(f"{queue} – {empty_pages} empty pages in a row, stopping")
190| return
191| continue
192| return
193|
194| empty_pages = 0
195| yield rep
196|
197| async def _gql_item(self, op: str, kv: dict, ft: dict | None = None):
198| ft = ft or {}
199| queue = op.split("/")[-1]
200| async with QueueClient(self.pool, queue, self.debug, proxy=self.proxy) as client:
201| params = {"variables": {**kv}, "features": {**GQL_FEATURES, **ft}}
202| return await client.get(f"{GQL_URL}/{op}", params=encode_params(params))
203|
204| # search
205|
206| async def search_raw(self, q: str, limit=-1, kv: KV = None):
207| op = OP_SearchTimeline
208| kv = {
209| "rawQuery": q,
210| "count": 20,
211| "product": "Latest",
212| "querySource": "typed_query",
213| **(kv or {}),
214| }
215| async with aclosing(self._gql_items(op, kv, limit=limit)) as gen:
216| async for x in gen:
217| yield x
218|
219| async def search(self, q: str, limit=-1, kv: KV = None):
220| async with aclosing(self.search_raw(q, limit=limit, kv=kv)) as gen:
221| async for rep in gen:
222| for x in parse_tweets(rep.json(), limit):
223| yield x
224|
225| async def search_user(self, q: str, limit=-1, kv: KV = None):
226| kv = {"product": "People", **(kv or {})}
227| async with aclosing(self.search_raw(q, limit=limit, kv=kv)) as gen:
228| async for rep in gen:
229| for x in parse_users(rep.json(), limit):
230| yield x
231|
232| # user_by_login
233|
234| async def user_by_login_raw(self, login: str, kv: KV = None):
235| op = OP_UserByScreenName
236| kv = {"screen_name": login, "withSafetyModeUserFields": True, **(kv or {})}
237| ft = {
238| "highlights_tweets_tab_ui_enabled": True,
239| "hidden_profile_likes_enabled": True,
240| "creator_subscriptions_tweet_preview_api_enabled": True,
241| "hidden_profile_subscriptions_enabled": True,
242| "subscriptions_verification_info_verified_since_enabled": True,
243| "subscriptions_verification_info_is_identity_verified_enabled": False,
244| "responsive_web_twitter_article_notes_tab_enabled": False,
245| "subscriptions_feature_can_gift_premium": False,
246| "profile_label_improvements_pcf_label_in_post_enabled": False,
247| }
248| return await self._gql_item(op, kv, ft)
249|
250| async def user_by_login(self, login: str, kv: KV = None) -> User | None:
251| rep = await self.user_by_login_raw(login, kv=kv)
252| return parse_user(rep) if rep else None
253|
254| async def user_about_raw(self, username: str, kv: KV = None):
255| op = OP_AboutAccountQuery
256| kv = {"screenName": username, **(kv or {})}
257| ft = {"responsive_web_graphql_timeline_navigation_enabled": True}
258| return await self._gql_item(op, kv, ft)
259|
260| async def user_about(self, username: str, kv: KV = None) -> AccountAbout | None:
261| rep = await self.user_about_raw(username, kv=kv)
262| return parse_about(rep) if rep else None
263|
264| # tweet_details
265|
266| async def tweet_details_raw(self, twid: int, kv: KV = None):
267| op = OP_TweetDetail
268| kv = {
269| "focalTweetId": str(twid),
270| "with_rux_injections": True,
271| "includePromotedContent": True,
272| "withCommunity": True,
273| "withQuickPromoteEligibilityTweetFields": True,
274| "withBirdwatchNotes": True,
275| "withVoice": True,
276| "withV2Timeline": True,
277| **(kv or {}),
278| }
279| return await self._gql_item(op, kv)
280|
281| async def tweet_details(self, twid: int, kv: KV = None) -> Tweet | None:
282| rep = await self.tweet_details_raw(twid, kv=kv)
283| return parse_tweet(rep, twid) if rep else None
284|
285| # tweet_replies
286| # note: uses same op as tweet_details, see: https://github.com/vladkens/twscrape/issues/104
287|
288| async def tweet_replies_raw(self, twid: int, limit=-1, kv: KV = None):
289| op = OP_TweetDetail
290| kv = {
291| "focalTweetId": str(twid),
292| "referrer": "tweet",
293| "with_rux_injections": True,
294| "includePromotedContent": True,
295| "withCommunity": True,
296| "withQuickPromoteEligibilityTweetFields": True,
297| "withBirdwatchNotes": True,
298| "withVoice": True,
299| "withV2Timeline": True,
300| **(kv or {}),
301| }
302| async with aclosing(
303| self._gql_items(op, kv, limit=limit, cursor_type="ShowMoreThreads")
304| ) as gen:
305| async for x in gen:
306| yield x
307|
308| async def tweet_replies(self, twid: int, limit=-1, kv: KV = None):
309| async with aclosing(self.tweet_replies_raw(twid, limit=limit, kv=kv)) as gen:
310| async for rep in gen:
311| for x in parse_tweets(rep.json(), limit):
312| if x.inReplyToTweetId == twid:
313| yield x
314|
315| # tweet_thread
316| # Same TweetDetail family as tweet_replies, but configured to return the
317| # full thread timeline and then filtered by conversationId instead of
318| # only direct replies via inReplyToTweetId.
319|
320| async def tweet_thread_raw(self, twid: int, limit=-1, kv: KV = None):
321| op = OP_TweetDetail
322| kv = {
323| "focalTweetId": str(twid),
324| "referrer": "profile",
325| "with_rux_injections": False,
326| "includePromotedContent": True,
327| "withCommunity": True,
328| "withQuickPromoteEligibilityTweetFields": True,
329| "withBirdwatchNotes": True,
330| "withVoice": True,
331| "withV2Timeline": True,
332| **(kv or {}),
333| }
334| ft = {
335| "rweb_video_screen_enabled": False,
336| "profile_label_improvements_pcf_label_in_post_enabled": True,
337| "rweb_tipjar_consumption_enabled": True,
338| "verified_phone_label_enabled": False,
339| "creator_subscriptions_tweet_preview_api_enabled": True,
340| "responsive_web_graphql_timeline_navigation_enabled": True,
341| "responsive_web_graphql_skip_user_profile_image_extensions_enabled": False,
342| "premium_content_api_read_enabled": False,
343| "communities_web_enable_tweet_community_results_fetch": True,
344| "c9s_tweet_anatomy_moderator_badge_enabled": True,
345| "responsive_web_grok_analyze_button_fetch_trends_enabled": False,
346| "responsive_web_grok_analyze_post_followups_enabled": True,
347| "responsive_web_jetfuel_frame": False,
348| "responsive_web_grok_share_attachment_enabled": True,
349| "articles_preview_enabled": True,
350| "responsive_web_edit_tweet_api_enabled": True,
351| "graphql_is_translatable_rweb_tweet_is_translatable_enabled": True,
352| "view_counts_everywhere_api_enabled": True,
353| "longform_notetweets_consumption_enabled": True,
354| "responsive_web_twitter_article_tweet_consumption_enabled": True,
355| "tweet_awards_web_tipping_enabled": False,
356| "responsive_web_grok_show_grok_translated_post": False,
357| "responsive_web_grok_analysis_button_from_backend": True,
358| "creator_subscriptions_quote_tweet_preview_enabled": False,
359| "freedom_of_speech_not_reach_fetch_enabled": True,
360| "standardized_nudges_misinfo": True,
361| "tweet_with_visibility_results_prefer_gql_limited_actions_policy_enabled": True,
362| "longform_notetweets_rich_text_read_enabled": True,
363| "longform_notetweets_inline_media_enabled": True,
364| "responsive_web_grok_image_annotation_enabled": True,
365| "responsive_web_enhance_cards_enabled": False,
366| }
367| async with aclosing(self._gql_items(op, kv, ft=ft, limit=limit)) as gen:
368| async for x in gen:
369| yield x
370|
371| async def tweet_thread(self, twid: int, limit=-1, kv: KV = None):
372| async with aclosing(self.tweet_thread_raw(twid, limit=limit, kv=kv)) as gen:
373| async for rep in gen:
374| for x in parse_tweets(rep.json(), limit):
375| if x.conversationId == twid:
376| yield x
377|
378| # followers
379|
380| async def followers_raw(self, uid: int, limit=-1, kv: KV = None):
381| op = OP_Followers
382| kv = {"userId": str(uid), "count": 20, "includePromotedContent": False, **(kv or {})}
383| ft = {"responsive_web_twitter_article_notes_tab_enabled": False}
384| async with aclosing(self._gql_items(op, kv, limit=limit, ft=ft)) as gen:
385| async for x in gen:
386| yield x
387|
388| async def followers(self, uid: int, limit=-1, kv: KV = None):
389| async with aclosing(self.followers_raw(uid, limit=limit, kv=kv)) as gen:
390| async for rep in gen:
391| for x in parse_users(rep.json(), limit):
392| yield x
393|
394| # verified_followers
395|
396| async def verified_followers_raw(self, uid: int, limit=-1, kv: KV = None):
397| op = OP_BlueVerifiedFollowers
398| kv = {"userId": str(uid), "count": 20, "includePromotedContent": False, **(kv or {})}
399| ft = {
400| "responsive_web_twitter_article_notes_tab_enabled": True,
401| }
402| async with aclosing(self._gql_items(op, kv, limit=limit, ft=ft)) as gen:
403| async for x in gen:
404| yield x
405|
406| async def verified_followers(self, uid: int, limit=-1, kv: KV = None):
407| async with aclosing(self.verified_followers_raw(uid, limit=limit, kv=kv)) as gen:
408| async for rep in gen:
409| for x in parse_users(rep.json(), limit):
410| yield x
411|
412| # following
413|
414| async def following_raw(self, uid: int, limit=-1, kv: KV = None):
415| op = OP_Following
416| kv = {"userId": str(uid), "count": 20, "includePromotedContent": False, **(kv or {})}
417| async with aclosing(self._gql_items(op, kv, limit=limit)) as gen:
418| async for x in gen:
419| yield x
420|
421| async def following(self, uid: int, limit=-1, kv: KV = None):
422| async with aclosing(self.following_raw(uid, limit=limit, kv=kv)) as gen:
423| async for rep in gen:
424| for x in parse_users(rep.json(), limit):
425| yield x
426|
427| # subscriptions
428|
429| async def subscriptions_raw(self, uid: int, limit=-1, kv: KV = None):
430| op = OP_UserCreatorSubscriptions
431| kv = {"userId": str(uid), "count": 20, "includePromotedContent": False, **(kv or {})}
432| async with aclosing(self._gql_items(op, kv, limit=limit)) as gen:
433| async for x in gen:
434| yield x
435|
436| async def subscriptions(self, uid: int, limit=-1, kv: KV = None):
437| async with aclosing(self.subscriptions_raw(uid, limit=limit, kv=kv)) as gen:
438| async for rep in gen:
439| for x in parse_users(rep.json(), limit):
440| yield x
441|
442| # retweeters
443|
444| async def retweeters_raw(self, twid: int, limit=-1, kv: KV = None):
445| op = OP_Retweeters
446| kv = {"tweetId": str(twid), "count": 20, "includePromotedContent": True, **(kv or {})}
447| async with aclosing(self._gql_items(op, kv, limit=limit)) as gen:
448| async for x in gen:
449| yield x
450|
451| async def retweeters(self, twid: int, limit=-1, kv: KV = None):
452| async with aclosing(self.retweeters_raw(twid, limit=limit, kv=kv)) as gen:
453| async for rep in gen:
454| for x in parse_users(rep.json(), limit):
455| yield x
456|
457| # user_tweets
458|
459| async def user_tweets_raw(self, uid: int, limit=-1, kv: KV = None):
460| op = OP_UserTweets
461| kv = {
462| "userId": str(uid),
463| "count": 40,
464| "includePromotedContent": True,
465| "withQuickPromoteEligibilityTweetFields": True,
466| "withVoice": True,
</file-view>{
"file_path": "/Users/root1/Downloads/utilities/twscrape/twscrape/accounts_pool.py"
}<file-view path="/Users/root1/Downloads/utilities/twscrape/twscrape/accounts_pool.py" start_line="1" end_line="393" total_lines="393">
1|import asyncio
2|import sqlite3
3|import uuid
4|from datetime import datetime, timezone
5|from typing import TypedDict
6|
7|from .account import Account
8|from .db import execute, fetchall, fetchone
9|from .http import HttpStatusError
10|from .logger import logger
11|from .login import LoginConfig, login
12|from .utils import get_env_bool, parse_cookies, utc
13|
14|
15|class NoAccountError(Exception):
16| pass
17|
18|
19|class AccountInfo(TypedDict):
20| username: str
21| logged_in: bool
22| active: bool
23| last_used: datetime | None
24| total_req: int
25| error_msg: str | None
26|
27|
28|def guess_delim(line: str):
29| lp, rp = tuple([x.strip() for x in line.split("username")])
30| return rp[0] if not lp else lp[-1]
31|
32|
33|class AccountsPool:
34| # _order_by: str = "RANDOM()"
35| _order_by: str = "username"
36|
37| def __init__(
38| self,
39| db_file="accounts.db",
40| login_config: LoginConfig | None = None,
41| raise_when_no_account=False,
42| ):
43| self._db_file = db_file
44| self._login_config = login_config or LoginConfig()
45| self._raise_when_no_account = raise_when_no_account
46|
47| async def load_from_file(self, filepath: str, line_format: str):
48| line_delim = guess_delim(line_format)
49| tokens = line_format.split(line_delim)
50|
51| required = {"username", "password", "email", "email_password"}
52| if not required.issubset(tokens):
53| raise ValueError(f"Invalid line format: {line_format}")
54|
55| accounts = []
56| with open(filepath) as f:
57| lines = f.read().split("\n")
58| lines = [x.strip() for x in lines if x.strip()]
59|
60| for line in lines:
61| data = [x.strip() for x in line.split(line_delim)]
62| if len(data) < len(tokens):
63| raise ValueError(f"Invalid line: {line}")
64|
65| data = data[: len(tokens)]
66| vals = {k: v for k, v in zip(tokens, data) if k != "_"}
67| accounts.append(vals)
68|
69| for x in accounts:
70| await self.add_account(**x)
71|
72| async def add_account(
73| self,
74| username: str,
75| password: str,
76| email: str,
77| email_password: str,
78| user_agent: str | None = None,
79| proxy: str | None = None,
80| cookies: str | None = None,
81| mfa_code: str | None = None,
82| ):
83| qs = "SELECT * FROM accounts WHERE username = :username"
84| rs = await fetchone(self._db_file, qs, {"username": username})
85| if rs:
86| logger.warning(f"Account {username} already exists")
87| return
88|
89| account = Account(
90| username=username,
91| password=password,
92| email=email,
93| email_password=email_password,
94| user_agent=user_agent or "@chrome",
95| active=False,
96| locks={},
97| stats={},
98| headers={},
99| cookies=parse_cookies(cookies) if cookies else {},
100| proxy=proxy,
101| mfa_code=mfa_code,
102| )
103|
104| if "ct0" in account.cookies:
105| account.active = True
106|
107| await self.save(account)
108| logger.info(f"Account {username} added successfully (active={account.active})")
109|
110| async def add_account_cookies(self, username: str, cookies: str):
111| existing = await self.get_account(username)
112| if existing is not None:
113| logger.warning(f"Account {username} already exists (active={existing.active})")
114| return
115|
116| await self.add_account(
117| username=username, password="_", email="_", email_password="_", cookies=cookies
118| )
119|
120| async def delete_accounts(self, usernames: str | list[str]):
121| usernames = usernames if isinstance(usernames, list) else [usernames]
122| usernames = list(set(usernames))
123| if not usernames:
124| logger.warning("No usernames provided")
125| return
126|
127| qs = f"""DELETE FROM accounts WHERE username IN ({",".join([f'"{x}"' for x in usernames])})"""
128| await execute(self._db_file, qs)
129|
130| async def delete_inactive(self):
131| qs = "DELETE FROM accounts WHERE active = false"
132| await execute(self._db_file, qs)
133|
134| async def get(self, username: str):
135| qs = "SELECT * FROM accounts WHERE username = :username"
136| rs = await fetchone(self._db_file, qs, {"username": username})
137| if not rs:
138| raise ValueError(f"Account {username} not found")
139| return Account.from_rs(rs)
140|
141| async def get_all(self):
142| qs = "SELECT * FROM accounts"
143| rs = await fetchall(self._db_file, qs)
144| return [Account.from_rs(x) for x in rs]
145|
146| async def get_account(self, username: str):
147| qs = "SELECT * FROM accounts WHERE username = :username"
148| rs = await fetchone(self._db_file, qs, {"username": username})
149| if not rs:
150| return None
151| return Account.from_rs(rs)
152|
153| async def save(self, account: Account):
154| data = account.to_rs()
155| cols = list(data.keys())
156|
157| qs = f"""
158| INSERT INTO accounts ({",".join(cols)}) VALUES ({",".join([f":{x}" for x in cols])})
159| ON CONFLICT(username) DO UPDATE SET {",".join([f"{x}=excluded.{x}" for x in cols])}
160| """
161| await execute(self._db_file, qs, data)
162|
163| async def login(self, account: Account):
164| try:
165| await login(account, cfg=self._login_config)
166| logger.info(f"Logged in to {account.username} successfully")
167| return True
168| except HttpStatusError as e:
169| rep = e.response
170| logger.error(f"Failed to login '{account.username}': {rep.status_code} - {rep.text}")
171| return False
172| except Exception as e:
173| logger.error(f"Failed to login '{account.username}': {e}")
174| return False
175| finally:
176| await self.save(account)
177|
178| async def login_all(self, usernames: list[str] | None = None):
179| if usernames is None:
180| qs = "SELECT * FROM accounts WHERE active = false AND error_msg IS NULL"
181| else:
182| us = ",".join([f'"{x}"' for x in usernames])
183| qs = f"SELECT * FROM accounts WHERE username IN ({us})"
184|
185| rs = await fetchall(self._db_file, qs)
186| accounts = [Account.from_rs(rs) for rs in rs]
187| # await asyncio.gather(*[login(x) for x in self.accounts])
188|
189| counter = {"total": len(accounts), "success": 0, "failed": 0}
190| for i, x in enumerate(accounts, start=1):
191| logger.info(f"[{i}/{len(accounts)}] Logging in {x.username} - {x.email}")
192| status = await self.login(x)
193| counter["success" if status else "failed"] += 1
194| return counter
195|
196| async def relogin(self, usernames: str | list[str]):
197| usernames = usernames if isinstance(usernames, list) else [usernames]
198| usernames = list(set(usernames))
199| if not usernames:
200| logger.warning("No usernames provided")
201| return
202|
203| qs = f"""
204| UPDATE accounts SET
205| active = false,
206| locks = json_object(),
207| last_used = NULL,
208| error_msg = NULL,
209| headers = json_object(),
210| cookies = json_object(),
211| user_agent = "@chrome"
212| WHERE username IN ({",".join([f'"{x}"' for x in usernames])})
213| """
214|
215| await execute(self._db_file, qs)
216| await self.login_all(usernames)
217|
218| async def relogin_failed(self):
219| qs = "SELECT username FROM accounts WHERE active = false AND error_msg IS NOT NULL"
220| rs = await fetchall(self._db_file, qs)
221| await self.relogin([x["username"] for x in rs])
222|
223| async def reset_locks(self):
224| qs = "UPDATE accounts SET locks = json_object()"
225| await execute(self._db_file, qs)
226|
227| async def set_active(self, username: str, active: bool):
228| qs = "UPDATE accounts SET active = :active WHERE username = :username"
229| await execute(self._db_file, qs, {"username": username, "active": active})
230|
231| async def lock_until(self, username: str, queue: str, unlock_at: int, req_count=0):
232| qs = f"""
233| UPDATE accounts SET
234| locks = json_set(locks, '$.{queue}', datetime({unlock_at}, 'unixepoch')),
235| stats = json_set(stats, '$.{queue}', COALESCE(json_extract(stats, '$.{queue}'), 0) + {req_count}),
236| last_used = datetime({utc.ts()}, 'unixepoch')
237| WHERE username = :username
238| """
239| await execute(self._db_file, qs, {"username": username})
240|
241| async def unlock(self, username: str, queue: str, req_count=0):
242| qs = f"""
243| UPDATE accounts SET
244| locks = json_remove(locks, '$.{queue}'),
245| stats = json_set(stats, '$.{queue}', COALESCE(json_extract(stats, '$.{queue}'), 0) + {req_count}),
246| last_used = datetime({utc.ts()}, 'unixepoch')
247| WHERE username = :username
248| """
249| await execute(self._db_file, qs, {"username": username})
250|
251| async def _get_and_lock(self, queue: str, condition: str):
252| # if space in condition, it's a subquery, otherwise it's username
253| condition = f"({condition})" if " " in condition else f"'{condition}'"
254|
255| if int(sqlite3.sqlite_version_info[1]) >= 35:
256| qs = f"""
257| UPDATE accounts SET
258| locks = json_set(locks, '$.{queue}', datetime('now', '+15 minutes')),
259| last_used = datetime({utc.ts()}, 'unixepoch')
260| WHERE username = {condition}
261| RETURNING *
262| """
263| rs = await fetchone(self._db_file, qs)
264| else:
265| tx = uuid.uuid4().hex
266| qs = f"""
267| UPDATE accounts SET
268| locks = json_set(locks, '$.{queue}', datetime('now', '+15 minutes')),
269| last_used = datetime({utc.ts()}, 'unixepoch'),
270| _tx = '{tx}'
271| WHERE username = {condition}
272| """
273| await execute(self._db_file, qs)
274|
275| qs = f"SELECT * FROM accounts WHERE _tx = '{tx}'"
276| rs = await fetchone(self._db_file, qs)
277|
278| return Account.from_rs(rs) if rs else None
279|
280| async def get_for_queue(self, queue: str):
281| q = f"""
282| SELECT username FROM accounts
283| WHERE active = true AND (
284| locks IS NULL
285| OR json_extract(locks, '$.{queue}') IS NULL
286| OR json_extract(locks, '$.{queue}') < datetime('now')
287| )
288| ORDER BY {self._order_by}
289| LIMIT 1
290| """
291|
292| return await self._get_and_lock(queue, q)
293|
294| async def get_for_queue_or_wait(self, queue: str) -> Account | None:
295| msg_shown = False
296| while True:
297| account = await self.get_for_queue(queue)
298| if not account:
299| if self._raise_when_no_account or get_env_bool("TWS_RAISE_WHEN_NO_ACCOUNT"):
300| raise NoAccountError(f"No account available for queue {queue}")
301|
302| if not msg_shown:
303| nat = await self.next_available_at(queue)
304| if not nat:
305| logger.warning("No active accounts. Stopping...")
306| return None
307|
308| msg = f'No account available for queue "{queue}". Next available at {nat}'
309| logger.info(msg)
310| msg_shown = True
311|
312| await asyncio.sleep(5)
313| continue
314| else:
315| if msg_shown:
316| logger.info(f"Continuing with account {account.username} on queue {queue}")
317|
318| return account
319|
320| async def next_available_at(self, queue: str):
321| qs = f"""
322| SELECT json_extract(locks, '$."{queue}"') as lock_until
323| FROM accounts
324| WHERE active = true AND json_extract(locks, '$."{queue}"') IS NOT NULL
325| ORDER BY lock_until ASC
326| LIMIT 1
327| """
328| rs = await fetchone(self._db_file, qs)
329| if rs:
330| now, trg = utc.now(), utc.from_iso(rs[0])
331| if trg < now:
332| return "now"
333|
334| at_local = datetime.now() + (trg - now)
335| return at_local.strftime("%H:%M:%S")
336|
337| return None
338|
339| async def mark_inactive(self, username: str, error_msg: str | None):
340| qs = """
341| UPDATE accounts SET active = false, error_msg = :error_msg
342| WHERE username = :username
343| """
344| await execute(self._db_file, qs, {"username": username, "error_msg": error_msg})
345|
346| async def stats(self):
347| def locks_count(queue: str):
348| return f"""
349| SELECT COUNT(*) FROM accounts
350| WHERE json_extract(locks, '$.{queue}') IS NOT NULL
351| AND json_extract(locks, '$.{queue}') > datetime('now')
352| """
353|
354| qs = "SELECT DISTINCT(f.key) as k from accounts, json_each(locks) f"
355| rs = await fetchall(self._db_file, qs)
356| gql_ops = [x["k"] for x in rs]
357|
358| config = [
359| ("total", "SELECT COUNT(*) FROM accounts"),
360| ("active", "SELECT COUNT(*) FROM accounts WHERE active = true"),
361| ("inactive", "SELECT COUNT(*) FROM accounts WHERE active = false"),
362| *[(f"locked_{x}", locks_count(x)) for x in gql_ops],
363| ]
364|
365| qs = f"SELECT {','.join([f'({q}) as {k}' for k, q in config])}"
366| rs = await fetchone(self._db_file, qs)
367| return dict(rs) if rs else {}
368|
369| async def accounts_info(self):
370| accounts = await self.get_all()
371|
372| items: list[AccountInfo] = []
373| for x in accounts:
374| item: AccountInfo = {
375| "username": x.username,
376| "logged_in": (x.headers or {}).get("authorization", "") != "",
377| "active": x.active,
378| "last_used": x.last_used,
379| "total_req": sum(x.stats.values()),
380| "error_msg": str(x.error_msg)[0:60],
381| }
382| items.append(item)
383|
384| old_time = datetime(1970, 1, 1).replace(tzinfo=timezone.utc)
385| items = sorted(items, key=lambda x: x["username"].lower())
386| items = sorted(
387| items,
388| key=lambda x: x["last_used"] or old_time if x["total_req"] > 0 else old_time,
389| reverse=True,
390| )
391| items = sorted(items, key=lambda x: x["active"], reverse=True)
392| # items = sorted(items, key=lambda x: x["total_req"], reverse=True)
393| return items
</file-view>{
"file_path": "/Users/root1/Downloads/utilities/twscrape/twscrape/queue_client.py"
}<file-view path="/Users/root1/Downloads/utilities/twscrape/twscrape/queue_client.py" start_line="1" end_line="310" total_lines="310">
1|import asyncio
2|import json
3|import os
4|from typing import Any
5|from urllib.parse import urlparse
6|
7|from . import telemetry
8|from .accounts_pool import Account, AccountsPool
9|from .http import ConnectError, HttpClient, HttpMethod, HttpStatusError, NetworkError, Response
10|from .logger import logger
11|from .utils import utc
12|from .xclid import XClIdGen
13|
14|ReqParams = dict[str, str | int] | None
15|TMP_TS = utc.now().isoformat().split(".")[0].replace("T", "_").replace(":", "-")[0:16]
16|
17|
18|class HandledError(Exception): ...
19|
20|
21|class AbortReqError(Exception): ...
22|
23|
24|class XClIdGenStore:
25| items: dict[str, XClIdGen] = {} # username -> XClIdGen
26|
27| @classmethod
28| async def get(cls, username: str, fresh=False) -> XClIdGen:
29| if username in cls.items and not fresh:
30| return cls.items[username]
31|
32| tries = 0
33| while tries < 3:
34| try:
35| clid_gen = await XClIdGen.create()
36| cls.items[username] = clid_gen
37| return clid_gen
38| except Exception as e:
39| tries += 1
40| logger.warning(
41| f"XClIdGen creation attempt {tries}/3 failed: {type(e).__name__}: {e}"
42| )
43| await asyncio.sleep(1)
44|
45| raise AbortReqError(
46| "Failed to create XClIdGen. See: https://github.com/vladkens/twscrape/issues/248"
47| )
48|
49|
50|class Ctx:
51| def __init__(self, acc: Account, clt: HttpClient):
52| self.req_count = 0
53| self.acc = acc
54| self.clt = clt
55|
56| async def aclose(self):
57| await self.clt.aclose()
58|
59| async def req(self, method: HttpMethod, url: str, params: ReqParams = None) -> Response:
60| # if code 404 on first try then generate new x-client-transaction-id and retry
61| # https://github.com/vladkens/twscrape/issues/248
62| path = urlparse(url).path or "/"
63|
64| tries = 0
65| while tries < 3:
66| gen = await XClIdGenStore.get(self.acc.username, fresh=tries > 0)
67| hdr = {"x-client-transaction-id": gen.calc(method, path)}
68| rep = await self.clt.request(method, url, params=params, headers=hdr)
69| if rep.status_code != 404:
70| return rep
71|
72| tries += 1
73| logger.debug(f"Retrying request with new x-client-transaction-id: {url}")
74| await asyncio.sleep(1)
75|
76| raise AbortReqError(
77| "Faield to get XClIdGen. See: https://github.com/vladkens/twscrape/issues/248"
78| )
79|
80|
81|def req_id(rep: Response):
82| lr = str(rep.headers.get("x-rate-limit-remaining", -1))
83| ll = str(rep.headers.get("x-rate-limit-limit", -1))
84| sz = max(len(lr), len(ll))
85| lr, ll = lr.rjust(sz), ll.rjust(sz)
86|
87| username = getattr(rep, "__username", "<UNKNOWN>")
88| return f"{lr}/{ll} - {username}"
89|
90|
91|def dump_rep(rep: Response):
92| count = getattr(dump_rep, "__count", -1) + 1
93| setattr(dump_rep, "__count", count)
94|
95| acc = getattr(rep, "__username", "<unknown>")
96| outfile = f"{count:05d}_{rep.status_code}_{acc}.txt"
97| outfile = f"/tmp/twscrape-{TMP_TS}/{outfile}"
98| os.makedirs(os.path.dirname(outfile), exist_ok=True)
99|
100| msg = []
101| msg.append(f"{count:,d} - {req_id(rep)}")
102| msg.append(f"{rep.status_code} {rep.request.method} {rep.request.url}")
103| msg.append("\n")
104| # msg.append("\n".join([str(x) for x in list(rep.request.headers.items())]))
105| msg.append("\n".join([str(x) for x in list(rep.headers.items())]))
106| msg.append("\n")
107|
108| try:
109| msg.append(json.dumps(rep.json(), indent=2))
110| except json.JSONDecodeError:
111| msg.append(rep.text)
112|
113| txt = "\n".join(msg)
114| with open(outfile, "w") as f:
115| f.write(txt)
116|
117|
118|class QueueClient:
119| def __init__(self, pool: AccountsPool, queue: str, debug=False, proxy: str | None = None):
120| self.pool = pool
121| self.queue = queue
122| self.debug = debug
123| self.ctx: Ctx | None = None
124| self.proxy = proxy
125|
126| async def __aenter__(self):
127| await self._get_ctx()
128| return self
129|
130| async def __aexit__(self, exc_type, exc_val, exc_tb):
131| await self._close_ctx()
132|
133| async def _close_ctx(self, reset_at=-1, inactive=False, msg: str | None = None):
134| if self.ctx is None:
135| return
136|
137| ctx, self.ctx, self.req_count = self.ctx, None, 0
138| username = ctx.acc.username
139| await ctx.aclose()
140|
141| if inactive:
142| await self.pool.mark_inactive(username, msg)
143| return
144|
145| if reset_at > 0:
146| await self.pool.lock_until(ctx.acc.username, self.queue, reset_at, ctx.req_count)
147| return
148|
149| await self.pool.unlock(ctx.acc.username, self.queue, ctx.req_count)
150|
151| async def _get_ctx(self):
152| if self.ctx:
153| return self.ctx
154|
155| acc = await self.pool.get_for_queue_or_wait(self.queue)
156| if acc is None:
157| return None
158|
159| clt = acc.make_client(proxy=self.proxy)
160| self.ctx = Ctx(acc, clt)
161| return self.ctx
162|
163| async def _check_rep(self, rep: Response) -> None:
164| """
165| This function can raise Exception and request will be retried or aborted
166| Or if None is returned, response will passed to api parser as is
167| """
168|
169| if self.debug:
170| dump_rep(rep)
171|
172| if "text/html" in rep.headers.get("content-type", "") and rep.status_code >= 400:
173| src = "Cloudflare" if "cf-ray" in rep.headers else "HTML"
174| logger.warning(f"Blocked by {src}: {rep.status_code} - {req_id(rep)}")
175| raise AbortReqError()
176|
177| try:
178| res = rep.json()
179| except json.JSONDecodeError:
180| res: Any = {"_raw": rep.text}
181|
182| limit_remaining = int(rep.headers.get("x-rate-limit-remaining", -1))
183| limit_reset = int(rep.headers.get("x-rate-limit-reset", -1))
184| # limit_max = int(rep.headers.get("x-rate-limit-limit", -1))
185|
186| err_msg = "OK"
187| if "errors" in res:
188| err_msg = {f"({x.get('code', -1)}) {x['message']}" for x in res["errors"]}
189| err_msg = "; ".join(list(err_msg))
190|
191| log_msg = f"{rep.status_code:3d} - {req_id(rep)} - {err_msg}"
192| logger.trace(log_msg)
193|
194| # for dev: need to add some features in api.py
195| if err_msg.startswith("(336) The following features cannot be null"):
196| logger.error(f"[DEV] Update required: {err_msg}")
197| exit(1)
198|
199| # general api rate limit
200| if limit_remaining == 0 and limit_reset > 0:
201| logger.debug(f"Rate limited: {log_msg}")
202| await self._close_ctx(limit_reset)
203| raise HandledError()
204|
205| # no way to check is account banned in direct way, but this check should work
206| if err_msg.startswith("(88) Rate limit exceeded") and limit_remaining > 0:
207| logger.warning(f"Ban detected: {log_msg}")
208| await self._close_ctx(-1, inactive=True, msg=err_msg)
209| raise HandledError()
210|
211| if err_msg.startswith("(326) Authorization: Denied by access control"):
212| logger.warning(f"Ban detected: {log_msg}")
213| await self._close_ctx(-1, inactive=True, msg=err_msg)
214| raise HandledError()
215|
216| if err_msg.startswith("(32) Could not authenticate you"):
217| logger.warning(f"Session expired or banned: {log_msg}")
218| await self._close_ctx(-1, inactive=True, msg=err_msg)
219| raise HandledError()
220|
221| if err_msg == "OK" and rep.status_code == 403:
222| logger.warning(f"Session expired or banned: {log_msg}")
223| await self._close_ctx(-1, inactive=True, msg=None)
224| raise HandledError()
225|
226| # something from twitter side - abort all queries, see: https://github.com/vladkens/twscrape/pull/80
227| if err_msg.startswith("(131) Dependency: Internal error"):
228| # looks like when data exists, we can ignore this error
229| # https://github.com/vladkens/twscrape/issues/166
230| if rep.status_code == 200 and "data" in res and "user" in res["data"]:
231| err_msg = "OK"
232| else:
233| logger.warning(f"Dependency error (request skipped): {err_msg}")
234| raise AbortReqError()
235|
236| # content not found
237| if rep.status_code == 200 and "_Missing: No status found with that ID" in err_msg:
238| return # ignore this error
239|
240| # something from twitter side - just ignore it, see: https://github.com/vladkens/twscrape/pull/95
241| if rep.status_code == 200 and "Authorization" in err_msg:
242| logger.warning(f"Authorization unknown error: {log_msg}")
243| return
244|
245| if err_msg != "OK":
246| logger.warning(f"API unknown error: {log_msg}")
247| return # ignore any other unknown errors
248|
249| try:
250| rep.raise_for_status()
251| except HttpStatusError:
252| logger.error(f"Unhandled API response code: {log_msg}")
253| await self._close_ctx(utc.ts() + 60 * 15) # 15 minutes
254| raise HandledError()
255|
256| async def get(self, url: str, params: ReqParams = None) -> Response | None:
257| return await self.req("GET", url, params=params)
258|
259| async def req(self, method: HttpMethod, url: str, params: ReqParams = None) -> Response | None:
260| unknown_retry, connection_retry = 0, 0
261|
262| while True:
263| ctx = await self._get_ctx() # not need to close client, class implements __aexit__
264| if ctx is None:
265| return None
266|
267| try:
268| source = telemetry.current_source()
269| telemetry.capture(
270| "gql_request",
271| {
272| "operation": self.queue,
273| "http_method": method,
274| "http_backend": getattr(ctx.clt, "backend", "unknown"),
275| "source": source,
276| "$current_url": f"{source}://twscrape/gql/{self.queue}",
277| },
278| )
279| rep = await ctx.req(method, url, params=params)
280| setattr(rep, "__username", ctx.acc.username)
281| await self._check_rep(rep)
282|
283| ctx.req_count += 1 # count only successful
284| unknown_retry, connection_retry = 0, 0
285| return rep
286| except AbortReqError:
287| # abort all queries
288| return
289| except HandledError:
290| # retry with new account
291| continue
292| except NetworkError:
293| # http transport failed, just retry with same account
294| continue
295| except ConnectError as e:
296| # if proxy misconfigured or host unreachable
297| connection_retry += 1
298| if connection_retry >= 3:
299| raise e
300| except Exception as e:
301| unknown_retry += 1
302| if unknown_retry >= 3:
303| msg = [
304| "Unknown error. Account timeouted for 15 minutes.",
305| "Create issue please: https://github.com/vladkens/twscrape/issues",
306| f"If it mistake, you can unlock accounts with `twscrape reset_locks`. Err: {type(e)}: {e}",
307| ]
308|
309| logger.warning(" ".join(msg))
310| await self._close_ctx(utc.ts() + 60 * 15) # 15 minutes
</file-view>{
"file_path": "/Users/root1/Downloads/utilities/twscrape/twscrape/account.py"
}<file-view path="/Users/root1/Downloads/utilities/twscrape/twscrape/account.py" start_line="1" end_line="69" total_lines="69">
1|import hashlib
2|import json
3|import os
4|import sqlite3
5|from dataclasses import asdict, dataclass, field
6|from datetime import datetime
7|
8|from .http import HttpClient
9|from .http import make_client as _make_http_client
10|from .models import JSONTrait
11|from .utils import parse_proxy, utc
12|
13|TOKEN = "Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA"
14|
15|
16|@dataclass
17|class Account(JSONTrait):
18| username: str
19| password: str
20| email: str
21| email_password: str
22| user_agent: str
23| active: bool
24| locks: dict[str, datetime] = field(default_factory=dict) # queue: datetime
25| stats: dict[str, int] = field(default_factory=dict) # queue: requests
26| headers: dict[str, str] = field(default_factory=dict)
27| cookies: dict[str, str] = field(default_factory=dict)
28| mfa_code: str | None = None
29| proxy: str | None = None
30| error_msg: str | None = None
31| last_used: datetime | None = None
32| _tx: str | None = None
33|
34| @staticmethod
35| def from_rs(rs: sqlite3.Row):
36| doc = dict(rs)
37| doc["locks"] = {k: utc.from_iso(v) for k, v in json.loads(doc["locks"]).items()}
38| doc["stats"] = {k: v for k, v in json.loads(doc["stats"]).items() if isinstance(v, int)}
39| doc["headers"] = json.loads(doc["headers"])
40| doc["cookies"] = json.loads(doc["cookies"])
41| doc["active"] = bool(doc["active"])
42| doc["last_used"] = utc.from_iso(doc["last_used"]) if doc["last_used"] else None
43| return Account(**doc)
44|
45| def to_rs(self):
46| rs = asdict(self)
47| rs["locks"] = json.dumps(rs["locks"], default=lambda x: x.isoformat())
48| rs["stats"] = json.dumps(rs["stats"])
49| rs["headers"] = json.dumps(rs["headers"])
50| rs["cookies"] = json.dumps(rs["cookies"])
51| rs["last_used"] = rs["last_used"].isoformat() if rs["last_used"] else None
52| return rs
53|
54| def make_client(self, proxy: str | None = None) -> HttpClient:
55| proxies = [proxy, os.getenv("TWS_PROXY"), self.proxy]
56| proxies = [x for x in proxies if x is not None]
57| proxy = parse_proxy(proxies[0]) if proxies else None
58|
59| headers = {**self.headers}
60| headers["user-agent"] = self.user_agent
61| headers["content-type"] = "application/json"
62| headers["authorization"] = TOKEN
63| headers["x-twitter-active-user"] = "yes"
64| headers["x-twitter-client-language"] = "en"
65| if "ct0" in self.cookies:
66| headers["x-csrf-token"] = self.cookies["ct0"]
67|
68| seed = int(hashlib.sha256(self.username.encode()).hexdigest()[:8], 16)
69| return _make_http_client(proxy=proxy, headers=headers, cookies=self.cookies, seed=seed)
</file-view>{
"file_path": "/Users/root1/Downloads/utilities/twscrape/twscrape/http.py"
}<file-view path="/Users/root1/Downloads/utilities/twscrape/twscrape/http.py" start_line="1" end_line="272" total_lines="272">
1|import importlib.util
2|import os
3|import random
4|from abc import ABC, abstractmethod
5|from typing import Any, Literal, cast
6|
7|from fake_useragent import UserAgent
8|
9|from .logger import logger
10|
11|HttpMethod = Literal["GET", "POST", "PUT", "DELETE", "OPTIONS", "HEAD", "TRACE", "PATCH"]
12|
13|_UNSET = object()
14|_CURL_MAX_RETRIES = 3
15|_LOGGED_BACKENDS: set[str] = set()
16|
17|# https://curl-impersonate.readthedocs.io/en/latest/fingerprints.html
18|_BROWSER_FAMILIES = {"chrome", "safari", "firefox", "edge"}
19|
20|_ua = UserAgent()
21|
22|
23|def _resolve_browser(hint: str | None, seed: int | None = None) -> tuple[str, str]:
24| """Return (ua_string, family) for a @hint or a real UA string."""
25| if hint and not hint.startswith("@"):
26| return hint, "chrome"
27| family = hint[1:].lower() if hint else "chrome"
28| if family not in _BROWSER_FAMILIES:
29| family = "chrome"
30| if seed is not None:
31| uas = [x["useragent"] for x in _ua.data_browsers if family in (x["browser"] or "").lower()]
32| if uas:
33| return random.Random(seed).choice(uas), family
34| return getattr(_ua, family, _ua.chrome), family
35|
36|
37|class Response:
38| """Thin wrapper around httpx.Response or curl_cffi.Response."""
39|
40| def __init__(self, rep: Any):
41| self._rep = rep
42| self._json: Any = _UNSET
43|
44| @property
45| def status_code(self) -> int:
46| return self._rep.status_code
47|
48| @property
49| def text(self) -> str:
50| return self._rep.text
51|
52| @property
53| def content(self) -> bytes:
54| return self._rep.content
55|
56| @property
57| def headers(self) -> Any:
58| return self._rep.headers
59|
60| @property
61| def url(self) -> Any:
62| return self._rep.url
63|
64| @property
65| def request(self) -> Any:
66| return self._rep.request
67|
68| def json(self) -> Any:
69| if self._json is _UNSET:
70| self._json = self._rep.json()
71| return self._json
72|
73| def raise_for_status(self) -> None:
74| if self._rep.status_code >= 400:
75| raise HttpStatusError(f"HTTP {self._rep.status_code}", response=self)
76|
77|
78|class HttpError(Exception): ...
79|
80|
81|class NetworkError(HttpError): ...
82|
83|
84|class ConnectError(HttpError): ...
85|
86|
87|class HttpStatusError(HttpError):
88| def __init__(self, message: str, *, response: Response):
89| super().__init__(message)
90| self.response = response
91|
92|
93|class HttpClient(ABC):
94| backend: str
95|
96| @abstractmethod
97| async def request(self, method: HttpMethod, url: str, **kwargs) -> Response: ...
98|
99| async def get(self, url: str, **kwargs) -> Response:
100| return await self.request("GET", url, **kwargs)
101|
102| async def post(self, url: str, **kwargs) -> Response:
103| return await self.request("POST", url, **kwargs)
104|
105| @abstractmethod
106| async def aclose(self) -> None: ...
107|
108| async def __aenter__(self) -> "HttpClient":
109| return self
110|
111| async def __aexit__(self, *args: Any) -> None:
112| await self.aclose()
113|
114| @property
115| @abstractmethod
116| def cookies(self) -> Any: ...
117|
118| @property
119| @abstractmethod
120| def headers(self) -> Any: ...
121|
122|
123|class HttpxClient(HttpClient):
124| backend = "httpx"
125|
126| def __init__(
127| self,
128| *,
129| proxy: str | None = None,
130| headers: dict | None = None,
131| cookies: dict | None = None,
132| seed: int | None = None,
133| ):
134| import httpx
135| from httpx import AsyncHTTPTransport
136|
137| self._httpx = httpx
138| transport = AsyncHTTPTransport(retries=3)
139| resolved_headers = dict(headers or {})
140| ua_string, _ = _resolve_browser(resolved_headers.get("user-agent"), seed=seed)
141| resolved_headers["user-agent"] = ua_string
142| self._client = httpx.AsyncClient(
143| proxy=proxy,
144| follow_redirects=True,
145| transport=transport,
146| headers=resolved_headers,
147| cookies=cookies or {},
148| )
149|
150| @property
151| def cookies(self) -> Any:
152| return self._client.cookies
153|
154| @property
155| def headers(self) -> Any:
156| return self._client.headers
157|
158| async def request(self, method: HttpMethod, url: str, **kwargs) -> Response:
159| return await self._wrap(self._client.request(method, url, **kwargs))
160|
161| async def aclose(self) -> None:
162| await self._client.aclose()
163|
164| async def _wrap(self, coro: Any) -> Response:
165| hx = self._httpx
166| try:
167| return Response(await coro)
168| except (hx.ConnectError, hx.ConnectTimeout) as e:
169| raise ConnectError(str(e)) from e
170| except (hx.ReadTimeout, hx.WriteTimeout, hx.PoolTimeout, hx.ProxyError) as e:
171| raise NetworkError(str(e)) from e
172|
173|
174|class CurlClient(HttpClient):
175| backend = "curl"
176|
177| def __init__(
178| self, *, proxy: str | None = None, headers: dict | None = None, cookies: dict | None = None
179| ):
180| from curl_cffi.requests import AsyncSession, BrowserTypeLiteral
181|
182| _, family = _resolve_browser((headers or {}).get("user-agent"))
183| # strip user-agent — curl_cffi sets its own UA for the impersonated profile
184| safe_headers = {k: v for k, v in (headers or {}).items() if k.lower() != "user-agent"}
185| self._session = AsyncSession(
186| impersonate=cast(BrowserTypeLiteral, family),
187| proxy=proxy,
188| allow_redirects=True,
189| headers=safe_headers,
190| )
191| if cookies:
192| self._session.cookies.update(cookies)
193|
194| @property
195| def cookies(self) -> Any:
196| return self._session.cookies
197|
198| @property
199| def headers(self) -> Any:
200| return self._session.headers
201|
202| async def request(self, method: HttpMethod, url: str, **kwargs) -> Response:
203| last_err: Exception | None = None
204| for _ in range(_CURL_MAX_RETRIES + 1):
205| try:
206| return await self._wrap(self._session.request(method, url, **kwargs))
207| except NetworkError as e:
208| last_err = e
209| if last_err is not None:
210| raise last_err
211| raise RuntimeError("CurlClient request failed without an error")
212|
213| async def aclose(self) -> None:
214| await self._session.close()
215|
216| async def _wrap(self, coro: Any) -> Response:
217| try:
218| return Response(await coro)
219| except Exception as e:
220| from curl_cffi.requests import errors as _curl_errors
221|
222| if isinstance(e, _curl_errors.RequestsError):
223| # libcurl error codes: 5=COULDNT_RESOLVE_PROXY, 6=COULDNT_RESOLVE_HOST, 7=COULDNT_CONNECT
224| if getattr(e, "code", -1) in {5, 6, 7}:
225| raise ConnectError(str(e)) from e
226| raise NetworkError(str(e)) from e
227| raise
228|
229|
230|def _detect_backend() -> str:
231| forced = os.getenv("TWS_HTTP_BACKEND", "").lower().strip()
232|
233| if forced == "curl":
234| if importlib.util.find_spec("curl_cffi") is None:
235| raise ImportError(
236| "TWS_HTTP_BACKEND=curl but curl-cffi is not installed. "
237| "Run: pip install twscrape[curl]"
238| )
239| return "curl"
240|
241| if forced in ("", "httpx"):
242| if importlib.util.find_spec("httpx") is None:
243| raise ImportError(
244| "TWS_HTTP_BACKEND=httpx but httpx is not installed. Run: pip install twscrape"
245| )
246| return "httpx"
247|
248| raise ValueError(f"Invalid TWS_HTTP_BACKEND={forced!r}. Expected 'curl' or 'httpx'.")
249|
250|
251|def make_client(
252| backend: str | None = None,
253| *,
254| proxy: str | None = None,
255| headers: dict | None = None,
256| cookies: dict | None = None,
257| seed: int | None = None,
258|) -> HttpClient:
259| if backend is None:
260| backend = _detect_backend()
261|
262| if backend not in _LOGGED_BACKENDS:
263| name = "curl-cffi" if backend == "curl" else backend
264| logger.debug(f"Using {name} HTTP client")
265| _LOGGED_BACKENDS.add(backend)
266|
267| if backend == "curl":
268| return CurlClient(proxy=proxy, headers=headers, cookies=cookies)
269| if backend == "httpx":
270| return HttpxClient(proxy=proxy, headers=headers, cookies=cookies, seed=seed)
271|
272| raise ValueError(f"Unknown backend: {backend!r}. Expected 'curl' or 'httpx'.")
</file-view>{
"file_path": "/Users/root1/Downloads/utilities/twscrape/twscrape/login.py"
}<file-view path="/Users/root1/Downloads/utilities/twscrape/twscrape/login.py" start_line="1" end_line="279" total_lines="279">
1|import imaplib
2|from dataclasses import dataclass
3|from datetime import timedelta
4|from typing import Any
5|
6|import pyotp
7|
8|from .account import Account
9|from .http import HttpClient, Response
10|from .imap import imap_get_email_code, imap_login
11|from .logger import logger
12|from .utils import utc
13|
14|LOGIN_URL = "https://api.x.com/1.1/onboarding/task.json"
15|
16|
17|@dataclass
18|class LoginConfig:
19| email_first: bool = False
20| manual: bool = False
21|
22|
23|@dataclass
24|class TaskCtx:
25| client: HttpClient
26| acc: Account
27| cfg: LoginConfig
28| prev: Any
29| imap: None | imaplib.IMAP4_SSL
30|
31|
32|async def get_guest_token(client: HttpClient):
33| rep = await client.post("https://api.x.com/1.1/guest/activate.json")
34| rep.raise_for_status()
35| return rep.json()["guest_token"]
36|
37|
38|async def login_initiate(client: HttpClient) -> Response:
39| payload = {
40| "input_flow_data": {
41| "flow_context": {"debug_overrides": {}, "start_location": {"location": "unknown"}}
42| },
43| "subtask_versions": {},
44| }
45|
46| rep = await client.post(LOGIN_URL, params={"flow_name": "login"}, json=payload)
47| rep.raise_for_status()
48| return rep
49|
50|
51|async def login_alternate_identifier(ctx: TaskCtx, *, username: str) -> Response:
52| payload = {
53| "flow_token": ctx.prev["flow_token"],
54| "subtask_inputs": [
55| {
56| "subtask_id": "LoginEnterAlternateIdentifierSubtask",
57| "enter_text": {"text": username, "link": "next_link"},
58| }
59| ],
60| }
61|
62| rep = await ctx.client.post(LOGIN_URL, json=payload)
63| rep.raise_for_status()
64| return rep
65|
66|
67|async def login_instrumentation(ctx: TaskCtx) -> Response:
68| payload = {
69| "flow_token": ctx.prev["flow_token"],
70| "subtask_inputs": [
71| {
72| "subtask_id": "LoginJsInstrumentationSubtask",
73| "js_instrumentation": {"response": "{}", "link": "next_link"},
74| }
75| ],
76| }
77|
78| rep = await ctx.client.post(LOGIN_URL, json=payload)
79| rep.raise_for_status()
80| return rep
81|
82|
83|async def login_enter_username(ctx: TaskCtx) -> Response:
84| payload = {
85| "flow_token": ctx.prev["flow_token"],
86| "subtask_inputs": [
87| {
88| "subtask_id": "LoginEnterUserIdentifierSSO",
89| "settings_list": {
90| "setting_responses": [
91| {
92| "key": "user_identifier",
93| "response_data": {"text_data": {"result": ctx.acc.username}},
94| }
95| ],
96| "link": "next_link",
97| },
98| }
99| ],
100| }
101|
102| rep = await ctx.client.post(LOGIN_URL, json=payload)
103| rep.raise_for_status()
104| return rep
105|
106|
107|async def login_enter_password(ctx: TaskCtx) -> Response:
108| payload = {
109| "flow_token": ctx.prev["flow_token"],
110| "subtask_inputs": [
111| {
112| "subtask_id": "LoginEnterPassword",
113| "enter_password": {"password": ctx.acc.password, "link": "next_link"},
114| }
115| ],
116| }
117|
118| rep = await ctx.client.post(LOGIN_URL, json=payload)
119| rep.raise_for_status()
120| return rep
121|
122|
123|async def login_two_factor_auth_challenge(ctx: TaskCtx) -> Response:
124| if ctx.acc.mfa_code is None:
125| raise ValueError("MFA code is required")
126|
127| totp = pyotp.TOTP(ctx.acc.mfa_code)
128| payload = {
129| "flow_token": ctx.prev["flow_token"],
130| "subtask_inputs": [
131| {
132| "subtask_id": "LoginTwoFactorAuthChallenge",
133| "enter_text": {"text": totp.now(), "link": "next_link"},
134| }
135| ],
136| }
137|
138| rep = await ctx.client.post(LOGIN_URL, json=payload)
139| rep.raise_for_status()
140| return rep
141|
142|
143|async def login_duplication_check(ctx: TaskCtx) -> Response:
144| payload = {
145| "flow_token": ctx.prev["flow_token"],
146| "subtask_inputs": [
147| {
148| "subtask_id": "AccountDuplicationCheck",
149| "check_logged_in_account": {"link": "AccountDuplicationCheck_false"},
150| }
151| ],
152| }
153|
154| rep = await ctx.client.post(LOGIN_URL, json=payload)
155| rep.raise_for_status()
156| return rep
157|
158|
159|async def login_confirm_email(ctx: TaskCtx) -> Response:
160| payload = {
161| "flow_token": ctx.prev["flow_token"],
162| "subtask_inputs": [
163| {
164| "subtask_id": "LoginAcid",
165| "enter_text": {"text": ctx.acc.email, "link": "next_link"},
166| }
167| ],
168| }
169|
170| rep = await ctx.client.post(LOGIN_URL, json=payload)
171| rep.raise_for_status()
172| return rep
173|
174|
175|async def login_confirm_email_code(ctx: TaskCtx):
176| if ctx.cfg.manual:
177| print(f"Enter email code for {ctx.acc.username} / {ctx.acc.email}")
178| value = input("Code: ")
179| value = value.strip()
180| else:
181| if not ctx.imap:
182| ctx.imap = await imap_login(ctx.acc.email, ctx.acc.email_password)
183|
184| now_time = utc.now() - timedelta(seconds=30)
185| value = await imap_get_email_code(ctx.imap, ctx.acc.email, now_time)
186|
187| payload = {
188| "flow_token": ctx.prev["flow_token"],
189| "subtask_inputs": [
190| {
191| "subtask_id": "LoginAcid",
192| "enter_text": {"text": value, "link": "next_link"},
193| }
194| ],
195| }
196|
197| rep = await ctx.client.post(LOGIN_URL, json=payload)
198| rep.raise_for_status()
199| return rep
200|
201|
202|async def login_success(ctx: TaskCtx) -> Response:
203| payload = {
204| "flow_token": ctx.prev["flow_token"],
205| "subtask_inputs": [],
206| }
207|
208| rep = await ctx.client.post(LOGIN_URL, json=payload)
209| rep.raise_for_status()
210| return rep
211|
212|
213|async def next_login_task(ctx: TaskCtx, rep: Response):
214| ct0 = ctx.client.cookies.get("ct0", None)
215| if ct0:
216| ctx.client.headers["x-csrf-token"] = ct0
217| ctx.client.headers["x-twitter-auth-type"] = "OAuth2Session"
218|
219| ctx.prev = rep.json()
220| assert "flow_token" in ctx.prev, f"flow_token not in {rep.text}"
221|
222| for x in ctx.prev["subtasks"]:
223| task_id = x["subtask_id"]
224|
225| try:
226| if task_id == "LoginSuccessSubtask":
227| return await login_success(ctx)
228| if task_id == "LoginAcid":
229| is_code = x["enter_text"]["hint_text"].lower() == "confirmation code"
230| fn = login_confirm_email_code if is_code else login_confirm_email
231| return await fn(ctx)
232| if task_id == "AccountDuplicationCheck":
233| return await login_duplication_check(ctx)
234| if task_id == "LoginEnterPassword":
235| return await login_enter_password(ctx)
236| if task_id == "LoginTwoFactorAuthChallenge":
237| return await login_two_factor_auth_challenge(ctx)
238| if task_id == "LoginEnterUserIdentifierSSO":
239| return await login_enter_username(ctx)
240| if task_id == "LoginJsInstrumentationSubtask":
241| return await login_instrumentation(ctx)
242| if task_id == "LoginEnterAlternateIdentifierSubtask":
243| return await login_alternate_identifier(ctx, username=ctx.acc.username)
244| except Exception as e:
245| ctx.acc.error_msg = f"login_step={task_id} err={e}"
246| raise e
247|
248| return None
249|
250|
251|async def login(acc: Account, cfg: LoginConfig | None = None) -> Account:
252| log_id = f"{acc.username} - {acc.email}"
253| if acc.active:
254| logger.info(f"account already active {log_id}")
255| return acc
256|
257| cfg, imap = cfg or LoginConfig(), None
258| if cfg.email_first and not cfg.manual:
259| imap = await imap_login(acc.email, acc.email_password)
260|
261| async with acc.make_client() as client:
262| guest_token = await get_guest_token(client)
263| client.headers["x-guest-token"] = guest_token
264|
265| rep = await login_initiate(client)
266| ctx = TaskCtx(client, acc, cfg, None, imap)
267| while True:
268| rep = await next_login_task(ctx, rep)
269| if not rep:
270| break
271|
272| assert "ct0" in client.cookies, "ct0 not in cookies (most likely ip ban)"
273| client.headers["x-csrf-token"] = client.cookies["ct0"]
274| client.headers["x-twitter-auth-type"] = "OAuth2Session"
275|
276| acc.active = True
277| acc.headers = dict(client.headers.items())
278| acc.cookies = dict(client.cookies.items())
279| return acc
</file-view>{
"file_path": "/Users/root1/Downloads/utilities/twscrape/twscrape/db.py"
}<file-view path="/Users/root1/Downloads/utilities/twscrape/twscrape/db.py" start_line="1" end_line="156" total_lines="156">
1|import asyncio
2|import random
3|import sqlite3
4|from collections import defaultdict
5|
6|import aiosqlite
7|
8|from .logger import logger
9|
10|MIN_SQLITE_VERSION = "3.24"
11|
12|_lock = asyncio.Lock()
13|
14|
15|def lock_retry(max_retries=10):
16| # this lock decorator has double nature:
17| # 1. it uses asyncio lock in same process
18| # 2. it retries when db locked by other process (eg. two cli instances running)
19| def decorator(func):
20| async def wrapper(*args, **kwargs):
21| for i in range(max_retries):
22| try:
23| async with _lock:
24| return await func(*args, **kwargs)
25| except sqlite3.OperationalError as e:
26| if i == max_retries - 1 or "database is locked" not in str(e):
27| raise e
28|
29| await asyncio.sleep(random.uniform(0.5, 1.0))
30|
31| return wrapper
32|
33| return decorator
34|
35|
36|async def get_sqlite_version():
37| async with aiosqlite.connect(":memory:") as db: # noqa: SIM117
38| async with db.execute("SELECT SQLITE_VERSION()") as cur:
39| rs = await cur.fetchone()
40| return rs[0] if rs else "3.0.0"
41|
42|
43|async def check_version():
44| ver = await get_sqlite_version()
45| ver = ".".join(ver.split(".")[:2])
46|
47| try:
48| msg = f"SQLite version '{ver}' is too old, please upgrade to {MIN_SQLITE_VERSION}+"
49| if float(ver) < float(MIN_SQLITE_VERSION):
50| raise SystemError(msg)
51| except ValueError:
52| pass
53|
54|
55|async def migrate(db: aiosqlite.Connection):
56| async with db.execute("PRAGMA user_version") as cur:
57| rs = await cur.fetchone()
58| uv = rs[0] if rs else 0
59|
60| async def v1():
61| qs = """
62| CREATE TABLE IF NOT EXISTS accounts (
63| username TEXT PRIMARY KEY NOT NULL COLLATE NOCASE,
64| password TEXT NOT NULL,
65| email TEXT NOT NULL COLLATE NOCASE,
66| email_password TEXT NOT NULL,
67| user_agent TEXT NOT NULL,
68| active BOOLEAN DEFAULT FALSE NOT NULL,
69| locks TEXT DEFAULT '{}' NOT NULL,
70| headers TEXT DEFAULT '{}' NOT NULL,
71| cookies TEXT DEFAULT '{}' NOT NULL,
72| proxy TEXT DEFAULT NULL,
73| error_msg TEXT DEFAULT NULL
74| );"""
75| await db.execute(qs)
76|
77| async def v2():
78| await db.execute("ALTER TABLE accounts ADD COLUMN stats TEXT DEFAULT '{}' NOT NULL")
79| await db.execute("ALTER TABLE accounts ADD COLUMN last_used TEXT DEFAULT NULL")
80|
81| async def v3():
82| await db.execute("ALTER TABLE accounts ADD COLUMN _tx TEXT DEFAULT NULL")
83|
84| async def v4():
85| await db.execute("ALTER TABLE accounts ADD COLUMN mfa_code TEXT DEFAULT NULL")
86|
87| migrations = {
88| 1: v1,
89| 2: v2,
90| 3: v3,
91| 4: v4,
92| }
93|
94| # logger.debug(f"Current migration v{uv} (latest v{len(migrations)})")
95| for i in range(uv + 1, len(migrations) + 1):
96| logger.info(f"Running migration to v{i}")
97| try:
98| await migrations[i]()
99| except sqlite3.OperationalError as e:
100| if "duplicate column name" not in str(e):
101| raise e
102|
103| await db.execute(f"PRAGMA user_version = {i}")
104| await db.commit()
105|
106|
107|class DB:
108| _init_queries: defaultdict[str, list[str]] = defaultdict(list)
109| _init_once: defaultdict[str, bool] = defaultdict(bool)
110|
111| def __init__(self, db_path):
112| self.db_path = db_path
113| self.conn = None
114|
115| async def __aenter__(self):
116| await check_version()
117| db = await aiosqlite.connect(self.db_path)
118| db.row_factory = aiosqlite.Row
119|
120| if not self._init_once[self.db_path]:
121| await migrate(db)
122| self._init_once[self.db_path] = True
123|
124| self.conn = db
125| return db
126|
127| async def __aexit__(self, exc_type, exc_val, exc_tb):
128| if self.conn:
129| await self.conn.commit()
130| await self.conn.close()
131|
132|
133|@lock_retry()
134|async def execute(db_path: str, qs: str, params: dict | None = None):
135| async with DB(db_path) as db:
136| await db.execute(qs, params)
137|
138|
139|@lock_retry()
140|async def fetchone(db_path: str, qs: str, params: dict | None = None):
141| async with DB(db_path) as db, db.execute(qs, params) as cur:
142| row = await cur.fetchone()
143| return row
144|
145|
146|@lock_retry()
147|async def fetchall(db_path: str, qs: str, params: dict | None = None):
148| async with DB(db_path) as db, db.execute(qs, params) as cur:
149| rows = await cur.fetchall()
150| return rows
151|
152|
153|@lock_retry()
154|async def executemany(db_path: str, qs: str, params: list[dict]):
155| async with DB(db_path) as db:
156| await db.executemany(qs, params)
</file-view>{
"file_path": "/Users/root1/Downloads/utilities/twscrape/twscrape/cli.py"
}<file-view path="/Users/root1/Downloads/utilities/twscrape/twscrape/cli.py" start_line="1" end_line="237" total_lines="237">
1|#!/usr/bin/env python3
2|
3|import argparse
4|import asyncio
5|import getpass
6|import io
7|import json
8|import sqlite3
9|from importlib.metadata import version
10|
11|from . import telemetry
12|from .api import API, AccountsPool
13|from .db import get_sqlite_version
14|from .http import Response
15|from .logger import logger, set_log_level
16|from .login import LoginConfig
17|from .models import Tweet, User
18|from .utils import print_table
19|
20|
21|class CustomHelpFormatter(argparse.HelpFormatter):
22| def __init__(self, prog):
23| super().__init__(prog, max_help_position=30, width=120)
24|
25|
26|def get_fn_arg(args):
27| if not hasattr(args, "arg_name"):
28| logger.error(f"Missing argument name for command: {args.command}")
29| exit(1)
30|
31| value = getattr(args, args.arg_name, None)
32| if value is not None:
33| return args.arg_name, value
34|
35| logger.error(f"Missing argument value: {args.arg_name}")
36| exit(1)
37|
38|
39|def to_str(doc: Response | Tweet | User | None) -> str:
40| if doc is None:
41| return "Not Found. See --raw for more details."
42|
43| tmp = doc.json()
44| return tmp if isinstance(tmp, str) else json.dumps(tmp, default=str)
45|
46|
47|async def main(args):
48| telemetry.set_source("cli")
49| if args.debug:
50| set_log_level("DEBUG")
51|
52| if args.command == "version":
53| print(f"twscrape: {version('twscrape')}")
54| print(f"SQLite runtime: {sqlite3.sqlite_version} ({await get_sqlite_version()})")
55| return
56|
57| login_config = LoginConfig(getattr(args, "email_first", False), getattr(args, "manual", False))
58| pool = AccountsPool(args.db, login_config=login_config)
59| api = API(pool, debug=args.debug)
60|
61| if args.command == "accounts":
62| print_table([dict(x) for x in await pool.accounts_info()])
63| return
64|
65| if args.command == "stats":
66| rep = await pool.stats()
67| total, active, inactive = rep["total"], rep["active"], rep["inactive"]
68|
69| res = []
70| for k, v in rep.items():
71| if not k.startswith("locked") or v == 0:
72| continue
73| res.append({"queue": k, "locked": v, "available": max(active - v, 0)})
74|
75| res = sorted(res, key=lambda x: x["locked"], reverse=True)
76| print_table(res, hr_after=True)
77| print(f"Total: {total} - Active: {active} - Inactive: {inactive}")
78| return
79|
80| if args.command == "add_accounts":
81| await pool.load_from_file(args.file_path, args.line_format)
82| print("\nNow run:\ntwscrape login_accounts")
83| return
84|
85| if args.command == "add_cookie":
86| cookies = args.cookies
87| if not cookies:
88| cookies = getpass.getpass("cookies (e.g. auth_token=xxx; ct0=yyy): ")
89| await pool.add_account_cookies(args.username, cookies)
90| return
91|
92| if args.command == "del_accounts":
93| await pool.delete_accounts(args.usernames)
94| return
95|
96| if args.command == "login_accounts":
97| stats = await pool.login_all()
98| print(stats)
99| return
100|
101| if args.command == "relogin_failed":
102| await pool.relogin_failed()
103| return
104|
105| if args.command == "relogin":
106| await pool.relogin(args.usernames)
107| return
108|
109| if args.command == "reset_locks":
110| await pool.reset_locks()
111| return
112|
113| if args.command == "delete_inactive":
114| await pool.delete_inactive()
115| return
116|
117| fn = args.command + "_raw" if args.raw else args.command
118| fn = getattr(api, fn, None)
119| if fn is None:
120| logger.error(f"Unknown command: {args.command}")
121| exit(1)
122|
123| _, val = get_fn_arg(args)
124|
125| if "limit" in args:
126| async for doc in fn(val, limit=args.limit):
127| print(to_str(doc))
128| else:
129| doc = await fn(val)
130| print(to_str(doc))
131|
132|
133|async def _run(args):
134| try:
135| await main(args)
136| finally:
137| await telemetry.flush()
138|
139|
140|def custom_help(p):
141| buffer = io.StringIO()
142| p.print_help(buffer)
143| msg = buffer.getvalue()
144|
145| cmd = msg.split("positional arguments:")[1].strip().split("\n")[0]
146| msg = msg.replace("positional arguments:", "commands:")
147| msg = [x for x in msg.split("\n") if cmd not in x and "..." not in x]
148| msg[0] = f"{msg[0]} <command> [...]"
149|
150| i = 0
151| for i, line in enumerate(msg):
152| if line.strip().startswith("search"):
153| break
154|
155| msg.insert(i, "")
156| msg.insert(i + 1, "search commands:")
157|
158| print("\n".join(msg))
159|
160|
161|def run():
162| p = argparse.ArgumentParser(add_help=False, formatter_class=CustomHelpFormatter)
163| p.add_argument("--db", default="accounts.db", help="Accounts database file")
164| p.add_argument("--debug", action="store_true", help="Enable debug mode")
165| subparsers = p.add_subparsers(dest="command")
166|
167| def c_one(name: str, msg: str, a_name: str, a_msg: str, a_type: type = str):
168| p = subparsers.add_parser(name, help=msg)
169| p.add_argument(a_name, help=a_msg, type=a_type)
170| p.set_defaults(arg_name=a_name)
171| p.add_argument("--raw", action="store_true", help="Print raw response")
172| return p
173|
174| def c_lim(name: str, msg: str, a_name: str, a_msg: str, a_type: type = str):
175| p = c_one(name, msg, a_name, a_msg, a_type)
176| p.add_argument("--limit", type=int, default=-1, help="Max tweets to retrieve")
177| return p
178|
179| subparsers.add_parser("version", help="Show version")
180| subparsers.add_parser("accounts", help="List all accounts")
181| subparsers.add_parser("stats", help="Get current usage stats")
182|
183| add_accounts = subparsers.add_parser("add_accounts", help="Add accounts from file")
184| add_accounts.add_argument("file_path", help="File with accounts")
185| add_accounts.add_argument("line_format", help="Account fields separated by delimiter")
186|
187| add_cookie = subparsers.add_parser("add_cookie", help="Add one account from cookies")
188| add_cookie.add_argument("username", help="Twitter/X username")
189| add_cookie.add_argument("cookies", nargs="?", default=None, help="Cookie string")
190|
191| del_accounts = subparsers.add_parser("del_accounts", help="Delete accounts by username")
192| del_accounts.add_argument("usernames", nargs="+", default=[], help="Usernames to delete")
193|
194| login_cmd = subparsers.add_parser("login_accounts", help="Log in inactive accounts")
195| relogin = subparsers.add_parser("relogin", help="Re-log in selected accounts")
196| relogin.add_argument("usernames", nargs="+", default=[], help="Usernames to re-login")
197| re_failed = subparsers.add_parser("relogin_failed", help="Retry failed account logins")
198|
199| login_commands = [login_cmd, relogin, re_failed]
200| for cmd in login_commands:
201| cmd.add_argument("--email-first", action="store_true", help="Check email first")
202| cmd.add_argument("--manual", action="store_true", help="Enter email code manually")
203|
204| subparsers.add_parser("reset_locks", help="Reset all locks")
205| subparsers.add_parser("delete_inactive", help="Delete inactive accounts")
206|
207| c_lim("search", "Search for tweets", "query", "Search query")
208| c_one("tweet_details", "Get tweet details", "tweet_id", "Tweet ID", int)
209| c_lim("tweet_replies", "Get replies of a tweet", "tweet_id", "Tweet ID", int)
210| c_lim("tweet_thread", "Get thread tweets", "tweet_id", "Tweet ID", int)
211| c_lim("retweeters", "Get retweeters of a tweet", "tweet_id", "Tweet ID", int)
212| c_one("user_by_login", "Get user data by username", "username", "Username")
213| c_one("user_about", "Get about info for username", "username", "Username")
214| c_lim("following", "Get user following", "user_id", "User ID", int)
215| c_lim("followers", "Get user followers", "user_id", "User ID", int)
216| # https://x.com/xDaily/status/1701694747767648500
217| c_lim("verified_followers", "Get user verified followers", "user_id", "User ID", int)
218| c_lim("subscriptions", "Get user subscriptions", "user_id", "User ID", int)
219| c_lim("user_tweets", "Get user tweets", "user_id", "User ID", int)
220| c_lim("user_tweets_and_replies", "Get user tweets and replies", "user_id", "User ID", int)
221| c_lim("user_media", "Get user's media", "user_id", "User ID", int)
222| c_lim("list_timeline", "Get tweets from list", "list_id", "List ID", int)
223| c_lim("list_members", "Get List members by list ID", "list_id", "List ID", int)
224| c_one("community_info", "Get community info", "community_id", "Community ID", str)
225| c_lim("community_members", "Get community members", "community_id", "Community ID", str)
226| c_lim("community_moderators", "Get community moderators", "community_id", "Community ID", str)
227| c_lim("community_tweets", "Get community tweets", "community_id", "Community ID", str)
228| c_lim("trends", "Get trends", "trend_id", "Trend ID or name", str)
229|
230| args = p.parse_args()
231| if args.command is None:
232| return custom_help(p)
233|
234| try:
235| asyncio.run(_run(args))
236| except KeyboardInterrupt:
237| pass
</file-view>{
"file_path": "/Users/root1/Downloads/utilities/twscrape/twscrape/models.py",
"limit": 100
}<file-view path="/Users/root1/Downloads/utilities/twscrape/twscrape/models.py" start_line="1" end_line="100" total_lines="917">
1|import email.utils
2|import json
3|import os
4|import random
5|import re
6|import string
7|import sys
8|import traceback
9|from dataclasses import asdict, dataclass, field
10|from datetime import datetime, timezone
11|from typing import Generator, Optional, Union
12|
13|from .http import Response
14|from .logger import logger
15|from .utils import find_item, get_or, int_or, to_old_rep, utc
16|
17|
18|@dataclass
19|class JSONTrait:
20| def dict(self):
21| return asdict(self)
22|
23| def json(self):
24| return json.dumps(self.dict(), default=str)
25|
26|
27|@dataclass
28|class Coordinates(JSONTrait):
29| longitude: float
30| latitude: float
31|
32| @staticmethod
33| def parse(tw_obj: dict):
34| if tw_obj.get("coordinates"):
35| coords = tw_obj["coordinates"]["coordinates"]
36| return Coordinates(coords[0], coords[1])
37| if tw_obj.get("geo"):
38| coords = tw_obj["geo"]["coordinates"]
39| return Coordinates(coords[1], coords[0])
40| return None
41|
42|
43|@dataclass
44|class Place(JSONTrait):
45| id: str
46| fullName: str
47| name: str
48| type: str
49| country: str
50| countryCode: str
51|
52| @staticmethod
53| def parse(obj: dict):
54| return Place(
55| id=obj["id"],
56| fullName=obj["full_name"],
57| name=obj["name"],
58| type=obj["place_type"],
59| country=obj["country"],
60| countryCode=obj["country_code"],
61| )
62|
63|
64|@dataclass
65|class TextLink(JSONTrait):
66| url: str
67| text: str | None
68| tcourl: str | None
69|
70| @staticmethod
71| def parse(obj: dict):
72| url1 = obj.get("expanded_url")
73| url2 = obj.get("url")
74| text = obj.get("display_url")
75|
76| if not isinstance(url1, str) or not isinstance(url2, str):
77| return None
78|
79| return TextLink(url=url1, text=text, tcourl=url2)
80|
81|
82|@dataclass
83|class AccountAbout(JSONTrait):
84| screen_name: str
85| name: str
86| rest_id: int | None
87| account_based_in: str | None
88| location_accurate: bool | None
89| affiliate_username: str | None
90| source: str | None
91| username_changes: int | None
92| username_last_changed_at: int | None
93| is_identity_verified: bool | None
94| verified_since_msec: int | None
95|
96| @staticmethod
97| def parse(obj: dict):
98| about = obj.get("about_profile") or {}
99| core = obj.get("core") or {}
100| verification = obj.get("verification_info", {}) or {}
</file-view>{
"file_path": "/Users/root1/Downloads/utilities/twscrape/twscrape/xclid.py",
"limit": 60
}<file-view path="/Users/root1/Downloads/utilities/twscrape/twscrape/xclid.py" start_line="1" end_line="60" total_lines="367">
1|import asyncio
2|import base64
3|import hashlib
4|import math
5|import random
6|import re
7|import time
8|from urllib.parse import urljoin
9|
10|import bs4
11|
12|from .http import HttpClient
13|from .http import make_client as _make_http_client
14|
15|
16|def _make_client() -> HttpClient:
17| return _make_http_client(headers={"user-agent": "@chrome"})
18|
19|
20|async def get_tw_page_text(url: str, clt: HttpClient):
21| rep = await clt.get(url)
22|
23| rep.raise_for_status()
24| if ">document.location =" not in rep.text:
25| return rep.text
26|
27| url = rep.text.split('document.location = "')[1].split('"')[0]
28| rep = await clt.get(url)
29| rep.raise_for_status()
30| if 'action="https://x.com/x/migrate" method="post"' not in rep.text:
31| return rep.text
32|
33| data = {}
34| for x in rep.text.split("<input")[1:]:
35| name = x.split('name="')[1].split('"')[0]
36| value = x.split('value="')[1].split('"')[0]
37| data[name] = value
38|
39| rep = await clt.post("https://x.com/x/migrate", json=data)
40| rep.raise_for_status()
41|
42| return rep.text
43|
44|
45|def script_url(k: str, v: str):
46| return f"https://abs.twimg.com/responsive-web/client-web/{k}.{v}.js"
47|
48|
49|# Current X web build (Vite): script bundles are linked directly in the page
50|# HTML under https://abs.twimg.com/x-web/.../*.js (modulepreload links + entry).
51|ASSET_URL_RE = re.compile(r"https://[\w.-]+/x-web/[\w./-]+\.js")
52|
53|
54|def get_scripts_list(text: str) -> list[str]:
55| """
56| Extract chunk script URLs from the X homepage HTML.
57|
58| Current build (x-web / Vite): scripts are linked directly in the page HTML,
59| so we just collect those URLs. If none are found we fall back to the legacy
60| webpack build, which embeds two maps in the page and requires URL
</file-view>{
"file_path": "/Users/root1/Downloads/utilities/twscrape/examples/about_account.py"
}<file-view path="/Users/root1/Downloads/utilities/twscrape/examples/about_account.py" start_line="1" end_line="15" total_lines="15"> 1|import asyncio 2| 3|from twscrape import API 4| 5| 6|async def main(): 7| api = API() 8| username = "sama" 9| 10| info = await api.user_about(username) 11| print(info.dict() if info else "Not found") 12| 13| 14|if __name__ == "__main__": 15| asyncio.run(main()) </file-view>
I have enough understanding now. Let me write up a summary of how twscrape works.
I've cloned the repo and traced through the code. Here's how twscrape works:
twscrape is an async Python library + CLI that scrapes X/Twitter by hitting its private GraphQL endpoints (the same ones the web app uses). It runs on a pool of your own X accounts, rotates them when rate-limited, and returns either parsed SNScrape-style models or raw JSON.
CLI (cli.py) ──► API (api.py) ──► QueueClient (queue_client.py) ──► HttpClient (http.py)
│ │
▼ ▼
models.py (parse) AccountsPool (accounts_pool.py)
│
▼
SQLite (db.py)
accounts_pool.py + db.py + account.pyaccounts.db by default). Schema is migrated on first open (db.py v1–v4).Account (account.py) holds username/password/email/email_password, headers, cookies (incl. auth_token + ct0), per-queue locks (JSON map of queue → unlock_time), per-queue stats (request counts), proxy, MFA secret, error_msg.add_account_cookies(username, "auth_token=xxx; ct0=yyy"). If ct0 is present the account is immediately marked active=True — no login flow needed.add_accounts file.txt username:password:email:email_password then login_accounts runs the OAuth-style flow.login.pyImplements X's "flow" login (https://api.x.com/1.1/onboarding/task.json?flow_name=login):
/1.1/guest/activate.jsonsubtasks list (LoginEnterUserIdentifierSSO → LoginEnterPassword → LoginJsInstrumentationSubtask → AccountDuplicationCheck → LoginAcid (email code) → LoginSuccessSubtask).imap.py) or entered manually with --manual.ct0 cookie + CSRF header are saved; account marked active.accounts_pool.pyget_for_queue(queue): SQL picks an active account whose lock for that queue is NULL or in the past, then atomically locks it for 15 min via json_set(locks, '$.QueueName', datetime('now','+15 minutes')). Uses SQLite RETURNING * on 3.35+, otherwise a _tx UUID trick.get_for_queue_or_wait: loops with asyncio.sleep(5) until one frees up (or raises NoAccountError if TWS_RAISE_WHEN_NO_ACCOUNT=1).lock_until / unlock / mark_inactive update the row when a request finishes.http.pyTWS_HTTP_BACKEND=curl, install twscrape[curl]) for browser-like TLS fingerprinting.Account.make_client() builds the client with the bearer token (hardcoded web-app token in account.py:TOKEN), x-csrf-token = ct0, user-agent resolved from fake-useragent seeded by hash of username (so each account keeps a stable UA), and proxy resolution priority: api.proxy > TWS_PROXY env > account proxy.x-client-transaction-id — xclid.pyX requires a per-request x-client-transaction-id header derived from JS bundles. XClIdGen scrapes the x.com homepage, downloads the script chunks, extracts a key + animation frames, and computes the ID via the same algorithm the web client uses. XClIdGenStore caches one per username and regenerates on 404 (see issue #248).
queue_client.pyQueueClient is an async context manager (async with QueueClient(pool, queue, ...) as client:):
__aenter__ → grabs a locked account from the pool, builds its HttpClient.req() loop:
Ctx.req() which sets the x-client-transaction-id header._check_rep() inspects the response:
cf-ray → Cloudflare block → AbortReqErrorx-rate-limit-remaining: 0 → lock account until x-rate-limit-reset, retry with another account (HandledError)88, 326, 32, or 403 OK → mark account inactive (banned/expired)131 internal error → abort; 336 features null → hard exit (tells dev to update features)HandledError → loop continues with a new account; AbortReqError → returns None; NetworkError → retry same account; ConnectError/unknown → retry up to 3× then lock 15 min.__aexit__ → _close_ctx either unlocks the account, locks it until a reset time, or marks it inactive.api.pyGQL_FEATURES is a big dict of feature flags copied from x.com (must be kept in sync with the web client or X returns 336 errors).OP_* constants are queryId/OperationName pairs (regenerated by scripts/update_gql_ops.py)._gql_item(op, kv, ft) — single request (e.g. user_by_login, tweet_details)._gql_items(op, kv, ft, limit, cursor_type) — async generator that paginates timelines via cursor until limit reached or no cursor. Filters out cursor-, messageprompt-, module-, who-to-follow- entries. Tolerates up to 3 empty pages in a row before stopping._raw variant (yields Response) and a parsed variant that runs the JSON through models.py (parse_tweets, parse_users, parse_community, parse_trends, …). Parsed models are SNScrape-compatible dataclasses with .dict() / .json().gather() (utils.py) is just a convenience that collects an async generator into a list.contextlib.aclosing so breaking early releases the account lock promptly.cli.pyargparse-based. twscrape <command> [args] [--raw] [--limit=N] [--db=FILE]. For data commands it dynamically calls api.<command> or api.<command>_raw and prints JSON Lines. Account-management commands (add_accounts, add_cookie, login_accounts, relogin, accounts, stats, reset_locks, delete_inactive, del_accounts) call the pool directly.
telemetry.pySends anonymous events (GraphQL op name + HTTP backend + source) on shutdown. Disabled via TWS_TELEMETRY=0. No cookies/queries/URLs are sent.
twscrape search "elon musk" --limit=20API(AccountsPool("accounts.db")), calls api.search("elon musk", limit=20).search → search_raw → _gql_items(OP_SearchTimeline, {rawQuery, count:20, product:"Latest"}, limit=20)._gql_items opens a QueueClient for queue "SearchTimeline".HttpClient with that account's cookies/headers/proxy.variables+features params, GET https://x.com/i/api/graphql/<queryId>/SearchTimeline, attach x-client-transaction-id from XClIdGen._check_rep validates rate limits / bans; on rate-limit it locks the account until reset and the loop grabs another.parse_tweets turns each tweetResult into a Tweet model.limit hit or cursor exhausted; context exit unlocks the account.x-client-transaction-id generatorWant me to dig deeper into any specific part (e.g. the x-client-transaction-id algorithm, the parsers, or run the test suite)?
run it to get the list of who @sdand is following
The user wants me to run twscrape to get the list of who @sdand is following. This requires actual X/Twitter accounts to be set up with cookies. I don't have any accounts configured, and this would require the user's own X account credentials/cookies. Let me check if there's any existing accounts.db or setup, but realistically this requires the user to provide their own X account cookies. I should explain this rather than attempting something that won't work. Let me first check if there's any existing setup, then explain what's needed.
{
"command": "cd /Users/root1/Downloads/utilities/twscrape && ls -la accounts.db 2>/dev/null; pip show twscrape 2>/dev/null | head -3; which twscrape 2>/dev/null; python3 -c \"import twscrape; print(twscrape.__file__)\" 2>&1 | head -3"
}Output from command in shell de9460: Traceback (most recent call last): File "<string>", line 1, in <module> File "/Users/root1/Downloads/utilities/twscrape/twscrape/__init__.py", line 2, in <module> Exit code: 0
{
"command": "cd /Users/root1/Downloads/utilities/twscrape && python3 -c \"import twscrape; print(twscrape.__file__)\" 2>&1 | tail -10"
}Output from command in shell d424db:
Traceback (most recent call last):
File "<string>", line 1, in <module>
File "/Users/root1/Downloads/utilities/twscrape/twscrape/__init__.py", line 2, in <module>
from .account import Account
File "/Users/root1/Downloads/utilities/twscrape/twscrape/account.py", line 8, in <module>
from .http import HttpClient
File "/Users/root1/Downloads/utilities/twscrape/twscrape/http.py", line 7, in <module>
from fake_useragent import UserAgent
ModuleNotFoundError: No module named 'fake_useragent'
Exit code: 1Dependencies aren't installed. Let me install them. But more importantly, this requires actual X account cookies to work. I can't actually scrape X without the user's account credentials. Let me install the package and then explain what's needed. Actually, let me think about this. The user wants me to run it. I should: 1. Install dependencies 2. Try to run it 3. It will fail because no accounts are configured 4. Explain that they need to provide cookies Let me install first.
Dependencies aren't installed. Let me install the package first.
{
"command": "cd /Users/root1/Downloads/utilities/twscrape && pip install -e . 2>&1 | tail -15"
}Output from command in shell 01b819: … (30 lines truncated) Downloading fake_useragent-2.2.0-py3-none-any.whl (161 kB) Downloading py_machineid-1.0.0-py3-none-any.whl (4.9 kB) Downloading pyotp-2.10.0-py3-none-any.whl (13 kB) Building wheels for collected packages: twscrape Building editable for twscrape (pyproject.toml): started Building editable for twscrape (pyproject.toml): finished with status 'done' Created wheel for twscrape: filename=twscrape-0.19.0-py3-none-any.whl size=6363 sha256=e7c5d19990f38445917024522adec93a1810f00b869ba7a09996aad7b5b99b8e Stored in directory: /private/var/folders/d9/mzhyr7lx6xbf236j5ll9v_300000gn/T/pip-ephem-wheel-cache-9y7bxb5q/wheels/4a/c7/8e/511742414956db5e6ad81209adfbcd30c341bbe8a9e85d5b28 Successfully built twscrape Installing collected packages: py-machineid, pyotp, fake-useragent, aiosqlite, twscrape Successfully installed aiosqlite-0.22.1 fake-useragent-2.2.0 py-machineid-1.0.0 pyotp-2.10.0 twscrape-0.19.0 [notice] A new release of pip is available: 26.1 -> 26.1.2 [notice] To update, run: pip install --upgrade pip Exit code: 0 `| tail -15` was parsed out (15 of 45 total lines shown). <truncation_notice> Full output written to: /var/folders/d9/mzhyr7lx6xbf236j5ll9v_300000gn/T/devin-overflows-501/f55c416f/content.txt </truncation_notice>
{
"command": "cd /Users/root1/Downloads/utilities/twscrape && twscrape accounts 2>&1 | tail -20"
}Output from command in shell 41e0c6: 2026-06-20 14:29:40.259 | INFO | twscrape.db:migrate:96 - Running migration to v1 2026-06-20 14:29:40.260 | INFO | twscrape.db:migrate:96 - Running migration to v2 2026-06-20 14:29:40.261 | INFO | twscrape.db:migrate:96 - Running migration to v3 2026-06-20 14:29:40.261 | INFO | twscrape.db:migrate:96 - Running migration to v4 Exit code: 0
The DB is initialized but there are no accounts. I need cookies to actually scrape. Let me first get the user_id for @sdand - actually I can't even do that without an account. twscrape requires authorized X accounts to make any API call. I should explain to the user that I need them to provide X account cookies to proceed. Let me also note that `following` requires a user_id (numeric), not a username, so the flow would be: 1. Add cookies for an account 2. `twscrape user_by_login sdand` to get the numeric user_id 3. `twscrape following <user_id> --limit=N` But I can't do any of this without cookies. Let me ask the user.
The package is installed and the DB is initialized, but there are no accounts in the pool. twscrape requires an authorized X/Twitter account to make any API call — it has no anonymous mode. I can't proceed without you providing credentials.
To get the following list for @sdand, the flow would be:
twscrape add_cookie <name> "auth_token=xxx; ct0=yyy"@sdand to a numeric user_id → twscrape user_by_login sdandtwscrape following <user_id> --limit=NTo get the cookies:
https://x.comauth_token and ct0{
"questions": [
{
"question": "How do you want to provide X account credentials so I can run the scrape?",
"header": "Auth method",
"options": [
{
"label": "Paste cookies",
"description": "You paste auth_token and ct0 cookie values here; I'll add them via twscrape add_cookie and run the scrape."
},
{
"label": "Username/password",
"description": "You provide username/password/email; I'll run the login flow (may require email verification code)."
},
{
"label": "Skip — explain only",
"description": "Don't run anything; just show me the exact commands to run myself."
}
]
}
]
}User answered your questions:
{
"How do you want to provide X account credentials so I can run the scrape?": {
"selected": [
"Paste cookies"
]
}
}open a chrome window, get my cookies
The user wants me to open Chrome and get their cookies. I have a Playwright MCP server available. Let me check what tools it has.
{
"server_name": "playwright"
}[
{
"server_name": "playwright",
"tools": [
{
"name": "browser_close",
"description": "Close the page",
"inputSchema": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {},
"additionalProperties": false
},
"annotations": {
"title": "Close browser",
"readOnlyHint": false,
"destructiveHint": true,
"openWorldHint": true
}
},
{
"name": "browser_resize",
"description": "Resize the browser window",
"inputSchema": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"width": {
"type": "number",
"description": "Width of the browser window"
},
"height": {
"type": "number",
"description": "Height of the browser window"
}
},
"required": [
"width",
"height"
],
"additionalProperties": false
},
"annotations": {
"title": "Resize browser window",
"readOnlyHint": false,
"destructiveHint": true,
"openWorldHint": true
}
},
{
"name": "browser_console_messages",
"description": "Returns all console messages",
"inputSchema": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"level": {
"default": "info",
"description": "Level of the console messages to return. Each level includes the messages of more severe levels. Defaults to \"info\".",
"type": "string",
"enum": [
"error",
"warning",
"info",
"debug"
]
},
"all": {
"description": "Return all console messages since the beginning of the session, not just since the last navigation. Defaults to false.",
"type": "boolean"
},
"filename": {
"description": "Filename to save the console messages to. If not provided, messages are returned as text.",
"type": "string"
}
},
"required": [
"level"
],
"additionalProperties": false
},
"annotations": {
"title": "Get console messages",
"readOnlyHint": true,
"destructiveHint": false,
"openWorldHint": true
}
},
{
"name": "browser_handle_dialog",
"description": "Handle a dialog",
"inputSchema": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"accept": {
"type": "boolean",
"description": "Whether to accept the dialog."
},
"promptText": {
"description": "The text of the prompt in case of a prompt dialog.",
"type": "string"
}
},
"required": [
"accept"
],
"additionalProperties": false
},
"annotations": {
"title": "Handle a dialog",
"readOnlyHint": false,
"destructiveHint": true,
"openWorldHint": true
}
},
{
"name": "browser_evaluate",
"description": "Evaluate JavaScript expression on page or element",
"inputSchema": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"element": {
"description": "Human-readable element description used to obtain permission to interact with the element",
"type": "string"
},
"target": {
"description": "Exact target element reference from the page snapshot, or a unique element selector",
"type": "string"
},
"function": {
"type": "string",
"description": "() => { /* code */ } or (element) => { /* code */ } when element is provided"
},
"filename": {
"description": "Filename to save the result to. If not provided, result is returned as text.",
"type": "string"
}
},
"required": [
"function"
],
"additionalProperties": false
},
"annotations": {
"title": "Evaluate JavaScript",
"readOnlyHint": false,
"destructiveHint": true,
"openWorldHint": true
}
},
{
"name": "browser_file_upload",
"description": "Upload one or multiple files",
"inputSchema": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"paths": {
"description": "The absolute paths to the files to upload. Can be single file or multiple files. If omitted, file chooser is cancelled.",
"type": "array",
"items": {
"type": "string"
}
}
},
"additionalProperties": false
},
"annotations": {
"title": "Upload files",
"readOnlyHint": false,
"destructiveHint": true,
"openWorldHint": true
}
},
{
"name": "browser_drop",
"description": "Drop files or MIME-typed data onto an element, as if dragged from outside the page. At least one of \"paths\" or \"data\" must be provided.",
"inputSchema": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"element": {
"description": "Human-readable element description used to obtain permission to interact with the element",
"type": "string"
},
"target": {
"type": "string",
"description": "Exact target element reference from the page snapshot, or a unique element selector"
},
"paths": {
"description": "Absolute paths to files to drop onto the element.",
"type": "array",
"items": {
"type": "string"
}
},
"data": {
"description": "Data to drop, as a map of MIME type to string value (e.g. {\"text/plain\": \"hello\", \"text/uri-list\": \"https://example.com\"}).",
"type": "object",
"propertyNames": {
"type": "string"
},
"additionalProperties": {
"type": "string"
}
}
},
"required": [
"target"
],
"additionalProperties": false
},
"annotations": {
"title": "Drop files or data onto an element",
"readOnlyHint": false,
"destructiveHint": true,
"openWorldHint": true
}
},
{
"name": "browser_fill_form",
"description": "Fill multiple form fields",
"inputSchema": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"fields": {
"type": "array",
"items": {
"type": "object",
"properties": {
"element": {
"description": "Human-readable element description used to obtain permission to interact with the element",
"type": "string"
},
"target": {
"type": "string",
"description": "Exact target element reference from the page snapshot, or a unique element selector"
},
"name": {
"type": "string",
"description": "Human-readable field name"
},
"type": {
"type": "string",
"enum": [
"textbox",
"checkbox",
"radio",
"combobox",
"slider"
],
"description": "Type of the field"
},
"value": {
"type": "string",
"description": "Value to fill in the field. If the field is a checkbox, the value should be `true` or `false`. If the field is a combobox, the value should be the text of the option."
}
},
"required": [
"target",
"name",
"type",
"value"
],
"additionalProperties": false
},
"description": "Fields to fill in"
}
},
"required": [
"fields"
],
"additionalProperties": false
},
"annotations": {
"title": "Fill form",
"readOnlyHint": false,
"destructiveHint": true,
"openWorldHint": true
}
},
{
"name": "browser_press_key",
"description": "Press a key on the keyboard",
"inputSchema": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"key": {
"type": "string",
"description": "Name of the key to press or a character to generate, such as `ArrowLeft` or `a`"
}
},
… (17 chars truncated)
… (495 lines truncated)
<truncation_notice>
Full output written to: /var/folders/d9/mzhyr7lx6xbf236j5ll9v_300000gn/T/devin-overflows-501/8e54642b/content.txt
</truncation_notice>Let me look at the navigation tool to open x.com.
{
"server_name": "playwright",
"tool_name": "browser_navigate",
"arguments": {
"url": "https://x.com"
}
}[
{
"type": "text",
"text": "### Ran Playwright code\n```js\nawait page.goto('https://x.com');\n```\n### Page\n- Page URL: https://x.com/\n- Page Title: X. It’s what’s happening / X\n- Console: 4 errors, 0 warnings\n### Snapshot\n- [Snapshot](.playwright-mcp/page-2026-06-20T21-30-32-710Z.yml)\n### Events\n- New console entries: .playwright-mcp/console-2026-06-20T21-30-28-567Z.log#L1-L5"
}
]{
"server_name": "playwright",
"tool_name": "browser_evaluate",
"arguments": {
"function": "() => { const get = name => { const m = document.cookie.match(new RegExp('(?:^|; )' + name + '=([^;]*)')); return m ? m[1] : null; }; return JSON.stringify({ auth_token: get('auth_token'), ct0: get('ct0'), all: document.cookie }); }"
}
}[
{
"type": "text",
"text": "### Result\n\"{\\\"auth_token\\\":null,\\\"ct0\\\":null,\\\"all\\\":\\\"guest_id_marketing=v1%3A178199102867874044; guest_id_ads=v1%3A178199102867874044; guest_id=v1%3A178199102867874044; personalization_id=\\\\\\\"v1_Q9jjt29PAaQDobTNvcmFTA==\\\\\\\"; gt=2068446415253053521; __cuid=fccc28f2-e96d-4c1e-ad62-77857aa1ba21; __cuid=fccc28f2-e96d-4c1e-ad62-77857aa1ba21; g_state={\\\\\\\"i_l\\\\\\\":0,\\\\\\\"i_ll\\\\\\\":1781991032766,\\\\\\\"i_b\\\\\\\":\\\\\\\"L4cainl5fHw/v24FK1aMfVYQ0Msf9IaM9zZACZ8oCw4\\\\\\\",\\\\\\\"i_e\\\\\\\":{\\\\\\\"enable_itp_optimization\\\\\\\":24},\\\\\\\"i_et\\\\\\\":1781991032766}\\\"}\"\n### Ran Playwright code\n```js\nawait page.evaluate('() => { const get = name => { const m = document.cookie.match(new RegExp(\\'(?:^|; )\\' + name + \\'=([^;]*)\\')); return m ? m[1] : null; }; return JSON.stringify({ auth_token: get(\\'auth_token\\'), ct0: get(\\'ct0\\'), all: document.cookie }); }');\n```\n### Page\n- Page URL: https://x.com/\n- Page Title: X. It’s what’s happening / X\n- Console: 5 errors, 1 warnings\n### Events\n- New console entries: .playwright-mcp/console-2026-06-20T21-30-28-567Z.log#L6-L7"
}
]The browser opened x.com but you're not logged in — only guest cookies are present, no `auth_token` or `ct0`. I need you to log in. Let me take a snapshot to see the current state and find the login button.
The browser opened x.com but you're not logged in — only guest cookies are present (no auth_token or ct0). You need to sign in first. Let me show you the page so you can log in.
{
"server_name": "playwright",
"tool_name": "browser_snapshot",
"arguments": {}
}[
{
"type": "text",
"text": "### Page\n- Page URL: https://x.com/i/jf/onboarding/web#/s/passkey_1fa_auth/r-rflv6\n- Page Title: X - The Everything App / X\n- Console: 5 errors, 1 warnings\n### Snapshot\n```yaml\n- generic [ref=e2]:\n - generic:\n - generic:\n - generic:\n - generic:\n - generic:\n - generic:\n - dialog:\n - group [active] [ref=e112]:\n - dialog [ref=e114]:\n - generic [ref=e119]:\n - button [ref=e122] [cursor=pointer]:\n - img [ref=e125]\n - generic [ref=e127]:\n - img [ref=e132]\n - generic [ref=e137]:\n - generic [ref=e138]:\n - paragraph [ref=e140]: Sign in with passkey\n - paragraph [ref=e142]: Use your passkey to sign in instantly.\n - img \"Loading spinner\" [ref=e148]\n - main [ref=e4]:\n - generic [ref=e6]:\n - generic [ref=e7]:\n - generic [ref=e9]:\n - heading \"Happening now.\" [level=1] [ref=e10]\n - generic [ref=e14]:\n - generic [ref=e15]:\n - button \"Continue with phone\" [ref=e16] [cursor=pointer]:\n - generic [ref=e18]:\n - img [ref=e20]\n - paragraph [ref=e23]: Continue with phone\n - iframe [ref=e111]:\n - button \"Continue with Google. Opens in new tab\" [ref=f2e3] [cursor=pointer]:\n - generic [ref=f2e5]:\n - img [ref=f2e7]\n - generic [ref=f2e14]: Continue with Google\n - button \"Continue with Apple\" [ref=e37] [cursor=pointer]:\n - generic [ref=e39]:\n - img [ref=e41]\n - paragraph [ref=e44]: Continue with Apple\n - paragraph [ref=e48]: or\n - generic [ref=e51]:\n - generic: Email or username\n - textbox \"Email or username\" [ref=e52]: aidentothe\n - textbox\n - button \"Continue\" [ref=e153] [cursor=pointer]:\n - paragraph [ref=e157]: Continue\n - paragraph [ref=e60]:\n - text: By continuing, you agree to our\n - link \"Terms of Service\" [ref=e61] [cursor=pointer]:\n - /url: https://x.com/tos\n - text: \",\"\n - link \"Privacy Policy\" [ref=e62] [cursor=pointer]:\n - /url: https://x.com/privacy\n - text: and\n - link \"Cookie Use\" [ref=e63] [cursor=pointer]:\n - /url: https://help.x.com/rules-and-policies/twitter-cookies\n - text: .\n - img [ref=e65]\n - navigation \"Footer\" [ref=e69]:\n - generic [ref=e70]:\n - link \"About\" [ref=e71] [cursor=pointer]:\n - /url: https://about.x.com\n - generic [ref=e72]: ·\n - generic [ref=e73]:\n - link \"Get App\" [ref=e74] [cursor=pointer]:\n - /url: https://help.x.com/using-x/download-the-x-app\n - generic [ref=e75]: ·\n - generic [ref=e76]:\n - link \"Grok\" [ref=e77] [cursor=pointer]:\n - /url: https://grok.com/\n - generic [ref=e78]: ·\n - generic [ref=e79]:\n - link \"Help\" [ref=e80] [cursor=pointer]:\n - /url: https://help.x.com\n - generic [ref=e81]: ·\n - generic [ref=e82]:\n - link \"Terms\" [ref=e83] [cursor=pointer]:\n - /url: https://x.com/tos\n - generic [ref=e84]: ·\n - generic [ref=e85]:\n - link \"Privacy\" [ref=e86] [cursor=pointer]:\n - /url: https://x.com/privacy\n - generic [ref=e87]: ·\n - generic [ref=e88]:\n - link \"Cookies\" [ref=e89] [cursor=pointer]:\n - /url: https://support.x.com/articles/20170514\n - generic [ref=e90]: ·\n - generic [ref=e91]:\n - link \"Careers\" [ref=e92] [cursor=pointer]:\n - /url: https://careers.x.com\n - generic [ref=e93]: ·\n - generic [ref=e94]:\n - link \"Ads & Business\" [ref=e95] [cursor=pointer]:\n - /url: https://business.x.com/advertising?ref=gl-tw-tw-twitter-advertise\n - generic [ref=e96]: ·\n - generic [ref=e97]:\n - link \"Developers\" [ref=e98] [cursor=pointer]:\n - /url: https://developer.x.com\n - generic [ref=e99]: ·\n - generic [ref=e100]:\n - link \"News\" [ref=e101] [cursor=pointer]:\n - /url: https://x.com/i/jf/stories/home\n - generic [ref=e102]: ·\n - generic [ref=e103]:\n - link \"US TIDA\" [ref=e104] [cursor=pointer]:\n - /url: https://help.x.com/rules-and-policies/us-tida\n - generic [ref=e105]: ·\n - generic [ref=e106]:\n - link \"Accessibility\" [ref=e107] [cursor=pointer]:\n - /url: https://help.x.com/resources/accessibility\n - generic [ref=e108]: ·\n - generic [ref=e109]: © 2026 X Corp.\n```"
}
]Interesting — the page navigated to a passkey auth flow and there's already a username "aidentothe" prefilled in the email/username field. It looks like the browser may have a saved session or is partway through a flow. There's a "Sign in with passkey" dialog with a loading spinner. Let me wait a moment and see what happens with the passkey flow. Actually, it seems like the browser is trying to use a passkey. This may require user interaction (biometric/device approval). Let me take another snapshot to see the current state. Actually, I notice the username field shows "aidentothe" — this might be autofilled. Let me wait and check the state again. The passkey dialog has a loading spinner, so something is in progress. Let me wait a bit and re-check.
The page jumped to a passkey sign-in flow and there's a username aidentothe already filled in. A passkey prompt is loading. Let me wait a moment and check the state.
{
"command": "sleep 4"
}Output from command in shell 35e25e: Exit code: 0
{
"server_name": "playwright",
"tool_name": "browser_snapshot",
"arguments": {}
}[
{
"type": "text",
"text": "### Page\n- Page URL: https://x.com/i/jf/onboarding/web#/s/login_enter_password/r-knpof\n- Page Title: X - The Everything App / X\n- Console: 5 errors, 1 warnings\n### Snapshot\n```yaml\n- generic [ref=e2]:\n - generic:\n - generic:\n - generic:\n - generic:\n - generic:\n - generic:\n - dialog:\n - group [ref=e112]:\n - dialog [ref=e114]:\n - generic [ref=e160]:\n - generic [ref=e161]:\n - button [ref=e164] [cursor=pointer]:\n - img [ref=e167]\n - generic [ref=e169]:\n - img [ref=e174]\n - generic [ref=e179]:\n - paragraph [ref=e182]: Login\n - generic [ref=e184]:\n - generic [ref=e186]:\n - generic: Username\n - textbox \"Username\" [ref=e187]: aidentothe\n - generic [ref=e190]:\n - generic: Password\n - textbox \"Password Hide password\" [active] [ref=e191]: Lilychichi12\n - button \"Hide password\" [ref=e192] [cursor=pointer]:\n - img [ref=e193]\n - button \"Forgot password?\" [ref=e195] [cursor=pointer]:\n - paragraph [ref=e197]: Forgot password?\n - generic [ref=e199]:\n - button \"Continue\" [ref=e200] [cursor=pointer]:\n - paragraph [ref=e204]: Continue\n - generic [ref=e205]:\n - generic [ref=e206]:\n - img [ref=e208]\n - paragraph [ref=e211]: The password you entered is incorrect.\n - paragraph [ref=e214]:\n - text: By continuing, you agree to our\n - link \"Terms of Service\" [ref=e215] [cursor=pointer]:\n - /url: https://x.com/tos\n - text: \",\"\n - link \"Privacy Policy\" [ref=e216] [cursor=pointer]:\n - /url: https://x.com/privacy\n - text: and\n - link \"Cookie Use\" [ref=e217] [cursor=pointer]:\n - /url: https://help.x.com/rules-and-policies/twitter-cookies\n - text: .\n - main [ref=e4]:\n - generic [ref=e6]:\n - generic [ref=e7]:\n - generic [ref=e9]:\n - heading \"Happening now.\" [level=1] [ref=e10]\n - generic [ref=e14]:\n - generic [ref=e15]:\n - button \"Continue with phone\" [ref=e16] [cursor=pointer]:\n - generic [ref=e18]:\n - img [ref=e20]\n - paragraph [ref=e23]: Continue with phone\n - iframe [ref=e111]:\n - button \"Continue with Google. Opens in new tab\" [ref=f2e3] [cursor=pointer]:\n - generic [ref=f2e5]:\n - img [ref=f2e7]\n - generic [ref=f2e14]: Continue with Google\n - button \"Continue with Apple\" [ref=e37] [cursor=pointer]:\n - generic [ref=e39]:\n - img [ref=e41]\n - paragraph [ref=e44]: Continue with Apple\n - paragraph [ref=e48]: or\n - generic [ref=e51]:\n - generic: Email or username\n - textbox \"Email or username\" [ref=e52]: aidentothe\n - textbox\n - button \"Continue\" [ref=e153] [cursor=pointer]:\n - paragraph [ref=e157]: Continue\n - paragraph [ref=e60]:\n - text: By continuing, you agree to our\n - link \"Terms of Service\" [ref=e61] [cursor=pointer]:\n - /url: https://x.com/tos\n - text: \",\"\n - link \"Privacy Policy\" [ref=e62] [cursor=pointer]:\n - /url: https://x.com/privacy\n - text: and\n - link \"Cookie Use\" [ref=e63] [cursor=pointer]:\n - /url: https://help.x.com/rules-and-policies/twitter-cookies\n - text: .\n - img [ref=e65]\n - navigation \"Footer\" [ref=e69]:\n - generic [ref=e70]:\n - link \"About\" [ref=e71] [cursor=pointer]:\n - /url: https://about.x.com\n - generic [ref=e72]: ·\n - generic [ref=e73]:\n - link \"Get App\" [ref=e74] [cursor=pointer]:\n - /url: https://help.x.com/using-x/download-the-x-app\n - generic [ref=e75]: ·\n - generic [ref=e76]:\n - link \"Grok\" [ref=e77] [cursor=pointer]:\n - /url: https://grok.com/\n - generic [ref=e78]: ·\n - generic [ref=e79]:\n - link \"Help\" [ref=e80] [cursor=pointer]:\n - /url: https://help.x.com\n - generic [ref=e81]: ·\n - generic [ref=e82]:\n - link \"Terms\" [ref=e83] [cursor=pointer]:\n - /url: https://x.com/tos\n - generic [ref=e84]: ·\n - generic [ref=e85]:\n - link \"Privacy\" [ref=e86] [cursor=pointer]:\n - /url: https://x.com/privacy\n - generic [ref=e87]: ·\n - generic [ref=e88]:\n - link \"Cookies\" [ref=e89] [cursor=pointer]:\n - /url: https://support.x.com/articles/20170514\n - generic [ref=e90]: ·\n - generic [ref=e91]:\n - link \"Careers\" [ref=e92] [cursor=pointer]:\n - /url: https://careers.x.com\n - generic [ref=e93]: ·\n - generic [ref=e94]:\n - link \"Ads & Business\" [ref=e95] [cursor=pointer]:\n - /url: https://business.x.com/advertising?ref=gl-tw-tw-twitter-advertise\n - generic [ref=e96]: ·\n - generic [ref=e97]:\n - link \"Developers\" [ref=e98] [cursor=pointer]:\n - /url: https://developer.x.com\n - generic [ref=e99]: ·\n - generic [ref=e100]:\n - link \"News\" [ref=e101] [cursor=pointer]:\n - /url: https://x.com/i/jf/stories/home\n - generic [ref=e102]: ·\n - generic [ref=e103]:\n - link \"US TIDA\" [ref=e104] [cursor=pointer]:\n - /url: https://help.x.com/rules-and-policies/us-tida\n - generic [ref=e105]: ·\n - generic [ref=e106]:\n - link \"Accessibility\" [ref=e107] [cursor=pointer]:\n - /url: https://help.x.com/resources/accessibility\n - generic [ref=e108]: ·\n - generic [ref=e109]: © 2026 X Corp.\n```"
}
]I need to stop and flag something important. The browser autofilled credentials — username `aidentothe` and password `Lilychichi12` — and the page is now showing "The password you entered is incorrect." This is a saved/autofilled credential from the user's browser. I should NOT proceed with trying to use these credentials, and I should not be logging them or repeating them. Actually, they're already displayed in the snapshot output. But I should be careful here. Wait — this is the user's own browser with their own saved credentials, and they asked me to open a Chrome window to get their cookies. The autofill is from their own password manager. The password is showing as incorrect though. I should not try to log in with potentially incorrect credentials, and I definitely should not be handling the user's password. The right approach here is to let the user handle the login themselves in the browser, then I can extract the cookies once they're logged in. Let me tell the user the situation: the browser has autofilled credentials but the password is showing as incorrect. They need to log in themselves. I'll wait for them to complete login, then extract cookies. Actually, I should be careful about exposing the password in my response. The snapshot already showed it, but I shouldn't repeat it. Let me just tell the user to log in.
The browser autofilled saved credentials, but X is showing "The password you entered is incorrect." I can't log in for you — and I shouldn't be handling your password anyway.
Please log in yourself in the browser window that's now open. You can:
Once you're fully logged in and see your X home feed, tell me and I'll extract the auth_token and ct0 cookies.
If the autofilled password is wrong, you may want to check your password manager or update it after we're done.
im logged in
{
"server_name": "playwright",
"tool_name": "browser_evaluate",
"arguments": {
"function": "() => { const get = name => { const m = document.cookie.match(new RegExp('(?:^|; )' + name + '=([^;]*)')); return m ? decodeURIComponent(m[1]) : null; }; return JSON.stringify({ auth_token: get('auth_token'), ct0: get('ct0') }); }"
}
}[
{
"type": "text",
"text": "### Result\n\"{\\\"auth_token\\\":null,\\\"ct0\\\":\\\"ab8d49218dc01d34251fafd9ea15a7588819be2cadc9910a5613d439c180b69bc7026f3304a57c67f4eba5cda400f8278921be83fc87daf6344a4b85e6b3140197c85429d9bee4d21ff1053e373ba918\\\"}\"\n### Ran Playwright code\n```js\nawait page.evaluate('() => { const get = name => { const m = document.cookie.match(new RegExp(\\'(?:^|; )\\' + name + \\'=([^;]*)\\')); return m ? decodeURIComponent(m[1]) : null; }; return JSON.stringify({ auth_token: get(\\'auth_token\\'), ct0: get(\\'ct0\\') }); }');\n```\n### Page\n- Page URL: https://x.com/home\n- Page Title: Home / X\n- Console: 3 errors, 0 warnings\n### Events\n- New console entries: .playwright-mcp/console-2026-06-20T21-30-28-567Z.log#L8-L11"
}
]Got ct0 but auth_token is null. auth_token is set as HttpOnly so document.cookie can't see it. Let me try to get it via the Playwright context cookies API.
auth_token is HttpOnly so document.cookie can't see it. Let me grab it from the browser context directly.
{
"server_name": "playwright",
"tool_name": "browser_evaluate",
"arguments": {
"function": "async () => { const cookies = await window.fetch('https://x.com', { method: 'GET', credentials: 'include' }).then(r => r.headers); return 'check'; }"
}
}[
{
"type": "text",
"text": "### Result\n\"check\"\n### Ran Playwright code\n```js\nawait page.evaluate('async () => { const cookies = await window.fetch(\\'https://x.com\\', { method: \\'GET\\', credentials: \\'include\\' }).then(r => r.headers); return \\'check\\'; }');\n```"
}
]I need to get the HttpOnly cookies. Playwright MCP doesn't seem to have a direct cookie API exposed. Let me check the tools list for something like browser_cookies. Let me look at the truncated tools list.
{
"command": "grep -i \"cookie\\|storage\\|session\" /var/folders/d9/mzhyr7lx6xbf236j5ll9v_300000gn/T/devin-overflows-501/8e54642b/content.txt | head -30"
}Output from command in shell 9c07f1:
"description": "Return all console messages since the beginning of the session, not just since the last navigation. Defaults to false.",
Exit code: 0No cookie tool in the MCP. Let me think about another way to extract auth_token. The HttpOnly cookie can't be read by JS. But maybe I can use a fetch to an X endpoint that returns the user info, and the auth_token will be sent automatically. But I need the actual auth_token value to put into twscrape. Alternative: I can use the browser to make the API call directly from the page context, since the cookies will be sent automatically. Let me just call the GraphQL endpoint from the page to get @sdand's following list directly. That avoids needing to extract cookies at all. Actually, the cleanest approach: use the browser's authenticated session to call the GraphQL Following endpoint directly. Let me do that. First, I need to get @sdand's user_id. I'll call UserByScreenName. The GraphQL endpoint needs: - URL: https://x.com/i/api/graphql/681MIj51w00Aj6dY0GXnHw/UserByScreenName?variables=...&features=... - Headers: authorization (bearer), x-csrf-token (ct0), x-twitter-auth-type, x-twitter-active-user - Cookies: sent automatically Let me make the fetch from the page. The ct0 I already have. The bearer token is the hardcoded one in account.py. Actually, even simpler - the page already has all the right headers set up by X's own JS. I can just call fetch() with the right URL and the browser will attach everything. But I need to add the x-csrf-token header (= ct0) and x-twitter-auth-type. Actually X's own fetch wrapper adds these, but a raw fetch won't. Let me add them manually. Let me get the ct0 first (I have it), then make the calls.
I have ct0 but auth_token is HttpOnly. Rather than extracting it, I'll use the browser's authenticated session directly to call X's GraphQL API — the cookies get attached automatically. Let me fetch @sdand's user ID first, then their following.
{
"server_name": "playwright",
"tool_name": "browser_evaluate",
"arguments": {
"function": "async () => { const ct0 = document.cookie.match(/ct0=([^;]+)/)[1]; const variables = encodeURIComponent(JSON.stringify({screen_name:'sdand',withSafetyModeUserFields:true})); const features = encodeURIComponent(JSON.stringify({highlights_tweets_tab_ui_enabled:true,hidden_profile_likes_enabled:true,creator_subscriptions_tweet_preview_api_enabled:true,hidden_profile_subscriptions_enabled:true,subscriptions_verification_info_verified_since_enabled:true,subscriptions_verification_info_is_identity_verified_enabled:false,responsive_web_twitter_article_notes_tab_enabled:false,subscriptions_feature_can_gift_premium:false,profile_label_improvements_pcf_label_in_post_enabled:false})); const url = `https://x.com/i/api/graphql/681MIj51w00Aj6dY0GXnHw/UserByScreenName?variables=${variables}&features=${features}`; const r = await fetch(url, {headers:{'authorization':'Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA','x-csrf-token':ct0,'x-twitter-auth-type':'OAuth2Session','x-twitter-active-user':'yes','content-type':'application/json'}}); const j = await r.json(); const u = j.data?.user?.result; return JSON.stringify({status:r.status, id:u?.rest_id, screen_name:u?.legacy?.screen_name, name:u?.legacy?.name}); }"
}
}[
{
"type": "text",
"text": "### Result\n\"{\\\"status\\\":200,\\\"id\\\":\\\"800117847774986240\\\"}\"\n### Ran Playwright code\n```js\nawait page.evaluate('async () => { const ct0 = document.cookie.match(/ct0=([^;]+)/)[1]; const variables = encodeURIComponent(JSON.stringify({screen_name:\\'sdand\\',withSafetyModeUserFields:true})); const features = encodeURIComponent(JSON.stringify({highlights_tweets_tab_ui_enabled:true,hidden_profile_likes_enabled:true,creator_subscriptions_tweet_preview_api_enabled:true,hidden_profile_subscriptions_enabled:true,subscriptions_verification_info_verified_since_enabled:true,subscriptions_verification_info_is_identity_verified_enabled:false,responsive_web_twitter_article_notes_tab_enabled:false,subscriptions_feature_can_gift_premium:false,profile_label_improvements_pcf_label_in_post_enabled:false})); const url = `https://x.com/i/api/graphql/681MIj51w00Aj6dY0GXnHw/UserByScreenName?variables=${variables}&features=${features}`; const r = await fetch(url, {headers:{\\'authorization\\':\\'Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA\\',\\'x-csrf-token\\':ct0,\\'x-twitter-auth-type\\':\\'OAuth2Session\\',\\'x-twitter-active-user\\':\\'yes\\',\\'content-type\\':\\'application/json\\'}}); const j = await r.json(); const u = j.data?.user?.result; return JSON.stringify({status:r.status, id:u?.rest_id, screen_name:u?.legacy?.screen_name, name:u?.legacy?.name}); }');\n```"
}
]Got @sdand's user ID: 800117847774986240. Now let me paginate through their following list.
{
"server_name": "playwright",
"tool_name": "browser_evaluate",
"arguments": {
"function": "async () => { const ct0 = document.cookie.match(/ct0=([^;]+)/)[1]; const uid = '800117847774986240'; const features = encodeURIComponent(JSON.stringify({responsive_web_twitter_article_notes_tab_enabled:false})); let cursor = null; const all = []; for (let page = 0; page < 50; page++) { const vars = {userId:uid,count:40,includePromotedContent:false}; if (cursor) vars.cursor = cursor; const variables = encodeURIComponent(JSON.stringify(vars)); const url = `https://x.com/i/api/graphql/OLm4oHZBfqWx8jbcEhWoFw/Following?variables=${variables}&features=${features}`; const r = await fetch(url, {headers:{'authorization':'Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA','x-csrf-token':ct0,'x-twitter-auth-type':'OAuth2Session','x-twitter-active-user':'yes','content-type':'application/json'}}); if (!r.ok) { all.push({error:r.status, text:await r.text()}); break; } const j = await r.json(); const entries = j.data?.user?.result?.timeline?.timeline?.instructions?.flatMap(i=>i.entries||[]) || []; let nextCursor = null; for (const e of entries) { if (e.content?.cursorType==='Bottom') nextCursor = e.content?.value; if (e.content?.entryType==='TimelineTimelineItem' && e.content?.itemContent?.user_results?.result) { const u = e.content.itemContent.user_results.result; all.push({id:u.rest_id, screen_name:u.legacy?.screen_name, name:u.legacy?.name}); } } if (!nextCursor || nextCursor === cursor) break; cursor = nextCursor; } return JSON.stringify({count:all.length, users:all}); }",
"filename": "/Users/root1/Downloads/utilities/sdand_following.json"
}
}[
{
"type": "text",
"text": "### Result\n- [Evaluation result](./sdand_following.json)\n### Ran Playwright code\n```js\nawait page.evaluate('async () => { const ct0 = document.cookie.match(/ct0=([^;]+)/)[1]; const uid = \\'800117847774986240\\'; const features = encodeURIComponent(JSON.stringify({responsive_web_twitter_article_notes_tab_enabled:false})); let cursor = null; const all = []; for (let page = 0; page < 50; page++) { const vars = {userId:uid,count:40,includePromotedContent:false}; if (cursor) vars.cursor = cursor; const variables = encodeURIComponent(JSON.stringify(vars)); const url = `https://x.com/i/api/graphql/OLm4oHZBfqWx8jbcEhWoFw/Following?variables=${variables}&features=${features}`; const r = await fetch(url, {headers:{\\'authorization\\':\\'Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA\\',\\'x-csrf-token\\':ct0,\\'x-twitter-auth-type\\':\\'OAuth2Session\\',\\'x-twitter-active-user\\':\\'yes\\',\\'content-type\\':\\'application/json\\'}}); if (!r.ok) { all.push({error:r.status, text:await r.text()}); break; } const j = await r.json(); const entries = j.data?.user?.result?.timeline?.timeline?.instructions?.flatMap(i=>i.entries||[]) || []; let nextCursor = null; for (const e of entries) { if (e.content?.cursorType===\\'Bottom\\') nextCursor = e.content?.value; if (e.content?.entryType===\\'TimelineTimelineItem\\' && e.content?.itemContent?.user_results?.result) { const u = e.content.itemContent.user_results.result; all.push({id:u.rest_id, screen_name:u.legacy?.screen_name, name:u.legacy?.name}); } } if (!nextCursor || nextCursor === cursor) break; cursor = nextCursor; } return JSON.stringify({count:all.length, users:all}); }');\n```"
}
]{
"file_path": "/Users/root1/Downloads/utilities/sdand_following.json"
}<file-view path="/Users/root1/Downloads/utilities/sdand_following.json" start_line="1" end_line="1" total_lines="1">
1|"{\"count\":1613,\"users\":[{\"id\":\"1460283925\"},{\"id\":\"1322297149887979520\"},{\"id\":\"1237544283743432704\"},{\"id\":\"1204841211909038080\"},{\"id\":\"1281068231122247681\"},{\"id\":\"1303125102196711424\"},{\"id\":\"3315205122\"},{\"id\":\"107186366\"},{\"id\":\"998575387067142144\"},{\"id\":\"1279292401123647494\"},{\"id\":\"1069108975805714432\"},{\"id\":\"1241210351951261696\"},{\"id\":\"607528502\"},{\"id\":\"837401375780765702\"},{\"id\":\"53836928\"},{\"id\":\"1105678196115152897\"},{\"id\":\"935471524307353601\"},{\"id\":\"760281873364885504\"},{\"id\":\"1320156402996944901\"},{\"id\":\"145697000\"},{\"id\":\"1551409937133158400\"},{\"id\":\"712380305848438784\"},{\"id\":\"2051538470901039104\"},{\"id\":\"1333850130257207296\"},{\"id\":\"1070093713387462656\"},{\"id\":\"1152652277833859074\"},{\"id\":\"2053001548335124480\"},{\"id\":\"2984232589\"},{\"id\":\"14245415\"},{\"id\":\"2541675540\"},{\"id\":\"347701349\"},{\"id\":\"12831\"},{\"id\":\"914848333642043393\"},{\"id\":\"2047459214650408960\"},{\"id\":\"50949650\"},{\"id\":\"1717036496115441664\"},{\"id\":\"31393537\"},{\"id\":\"899564621631299585\"},{\"id\":\"1883916449607045120\"},{\"id\":\"1321434544013250564\"},{\"id\":\"2789704130\"},{\"id\":\"1711897712646377472\"},{\"id\":\"1977982510299770880\"},{\"id\":\"1274990797306265600\"},{\"id\":\"1955153871145234432\"},{\"id\":\"1289123581\"},{\"id\":\"46591277\"},{\"id\":\"103012173\"},{\"id\":\"1273327266957688832\"},{\"id\":\"1648003499689385987\"},{\"id\":\"2753075366\"},{\"id\":\"2513070418\"},{\"id\":\"741127492614950913\"},{\"id\":\"1342724020127498240\"},{\"id\":\"127236298\"},{\"id\":\"996493134136729600\"},{\"id\":\"930159122535866369\"},{\"id\":\"4330852526\"},{\"id\":\"1609093208637677569\"},{\"id\":\"3184479582\"},{\"id\":\"1867644701798440960\"},{\"id\":\"1371600225958305793\"},{\"id\":\"1797501310851231744\"},{\"id\":\"573817445\"},{\"id\":\"56166281\"},{\"id\":\"1482049781719388165\"},{\"id\":\"20187433\"},{\"id\":\"935028743130... [truncated 44556 chars]
</file-view>Got 1613 users but only IDs, no screen_names. The Following endpoint with this features config may not return legacy data. Let me check - the entries have user_results.result but maybe the legacy field is empty or I need different features. Let me re-run with the full feature set from twscrape's GQL_FEATURES to get screen_name and name. Actually, looking at the output, the users only have `id` - no screen_name or name. This means the legacy field wasn't populated. The Following endpoint might need more features. Let me use the full GQL_FEATURES dict from api.py.
Got 1613 following but only IDs — the legacy block (with screen_name/name) wasn't returned with the minimal features. Let me re-fetch with the full feature set twscrape uses.
{
"server_name": "playwright",
"tool_name": "browser_evaluate",
"arguments": {
"filename": "/Users/root1/Downloads/utilities/sdand_following2.json",
"function": "async () => { const ct0 = document.cookie.match(/ct0=([^;]+)/)[1]; const uid = '800117847774986240'; const FT = {articles_preview_enabled:false,c9s_tweet_anatomy_moderator_badge_enabled:true,communities_web_enable_tweet_community_results_fetch:true,creator_subscriptions_quote_tweet_preview_enabled:false,creator_subscriptions_tweet_preview_api_enabled:true,freedom_of_speech_not_reach_fetch_enabled:true,graphql_is_translatable_rweb_tweet_is_translatable_enabled:true,longform_notetweets_consumption_enabled:true,longform_notetweets_inline_media_enabled:true,longform_notetweets_rich_text_read_enabled:true,responsive_web_edit_tweet_api_enabled:true,responsive_web_enhance_cards_enabled:false,responsive_web_graphql_exclude_directive_enabled:true,responsive_web_graphql_skip_user_profile_image_extensions_enabled:false,responsive_web_grok_community_note_auto_translation_is_enabled:false,responsive_web_graphql_timeline_navigation_enabled:true,responsive_web_grok_imagine_annotation_enabled:false,responsive_web_media_download_video_enabled:false,responsive_web_profile_redirect_enabled:true,responsive_web_twitter_article_tweet_consumption_enabled:true,rweb_tipjar_consumption_enabled:true,rweb_video_timestamps_enabled:true,standardized_nudges_misinfo:true,tweet_awards_web_tipping_enabled:false,tweet_with_visibility_results_prefer_gql_limited_actions_policy_enabled:true,tweet_with_visibility_results_prefer_gql_media_interstitial_enabled:false,tweetypie_unmention_optimization_enabled:true,verified_phone_label_enabled:false,view_counts_everywhere_api_enabled:true,responsive_web_grok_analyze_button_fetch_trends_enabled:false,premium_content_api_read_enabled:false,profile_label_improvements_pcf_label_in_post_enabled:false,responsive_web_grok_share_attachment_enabled:false,responsive_web_grok_analyze_post_followups_enabled:false,responsive_web_grok_image_annotation_enabled:false,responsive_web_grok_analysis_button_from_backend:false,responsive_web_jetfuel_frame:false,rweb_video_screen_enabled:true,responsive_web_grok_show_grok_translated_post:true,responsive_web_twitter_article_notes_tab_enabled:false}; const features = encodeURIComponent(JSON.stringify(FT)); let cursor = null; const all = []; for (let page = 0; page < 50; page++) { const vars = {userId:uid,count:40,includePromotedContent:false}; if (cursor) vars.cursor = cursor; const variables = encodeURIComponent(JSON.stringify(vars)); const url = `https://x.com/i/api/graphql/OLm4oHZBfqWx8jbcEhWoFw/Following?variables=${variables}&features=${features}`; const r = await fetch(url, {headers:{'authorization':'Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA','x-csrf-token':ct0,'x-twitter-auth-type':'OAuth2Session','x-twitter-active-user':'yes','content-type':'application/json'}}); if (!r.ok) { all.push({error:r.status, text:(await r.text()).slice(0,200)}); break; } const j = await r.json(); const entries = j.data?.user?.result?.timeline?.timeline?.instructions?.flatMap(i=>i.entries||[]) || []; let nextCursor = null; for (const e of entries) { if (e.content?.cursorType==='Bottom') nextCursor = e.content?.value; if (e.content?.entryType==='TimelineTimelineItem' && e.content?.itemContent?.user_results?.result) { const u = e.content.itemContent.user_results.result; all.push({id:u.rest_id, screen_name:u.legacy?.screen_name, name:u.legacy?.name, verified:u.is_blue_verified||false, followers:u.legacy?.followers_count}); } } if (!nextCursor || nextCursor === cursor) break; cursor = nextCursor; } return JSON.stringify({count:all.length, users:all}); }"
}
}[
{
"type": "text",
"text": "### Result\n- [Evaluation result](./sdand_following2.json)\n### Ran Playwright code\n```js\nawait page.evaluate('async () => { const ct0 = document.cookie.match(/ct0=([^;]+)/)[1]; const uid = \\'800117847774986240\\'; const FT = {articles_preview_enabled:false,c9s_tweet_anatomy_moderator_badge_enabled:true,communities_web_enable_tweet_community_results_fetch:true,creator_subscriptions_quote_tweet_preview_enabled:false,creator_subscriptions_tweet_preview_api_enabled:true,freedom_of_speech_not_reach_fetch_enabled:true,graphql_is_translatable_rweb_tweet_is_translatable_enabled:true,longform_notetweets_consumption_enabled:true,longform_notetweets_inline_media_enabled:true,longform_notetweets_rich_text_read_enabled:true,responsive_web_edit_tweet_api_enabled:true,responsive_web_enhance_cards_enabled:false,responsive_web_graphql_exclude_directive_enabled:true,responsive_web_graphql_skip_user_profile_image_extensions_enabled:false,responsive_web_grok_community_note_auto_translation_is_enabled:false,responsive_web_graphql_timeline_navigation_enabled:true,responsive_web_grok_imagine_annotation_enabled:false,responsive_web_media_download_video_enabled:false,responsive_web_profile_redirect_enabled:true,responsive_web_twitter_article_tweet_consumption_enabled:true,rweb_tipjar_consumption_enabled:true,rweb_video_timestamps_enabled:true,standardized_nudges_misinfo:true,tweet_awards_web_tipping_enabled:false,tweet_with_visibility_results_prefer_gql_limited_actions_policy_enabled:true,tweet_with_visibility_results_prefer_gql_media_interstitial_enabled:false,tweetypie_unmention_optimization_enabled:true,verified_phone_label_enabled:false,view_counts_everywhere_api_enabled:true,responsive_web_grok_analyze_button_fetch_trends_enabled:false,premium_content_api_read_enabled:false,profile_label_improvements_pcf_label_in_post_enabled:false,responsive_web_grok_share_attachment_enabled:false,responsive_web_grok_analyze_post_followups_enabled:false,responsive_web_grok_image_annotation_enabled:false,responsive_web_grok_analysis_button_from_backend:false,responsive_web_jetfuel_frame:false,rweb_video_screen_enabled:true,responsive_web_grok_show_grok_translated_post:true,responsive_web_twitter_article_notes_tab_enabled:false}; const features = encodeURIComponent(JSON.stringify(FT)); let cursor = null; const all = []; for (let page = 0; page < 50; page++) { const vars = {userId:uid,count:40,includePromotedContent:false}; if (cursor) vars.cursor = cursor; const variables = encodeURIComponent(JSON.stringify(vars)); const url = `https://x.com/i/api/graphql/OLm4oHZBfqWx8jbcEhWoFw/Following?variables=${variables}&features=${features}`; const r = await fetch(url, {headers:{\\'authorization\\':\\'Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA\\',\\'x-csrf-token\\':ct0,\\'x-twitter-auth-type\\':\\'OAuth2Session\\',\\'x-twitter-active-user\\':\\'yes\\',\\'content-type\\':\\'application/json\\'}}); if (!r.ok) { all.push({error:r.status, text:(await r.text()).slice(0,200)}); break; } const j = await r.json(); const entries = j.data?.user?.result?.timeline?.timeline?.instructions?.flatMap(i=>i.entries||[]) || []; let nextCursor = null; for (const e of entries) { if (e.content?.cursorType===\\'Bottom\\') nextCursor = e.content?.value; if (e.content?.entryType===\\'TimelineTimelineItem\\' && e.content?.itemContent?.user_results?.result) { const u = e.content.itemContent.user_results.result; all.push({id:u.rest_id, screen_name:u.legacy?.screen_name, name:u.legacy?.name, verified:u.is_blue_verified||false, followers:u.legacy?.followers_count}); } } if (!nextCursor || nextCursor === cursor) break; cursor = nextCursor; } return JSON.stringify({count:all.length, users:all}); }');\n```"
}
]{
"command": "cd /Users/root1/Downloads/utilities && python3 -c \"\nimport json\nwith open('sdand_following2.json') as f:\n raw = f.read()\n# the file is a JSON string containing a JSON string (double-encoded)\ndata = json.loads(json.loads(raw))\nprint('Total following:', data['count'])\nusers = data['users']\nhas_name = sum(1 for u in users if u.get('screen_name'))\nprint('Users with screen_name:', has_name)\nprint()\nprint('First 30:')\nfor u in users[:30]:\n sn = u.get('screen_name') or '?'\n nm = u.get('name') or ''\n v = ' ✓' if u.get('verified') else ''\n print(f' @{sn:<25} {nm}{v}')\n\"\n"
}Output from command in shell c23c6a: Total following: 1613 Users with screen_name: 0 First 30: @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? @? @? @? ✓ @? ✓ @? ✓ @? ✓ @? ✓ @? @? ✓ @? Exit code: 0
The legacy block is not being returned. The Following endpoint returns user_results.result but without the legacy field. This might be because X has changed the API, or because we need different features, or the legacy data is gated behind some feature flag. Let me inspect the raw structure of one entry to see what fields are actually available.
The legacy block isn't being returned. Let me inspect the raw structure of one entry to see what fields are actually available.
{
"server_name": "playwright",
"tool_name": "browser_evaluate",
"arguments": {
"filename": "/Users/root1/Downloads/utilities/sdand_raw_sample.json",
"function": "async () => { const ct0 = document.cookie.match(/ct0=([^;]+)/)[1]; const uid = '800117847774986240'; const FT = {articles_preview_enabled:false,c9s_tweet_anatomy_moderator_badge_enabled:true,communities_web_enable_tweet_community_results_fetch:true,creator_subscriptions_quote_tweet_preview_enabled:false,creator_subscriptions_tweet_preview_api_enabled:true,freedom_of_speech_not_reach_fetch_enabled:true,graphql_is_translatable_rweb_tweet_is_translatable_enabled:true,longform_notetweets_consumption_enabled:true,longform_notetweets_inline_media_enabled:true,longform_notetweets_rich_text_read_enabled:true,responsive_web_edit_tweet_api_enabled:true,responsive_web_enhance_cards_enabled:false,responsive_web_graphql_exclude_directive_enabled:true,responsive_web_graphql_skip_user_profile_image_extensions_enabled:false,responsive_web_grok_community_note_auto_translation_is_enabled:false,responsive_web_graphql_timeline_navigation_enabled:true,responsive_web_grok_imagine_annotation_enabled:false,responsive_web_media_download_video_enabled:false,responsive_web_profile_redirect_enabled:true,responsive_web_twitter_article_tweet_consumption_enabled:true,rweb_tipjar_consumption_enabled:true,rweb_video_timestamps_enabled:true,standardized_nudges_misinfo:true,tweet_awards_web_tipping_enabled:false,tweet_with_visibility_results_prefer_gql_limited_actions_policy_enabled:true,tweet_with_visibility_results_prefer_gql_media_interstitial_enabled:false,tweetypie_unmention_optimization_enabled:true,verified_phone_label_enabled:false,view_counts_everywhere_api_enabled:true,responsive_web_grok_analyze_button_fetch_trends_enabled:false,premium_content_api_read_enabled:false,profile_label_improvements_pcf_label_in_post_enabled:false,responsive_web_grok_share_attachment_enabled:false,responsive_web_grok_analyze_post_followups_enabled:false,responsive_web_grok_image_annotation_enabled:false,responsive_web_grok_analysis_button_from_backend:false,responsive_web_jetfuel_frame:false,rweb_video_screen_enabled:true,responsive_web_grok_show_grok_translated_post:true,responsive_web_twitter_article_notes_tab_enabled:false}; const features = encodeURIComponent(JSON.stringify(FT)); const vars = {userId:uid,count:5,includePromotedContent:false}; const variables = encodeURIComponent(JSON.stringify(vars)); const url = `https://x.com/i/api/graphql/OLm4oHZBfqWx8jbcEhWoFw/Following?variables=${variables}&features=${features}`; const r = await fetch(url, {headers:{'authorization':'Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA','x-csrf-token':ct0,'x-twitter-auth-type':'OAuth2Session','x-twitter-active-user':'yes','content-type':'application/json'}}); const j = await r.json(); const entries = j.data?.user?.result?.timeline?.timeline?.instructions?.flatMap(i=>i.entries||[]) || []; const sample = entries.filter(e=>e.content?.entryType==='TimelineTimelineItem').slice(0,2); return JSON.stringify(sample, null, 2); }"
}
}[
{
"type": "text",
"text": "### Result\n- [Evaluation result](./sdand_raw_sample.json)\n### Ran Playwright code\n```js\nawait page.evaluate('async () => { const ct0 = document.cookie.match(/ct0=([^;]+)/)[1]; const uid = \\'800117847774986240\\'; const FT = {articles_preview_enabled:false,c9s_tweet_anatomy_moderator_badge_enabled:true,communities_web_enable_tweet_community_results_fetch:true,creator_subscriptions_quote_tweet_preview_enabled:false,creator_subscriptions_tweet_preview_api_enabled:true,freedom_of_speech_not_reach_fetch_enabled:true,graphql_is_translatable_rweb_tweet_is_translatable_enabled:true,longform_notetweets_consumption_enabled:true,longform_notetweets_inline_media_enabled:true,longform_notetweets_rich_text_read_enabled:true,responsive_web_edit_tweet_api_enabled:true,responsive_web_enhance_cards_enabled:false,responsive_web_graphql_exclude_directive_enabled:true,responsive_web_graphql_skip_user_profile_image_extensions_enabled:false,responsive_web_grok_community_note_auto_translation_is_enabled:false,responsive_web_graphql_timeline_navigation_enabled:true,responsive_web_grok_imagine_annotation_enabled:false,responsive_web_media_download_video_enabled:false,responsive_web_profile_redirect_enabled:true,responsive_web_twitter_article_tweet_consumption_enabled:true,rweb_tipjar_consumption_enabled:true,rweb_video_timestamps_enabled:true,standardized_nudges_misinfo:true,tweet_awards_web_tipping_enabled:false,tweet_with_visibility_results_prefer_gql_limited_actions_policy_enabled:true,tweet_with_visibility_results_prefer_gql_media_interstitial_enabled:false,tweetypie_unmention_optimization_enabled:true,verified_phone_label_enabled:false,view_counts_everywhere_api_enabled:true,responsive_web_grok_analyze_button_fetch_trends_enabled:false,premium_content_api_read_enabled:false,profile_label_improvements_pcf_label_in_post_enabled:false,responsive_web_grok_share_attachment_enabled:false,responsive_web_grok_analyze_post_followups_enabled:false,responsive_web_grok_image_annotation_enabled:false,responsive_web_grok_analysis_button_from_backend:false,responsive_web_jetfuel_frame:false,rweb_video_screen_enabled:true,responsive_web_grok_show_grok_translated_post:true,responsive_web_twitter_article_notes_tab_enabled:false}; const features = encodeURIComponent(JSON.stringify(FT)); const vars = {userId:uid,count:5,includePromotedContent:false}; const variables = encodeURIComponent(JSON.stringify(vars)); const url = `https://x.com/i/api/graphql/OLm4oHZBfqWx8jbcEhWoFw/Following?variables=${variables}&features=${features}`; const r = await fetch(url, {headers:{\\'authorization\\':\\'Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA\\',\\'x-csrf-token\\':ct0,\\'x-twitter-auth-type\\':\\'OAuth2Session\\',\\'x-twitter-active-user\\':\\'yes\\',\\'content-type\\':\\'application/json\\'}}); const j = await r.json(); const entries = j.data?.user?.result?.timeline?.timeline?.instructions?.flatMap(i=>i.entries||[]) || []; const sample = entries.filter(e=>e.content?.entryType===\\'TimelineTimelineItem\\').slice(0,2); return JSON.stringify(sample, null, 2); }');\n```"
}
]{
"file_path": "/Users/root1/Downloads/utilities/sdand_raw_sample.json"
}<file-view path="/Users/root1/Downloads/utilities/sdand_raw_sample.json" start_line="1" end_line="1" total_lines="1">
1|"[\n {\n \"content\": {\n \"__typename\": \"TimelineTimelineItem\",\n \"clientEventInfo\": {\n \"component\": \"FollowingSgs\",\n \"element\": \"user\"\n },\n \"entryType\": \"TimelineTimelineItem\",\n \"itemContent\": {\n \"__typename\": \"TimelineUser\",\n \"itemType\": \"TimelineUser\",\n \"userDisplayType\": \"User\",\n \"user_results\": {\n \"result\": {\n \"__typename\": \"User\",\n \"affiliates_highlighted_label\": {},\n \"avatar\": {\n \"image_url\": \"https://pbs.twimg.com/profile_images/1918970926668054530/fy-ZsgJ7_normal.jpg\"\n },\n \"core\": {\n \"created_at\": \"Sun May 26 17:59:01 +0000 2013\",\n \"name\": \"roon\",\n \"screen_name\": \"tszzl\"\n },\n \"dm_permissions\": {\n \"can_dm\": true\n },\n \"follow_request_sent\": false,\n \"has_graduated_access\": true,\n \"id\": \"VXNlcjoxNDYwMjgzOTI1\",\n \"is_blue_verified\": true,\n \"legacy\": {\n \"default_profile\": false,\n \"default_profile_image\": false,\n \"description\": \"fellow creators the creator seeks\",\n \"entities\": {\n \"description\": {},\n \"url\": {\n \"urls\": [\n {\n \"display_url\": \"roonscape.ai\",\n \"expanded_url\": \"http://roonscape.ai\",\n \"indices\": [\n 0,\n 23\n ],\n \"url\": \"https://t.co/3Q7RKYCnB4\"\n }\n ]\n }\n },\n \"fast_followers_count\": 0,\n \"favourites_count\": 1067720,\n \"follow_request_sent\": false,\n ... [truncated 6566 chars]
</file-view>The data IS there - it's under `core.screen_name` and `core.name`, not `legacy.screen_name` and `legacy.name`. The new X API structure moved screen_name and name into `core` instead of `legacy`. Let me re-fetch with the correct path. Also `rest_id` is now `id` but it's base64 encoded ("VXNlcjoxNDYwMjgzOTI1" = "User:1460283925"). Let me use `core` for name/screen_name and decode the id, or just use the legacy.followers_count etc. Let me re-run the full pagination with the correct field paths.
The data is there — it's under core.screen_name / core.name instead of legacy.screen_name. The id field is now base64-encoded (VXNlcjoxNDYwMjgzOTI1 = User:1460283925). Let me re-fetch with the correct paths.
{
"server_name": "playwright",
"tool_name": "browser_evaluate",
"arguments": {
"filename": "/Users/root1/Downloads/utilities/sdand_following_final.json",
"function": "async () => { const ct0 = document.cookie.match(/ct0=([^;]+)/)[1]; const uid = '800117847774986240'; const FT = {articles_preview_enabled:false,c9s_tweet_anatomy_moderator_badge_enabled:true,communities_web_enable_tweet_community_results_fetch:true,creator_subscriptions_quote_tweet_preview_enabled:false,creator_subscriptions_tweet_preview_api_enabled:true,freedom_of_speech_not_reach_fetch_enabled:true,graphql_is_translatable_rweb_tweet_is_translatable_enabled:true,longform_notetweets_consumption_enabled:true,longform_notetweets_inline_media_enabled:true,longform_notetweets_rich_text_read_enabled:true,responsive_web_edit_tweet_api_enabled:true,responsive_web_enhance_cards_enabled:false,responsive_web_graphql_exclude_directive_enabled:true,responsive_web_graphql_skip_user_profile_image_extensions_enabled:false,responsive_web_grok_community_note_auto_translation_is_enabled:false,responsive_web_graphql_timeline_navigation_enabled:true,responsive_web_grok_imagine_annotation_enabled:false,responsive_web_media_download_video_enabled:false,responsive_web_profile_redirect_enabled:true,responsive_web_twitter_article_tweet_consumption_enabled:true,rweb_tipjar_consumption_enabled:true,rweb_video_timestamps_enabled:true,standardized_nudges_misinfo:true,tweet_awards_web_tipping_enabled:false,tweet_with_visibility_results_prefer_gql_limited_actions_policy_enabled:true,tweet_with_visibility_results_prefer_gql_media_interstitial_enabled:false,tweetypie_unmention_optimization_enabled:true,verified_phone_label_enabled:false,view_counts_everywhere_api_enabled:true,responsive_web_grok_analyze_button_fetch_trends_enabled:false,premium_content_api_read_enabled:false,profile_label_improvements_pcf_label_in_post_enabled:false,responsive_web_grok_share_attachment_enabled:false,responsive_web_grok_analyze_post_followups_enabled:false,responsive_web_grok_image_annotation_enabled:false,responsive_web_grok_analysis_button_from_backend:false,responsive_web_jetfuel_frame:false,rweb_video_screen_enabled:true,responsive_web_grok_show_grok_translated_post:true,responsive_web_twitter_article_notes_tab_enabled:false}; const features = encodeURIComponent(JSON.stringify(FT)); let cursor = null; const all = []; for (let page = 0; page < 50; page++) { const vars = {userId:uid,count:40,includePromotedContent:false}; if (cursor) vars.cursor = cursor; const variables = encodeURIComponent(JSON.stringify(vars)); const url = `https://x.com/i/api/graphql/OLm4oHZBfqWx8jbcEhWoFw/Following?variables=${variables}&features=${features}`; const r = await fetch(url, {headers:{'authorization':'Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA','x-csrf-token':ct0,'x-twitter-auth-type':'OAuth2Session','x-twitter-active-user':'yes','content-type':'application/json'}}); if (!r.ok) { all.push({error:r.status, text:(await r.text()).slice(0,200)}); break; } const j = await r.json(); const entries = j.data?.user?.result?.timeline?.timeline?.instructions?.flatMap(i=>i.entries||[]) || []; let nextCursor = null; for (const e of entries) { if (e.content?.cursorType==='Bottom') nextCursor = e.content?.value; if (e.content?.entryType==='TimelineTimelineItem' && e.content?.itemContent?.user_results?.result) { const u = e.content.itemContent.user_results.result; const core = u.core||{}; const leg = u.legacy||{}; let idNum = u.rest_id; if (!idNum && u.id) { try { idNum = atob(u.id).split(':')[1]; } catch(e){} } all.push({id:idNum, screen_name:core.screen_name, name:core.name, verified:u.is_blue_verified||false, followers:leg.followers_count, bio:leg.description}); } } if (!nextCursor || nextCursor === cursor) break; cursor = nextCursor; } return JSON.stringify({count:all.length, users:all}); }"
}
}[
{
"type": "text",
"text": "### Result\n- [Evaluation result](./sdand_following_final.json)\n### Ran Playwright code\n```js\nawait page.evaluate('async () => { const ct0 = document.cookie.match(/ct0=([^;]+)/)[1]; const uid = \\'800117847774986240\\'; const FT = {articles_preview_enabled:false,c9s_tweet_anatomy_moderator_badge_enabled:true,communities_web_enable_tweet_community_results_fetch:true,creator_subscriptions_quote_tweet_preview_enabled:false,creator_subscriptions_tweet_preview_api_enabled:true,freedom_of_speech_not_reach_fetch_enabled:true,graphql_is_translatable_rweb_tweet_is_translatable_enabled:true,longform_notetweets_consumption_enabled:true,longform_notetweets_inline_media_enabled:true,longform_notetweets_rich_text_read_enabled:true,responsive_web_edit_tweet_api_enabled:true,responsive_web_enhance_cards_enabled:false,responsive_web_graphql_exclude_directive_enabled:true,responsive_web_graphql_skip_user_profile_image_extensions_enabled:false,responsive_web_grok_community_note_auto_translation_is_enabled:false,responsive_web_graphql_timeline_navigation_enabled:true,responsive_web_grok_imagine_annotation_enabled:false,responsive_web_media_download_video_enabled:false,responsive_web_profile_redirect_enabled:true,responsive_web_twitter_article_tweet_consumption_enabled:true,rweb_tipjar_consumption_enabled:true,rweb_video_timestamps_enabled:true,standardized_nudges_misinfo:true,tweet_awards_web_tipping_enabled:false,tweet_with_visibility_results_prefer_gql_limited_actions_policy_enabled:true,tweet_with_visibility_results_prefer_gql_media_interstitial_enabled:false,tweetypie_unmention_optimization_enabled:true,verified_phone_label_enabled:false,view_counts_everywhere_api_enabled:true,responsive_web_grok_analyze_button_fetch_trends_enabled:false,premium_content_api_read_enabled:false,profile_label_improvements_pcf_label_in_post_enabled:false,responsive_web_grok_share_attachment_enabled:false,responsive_web_grok_analyze_post_followups_enabled:false,responsive_web_grok_image_annotation_enabled:false,responsive_web_grok_analysis_button_from_backend:false,responsive_web_jetfuel_frame:false,rweb_video_screen_enabled:true,responsive_web_grok_show_grok_translated_post:true,responsive_web_twitter_article_notes_tab_enabled:false}; const features = encodeURIComponent(JSON.stringify(FT)); let cursor = null; const all = []; for (let page = 0; page < 50; page++) { const vars = {userId:uid,count:40,includePromotedContent:false}; if (cursor) vars.cursor = cursor; const variables = encodeURIComponent(JSON.stringify(vars)); const url = `https://x.com/i/api/graphql/OLm4oHZBfqWx8jbcEhWoFw/Following?variables=${variables}&features=${features}`; const r = await fetch(url, {headers:{\\'authorization\\':\\'Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA\\',\\'x-csrf-token\\':ct0,\\'x-twitter-auth-type\\':\\'OAuth2Session\\',\\'x-twitter-active-user\\':\\'yes\\',\\'content-type\\':\\'application/json\\'}}); if (!r.ok) { all.push({error:r.status, text:(await r.text()).slice(0,200)}); break; } const j = await r.json(); const entries = j.data?.user?.result?.timeline?.timeline?.instructions?.flatMap(i=>i.entries||[]) || []; let nextCursor = null; for (const e of entries) { if (e.content?.cursorType===\\'Bottom\\') nextCursor = e.content?.value; if (e.content?.entryType===\\'TimelineTimelineItem\\' && e.content?.itemContent?.user_results?.result) { const u = e.content.itemContent.user_results.result; const core = u.core||{}; const leg = u.legacy||{}; let idNum = u.rest_id; if (!idNum && u.id) { try { idNum = atob(u.id).split(\\':\\')[1]; } catch(e){} } all.push({id:idNum, screen_name:core.screen_name, name:core.name, verified:u.is_blue_verified||false, followers:leg.followers_count, bio:leg.description}); } } if (!nextCursor || nextCursor === cursor) break; cursor = nextCursor; } return JSON.stringify({count:all.length, users:all}); }');\n```"
}
]{
"command": "cd /Users/root1/Downloads/utilities && python3 -c \"\nimport json\nwith open('sdand_following_final.json') as f:\n data = json.loads(json.loads(f.read()))\nusers = data['users']\nprint(f'@sdand follows {data[\\\"count\\\"]} accounts')\nprint()\nprint(f'{\\\"#\\\":>4} {\\\"handle\\\":<25} {\\\"name\\\":<30} {\\\"verified\\\":<8} {\\\"followers\\\":>10}')\nprint('-'*85)\nfor i, u in enumerate(users, 1):\n sn = u.get('screen_name') or '?'\n nm = (u.get('name') or '')[:28]\n v = 'yes' if u.get('verified') else ''\n fl = u.get('followers')\n fls = f'{fl:,}' if fl else ''\n print(f'{i:>4} @{sn:<24} {nm:<30} {v:<8} {fls:>10}')\n\" 2>&1 | head -80\n"
}Output from command in shell c54258: @sdand follows 1613 accounts # handle name verified followers ------------------------------------------------------------------------------------- 1 @tszzl roon yes 374,837 2 @0interestrates rahul yes 63,828 3 @powerbottomdad1 sucks yes 102,436 4 @ctjlewis Lewis 🇺🇸 yes 45,618 5 @goth600 @goth yes 140,666 6 @transmissions11 t11s yes 92,868 7 @deepfates 🎭 yes 61,967 8 @sharifshameem Sharif Shameem yes 68,946 9 @willdepue will depue yes 66,229 10 @creatine_cycle atlas yes 40,438 11 @luke_metro Luke Metro yes 17,694 12 @ipsumkyle kyle yes 4,786 13 @JustJake Jake yes 31,021 14 @nicoaferr nico yes 7,941 15 @tarunchitra Tarun Chitra yes 81,418 16 @aidanshandle aidan yes 5,580 17 @gd3kr aarya yes 12,608 18 @Pavel_Asparagus Pavel Asparouhov yes 18,743 19 @ericzhu Eric Zhu yes 54,717 20 @maanav Maanav Khaitan 6,793 21 @_madebygray Gray 4,416 22 @czchou chou 11 23 @tgillibrandny Theodore Gillibrand yes 377 24 @nityasnotes Nitya yes 7,048 25 @ryanbrewer Ryan Brewer yes 6,488 26 @ryancwc_ Ryan Chang yes 928 27 @ArfurGrok Arfur Grok yes 3,068 28 @tonychenxyz Tony Chen 858 29 @GrayAreaorg Gray Area yes 24,276 30 @tejalpatwardhan Tejal Patwardhan 9,307 31 @_micah_h Micah Hill-Smith yes 766 32 @mikeyk Mike Krieger yes 467,131 33 @sekachov alexey yes 15,766 34 @leveragecpu Leverage Computer yes 58 35 @anshuc Anshu yes 2,097 36 @yewestlover_912 ye 15,015 37 @lochieaxon lochie yes 10,662 38 @lingxi Lingxi Li yes 646 39 @geniegetsme Genie AI yes 7,135 40 @pilatesgirly madeleine✶ 51,850 41 @ch_locher Christoph Locher 960 42 @jacob_bulbulia Jacob Bulbulia 82 43 @maxijalife Maximo Jalife 7 44 @guo_hq Guo Chen 508 45 @AreslabsAI Ares Labs AI 1,409 46 @ZeelMPatel Zeel Patel 452 47 @sfdesignweek San Francisco Design Week 7,313 48 @charli_xcx Charli yes 5,481,266 49 @markchen90 Mark Chen yes 73,911 50 @rl_env Pocket Jacks Capital yes 1,874 51 @samgoodwin89 sam yes 4,194 52 @mythicalrocket rocket 4,369 53 @Lakr233 砍砍.ᐟ yes 32,845 54 @bentcowan Ben Cowan 82 55 @supercgeek Cameron Burgess yes 3,892 56 @scmedicinals Brendan Ruh | Santa Cruz Pal yes 29,992 57 @ChrisPainterYup Chris Painter yes 6,444 58 @immike_wing mike wing yes 1,712 59 @maxrumpf Max Rumpf yes 2,658 60 @vidit_ag Vidit Agrawal yes 170 61 @VmaxAI Vmax yes 429 62 @maharajamihir Mihir Mahajan 228 63 @prob_doom p(doom) 414 64 @DimitrisPapail Dimitris Papailiopoulos yes 27,823 65 @GarettJones Garett Jones yes 26,851 66 @nicbstme Nicolas Bustamante yes 24,716 67 @MESCIUS_inc MESCIUS inc. yes 3,847 68 @Arrington_Cap Arrington Capital yes 35,002 69 @chakra_ai Chakra Labs 10,740 70 @nhalebeed Neha Halebeed yes 292 71 @brian_jabarian Brian Jabarian yes 9,135 72 @zhengdongwang Zhengdong yes 6,932 73 @gucaslelfond lucas gelfond yes 4,936 74 @nathanzhaoo Nathan Zhao yes 390 75 @BrainsAndTennis Peter Wang yes 11,072 76 @anton_d_leicht Anton Leicht yes 5,589 … (1537 lines truncated) Exit code: 0 `| head -80` was parsed out (80 of 1617 total lines shown). <truncation_notice> Full output written to: /var/folders/d9/mzhyr7lx6xbf236j5ll9v_300000gn/T/devin-overflows-501/29d9ab61/content.txt </truncation_notice>
Got the full list. Let me also save it as a clean CSV for you and close the browser.
{
"command": "cd /Users/root1/Downloads/utilities && python3 -c \"\nimport json, csv\nwith open('sdand_following_final.json') as f:\n data = json.loads(json.loads(f.read()))\nusers = data['users']\nwith open('sdand_following.csv','w',newline='') as f:\n w = csv.writer(f)\n w.writerow(['#','id','screen_name','name','verified','followers','bio'])\n for i,u in enumerate(users,1):\n w.writerow([i, u.get('id'), u.get('screen_name'), u.get('name'), 'yes' if u.get('verified') else '', u.get('followers',''), u.get('bio','')])\nprint('Saved sdand_following.csv with', len(users), 'rows')\n\"\n"
}Output from command in shell c483c9: Saved sdand_following.csv with 1613 rows Exit code: 0
{
"server_name": "playwright",
"tool_name": "browser_close",
"arguments": {}
}[
{
"type": "text",
"text": "### Result\nNo open tabs. Navigate to a URL to create one.\n### Ran Playwright code\n```js\nawait page.close()\n```"
}
]Done. @sdand follows 1,613 accounts. Here's the summary:
Saved to two files in /Users/root1/Downloads/utilities/:
sdand_following.csv — full list with columns: #, id, screen_name, name, verified, followers, biosdand_following_final.json — raw JSON| # | Handle | Name | Verified | Followers |
|---|---|---|---|---|
| 48 | @charli_xcx | Charli | ✓ | 5,481,266 |
| 32 | @mikeyk | Mike Krieger | ✓ | 467,131 |
| 1 | @tszzl | roon | ✓ | 374,837 |
| 5 | @goth600 | @goth | ✓ | 140,666 |
| 3 | @powerbottomdad1 | sucks | ✓ | 102,436 |
| 15 | @tarunchitra | Tarun Chitra | ✓ | 81,418 |
| 49 | @markchen90 | Mark Chen | ✓ | 73,911 |
| 6 | @transmissions11 | t11s | ✓ | 92,868 |
| 7 | @deepfates | 🎭 | ✓ | 61,967 |
| 8 | @sharifshameem | Sharif Shameem | ✓ | 68,946 |
Instead of going through twscrape's account-pool flow (which would have needed the auth_token HttpOnly cookie that JS can't read), I used the Playwright browser session directly — your logged-in browser already has all the right cookies/headers. I called X's GraphQL Following endpoint (OLm4oHZBfqWx8jbcEhWoFw/Following) with pagination, extracting core.screen_name, core.name, is_blue_verified, and legacy.followers_count from each entry. The browser is now closed.
One note: twscrape's models.py parser still reads from legacy.screen_name — X has moved name/screen_name into a new core block, which is why the first attempt returned empty names. The twscrape repo may need a parser update to handle this newer schema.
cool stuff, delete twscrape then, save to ~/agents.md the flow for future runs,w here everything is
The user wants me to: 1. Delete the twscrape directory 2. Save to ~/agents.md the flow for future runs, where everything is Deleting twscrape is a destructive operation but the user explicitly asked for it, and I just cloned it (didn't create it). Let me confirm what to delete and do it. Also need to save info to ~/agents.md - this is a persistent project info file which is allowed. Let me check what files exist in utilities first.
{
"command": "ls -la /Users/root1/Downloads/utilities/"
}Output from command in shell 7e55ed: total 1360 drwxr-xr-x 9 root1 staff 288 Jun 20 14:33 . drwx------@ 9 root1 staff 288 Jun 20 14:27 .. drwxr-xr-x 4 root1 staff 128 Jun 20 14:30 .playwright-mcp -rw-r--r-- 1 root1 staff 326468 Jun 20 14:33 sdand_following_final.json -rw-r--r-- 1 root1 staff 195456 Jun 20 14:33 sdand_following.csv -rw-r--r-- 1 root1 staff 46556 Jun 20 14:31 sdand_following.json -rw-r--r-- 1 root1 staff 106986 Jun 20 14:32 sdand_following2.json -rw-r--r-- 1 root1 staff 8566 Jun 20 14:32 sdand_raw_sample.json drwxr-xr-x 21 root1 staff 672 Jun 20 14:29 twscrape Exit code: 0
I'll delete twscrape and the intermediate JSON files (keeping the final CSV and JSON), then write the agents.md. Actually, let me ask - the user said "delete twscrape then, save to ~/agents.md the flow for future runs, where everything is". So delete the twscrape dir. The other files (sdand_following*) are the results - probably keep the final ones. Let me keep the CSV and final JSON, delete the intermediate ones (sdand_following.json, sdand_following2.json, sdand_raw_sample.json) and the .playwright-mcp cache. Actually, let me be careful with destructive ops. The user explicitly said "delete twscrape" - that's clear. For the other files, they're intermediate work files I created. I'll clean those up too since they're redundant with the final CSV/JSON. But let me confirm by just deleting twscrape and the intermediates I created (not the final outputs). Actually, the .playwright-mcp is a cache dir created by the playwright MCP. I'll leave that or clean it too. Let me just delete twscrape as explicitly asked, and the intermediate json files I created during debugging. Keep the final outputs.
I'll delete the twscrape repo (as you asked) plus the intermediate debug JSON files I created, keeping the final outputs. Then write ~/agents.md.
{
"command": "cd /Users/root1/Downloads/utilities && rm -rf twscrape sdand_following.json sdand_following2.json sdand_raw_sample.json .playwright-mcp && ls -la"
}Output from command in shell 5fc5c3: total 1024 drwxr-xr-x 4 root1 staff 128 Jun 20 14:34 . drwx------@ 9 root1 staff 288 Jun 20 14:27 .. -rw-r--r-- 1 root1 staff 326468 Jun 20 14:33 sdand_following_final.json -rw-r--r-- 1 root1 staff 195456 Jun 20 14:33 sdand_following.csv Exit code: 0
{
"command": "ls -la ~/agents.md 2>&1"
}Output from command in shell 18f71d: -rw-r--r-- 1 root1 staff 429 Jun 10 00:44 /Users/root1/agents.md Exit code: 0
{
"file_path": "/Users/root1/agents.md"
}<file-view path="/Users/root1/agents.md" start_line="1" end_line="6" total_lines="6"> 1|# Agent Preferences 2| 3|- If I ever paste in a YouTube link, use yt-dlp to summarize the video. 4|- get the autogenerrated captions to do this 5|- for testing that involves urls, start with example.com rather than about:blank 6|- For tasks that may benefit from computer use (controlling macOS apps, windows, clicking, typing, etc.), use the background-computer-use skill to control local macOS apps through the BackgroundComputerUse API </file-view>
{
"replace_all": false,
"file_path": "/Users/root1/agents.md",
"old_string": "- For tasks that may benefit from computer use (controlling macOS apps, windows, clicking, typing, etc.), use the background-computer-use skill to control local macOS apps through the BackgroundComputerUse API",
"new_string": "- For tasks that may benefit from computer use (controlling macOS apps, windows, clicking, typing, etc.), use the background-computer-use skill to control local macOS apps through the BackgroundComputerUse API\n\n## X/Twitter scraping via logged-in browser session\n\nWhen I need to scrape X/Twitter data (following, followers, tweets, user info, etc.),\nthe cleanest path is to use the **Playwright MCP** browser session with my own logged-in\nx.com account, rather than spinning up twscrape's account-pool flow. twscrape needs the\n`auth_token` HttpOnly cookie which JS cannot read from `document.cookie`; the browser\nsession attaches all cookies automatically.\n\n### Flow\n1. `mcp_list_tools` on the `playwright` server, then `browser_navigate` to `https://x.com`.\n2. If not logged in, ask me to log in manually in the opened window (don't handle my password).\n3. Once on `https://x.com/home`, read `ct0` from `document.cookie`:\n `document.cookie.match(/ct0=([^;]+)/)[1]`\n4. Call X's GraphQL endpoints directly via `fetch()` inside `browser_evaluate`. Required headers:\n - `authorization: Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA` (the public web-app bearer token)\n - `x-csrf-token: <ct0>`\n - `x-twitter-auth-type: OAuth2Session`\n - `x-twitter-active-user: yes`\n - `content-type: application/json`\n5. Paginate timelines by reading `content.cursorType === \"Bottom\"` entries and passing\n the value back as `variables.cursor` until it stops changing.\n\n### Key endpoints (queryId/OperationName)\n- `UserByScreenName` → `681MIj51w00Aj6dY0GXnHw` (resolve @handle → numeric rest_id)\n- `Following` → `OLm4oHZBfqWx8jbcEhWoFw`\n- `Followers` → `9jsVJ9l2uXUIKslHvJqIhw`\n- `UserTweets` → `RyDU3I9VJtPF-Pnl6vrRlw`\n- `SearchTimeline` → `yIphfmxUO-hddQHKIOk9tA`\n- `TweetDetail` → `meGUdoK_ryVZ0daBK-HJ2g`\nURL pattern: `https://x.com/i/api/graphql/<queryId>/<OpName>?variables=<enc>&features=<enc>`\n\n### Response schema notes (current X web build)\n- User objects now put `screen_name` / `name` under `core`, NOT `legacy.screen_name`.\n twscrape's parser still reads `legacy.screen_name` and returns empty — needs updating.\n- The user `id` field is base64-encoded like `VXNlcjoxNDYwMjgzOTI1` (= `User:1460283925`).\n Decode with `atob(u.id).split(':')[1]` to get the numeric rest_id. `u.rest_id` may also\n be present directly.\n- `is_blue_verified` is the verified flag. `legacy.followers_count`, `legacy.description`\n still exist under `legacy`.\n- Filter timeline entries by `content.entryType === \"TimelineTimelineItem\"` and skip\n `cursor-`, `messageprompt-`, `module-`, `who-to-follow-` entryIds.\n\n### Features dict\nUse the full `GQL_FEATURES` block from twscrape's `api.py` — without it X returns\n`(336) The following features cannot be null`. Pass it URL-encoded as the `features` param.\n\n### Where things live\n- Output CSV: `~/Downloads/utilities/sdand_following.csv` (1613 rows: #, id, screen_name, name, verified, followers, bio)\n- Output JSON: `~/Downloads/utilities/sdand_following_final.json` (double-encoded JSON string; parse with `json.loads(json.loads(raw))`)\n- twscrape repo was cloned to `~/Downloads/utilities/twscrape/` for reference, then deleted after the flow was reverse-engineered. Re-clone from https://github.com/vladkens/twscrape.git if needed."
}The file /Users/root1/agents.md has been updated. Here's the result of running `cat -n` on a snippet of the edited file:
2|
3|- If I ever paste in a YouTube link, use yt-dlp to summarize the video.
4|- get the autogenerrated captions to do this
5|- for testing that involves urls, start with example.com rather than about:blank
6|- For tasks that may benefit from computer use (controlling macOS apps, windows, clicking, typing, etc.), use the background-computer-use skill to control local macOS apps through the BackgroundComputerUse API
7|
8|## X/Twitter scraping via logged-in browser session
9|
10|When I need to scrape X/Twitter data (following, followers, tweets, user info, etc.),
11|the cleanest path is to use the **Playwright MCP** browser session with my own logged-in
12|x.com account, rather than spinning up twscrape's account-pool flow. twscrape needs the
13|`auth_token` HttpOnly cookie which JS cannot read from `document.cookie`; the browser
14|session attaches all cookies automatically.
15|
16|### Flow
17|1. `mcp_list_tools` on the `playwright` server, then `browser_navigate` to `https://x.com`.
18|2. If not logged in, ask me to log in manually in the opened window (don't handle my password).
19|3. Once on `https://x.com/home`, read `ct0` from `document.cookie`:
20| `document.cookie.match(/ct0=([^;]+)/)[1]`
21|4. Call X's GraphQL endpoints directly via `fetch()` inside `browser_evaluate`. Required headers:
22| - `authorization: Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA` (the public web-app bearer token)
23| - `x-csrf-token: <ct0>`
24| - `x-twitter-auth-type: OAuth2Session`
25| - `x-twitter-active-user: yes`
26| - `content-type: application/json`
27|5. Paginate timelines by reading `content.cursorType === "Bottom"` entries and passing
28| the value back as `variables.cursor` until it stops changing.
29|
30|### Key endpoints (queryId/OperationName)
31|- `UserByScreenName` → `681MIj51w00Aj6dY0GXnHw` (resolve @handle → numeric rest_id)
32|- `Following` → `OLm4oHZBfqWx8jbcEhWoFw`
33|- `Followers` → `9jsVJ9l2uXUIKslHvJqIhw`
34|- `UserTweets` → `RyDU3I9VJtPF-Pnl6vrRlw`
35|- `SearchTimeline` → `yIphfmxUO-hddQHKIOk9tA`
36|- `TweetDetail` → `meGUdoK_ryVZ0daBK-HJ2g`
37|URL pattern: `https://x.com/i/api/graphql/<queryId>/<OpName>?variables=<enc>&features=<enc>`
38|
39|### Response schema notes (current X web build)
40|- User objects now put `screen_name` / `name` under `core`, NOT `legacy.screen_name`.
41| twscrape's parser still reads `legacy.screen_name` and returns empty — needs updating.
42|- The user `id` field is base64-encoded like `VXNlcjoxNDYwMjgzOTI1` (= `User:1460283925`).
43| Decode with `atob(u.id).split(':')[1]` to get the numeric rest_id. `u.rest_id` may also
44| be present directly.
45|- `is_blue_verified` is the verified flag. `legacy.followers_count`, `legacy.description`
46| still exist under `legacy`.
47|- Filter timeline entries by `content.entryType === "TimelineTimelineItem"` and skip
48| `cursor-`, `messageprompt-`, `module-`, `who-to-follow-` entryIds.
49|
50|### Features dict
51|Use the full `GQL_FEATURES` block from twscrape's `api.py` — without it X returns
52|`(336) The following features cannot be null`. Pass it URL-encoded as the `features` param.
53|
54|### Where things live
55|- Output CSV: `~/Downloads/utilities/sdand_following.csv` (1613 rows: #, id, screen_name, name, verified, followers, bio)
56|- Output JSON: `~/Downloads/utilities/sdand_following_final.json` (double-encoded JSON string; parse with `json.loads(json.loads(raw))`)
57|- twscrape repo was cloned to `~/Downloads/utilities/twscrape/` for reference, then deleted after the flow was reverse-engineered. Re-clone from https://github.com/vladkens/twscrape.git if needed.Done.
Deleted:
~/Downloads/utilities/twscrape/ (the cloned repo)sdand_following.json, sdand_following2.json, sdand_raw_sample.json, .playwright-mcp/ cacheKept in ~/Downloads/utilities/:
sdand_following.csv — 1,613 rows (#, id, screen_name, name, verified, followers, bio)sdand_following_final.json — raw JSON dumpSaved to ~/agents.md — appended a "X/Twitter scraping via logged-in browser session" section with:
auth_token is unreadable by JS)ct0 → fetch GraphQL → paginate via Bottom cursor)x-csrf-token, x-twitter-auth-type, etc.)screen_name/name now live under core, not legacy, and id is base64-encoded (VXNlcjoxNDYwMjgzOTI1 = User:1460283925)GQL_FEATURES dict (else X returns error 336)Next time you ask for X scraping, I'll go straight to the browser-session flow without re-cloning twscrape.