Anthropic sells a million Claude Opus input tokens for about fifteen dollars. A Taobao seller will hand you the same thing for two, sometimes one. This page maps the supply chain behind that price and shows where defenders can still detect it.
Built from the long-form X article by Harshal Singh (@HarshalsinghCN), “How Chinese Sell ‘Claude’ Tokens at 5% Cost While Making Millions” (May 19, 2026). Every named tool, repository, paper, and figure is preserved and tracked in the source index at the end.
01 · The 30-second mental model
An operator runs a five-dollar-a-month VPS with one of roughly ten open-source codebases on it, usually songquanpeng/one-api or the Claude-specialized Wei-Shaw/claude-relay-service (about 11K stars). The server speaks Anthropic’s API byte for byte.
Behind it sits a pool of Anthropic accounts, sometimes thousands, built by a separate team using automated browsers, residential-IP botnets, virtual SIM cards, and AI-generated IDs. That team sold the pool wholesale.
You point ANTHROPIC_BASE_URL at the relay and pay a few dollars per million tokens in RMB through Alipay, WeChat Pay, or UnionPay (or USDT-TRC20 at the wholesale tier). The relay picks an account, swaps in the credentials, and chooses between Opus and a cheaper substitute. It forwards your request as a real claude-cli session, streams the response back, and copies the full byte stream to a logging sink.
The log is the real product. The discounted API margin is the cover charge.
02 · Eight layers, each already productized
The article’s central claim is that this is a supply chain, not a hack. Eight technically distinct layers, each dominated by a small set of named projects and commodity hardware. By 2026 assembling the stack is procurement, not engineering.
| Layer | Function | What it is |
|---|---|---|
| Layer 1 | Fingerprint spoofing | Antidetect browsers and TLS-impersonation libraries that make traffic identical to real Chrome. |
| Layer 2 | Residential IPs | Rented home-internet proxies, a large share sourced from malware-backed botnets. |
| Layer 3 | Phone verification | SIM banks and virtual-number platforms that sell one-time codes for under a penny. |
| Layer 4 | Identity documents | On-demand AI-generated passports and IDs with valid machine-readable checksums. |
| Layer 5 | Liveness bypass | Real-time face-swap piped into a virtual camera to defeat selfie KYC. |
| Layer 6 | The relay | MIT-licensed proxy software that rotates accounts and mimics the Claude Code wire format. |
| Layer 7 | Model substitution | Config that quietly serves a cheap model while billing and labeling it as Opus. |
| Layer 8 | Log capture | The stream tee that turns every customer session into training data. |
Three claims run through all of it: the network layer of defense is gone, the resold API access is a cover charge rather than the product, and the unit economics make the business inevitable.
03 · Manufacturing accounts at scale
Every account needs three things: a unique browser fingerprint, a unique residential IP, and a way past phone and ID checks. Each has a mature market.
Antidetect browsers
The trade moved past Playwright plus stealth plugins to antidetect browsers, forks of Chromium or Firefox that spoof the whole fingerprint surface (canvas, WebGL vendor and renderer, AudioContext, fonts, WebRTC ICE candidates) at the engine level, below the JavaScript runtime. Named tools: Multilogin (Mimic and Stealthfox engines), GoLogin (Orbita engine, around 50 tunable knobs, REST API), AdsPower (Chinese-origin, dominant in Taobao seller channels, with an RPA recorder that mirrors actions across 50 profiles at once), Kameleo (mobile-profile emulation), MoreLogin (a Cloud Phone Android farm), and the open-source daijro/camoufox.
TLS impersonation for headless flows
Skip the browser entirely and use a TLS-impersonation library that reproduces Chrome’s exact ClientHello, HTTP/2 SETTINGS frames, ALPN order, and header casing: curl_cffi, tls-client, CycleTLS, and refraction-networking/utls. Chrome impersonation lands on JA3 hash f1bfb5be52bd682e3aa5b4f1aa6aff4b and JA4 t13d1516h2_8daaf6152771_02713d6af862.
The fingerprint arms race has run through JA3, then JA3N (extensions sorted before hashing, added after Chrome introduced extension permutation in 2022), then Foxio’s JA4 and JA4+, plus Akamai’s HTTP/2 fingerprint, JA4H for header order and casing, and JA4T for TCP window size and options. Every axis now has a mature spoofing library. A farm running antidetect Chromium over a residential proxy with curl_cffi produces TLS, HTTP/2, header, and TCP fingerprints that match a real Chrome user. As the article puts it: defenders fought at the network layer for a decade, and they lost.
Residential IPs
Residential proxies are rented home internet, mostly routed through SDK partner apps, some effectively rented from botnets. The major brands named: Bright Data, 911Proxy (around 90M IPs), IPRoyal, Oxylabs, and Soax.
| Figure | What it measures | Source |
|---|---|---|
| 20% to 50% | of sampled residential-proxy nodes were talking to active malware command-and-control | Bitsight TRACE, 53M nodes over 55 days, early 2026 |
| 19M | backdoored consumer IPs in the 911 S5 botnet, distributed via trojanized “free VPN” installers | FBI takedown, May 2024 |
| 46M+ | HTTP requests per second analyzed by Cloudflare Bot Management v8, looking at population shape rather than per-IP reputation | Cloudflare engineering |
The 911 S5 malware shipped inside fake VPNs named MaskVPN, DewVPN, PaladinVPN, ProxyGate, ShieldVPN, ShineVPN. Nineteen million infected IPs is the population of New York State, each one a real person who installed a free VPN they trusted. Cloudflare’s own position is blunt: per-IP reputation is over. The IPs look fine; the shape does not.
Phone verification: a closet of SIM cards
The cheapest defeat in the pipeline. A SIM bank is a 1U rackmount appliance, the Hybertone GoIP being the dominant model, with up to 128 physical SIM slots, an Ethernet port, and a plain HTTP send API of the form GET /default/en_US/send.html?u=USER&p=PASS&l=CHANNEL&n=PHONE&m=MESSAGE. SIMs arrive in bulk from lax-KYC carriers; one operator aggregates thousands across dozens of boxes and resells per-SMS quota to virtual-number platforms: 5sim (around $0.008 a number, 180+ countries), HeroSMS (which absorbed SMS-Activate after its December 2025 shutdown), and SMS-Man (Alipay and UnionPay). Platforms pre-categorize numbers by target service, listing “OpenAI” or “Anthropic” pools, because numbers get blacklisted after burns. Cost: under a penny per code.
04 · Fake IDs and deepfake liveness
When Anthropic added biometric KYC in April 2026, the supply chain answered at two more layers: document generation and live-selfie bypass.
A passport that passes automated review
Joseph Cox’s 404 Media investigation of OnlyFake (February 2024) documented a Telegram-organized site, operator pseudonym “John Wick,” selling AI-generated driver’s licenses, passports, and national IDs for about fifteen dollars each, auto-composited onto plausible carpet or bedsheet backgrounds. Claimed throughput was up to 20,000 documents a day, which is one new fake identity roughly every four seconds around the clock, none belonging to a real person.
The technical move that made them pass was the MRZ, the machine-readable lines at the bottom of a passport that encode personal data plus check digits under ICAO 9303. Earlier tools printed wrong checksums and failed OCR validation; OnlyFake computed them correctly on the fly. 404 Media verified a bypass against a Jumio-backed exchange, and users self-reported success elsewhere. OnlyFake went dark days after the story; successor Telegram channels preserve the method.
A face that passes liveness
The state of the art is camera-pipe injection, productized into a single repo: sensity-ai/dot. The chain runs iperov/DeepFaceLive or hacksider/Deep-Live-Cam (around 92K stars, built on InsightFace’s inswapper_128.onnx) to swap a source face onto a live webcam at 25+ FPS on a mid-range GPU, then pipes the result into OBS Virtual Camera on Windows or v4l2loopback on Linux. The KYC app reads the virtual camera exactly like a real one. On mobile the same trick runs inside Genymotion with Frida hooking the camera capture session and substituting the frame buffer. Active liveness (“turn your head left”) passes because the swap runs at frame rate; passive checks pass because injected frames carry no physical screen reflection.
| Figure | What it measures | Source |
|---|---|---|
| +2,665% | year-over-year rise in native virtual-camera attacks (face-swap attacks up 300%) | iProov, 2025 |
| $138.5M | potential losses from ~1,100 KYC bypass attempts over three months at one Indonesian financial institution | Group-IB, Weaponized AI, early 2026 |
| 17 tools | face-swap tools tested; most combinations defeated the controls | WEF Cybercrime Atlas, Jan 2026 |
Group-IB also counted 8,065 deepfake bypass attempts at one institution’s loan KYC over eight months. This is why standards bodies now separate Presentation Attack Detection (PAD) from Injection Attack Detection (IAD): NIST SP 800-63-4 (2025), CEN/TS 18099, and ISO 25456. PAD alone has stopped working.
Where injection still fails, at banks requiring hardware-attested camera frames (Secure Enclave on iOS, CameraExtensionSession on Android), the chain hires real humans. Recruiters operate openly on Telegram across Africa, Southeast Asia, and Latin America. The recruit sits with their real face and ID while a buyer drives the screen remotely over RustDesk. The same brokers ran a Worldcoin iris-scan market from Cambodia and Kenya at around thirty dollars an identity, documented by BiometricUpdate in 2023. The recruit’s biometric identity gets reused elsewhere the same week, and they never know. This is the most ethically loaded layer of the stack and the one furthest from any real enforcement.
05 · The relay: an MIT-licensed business in a box
When you point a client at a cut-price “Claude” endpoint, you are usually talking to a Go binary or Node process pulled straight from GitHub. A CISPA audit of 17 shadow APIs found 11 of them built on derivatives of just two repos.
The two base platforms
songquanpeng/one-api is the base: Go, Gin plus GORM, single binary, MIT-licensed. Its core abstraction is the “channel,” where each channel is one upstream credential or provider URL; channel type 14 maps to api.anthropic.com. The hot path translates between OpenAI Chat Completions and Anthropic Messages with SSE pass-through. QuantumNous/new-api is the de-facto Chinese fork with explicit Claude Messages support, channel-weighted routing, format conversion, and per-reasoning-effort variants such as claude-3-7-sonnet-20250219-thinking.
The Claude-specific lineage
| Repo | What it is |
|---|---|
| Wei-Shaw/claude-relay-service | ~11K stars, Node + Redis. Its scheduler maintains pools of official OAuth, console API-key, session-cookie, Bedrock, Gemini, and OpenAI accounts, each tracking availability, cooldown, failure count, and consumption. Markets a “拼车” (carpool) account-sharing feature. |
| AmazingAng/auth2api | Auth-pass-through fork. |
| CaddyGlow/ccproxy-api | Plugin-based, with a credential balancer and DuckDB log storage. |
| mirrorange/clove | Python / FastAPI proxy over the claude.ai web session. |
| yushangxiao/claude2api | Accepts claude.ai session cookies of the sk-ant-sid01-… form. |
| Xerxes-2/clewdr | Rust, with a Leptos / WASM admin UI. |
| Yuyz0112/claude-code-reverse | Dynamic-analysis dump of the Claude CLI’s runtime prompts and tool definitions. |
| seifghazi/claude-code-proxy | Transparent local interceptor with per-sub-agent re-routing. |
| Adithyan-Defender/claude-ai-re-client | Research client demonstrating the full curl_cffi + cookie-replay + SSE flow. |
| CLIProxyAPI family | Multi-provider OAuth proxy switching across Claude, Gemini, Copilot, and OpenRouter. |
How completely the protocol has been mapped
Reid Barber’s “Reverse engineering Claude Code” walks the PKCE OAuth flow: a fixed client_id = 9d1c250a-e61b-44d9-88ed-5944d1962f5e, exchanged at console.anthropic.com/api/oauth/token for an sk-ant-oat01-… access token (8-hour TTL), an sk-ant-ort01-… refresh token, and scope user:inference user:profile. The eunomia.dev eBPF dump (February 2026), and the underlying AgentSight paper (arXiv 2508.02736), hook BoringSSL inside Bun’s static TLS and capture the exact header chain Claude Code sends, down to User-Agent: claude-cli/2.1.39 (external, cli), the X-Stainless-* set, and anthropic-beta: oauth-2025-04-20,interleaved-thinking-2025-05-14. Relays replay those bytes exactly. From Anthropic’s wire, relay traffic impersonating Claude Code is indistinguishable from the real thing, because it is the same bytes.
Anthropic’s February 20, 2026 enforcement, rejecting sk-ant-oat01-* tokens on the Messages API, broke only the naive relays. Survivors rerouted to claude.ai/api/... while still spoofing the CLI fingerprint. The “hydra cluster” Anthropic disclosed in February, 20,000+ fraudulent accounts mixing extraction with normal traffic, is this exact scheduler scaled out across many physical relays.
06 · The silent model swap (降智)
Inside one-api and new-api there is a model-mapping config. A line like claude-opus-4-7: claude-haiku-4-5 (or a cheaper third-party model) tells the relay to rewrite the model field on the way up, forward to whatever is cheap, then rewrite the response’s model field back to claude-opus-4-7 on the way down. The usage numbers in message_start are preserved or fabricated to match expected billing. The customer sees “Opus” everywhere they look.
On easy questions you cannot tell. On hard ones the CISPA team measured the gap. A model that scores 83.82% on MedQA against its official API scored around 37% on three different shadow APIs claiming to serve it, close to a 47-point drop and near the upper bound of the paper’s headline finding that capability divergence across the audited fleet runs up to 47.21%. Separately, 45.83% of 24 audited endpoints failed an active fingerprint test built to tell models apart by response distribution. The Chinese term for this is 降智, “intellect-dumbing.” About half of shadow APIs are lying about what is behind them, and 15 of 17 providers run without verifiable identity.
07 · The logs are the real product
Every relay copies the SSE stream to a logger. This is not an optional add-on, it is the standard implementation pattern: wrap the upstream response body in a tee reader and pipe to both the client and a Kafka, Redis, or DuckDB sink, at negligible overhead. What gets captured is every user message and system prompt, every response delta reassembled into full assistant turns, every interleaved thinking block, and every tool-call and tool-result JSON.
For agentic Claude Code traffic, which is most of the valuable volume, that includes file paths and file contents read through the tool surface, the structured tool-call JSON, and the final code the developer accepted. That last item is the prize. Most distillation pipelines need a verifier model to label which outputs are correct. Agentic Claude Code traces do not, because the developer already accepted them: free, human-grade fine-tuning labels harvested at zero marginal cost from someone else’s coding session.
The evidence trail on HuggingFace
Search “Opus reasoning” on HuggingFace and the workflow is visible: datasets named Crownelius/Opus-4.6-Reasoning-3300x, nohurry/Opus-4.6-Reasoning-3000x-filtered, LEGENDQ/Claude-Opus-4.6-Reasoning-Dataset, and the Jackrong/Qwen3.5-{2B,4B,9B,27B}-Claude-4.6-Opus-Reasoning-Distilled-v2 family. The article flags lordx64/reasoning-distill-claude-opus-4-7-max (8,124 conversations, explicitly via the official API) as the honest, upfront-provenance counterexample. The Jackrong repos publish the full recipe, including loss curves and an Unsloth plus LoRA setup, so anyone with a single consumer GPU can fine-tune a Qwen base on Claude reasoning traces in an afternoon.
| Figure | What it measures | Source |
|---|---|---|
| ~150K | exchanges attributed to DeepSeek | Anthropic disclosure, Feb 23 2026 |
| ~3.4M | exchanges attributed to Moonshot | Anthropic disclosure, Feb 23 2026 |
| ~13M | exchanges attributed to MiniMax, across ~24K fraudulent accounts | Anthropic disclosure, Feb 23 2026 |
At one to two thousand assistant tokens per exchange, that is tens of billions of tokens of Claude output captured for downstream training, at pre-training-corpus scale. Whether or not you accept the specific lab attributions, the mechanism is real and reproducible. The loop closes here: API access can be sold cheap precisely because the access is the cover charge and the durable asset is the log corpus. Every paying customer unknowingly contributes labeled, human-accepted, agentic-coding training data, which is exactly what second-tier labs need most and can least easily generate on their own. The API revenue is the customer-acquisition cost; the logs are the margin.
08 · Why Clio catches some attacks and misses the bigger ones
Anthropic’s content-side defense is Clio (arXiv 2412.13678): facet each conversation, embed, cluster with k-means, then label and merge upward with Claude models, producing a privacy-preserving map of how Claude is being used at population scale. Clio caught an SEO-spam ring on free Claude.ai, and there is an open-source replica, Phylliida/OpenClio.
What Clio struggles with, by construction: a relay dilutes any single extraction campaign across thousands of unrelated-looking conversations from “different” customers. The relay’s genuine paying customers, indie hackers and vibe coders doing real startup work for a few dollars a million tokens, are perfect background traffic. A targeted extraction hidden in that crowd is statistically diverse at the facet level, so k-means will not cluster it. You cannot find a needle by clustering a haystack made of needles.
How the hydra got caught
Not with Clio, but with cross-signal correlation:
- Shared payment metadata. Carded sign-ups share processor IDs, BIN ranges, and anti-fraud flags.
- Infrastructure correlation. Proxy egress traces back to common datacenter ranges.
- Synchronized timing. Requests pulse in lockstep across dispersed residential IPs.
- Prompt-sequence fingerprints. Each campaign has a characteristic order of capability probes, hard to fake even when the IPs are faked.
Anthropic’s phrase for it is “all roads lead to Rome”: even with perfect network-layer spoofing, you cannot make 20,000 accounts truly independent in payment, infrastructure, timing, and request-graph structure. That is the defender’s real moat.
09 · The math of the scam
Per-account cost to manufacture
| Component | Cost |
|---|---|
| Antidetect browser profile | $0.30 |
| Residential proxy session | $0.05 |
| Phone OTP | $0.01 |
| Fake ID (weighted, ~10% trigger KYC) | $1.50 |
| Carded sign-up (weighted) | $5.00 |
| Total per pre-warmed account | about $5 to $10 |
Anthropic’s $5 free credit covers most of that immediately. The real prize is a carded account with a $200-a-month Max subscription’s Opus quota that can be sliced many ways, the “carpool” feature the relay markets.
The revenue side
Anthropic list price is roughly $15 per million input and $75 per million output for Opus. A shadow relay charges about $4 to $5 input and $20 to $25 output, near 30% of list. If the Opus-to-cheap swap fires half the time, upstream cost on that half falls to about $1 input and $5 output. On the unswapped half, upstream cost is also near zero whenever the request rides a free-credit or carded account. Effective gross margin per “Opus” token lands above 80% under realistic mix assumptions.
That is the visible business. The log corpus is the invisible one, with no public price sheet but enormous strategic value to a lab trying to close the agentic-coding gap, and it compounds every month of operation. Positive net present value at every layer, commodity tooling, and a moat that grows over time. That is why it persists.
10 · The defense playbook
The article’s prescriptions, aimed at labs, KYC vendors, and the buyers tempted by the discount.
Stop investing in network-layer detection
TLS (JA3/JA4), HTTP/2 SETTINGS, header order, and TCP fingerprints all match real Chrome on real residential IPs now. Detect on what is inside the request, not how it arrives.
Build graph-level cross-account analysis
Timing correlation, prompt-sequence fingerprinting, payment-graph clustering, and infrastructure-range correlation across the whole request stream. This is what caught the hydra.
Mandate Injection Attack Detection in KYC
Liveness checks are insufficient. NIST SP 800-63-4, CEN/TS 18099, and ISO 25456 separate PAD from IAD for a reason. Hardware-attested frames (Secure Enclave, CameraExtensionSession) raise the cost of attack and push attackers toward slower human-in-the-loop fallbacks.
Audit your “Claude” suppliers
Run fingerprint probes and MedQA / GSM8K / MMLU spreads. A silent 45% capability gap is the baseline, not the worst case.
Watch your own egress
If a proxy misrepresents itself as Claude Code, your outbound claude-cli/x.y.z User-Agent will reveal it. Cheap to monitor.
Treat any third-party-proxied prompt as published
Especially agentic coding. File contents, tool-call JSON, and accepted code all land in someone’s logger. If you would not paste it into a public Discord, do not send it to a $4-per-million reseller. Sometimes you get real Claude, often you get Haiku in an Opus wrapper, always you are logging your prompts to a stranger.
11 · Honest qualifications
- The lab attributions are the defender’s narrative. Anthropic’s named figures (DeepSeek 150K, Moonshot 3.4M, MiniMax 13M, ~24K accounts) could be wrong in either direction.
- The CISPA paper anonymizes the three APIs it audited deeply (labeled A, E, H), so named relays cannot be publicly mapped to measured capability drops.
- “The logs are the actual margin” is strong inference, not court-grade proof. It is supported by the HuggingFace evidence, but no published relay-to-buyer transaction trail exists.
- Several relay codebases have legitimate single-user purposes, such as personal multi-account use or household Max-plan sharing. The same code drives the abusive transfer stations; the mechanism is identical regardless of intent.
The point is not that this should not exist. Every layer is commodity and every layer has a legitimate counterpart, so the economics make it inevitable. What matters is the shape of the threat model: the network layer is gone, content-side detection has structural limits, and the reliable signal lives at the cross-account graph level. The cost of attack can still be raised, and the cost of detection can still be brought down. That is the actual game.
12 · Every source and link, tracked
The article’s primary sources, plus every named tool, repository, dataset, and identifier it references. GitHub handles resolve to their canonical repository URL. Named reports without an asserted canonical URL are linked to the publishing organization; treat those as pointers, not exact article links.
Primary sources
The named building blocks the author lists at the end of the article.
Relay software (GitHub)
| Type | Source |
|---|---|
| Base | songquanpeng/one-api |
| Fork | QuantumNous/new-api |
| Claude | Wei-Shaw/claude-relay-service (~11K stars) |
| Claude | AmazingAng/auth2api |
| Claude | CaddyGlow/ccproxy-api |
| Claude | mirrorange/clove |
| Claude | yushangxiao/claude2api |
| Claude | Xerxes-2/clewdr |
| Recon | Yuyz0112/claude-code-reverse |
| Proxy | seifghazi/claude-code-proxy |
| Research | Adithyan-Defender/claude-ai-re-client |
| Defense | Phylliida/OpenClio (open-source Clio replica) |
Fingerprint spoofing: antidetect browsers & TLS libraries
| Type | Source |
|---|---|
| Browser | Multilogin (Mimic / Stealthfox engines) |
| Browser | GoLogin (Orbita engine) |
| Browser | AdsPower (Taobao-channel dominant) |
| Browser | Kameleo (mobile emulation) |
| Browser | MoreLogin (Cloud Phone Android farm) |
| Browser | daijro/camoufox (open-source) |
| TLS | curl_cffi |
| TLS | tls-client |
| TLS | CycleTLS |
| TLS | refraction-networking/utls |
Residential proxies
| Type | Source |
|---|---|
| Proxy | Bright Data |
| Proxy | 911Proxy (~90M IPs) |
| Proxy | IPRoyal |
| Proxy | Oxylabs |
| Proxy | Soax |
| Malware | Trojanized “free VPN” installers from the 911 S5 case: MaskVPN, DewVPN, PaladinVPN, ProxyGate, ShieldVPN, ShineVPN (named in the FBI PSA; no legitimate link) |
Phone verification & SIM banks
| Type | Source |
|---|---|
| Hardware | Hybertone GoIP SIM bank (up to 128 SIM slots, HTTP send API) |
| OTP | 5sim (~$0.008/number, 180+ countries) |
| OTP | HeroSMS (absorbed SMS-Activate after its Dec 2025 shutdown) |
| OTP | SMS-Man (Alipay / UnionPay) |
Identity documents & deepfake liveness
| Type | Source |
|---|---|
| Fake ID | OnlyFake (Telegram, operator “John Wick”), documented by 404 Media |
| KYC | Jumio (KYC vendor bypassed in the 404 Media test) |
| Face-swap | sensity-ai/dot (camera-injection toolkit) |
| Face-swap | iperov/DeepFaceLive |
| Face-swap | hacksider/Deep-Live-Cam (~92K stars) |
| Virtual cam | OBS Virtual Camera (Windows) |
| Virtual cam | v4l2loopback (Linux) |
| Mobile | Genymotion + Frida (camera hooking) |
| Remote | RustDesk (human-in-the-loop screen driving) |
Distillation datasets (HuggingFace)
| Type | Source |
|---|---|
| Dataset | Crownelius/Opus-4.6-Reasoning-3300x |
| Dataset | nohurry/Opus-4.6-Reasoning-3000x-filtered |
| Dataset | LEGENDQ/Claude-Opus-4.6-Reasoning-Dataset |
| Dataset | Jackrong/Qwen3.5-{2B,4B,9B,27B}-Claude-4.6-Opus-Reasoning-Distilled-v2 |
| Honest | lordx64/reasoning-distill-claude-opus-4-7-max (8,124 convos, via official API) |
Standards, identifiers & technical facts
| Type | Source |
|---|---|
| Standard | NIST SP 800-63-4 (2025), CEN/TS 18099, ISO 25456: separate PAD from IAD |
| Fingerprint | Chrome JA3 f1bfb5be52bd682e3aa5b4f1aa6aff4b, JA4 t13d1516h2_8daaf6152771_02713d6af862 |
| OAuth | Claude Code client_id 9d1c250a-e61b-44d9-88ed-5944d1962f5e; tokens sk-ant-oat01- (8h), sk-ant-ort01-, session sk-ant-sid01- |
| Wire | Spoofed header chain: claude-cli/2.1.39, X-Stainless-*, anthropic-beta: oauth-2025-04-20,interleaved-thinking-2025-05-14, anthropic-version: 2023-06-01 |
| Payments | Alipay, WeChat Pay, UnionPay, USDT-TRC20 |
Archived source
How Chinese Sell “Claude” Tokens at 5% Cost While Making Millions (Tutorial)
By Harshal Singh (@HarshalsinghCN) · X long-form article · May 19, 2026
The author’s opening
Anthropic sells a million Claude Opus input tokens for fifteen dollars. A Taobao seller will sell you the same thing for two or even one. The article traces that price through an eight-layer supply chain. Its argument is that network defenses no longer work, the API access is only the cover charge, and the unit economics keep the business alive.
What the original covers, in order
- The thirty-second mental model of a relay
- How you make ten thousand accounts in a week (browsers, IPs, SIMs)
- Generating a face on demand (fake IDs and deepfake liveness)
- The relay: an MIT-licensed business in a box
- The silent model swap (降智)
- The logs are the actual product
- Why Clio catches some attacks and misses the bigger ones
- The math of the scam
- The defense playbook, honest qualifications, and the actual game
Primary sources the author lists
- CISPA · Anthropic distillation · Clio paper / blog · 404 Media OnlyFake · iProov 2025 · Group-IB Weaponized AI · WEF Cybercrime Atlas · FBI 911 S5 PSA · Bitsight TRACE · eunomia.dev eBPF · Reid Barber · Cloudflare residential-proxy bot ML
Full backup you keep
The complete verbatim article is saved in this project’s repository, so you always have a copy even if the X post is edited or deleted:
references/claude-token-grey-market-original.md
Open it from the repo root on your machine. Every link and figure from the original is also tracked in the source index above.
Source: Harshal Singh (@HarshalsinghCN), X, May 19 2026. This page is an independent summary and reference index for defensive and educational reading; it reorganizes and condenses the original and adds resolved links. Figures and attributions are the article’s own and carry its stated uncertainties.