ultimate-browsing
code-yeongyu/oh-my-openagent · Agent Skill
Accesses web pages, including JS rendering, clicks, forms, screenshots, and WAF bypass.
Jump to install ↓Install options
The review-first prompt is selected by default. You can switch to a direct command.
Copy the prompt, paste it into Claude, Codex, or another assistant, and it will review the skill page and install it.
I'd like to install this Claude Code skill: https://github.com/code-yeongyu/oh-my-openagent/tree/dev/packages/shared-skills/skills/ultimate-browsing Please review the SKILL.md file, verify the skill is legitimate, then copy the skill folder into .claude/skills/ of my project.
A direct command skips the review — inspect the source first.
npx skills add https://github.com/code-yeongyu/oh-my-openagent/tree/dev/packages/shared-skills/skills/ultimate-browsing
Read SKILL.md and any companion files before deciding whether to install. Click a file to preview it.
SKILL.md contents
Original author text · read-onlyUltimate Browsing
Web access for everything a plain fetch cannot finish: a page that renders in JS, a click or a form, a screenshot, a login that must persist across pages, or a host that blocks generic fetchers (WAF / 403 / Cloudflare). Start at the cheapest tier that can do the job and climb only when it cannot:
Tier 1 — insane-search (headless extraction + WAF bypass) -> Tier 1.5 — agent-reach (platform-native APIs, esp. Chinese platforms) -> Tier 2 — a real browser: 2a a code-driven kernel browser, 2b Chrome stealth (CloakBrowser + agent-browser) when the page fights back.
PHASE 0 — ROUTE FIRST (MANDATORY)
User request
|
+- extract text/data from a URL --------------------- TIER 1 insane-search
+- URL blocked / 403 / Cloudflare / WAF ------------- TIER 1 insane-search
+- YouTube/Vimeo/TikTok subtitles or metadata ------- TIER 1 insane-search (yt-dlp)
+- read an article / blog / Reddit / HN / arXiv ----- TIER 1 insane-search
|
+- Chinese platform (xhs/douyin/weibo/bilibili/v2ex/wechat) TIER 1.5 agent-reach
+- podcast transcript / stock forum ----------------- TIER 1.5 agent-reach
+- Twitter feed / LinkedIn profile / GitHub via CLI - TIER 1.5 agent-reach
|
+- Tier 1/1.5 returned empty or partial ------------- TIER 2 2a kernel browser -> 2b stealth
+- click / fill form / scroll / interact ------------ TIER 2 2a kernel browser -> 2b stealth
+- screenshot / render / play video ----------------- TIER 2 2a kernel browser -> 2b stealth
+- login session across pages / inject cookies ------ TIER 2 2b Chrome stealth (profile + cookies)
+- test web app / QA / dogfood ---------------------- TIER 2 2a kernel browser -> 2b stealth
|
+- simple search query ------------------------------ NOT this skill (use web-search)
Read the matching reference before acting: references/insane-search/README.md, references/agent-reach/README.md, or references/chrome-stealth.md.
Tier 1 — insane-search (headless extraction)
When: content extraction, blocked-URL bypass, media metadata — no browser UI needed.
Why first: ~10x faster than a browser, no process spin-up; handles most "fetch this blocked page" requests via curl_cffi TLS impersonation, yt-dlp (1858 sites), official public APIs, mobile URL transforms, Phase-2.5 surrogate archives (Wayback / archive.today snapshots, provenance-tagged — see references/insane-search/cache-archive.md), a key-gated Jina Reader (JINA_API_KEY), and a Playwright real-Chrome fallback. The engine lives inside this skill at engine/ and is invoked as a module. Surrogate results are dated COPIES: a result whose provenance is snapshot must be reported with its snapshot_timestamp, never presented as the live page.
# Core command — auto-detects WAF, runs the full fetch grid (run from the skill dir):
python3 -m engine "https://example.com/blocked-page"
# add --selector "<CSS>" for positive-proof validation, --device auto|desktop|mobile,
# --trace to inspect every attempt, --json for machine-readable output.
# YouTube subtitles / metadata (no browser):
yt-dlp --write-sub --write-auto-sub --sub-lang "en,ko" --skip-download -o "/tmp/%(id)s" "<URL>"
# Reddit / HN / Bluesky / arXiv etc. use official public endpoints — see the Phase 0 index in
# references/insane-search/README.md (Twitter syndication, Reddit .json, HN Firebase, ...).
The full engine harness (rules R1-R7, the Phase 0 official-API index, the no-site-name rule, and the references/insane-search/*.md deep-dives for TLS, Playwright routing, Naver, media, etc.) is in references/insane-search/README.md. Read it before tuning the engine or adding a WAF profile.
Escalate to Tier 1.5 or Tier 2 when
- The target is a Chinese / social platform with a native reader -> Tier 1.5.
- insane-search returns empty/partial, or the page needs JS interaction, a screenshot, a persistent login, or media playback -> Tier 2.
Tier 1.5 — agent-reach (platform-native readers)
When: the target is a platform with a first-class API/CLI that beats generic fetching — especially Chinese platforms that stealth browsers still cannot reach cleanly. Several channels are zero-config (Douyin, V2EX, Reddit, RSS, YouTube); others need a one-time auth you supply via environment variables if you have access (JINA_API_KEY for Jina Reader — anonymous access is dead, see references/insane-search/jina.md; TWITTER_* for X; a transcription key for podcasts).
| Category | Platforms | Entry |
|---|---|---|
| social | xhs (Xiaohongshu), douyin, weibo, bilibili, V2EX, Reddit, Twitter/X | references/agent-reach/social.md |
| web | Jina Reader, WeChat articles, RSS | references/agent-reach/web.md |
| video | YouTube, Bilibili, podcast transcripts, Douyin video | references/agent-reach/video.md |
| career | LinkedIn | references/agent-reach/career.md |
| dev | GitHub (gh CLI) | references/agent-reach/dev.md |
| search | Exa AI | references/agent-reach/search.md |
mcporter call 'douyin.parse_douyin_video_info(url: "<URL>")' # douyin, zero-config
curl -s "https://r.jina.ai/https://weibo.com/<uid>/<pid>" # weibo via Jina
yt-dlp --dump-json "<bilibili-url>" # Bilibili (overseas: add --cookies-from-browser)
curl -s "https://www.v2ex.com/api/topics/hot.json" # V2EX public API
Routing table, per-platform auth (set TWITTER_* env vars, gh auth login, a transcription key — only if you have access), rate-limit notes, and known version quirks are in references/agent-reach/README.md.
Tier 2 — a real browser (real interaction)
When: real interaction is needed (clicks, forms, screenshots, video, persistent login), or Tier 1/1.5 failed.
Tier 2a — kernel browser (default)
Drive the page from the code cell you are already in, with no CLI process and no open CDP port. On a Bun >= 1.4 runtime that is new Bun.WebView() (navigate, click, type, evaluate, screenshot, raw cdp); elsewhere it is playwright-core/puppeteer-core against the local Chrome. Snapshots and screenshots come back in-process, so this is the cheapest way to answer "what does the page actually render".
await using view = new Bun.WebView({ width: 1280, height: 800 })
await view.navigate(url)
const title = await view.evaluate("document.title")
await view.screenshot({ path: "/tmp/page.png" })
Climb to 2b when the site detects automation (Turnstile, FingerprintJS, "unusual traffic"), when the flow needs a persistent logged-in profile or injected cookies, or when the kernel browser cannot reach the page at all.
Tier 2b — Chrome stealth (blocked or logged-in pages)
CloakBrowser is a stealth Chromium with source-level fingerprint patches that passes Cloudflare Turnstile, FingerprintJS, BrowserScan, and 30+ detectors; agent-browser is the CDP automation CLI that drives it. Both are runtime-installed tools (not vendored here). Full setup, version pins, launch flow, cookie login, and cross-platform notes are in references/chrome-stealth.md.
NEVER clear cookies, cache, or site data (Network.clearBrowserCookies, Storage.clearCookies, chrome.browsingData.remove, "clear browsing data") on the user's real/main browser profile — it wipes their logged-in state everywhere. If the task needs that profile's login state, clone the profile directory first (rsync -a <profile>/ <tmp-clone>/) and launch CloakBrowser / agent-browser with the clone as the user-data-dir; run any clearing on the clone only.
# 1. Launch CloakBrowser with CDP on :9242 (see chrome-stealth.md for install + venv).
# 2. CloakBrowser launches tabless — open the first tab via CDP before any agent-browser command:
curl -s -X PUT "http://127.0.0.1:9242/json/new?https://example.com"
# 3. Drive it with agent-browser over CDP:
agent-browser --cdp 9242 snapshot -i # interactive elements (@eN refs)
agent-browser --cdp 9242 click @e3
agent-browser --cdp 9242 screenshot out.png
agent-browser --cdp 9242 close
Cookie login (cross-platform)
scripts/extract_cookies.py reads cookies from a local Chromium-family or Firefox-family browser and optionally injects them into the running CDP session. It resolves browser profile paths and decrypts cookie values per-OS (macOS Keychain, Linux libsecret, Windows DPAPI):
# Extract cookies to a file:
mkdir -p ~/.local/state/omo-cookies
python3 scripts/extract_cookies.py --browser chrome --domain youtube.com --output ~/.local/state/omo-cookies/youtube.cookies.json
# Extract and inject into the running CDP session:
python3 scripts/extract_cookies.py --browser chrome --domain youtube.com --inject --cdp 9242
Cookie export files are written with owner-only 0600 permissions. Do not place live auth cookies in shared temp directories or commit them to a repo. Cookie injection sends values to CDP over stdin rather than argv. Cookies apply on next navigation — reload after injecting. Google services use fingerprint-bound tokens that may not transfer across browser profiles. Full detail in references/chrome-stealth.md.
Reference docs
| File | When to read |
|------|-------------|
| references/insane-search/README.md | Tier-1 engine harness (R1-R7, Phase 0 API index, no-site-name rule) + its *.md deep-dives |
| references/agent-reach/README.md | Tier-1.5 routing table, platform auth, per-category *.md |
| references/chrome-stealth.md | Tier-2 CloakBrowser + agent-browser install, CDP flow, version pins, cookie login |
Environment variables
CLOAK_CDP_PORT=9242 # CloakBrowser CDP port (default 9242)
AGENT_BROWSER_USER_AGENT="..." # override UA to hide HeadlessChrome
AGENT_BROWSER_HEADED=1 # show the browser window
# agent-reach auth: set the channel-specific env vars from each tool's docs only if you have access
# insane-search needs no env vars — it auto-installs deps on first run
Anti-patterns
- Do NOT launch Chrome stealth for plain text extraction — use Tier 1.
- Do NOT pass an
--init-scriptfor the webdriver flag — CloakBrowser already patches it at source; the only required override is--user-agent. - Do NOT run agent-browser before creating the first tab via
curl -X PUT .../json/new— CloakBrowser launches tabless. - Do NOT use vanilla Chrome when stealth is needed — always CloakBrowser.
- Do NOT forget to
closethe session when done. - Do NOT inject cookies without reloading the page.
- Do NOT hardcode site domains/selectors into
engine/**orwaf_profiles.yaml— runtime hints only (see the no-site-name rule in the insane-search reference).