Selected work

Case study / Content / browser automation / local tools

Web to Markdown

How can a rendered webpage become a document I can edit and keep?

A local web app and CLI that render public webpages with Chromium, convert the result to Markdown, and support editing and export.

Project type
software
Contribution
Built the conversion pipeline, local interface, CLI, and tests
Status
Local app and CLI
Context
Independent JavaScript project
Updated
Web to Markdown showing a completed Homebrew conversion with editable Markdown and rendered preview side by side
Local app, with the Homebrew page in Split view. Homebrew content belongs to its authors.

Key decision: Render through Chromium, then watch for substantive body text to settle within a bounded wait before extracting HTML.

Start here

Overview

The engineering problem

A page's initial HTML may omit content added by JavaScript. Extracting an article can also lose useful structure or capture surrounding navigation. Fetching user-supplied URLs introduces network access that needs explicit limits.

Contribution

I built the Chromium conversion pipeline, URL and network checks, local web interface, Markdown editor and preview, and CLI. I added automated fixtures for conversion, security boundaries, document editing, and export behaviour.

What exists now

The repository provides a local web app and CLI for converting a page into Markdown. The web app supports editing, a sanitized preview, an outline that follows edits, copying, and saving the resulting document.

Inspect the work

Engineering detail

Constraints

  • The application runs locally with Node.js and Playwright Chromium. The server listens on 127.0.0.1 and the project has no hosted service.
  • Conversion processes one public HTTP or HTTPS page per request. It does not sign in, crawl a site, or interact with consent banners.
  • Network rendering must have bounded time, requests, and content sizes. Those bounds can leave a page incomplete.

Architecture at a glance

The web interface sends URL conversions to an Express endpoint. The CLI uses the same conversion function. A Node fetch layer validates and pins public addresses before supplying responses to Chromium. The rendered HTML passes through Readability and Turndown. In the browser, editable Markdown feeds a Marked preview sanitized with DOMPurify.

Key decisions

Decision 01

When is a page ready to convert?

Fetching HTML alone misses JavaScript content, while waiting for network silence can stall on continuously active pages.

Selected
Render through Chromium, then watch for substantive body text to settle within a bounded wait before extracting HTML.
Alternatives
Convert the initial HTTP response without running JavaScript; Wait for all network activity to stop
Trade-off
Late content, blocked resources, and pages with little stable text may still be incomplete.
Consequence
JavaScript-rendered fixtures can be converted without requiring a quiet network.

Decision 02

What should happen when article extraction is uncertain?

Readability does not identify a useful article on every page.

Selected
Use article content when it contains enough paragraph text. Otherwise convert the body and return a visible warning. Also offer explicit Full page conversion.
Alternatives
Fail when article extraction finds no article; Always convert the complete body
Trade-off
The body fallback may include navigation and repeated page furniture.
Consequence
The user can inspect and edit the result rather than silently receiving an assumed clean article.

Decision 03

How should browser rendering access the network?

A public-looking URL can resolve or redirect to a nonpublic address, and pages can request additional resources.

Selected
Validate HTTP and HTTPS URLs, standard ports, and every DNS answer. Pin the selected public address for each Node request and recheck redirects and allowed browser subresources through the same fetch layer.
Alternatives
Validate only the first URL and let Chromium fetch subsequent resources directly
Trade-off
Blocking media, workers, WebSockets, and other browser features reduces compatibility. These checks do not establish universal safety against hostile pages.
Consequence
Network access has an inspectable boundary with shared request, byte, and time budgets.

Decision 04

How can editing and preview stay useful without loading remote content?

Converted or opened Markdown can contain active HTML and remote images. Export needs to reflect the current edits.

Selected
Keep the editable document in the browser, sanitize the rendered preview with DOMPurify, and represent remote images as text placeholders. Copy and Save as .md use the edited text.
Alternatives
Render raw HTML directly or preview remote images automatically
Trade-off
The preview is intentionally less visually complete than the source page.
Consequence
Source, Preview, and desktop Split views support review without the preview fetching remote images.

Reliability and quality

Automated tests
The test suite covers URL and mixed-DNS rejection, redirects, pinned lookup, conversion fixtures, API validation, CLI arguments and overwrite protection, preview sanitization, heading analysis, and Chromium UI interactions.
Rendering and editing fixtures
Browser tests exercise JavaScript content with ongoing network activity, failed content resources, editing, outline actions, preview, and save controls. Deterministic fixtures do not establish compatibility with arbitrary live websites.
CI definition
The workflow defines lint, formatting, automated tests, and a dependency audit on Node.js 22, 24, and 26. Live-site conversion is excluded from CI. A manual smoke script targets Homebrew separately.

Limits that shape conversion

The fetch layer (opens in a new tab) budgets 30 seconds, 80 requests, 2 MB per response, 12 MB of downloads, and at most five redirects per fetch. The converter also limits rendered HTML to 4 MB and Markdown output to 2 MB. The local server admits two active conversions. These are configured bounds, not measured performance results.

The converter (opens in a new tab) routes allowed document, script, and data requests through the checked fetch layer. It blocks image, media, font, WebSocket, and event-stream requests, disables workers, and warns when content requests fail. The settling check waits for repeated substantive text within a cap, so a page can still change after capture.

Two interfaces, one conversion path

The web app offers article and full-page modes, optional front matter, and choices for images and link destinations. Local Markdown files are read in the browser, with a 2 MB limit and a prompt before replacing unsaved edits. The preview implementation (opens in a new tab) keeps remote images as placeholders.

The CLI (opens in a new tab) writes Markdown to stdout or a selected file, refuses to overwrite an existing file, and reports failures through stderr and a nonzero exit code. Both interfaces reuse the conversion function, so rendering and extraction decisions stay in one place.

What next

Live-page compatibility and operational isolation would need separate evidence before offering this as a public service. The local workflow keeps the current deployment scope explicit.

Scope of the evidence

Limitations

  • This is a local application, not a hosted conversion service. Chromium must be installed and the machine needs access to the public target website.
  • Article extraction is heuristic. Body fallback may include navigation, while delayed content, blocked scripts, browser restrictions, and site bot controls can prevent a complete conversion.
  • The fetch layer reduces access to nonpublic destinations but is not a complete security guarantee for running untrusted browser code. Public deployment would need a separate isolation and operational review.
  • Tests use controlled fixtures and substituted dependencies. They do not prove universal website compatibility or security against every hostile input.
  • Preview sanitization applies to the local preview. Exported Markdown should still be treated as untrusted content by whichever application opens it.