ego (lite) is just a browser, ego is your personal agent across devices.
Join waitlist
URL to Markdown and rendered PDF conversion in a real browser

Free Webpage to Markdown converter and PDF exporter for your AI agent

Try:

Extract website content to clean Markdown or a rendered PDF using the browser session you are already signed into. Your AI agent converts or prints what you can actually see, including internal docs and subscribed articles, and keeps the source URL, checked-at time, and access status with each file.

Trusted by developers from

OpenAIAnthropicGoogleMetaNVIDIACursorPerplexity
SpaceXTeslaNotionFigmaStripeNetflixAirbnb

How to convert a webpage or HTML page to Markdown or PDF in 5 steps

For one public article, a free web converter or the browser's Print command may be faster and you should use it. This workflow is for the rest: pages that need the session you are already signed into, a list you want converted in one run, or a source-linked Markdown/PDF handoff. Give your Agent a bounded list of URLs, the output shape you want, and a rule to stop at anything that needs a human. ego (lite) supplies the visible browser, a separate Space per page, and your existing session for the pages that require one.

1

Install ego (lite) and choose an authorized browser context

Download ego (lite) for Mac, then import only the Chrome context you are authorized to use. That context is the whole point: it is what lets the Agent open an internal doc or a subscribed article as you, rather than as an anonymous server. The Agent works inside the context you choose and never enters credentials or works around a check.

2

Send the Markdown or PDF conversion prompt to your AI agent

Give your Agent a short prompt: the URLs, whether you want one file per page or a single combined document, and whether the output should be Markdown, PDF, or both. For PDF, specify paper size, orientation, margins, background graphics, and whether the page should be printed after a particular state appears. Keep the rule that the Agent stops when a page needs a login, payment, or consent decision you have not already made.

3

Watch each page render before it is converted or printed

Your Agent opens each URL in its own ego (lite) Space and waits for the requested rendered state. Content that a plain HTTP fetch would miss is present in the Markdown, while PDF output captures the page as it is visibly rendered at the chosen viewport. Headings, lists, tables, code blocks, links, and images map to Markdown; navigation, ads, and cookie banners are left out when you request the content export.

4

Convert a list of pages in parallel Spaces

Give each page its own Space and convert a whole list at once instead of pasting URLs one at a time. The pages progress side by side without mixing browser state, so a docs section or a reading list becomes one run. Your own tabs stay untouched while the Agent works.

5

Get source-linked Markdown files or PDFs

The Agent writes a .md file per page, a PDF per page, or both, with the source URL, page title, and checked-at time in the handoff. Markdown links and images are resolved to absolute URLs; PDFs preserve the rendered print layout as supported by the browser. Any page that did not fully open is listed with its status instead of being saved as a clean file of the wrong content.

Why use ego (lite) to convert a website to Markdown or PDF?

Hosted web converters are quick for public pages and blocked on everything else. Libraries like Turndown and markdownify convert HTML you already have, while browser print commands make a quick PDF. ego (lite) sits where those shortcuts fall short: a real browser holding your own session, several pages at once, and source-linked Markdown or PDF files that record what it actually rendered.

Finish this browser task 3.5× faster with ego (lite)

Use the same coding Agent for the conversion and whatever you do next with the Markdown. It renders each page in a real browser and returns files you can drop straight into your notes or an agent's context. In the task shown here, ego (lite) finished in 81.8 seconds, compared with 282.9 seconds for an agent browser. Actual timing varies by website, workflow, and network conditions.

Task-time comparison for ego (lite) and an agent browser during a webpage conversion task

Convert or print a whole list at once, not one URL at a time

Free web converters take one page per paste, and browser print dialogs make repeated archiving tedious. Here each page gets its own ego (lite) Space and the list runs in parallel, which is what makes a docs section or a reading list a single job. The scope stays your list: this is bounded conversion, not a site crawler.

Parallel ego (lite) Spaces converting separate web pages to Markdown

Keep multitasking while your Spaces render

Run several Markdown conversions or PDF prints in parallel, each in its own visible Space. Keep working in your own tabs, or open any Space when you want to check what a page rendered. If one page hits a login wall, a paywall, or a consent screen, the Agent pauses only that Space and hands it back to you—without interrupting the rest.

Reach the pages a hosted converter never can

This is the difference that matters. Import the Chrome context you choose, and the Agent can convert the internal wiki page, the shared doc, the staging environment, or the publication you subscribe to — the pages an anonymous server request returns a login screen for. It uses the access you already have and never creates new access.

ego (lite) Chrome context import for converting pages that require a signed-in session

What this Markdown converter can and cannot do

The scope is deliberately narrow: pages you can already open, converted in a browser you control, for a list you provide. It uses the access you have; it does not create access you do not have, and it is not a site crawler.

What your Agent can convert

Pages the current authorized session can already open.

  • Article and documentation pages, including client-rendered ones, converted after the page finishes rendering
  • Pages that need the session you are already signed into, such as an internal wiki, a shared doc, a staging environment, or a subscription you hold
  • Headings, paragraphs, lists, tables, code blocks with their language, emphasis, links, and images, with URLs resolved to absolute form
  • A bounded list of URLs converted in parallel, as separate .md files or one combined document
  • PDF files printed from the visible rendered page with recorded viewport, print settings, source URL, checked-at time, and access status
  • Front matter carrying source_url, title, checked_at, and access status for every file

What stays out of scope

No bypass, no crawling, no invented content.

  • Getting past a paywall, login wall, consent screen, or bot check, or entering credentials on your behalf
  • Crawling a whole site or following links beyond the URLs you list
  • Producing a file for a page that did not load; the status is recorded instead
  • Reading text out of images, screenshots, or video, or transcribing audio
  • A guarantee that interactive widgets, animations, embedded media, fonts, or complex print layouts match the live page exactly
  • Any right to republish what you converted — the source's copyright and terms still apply

Read the CommonMark specification

Convert or print the pages other tools return a login screen for

Give your Agent the URLs and the output shape you want. ego (lite) converts or prints each page in a visible browser holding your own session, then hands back source-linked Markdown files or PDFs that remember where they came from and what state the page was in.

Try the free Markdown and PDF converter

Webpage to Markdown converter FAQ

For one public page, paste the URL into a free converter such as Firecrawl's or Microlink's and download the result. For a page your session must already be able to open, or for a bounded list you want converted together, give the URLs to your coding Agent in ego (lite). It waits for each rendered page, keeps the main content while leaving out navigation and ads, and writes Markdown files with the source URL, title, checked-at time, and access status. The output is a conversion of what your authorized browser can see—not a paywall bypass, hidden-content extractor, or site-wide crawler.

For a single public page, paste it into a free web converter such as Firecrawl's or Microlink's — that is the fastest route and this workflow is not trying to replace it. For pages those tools cannot reach, install ego (lite) and give your coding Agent the URLs, the output shape you want, and a rule to stop at anything needing a human. The Agent opens each page in a visible browser Space using your own session, converts what rendered, and writes .md files that keep the source URL, title, and check time.

Because they fetch the page from their own servers, with no access to your cookies or session. Firecrawl's free tool states it directly: it works on publicly accessible pages, and content behind a login or paywall is not accessible without credentials. What you usually get back is not an error but clean Markdown of a login form or a paywall teaser. ego (lite) converts inside the browser context you authorized, so it sees the page you see, and it records the access status alongside the content.

ego (lite) is free to download for Mac, and the work is done by the coding Agent you already use, so the only cost is your Agent's tokens. There are no conversion credits or per-page charges in this workflow. Compare that to the usual meaning of free in this category: Microlink allows 25 conversions a day, Firecrawl's free plan runs on a monthly credit allowance and limits its free tool to one page at a time, and Jina AI's Reader offers a rate-limited free tier before paid plans. The trade-off here is setup: you need a Mac and a coding Agent.

Yes, and for one public page you should. A hosted converter or the r.jina.ai URL prefix will do it in seconds with nothing installed. Choose this workflow when the thing you actually need is different: a page that requires your session, a list of pages in one run, or files that record which page state was converted. If none of those apply, a web converter is the better tool and this page will not pretend otherwise.

If you already have the HTML, use a library: Turndown in JavaScript or TypeScript, markdownify or html2text in Python, or Pandoc on the command line. They are mature, free, and the right answer for HTML in hand. Their limit is that they convert what you give them — they do not go and render a page, and they have no access to your logged-in session. ego (lite) covers that earlier step: getting the correct rendered page out of a real browser. The two compose; use a library for HTML you have, this for pages you need to open first.

If you already have an HTML file, use your browser's print-to-PDF command or a server-side HTML-to-PDF library. For a webpage that must render in an authorized browser session, ask your coding Agent to open the URL in ego (lite), wait for the required state, and print the visible page with a defined paper size, orientation, margins, and background setting. The output is a PDF of that rendered state with source URL, checked-at time, and access status recorded. It does not bypass login walls or guarantee that interactive widgets, animations, fonts, or complex layouts reproduce exactly.

Yes, when the page is already openable in a browser session you are authorized to use. The Agent can print the visible page from its ego (lite) Space after you handle any required login, consent, CAPTCHA, or verification step. It never enters credentials or circumvents a control, and the PDF should retain the source URL, checked-at time, and access status so you can review what was archived. Check the source's terms and any sensitive data before sharing the file.

Yes. The Agent works in a real Chromium browser and converts the page after it has rendered, so content a plain HTTP request would miss is present in the Markdown. Some hosted converters also render pages now, so this is no longer a unique advantage on public URLs. It matters most in combination with your session: a client-rendered page that also requires a login is where anonymous fetchers fail on both counts.

Ask for it in the prompt. Turning a URL to a Markdown file is the default output: the Agent writes one .md per page into the location you specify, or a single combined document if you prefer, with source_url, title, and checked_at in the front matter. Because links and images are resolved to absolute URLs, the file still works after you move it into a notes vault or a repository. Pages that did not fully open appear in the run summary with their status rather than as files.

It converts a bounded list of URLs, each in its own Space, in parallel — a docs section or a reading list is one run. It is not a site crawler: it will not discover and follow links across a domain, and this page will not claim otherwise. If you need site-wide crawling with scheduling and retries, a crawler API such as Firecrawl's is the built-for-purpose category. If you need a supervised run over pages you chose, including ones that require your session, this fits better.

Tables become GitHub-flavored Markdown tables, code blocks become fenced blocks carrying the language when the page shows it, and emphasis, headings, and lists map to their Markdown equivalents. Links and image sources are rewritten to absolute URLs so they resolve outside the original site. You can ask for images as Markdown image links, as a separate list, or omitted. Text inside an image is not extracted; there is no OCR in this workflow.

Markdown carries the document's structure — headings, lists, tables, code — in far fewer tokens than the equivalent HTML, so the same content costs less context and its hierarchy survives chunking. That is why Markdown has become the common target format for feeding web content to models. What this workflow adds is provenance and reach: each file records the URL and time it came from, so you can re-check a retrieved chunk against its source, and internal documentation that a hosted fetcher cannot open can go into the pipeline too.

No, and no converter's is. Markdown has no equivalent for many things a page can do: interactive widgets, forms, embedded players, complex nested layouts, and heavy styling do not survive, and unusual table structures may need a manual pass. The Agent converts the main content and leaves out navigation, ads, and consent banners, which is normally what you want but is still a judgment about what counts as content. The front matter keeps the source URL so you can always go back and compare.

Converting a page you are authorized to view into another format for your own reading or research is ordinary use of your browser, and this workflow does nothing your browser cannot already do. Two things still apply: the source's copyright does not change because the format did, so republishing converted content is a separate question, and a site's terms may restrict automated access regardless of format. This is not legal advice. Keep runs bounded to pages you have legitimate access to, and never use this to get around a paywall.

Convert the main article or documentation element, keep its heading levels, lists, tables, and code fences, and remove navigation, ads, cookie banners, and repeated footer text. Preserve source_url, title, and checked_at in front matter, then chunk by headings with a small overlap and review a sample against the page. Markdown reduces structural noise, but it does not make a lossy extraction factual; provenance and spot checks still matter.

It can convert a page after it renders in the authorized browser session, including client-rendered documentation that an anonymous HTTP fetch misses. Cloudflare, CAPTCHA, login, paywall, and consent screens remain access controls: when one appears, the Agent stops and records the status instead of bypassing it. A visible browser improves reach for pages you are allowed to open; it does not promise access to a protected site.

Send the Agent only the target URLs and the main-content boundary, ask for conversion rather than a summary, and exclude navigation, ads, and repeated boilerplate. Return the Markdown file and a short status list instead of pasting the entire page back into the conversation. Keep source links and checked-at times so a reviewer can compare a chunk with the live page; no cleaner can guarantee that an ambiguous page was interpreted correctly without that check.

Yes, for fields the authorized page visibly exposes. Ask the Agent for a JSON schema such as name, SKU, displayed price, currency, availability, and source_url, require null or ‘not displayed’ for missing fields, and keep the rendered page as the evidence source. This avoids hand-written XPath for a one-off task, but it is not a guarantee that hidden metadata or an internal API is available or permitted.

A browser extension can be convenient for one visible page, and an authorized chat or document URL can be included in a bounded ego (lite) conversion run. ego (lite) does not claim a background export API for every ChatGPT, Claude, or Gemini conversation, and it will not crawl your history automatically. List the pages you want, keep each source and access status, and manually review sensitive conversation content before sharing the files.

Choose the specific documentation URLs or a bounded section, convert them to Markdown with headings and absolute links preserved, then upload the resulting files or add them to the assistant's supported knowledge source. A full-site crawl is a different, resource-intensive job; this workflow does not discover links beyond the URLs you provide. Keep the source URL and date in front matter so the assistant's answers can be traced back to the current docs.

If you can open the document in an authorized browser session, add its URL to the same bounded conversion prompt and ask for headings, tables, links, and checked-at metadata to be preserved. The Agent stops at a permission dialog, login, or export restriction; it does not use private APIs or copy documents a session cannot access. For an organization-wide export, use Feishu/Lark's official export or API with the workspace administrator's approval.