ego (lite) is just a browser, ego is your personal agent across devices.
Join waitlist
Website data extraction for spreadsheet workflows

Free Website data extractor to CSV or spreadsheets with your AI agent

Try:

Give your AI agent a list of pages and the columns you need. ego (lite) opens each page in a visible browser on your Mac, extracts rendered data, and exports a source-linked CSV or Markdown report you can open in Excel or Google Sheets.

Trusted by developers from

OpenAIAnthropicGoogleMetaNVIDIACursorPerplexity
SpaceXTeslaNotionFigmaStripeNetflixAirbnb

How to extract website data into a spreadsheet in 5 steps

Give your Agent the source pages, required columns, and a clear row limit. ego (lite) opens each page in a visible Space, keeps browser state separate, and pauses whenever a page needs your input. The result is CSV or Markdown that imports cleanly into a spreadsheet; the workflow does not pretend to be a native Google Sheets connector.

1

Install ego (lite) and choose an authorized browser context

Download ego (lite) for Mac, then choose only a Chrome context you are authorized to use. The Agent works with the same rendered pages you can open and stays inside that visible access boundary.

2

Define the source pages, fields, and limits

List the pages you want to process, the fields you need, and a maximum result count. Ask for stable column names, one row per item, source page, checked-at time, access status, deduplication, and CSV or Markdown output so the task stays bounded and reviewable.

3

Watch the Agent collect rendered data

The Agent opens each source page in its own Space and records the requested fields from the rendered document. It can resolve relative URLs and keep anchor text, link type, or other visible values without following destinations or inventing missing cells.

4

Keep each source in a separate Space

Give each page or page group its own Space and run the same output rules in parallel. Browser state stays independent, while you can open any Space to inspect progress, pause the task, or take over.

5

Review and open the spreadsheet-ready export

The Agent returns a CSV or Markdown report with one row per requested item. Review the values, column names, source page, check time, and access status before opening the CSV in Excel or importing it into Google Sheets. A spreadsheet import does not replace source review.

Why use ego (lite) for website data to spreadsheet workflows?

Copy-and-paste works for one public table, and an API works when a stable endpoint already exists. ego (lite) adds a reviewable browser workflow for rendered pages, authorized sessions, and bounded page batches where every spreadsheet row needs a source.

Keep extraction, cleanup, and export in one Agent workflow

The same coding Agent collects visible fields, resolves relative paths when requested, normalizes values, removes duplicates, and writes the final CSV or Markdown files. You avoid passing the data between multiple extensions and preserve the same columns from input to export.

Existing ego (lite) browser-research task-time comparison, shown as product context rather than a URL extraction guarantee

Run several page-level extractions side by side

Give each source page or page group its own Space and apply the same columns and deduplication rules. Every spreadsheet row keeps its source, and one blocked page does not stop the other Spaces from finishing.

Parallel ego (lite) Spaces extracting links from separate source pages

Multitask across parallel Spaces

Run several URL extraction tasks at the same time, each in its own ego (lite) Space. Switch between Spaces to monitor progress without interrupting the others, and take over only when one task needs a login, CAPTCHA, consent screen, or another human decision.

Use the browser context you already control

Import only the Chrome context you authorize. The Agent can work with the same rendered page you can open, including an authorized tool or dashboard, while keeping the extraction local to your Mac and every result tied to its source.

ego (lite) Chrome context import for an authorized page-level link extraction task

What this website data extractor can and cannot collect

This browser workflow extracts visible values from the source pages you provide and formats them for a spreadsheet. It does not become a complete site crawler, native Sheets integration, or data-quality guarantee unless you explicitly add those steps.

What your Agent can extract

Visible fields, link destinations, and useful context from the rendered source pages in your bounded list.

  • A schema of visible fields you name, such as product title, price, availability, author, date, or status, with one row per item
  • Destination URLs resolved from relative and absolute href values in the rendered document
  • Anchor text, rel, target, occurrence count, and internal, external, file, email, phone, or fragment classification
  • A source-linked CSV or Markdown report with stable columns, check time, and access status across several pages; the CSV can be opened in Excel or imported into Google Sheets

What stays out of scope by default

No unlimited crawl, hidden-page discovery, destination check, or access bypass.

  • A complete site inventory or orphan-page list from pages that were never linked or provided
  • URLs stored only in network traffic, canvas content, scripts, or unlinked text unless you add that source type
  • Native Google Sheets or Excel API synchronization, destination status, redirects, canonicals, or access bypasses unless you request a separate authorized step

Read the Robots Exclusion Protocol

Extract spreadsheet rows you can trace back to every page

Give your Agent the source pages and output columns. ego (lite) returns a reviewable CSV or Markdown table with visible browser work and a source page on every row.

Try free website data extraction

Website data, URL, and link extractor FAQ

Choose the pages, the columns, and a row limit, then ask your Agent to open each page in an ego (lite) Space and return one row per visible item. The export is a source-linked CSV or Markdown file you can open in Excel or import into Google Sheets. Keep source_url, checked_at, and access_status columns so every value can be reviewed; this workflow does not provide a native Sheets sync or invent fields that the page does not show.

ego (lite) does not claim a native Google Sheets connector. It writes a CSV or Markdown file locally, which you can import into a sheet after reviewing the rows. Google Sheets also has IMPORTHTML for public HTML tables and lists, while a visible browser workflow is useful when content is JavaScript-rendered, spread across several pages, or available only in an authorized session.

Start with the fields that answer your research question, then add source_url, source_title, checked_at, access_status, and an evidence or notes column. Keep one value per cell, use null or ‘not displayed’ for missing values, and define a stable key for deduplication. A schema makes a CSV useful in a spreadsheet and makes later refreshes comparable.

A URL extractor turns links from a page, HTML block, document, or text into a clean list. For a live web page, a useful result usually includes the resolved destination URL, anchor text, link type, and source page. With ego (lite), your coding Agent opens the live page in a visible browser, extracts the links from the rendered document, and returns a source-linked CSV or Markdown report.

Install ego (lite), give your Agent the page URLs, and choose the columns and output format you need. The Agent opens each page in its own Space, waits for it to render, resolves relative paths, classifies and deduplicates the links, and writes the final list. You can watch any Space and take over if a page asks for a human decision.

A URL extractor often means a text tool that finds URL-shaped strings wherever they appear. A link extractor or hyperlink extractor usually reads actual HTML anchors, so it can keep anchor text, rel attributes, targets, and internal-versus-external context. This ego (lite) workflow focuses on page-level hyperlinks by default, while the prompt can be expanded to include plain-text URLs when that is genuinely part of the task.

Not from one page alone. It can get the links available in the rendered document for every source page you list, but a complete website inventory may also require an XML sitemap, a bounded internal crawl, and another discovery source for orphan pages. The page keeps that distinction explicit instead of calling one extracted page a complete site map.

Yes, when the links become available in the rendered document that your authorized browser session can see. The Agent waits for the page's main content before collecting anchors. Links hidden in closed shadow roots, canvas drawings, network responses, or code are separate extraction sources and are not included unless you ask for them explicitly.

Yes. The Agent resolves each href against the final source-page URL, compares the destination origin with the source origin, and labels the row as internal or external. It can also keep file, email, phone, fragment, rel, target, anchor-text, and occurrence fields so the list is useful for an SEO or content audit.

Yes. Ask your Agent for CSV, Markdown, or both, with the same fields in each file. You can also request a short summary of source-page coverage, unique links, internal and external counts, and any page that needs your review.

Only when you are authorized to view the page and choose an ego (lite) browser context that already has the required session. The Agent never asks for or enters your credentials. If the session does not have access, or the page asks for a CAPTCHA, approval, or another human decision, the Agent stops and hands the visible Space back to you.

Not by default. Extracting a destination from a source page does not prove that the destination returns 200, redirects correctly, or has the expected canonical. Add a separate, bounded checking pass if you need those fields; that pass must open each destination and record its own check time and status instead of guessing from the source link.

ego (lite) is free to download for Mac. The workflow uses the coding Agent you already run, so its token usage and any provider costs still apply. It is designed for visible, supervised page batches rather than an unlimited unattended site crawl.

Open the page in a real browser, wait for the pricing content to render, and then extract the anchors from the visible document. Keep the final URL, page title, anchor text, and check time with each result. If the links exist only in a script, API response, or closed shadow root, name that as a separate authorized source instead of claiming the visible HTML contains it.

Record the URL and the block, then use an official sitemap, public feed, employer API, or a human review path that the site permits. Do not bypass a robot check with stealth settings, fingerprint changes, proxy rotation, or credential sharing. A blocked source should remain blocked in the report so a partial extraction is not mistaken for complete coverage.

You can extract contact details that are visibly published on pages you are authorized to process, but treat them as personal data. Keep the source URL and purpose, collect only the fields you need, avoid hidden or obfuscated values unless the page renders them for users, and follow privacy, consent, retention, and outreach rules. Never infer a person's contact information or send messages automatically.

Define the allowed domain, starting URLs, pagination pattern, maximum page count, and stop conditions before running. Follow only visible, authorized next-page links, deduplicate by canonicalized URL, and record each page's status. An email hidden in source code or an unlinked endpoint is a different extraction request; do not expand into an unlimited crawl or access-restricted content by assumption.

A single URL usually exposes one page, not every article in a blog. First extract the blog's visible article index or XML sitemap, cap the number of pages, then fetch each permitted article and produce a source-linked digest or Markdown bundle. Preserve titles, URLs, dates, and failures so the model can distinguish missing pages from pages that were never requested.

Use a generic hierarchy of signals: semantic article links and headings in the rendered page, pagination or load-more controls, then an XML sitemap when available. Normalize URLs against the page origin, retain anchor text and dates where shown, and ask for a review of ambiguous cards. A generic extractor handles common layouts, but no adapter-free method can guarantee every custom JavaScript pattern.

Yes, for a bounded list of pages you are allowed to capture. Give each URL its own Space, wait for a stated ready condition, set a consistent viewport, and save a timestamped filename with the source URL. Redact credentials and personal data, respect robots or site terms, and mark pages that fail or require a human decision instead of retrying without a limit.

A rendered-browser workflow can produce the DOM or page text after JavaScript runs, but ego (lite) is a local visible browser rather than a hosted HTML-rendering API. If you need an API, compare rendering timing, JavaScript support, caching, data retention, authentication, and per-request pricing, and verify that sending the target page to the provider is allowed. Keep source URL and render time in the output.