URL Extractor
Paste any text or HTML and get back a clean, deduplicated list of the links inside it.
Overview
Links hide in all kinds of text: page source, sitemaps pasted as text, exported chat logs, server logs, or a long email. This tool scans for anything that starts with http:// or https://, trims the punctuation that sentences leave stuck to the end of a link, removes duplicates, and lists what it found, one per line, in the order each link first appears.
Examples & Sample Data
Extracting links from a sentence
Read https://example.com/guide, then the FAQ (https://example.com/faq).
https://example.com/guide https://example.com/faq
How It Works
- Paste text or HTML containing one or more links into the box.
- Click Extract.
- Copy the deduplicated list of links found.
Common Use Cases
Auditing the links on a page
Paste a page's HTML source to list every absolute link it contains, then check them in bulk with a workflow.
Collecting links from a thread or log
Pull just the URLs out of a chat export, an email chain, or a server log without copying them by hand.
Tips & Best Practices
- Only absolute http and https links are found. Relative links such as /about have no host, so they can't be checked on their own.
- Chain it into the Broken Link Checker or Redirect Checker in Workflows to check each link it finds.
- Nothing you paste leaves your browser: extraction runs entirely client-side.
Frequently Asked Questions
No. The scan runs in your browser using JavaScript. Nothing you paste is uploaded anywhere.
Punctuation that ends a sentence, or a closing bracket around a link, is trimmed because it is almost never part of the address. A closing bracket is kept when the link itself opened one, as in many Wikipedia URLs.
Yes. Two spellings of the same address, such as a differently capitalised host name, are listed once, using the first spelling found.
Related Tools
Explore more high-performance utilities.