Web Data
https://api.deepinfo.com/v1/lookup/webdataLoads a web page in a real browser and returns everything Deepinfo extracts from it, in one call: detected technologies, page metadata (title, description, Open Graph tags, JSON-LD), the HTML source and visible text, links, scripts, trackers (Google Analytics, AdSense and Tag Manager IDs), e-mail addresses, favicons, robots.txt, and the HTTP response headers, cookies and redirects. Add screenshot=true to capture a screenshot as well.
Use it to profile a website in depth, for example to compare a look-alike domain's page with your own or to find the analytics IDs it shares with other sites. For technologies only, Technology is lighter; for an image only, use Screenshot.
A request takes several seconds (about 10 s in our tests).
Authentication
Send your API key in the apikey request header.
Query Parameters
| Parameter | Required | Description |
|---|---|---|
url | Required | Web page to analyze, e.g. https://www.deepinfo.com.Example https://www.deepinfo.com |
screenshot | Optional | Also capture a JPEG screenshot of the rendered page. The response then includes a screenshot object. Default false.Example true |
wait_ | Optional | Page load events to wait for before the page is parsed, comma-separated. Ignored when screenshot is true: the page is then always parsed after load. One or more of load, domcontentloaded, networkidle0, networkidle2. Default networkidle2,load,domcontentloaded.Example networkidle2,load,domcontentloaded |
max_ | Optional | Maximum size of html.source_code, in bytes. For a larger page, source_code holds a placeholder message instead of the HTML. Range 1–33554432. Default 8388608.Example 8388608 |
proxy | Optional | Proxy to load the page through, as a URL with an explicit port, e.g. http://user:password@host:8080. The proxy host must be a public IP address or an allowed domain. |
Response Fields
| Field | Description |
|---|---|
url | The requested URL |
technology. | One entry per detected technology |
technology. | Identifier of the technology |
technology. | Name of the technology |
technology. | What the technology is |
technology. | Categories it belongs to |
technology. | Detection confidence (0–100) |
technology. | Detected version; empty when the site does not reveal it |
technology. | The version as a number, or null |
technology. | CPE name |
technology. | Other CPE names of the technology |
technology. | The vendor's website |
technology. | Icon file name |
html. | Page title |
html. | Site name given in the page metadata |
html. | Meta description |
html. | Meta keywords |
html. | Language of the page |
html. | Other language versions the page declares |
html. | Character encoding |
html. | Whether the page asks search engines not to index it |
html. | Canonical URL |
html. | Open Graph and Twitter card tags, one entry per tag |
html. | Tag name |
html. | Tag value |
html. | Structured data blocks, in expanded JSON-LD form |
html. | The page HTML. For a page larger than max_file_size, a placeholder message instead of the HTML |
html. | SHA-256 hash of html.source_code |
html. | The visible text of the page |
html. | SHA-256 hash of html.content |
html. | The most frequent words of html.content |
html. | Host names on the site's own domain that the page links to |
html. | Links to other sites |
html. | Host names of those links |
html. | Registered domains of those links |
html. | Scripts the page loads |
html. | Iframes the page loads |
html. | Icons the page links to |
html. | Tracking IDs found in the page, one entry per tracker (e.g. Google Analytics) |
html. | Tracker name |
html. | IDs found for that tracker |
html. | E-mail addresses found in the page |
html. | E-mail addresses found in the page that are on the site's own domain |
html. | Whether the page tries to block inspection of its content |
favicon[] | One entry per icon |
favicon[]. | Icon URL |
favicon[]. | Hash of the icon |
robots_ | Content of the site's robots.txt |
robots_ | Hash of robots_ |
robots_ | The Disallow paths |
http. | HTTP response headers, one entry per header |
http. | Header name |
http. | Header value |
http. | Cookies the page set, one entry per cookie |
http. | Cookie name |
http. | Cookie value |
http. | Domain the cookie applies to |
http. | Path the cookie applies to |
http. | Expiry time |
http. | Whether the cookie is sent over HTTPS only |
http. | Whether the cookie is hidden from JavaScript |
http. | The cookie's SameSite attribute |
http. | The cookie's SameParty attribute |
http. | The cookie's Priority attribute |
http. | Size of the cookie, in bytes |
http. | Whether it is a session cookie |
http. | Each URL on the way to the final page |
http. | The URL |
http. | The HTTP status it returned |
screenshot | Only with screenshot=true: the capture, with the screenshot.* fields below |
screenshot. | The requested URL |
screenshot. | The URL the browser ended up on after redirects |
screenshot. | Link to the JPEG image: a signed link that expires after 7 days. null if the capture was skipped |
screenshot. | success when the page was loaded |
screenshot. | When the screenshot was taken (UTC) |
connection_ | success when the page was loaded |
version | Version of the result format (currently 1) |
check_ | When the lookup was performed (UTC) |
Response Schema
Inferred from examples Built from the 2 saved 2xx example responses: the fields they contain, with the types seen there. It is not a contract.
| Field | Type | Present | Example |
|---|---|---|---|
html | object | 2 of 2 | |
html. | object | 2 of 2 | |
html. | string | 2 of 2 | "Deepinfo" |
html. | string | 2 of 2 | "Deepinfo is a Continuous Threat Exp…" |
html. | string | 2 of 2 | "en" |
html. | array | 2 of 2 | |
html. | array | 2 of 2 | |
html. | boolean | 2 of 2 | false |
html. | string | 2 of 2 | "utf-8" |
html. | string | 2 of 2 | "https://www.deepinfo.com/" |
html. | string | 2 of 2 | "Deepinfo | Continuous Threat Exposu…" |
html. | array< | 2 of 2 | |
html. | string | 2 of 2 | "https://www.deepinfo.com/#organizat…" |
html. | array< | 2 of 2 | "http://schema.org/Organization" |
html. | array< | 2 of 2 | |
html. | string | 2 of 2 | "Deepinfo" |
html. | array< | 2 of 2 | |
html. | string | 2 of 2 | "og:site_name" |
html. | string | 2 of 2 | "Deepinfo" |
html. | array< | 2 of 2 | "www.deepinfo.com" |
html. | array< | 2 of 2 | "linkedin.com" |
html. | array< | 2 of 2 | "linkedin.com" |
html. | array< | 2 of 2 | "https://linkedin.com/company/deepin…" |
html. | array< | 2 of 2 | "https://www.deepinfo.com//static/js…" |
html. | array | 2 of 2 | |
html. | array | 2 of 2 | |
html. | array | 2 of 2 | |
html. | array | 2 of 2 | |
html. | array< | 2 of 2 | "https://www.deepinfo.com//static/im…" |
html. | boolean | 2 of 2 | false |
html. | string | 2 of 2 | "<!DOCTYPE html><html lang=\"en\"><hea…" |
html. | string | 2 of 2 | "ad93f88fec2753e10d83cad96f866788c00…" |
html. | string | 2 of 2 | "Skip to main content Platform PLATF…" |
html. | array< | 2 of 2 | "data" |
html. | string | 2 of 2 | "14ab6575f7a485db17770369046e7516de0…" |
http | object | 2 of 2 | |
http. | array< | 2 of 2 | |
http. | string | 2 of 2 | "https://www.deepinfo.com" |
http. | number | 2 of 2 | 200 |
http. | array< | 2 of 2 | |
http. | string | 2 of 2 | "intercom-device-id-xki6sc8r" |
http. | string | 2 of 2 | "<cookie value>" |
http. | string | 2 of 2 | ".www.deepinfo.com" |
http. | string | 2 of 2 | "/" |
http. | string | 2 of 2 | "Lax" |
http. | null | 2 of 2 | |
http. | number | 2 of 2 | 63 |
http. | string | 2 of 2 | "2027-06-20T15:18:37.000Z" |
http. | boolean | 2 of 2 | false |
http. | boolean | 2 of 2 | false |
http. | string | 2 of 2 | "Lax" |
http. | boolean | 2 of 2 | false |
http. | array< | 2 of 2 | |
http. | string | 2 of 2 | "content-type" |
http. | string | 2 of 2 | "text/html; charset=utf-8" |
url | string | 2 of 2 | "https://www.deepinfo.com" |
favicon | array< | 2 of 2 | |
favicon[]. | string | 2 of 2 | "https://www.deepinfo.com//static/im…" |
favicon[]. | string | 2 of 2 | "b742b1a1b669ce7b299449865e7cf766abd…" |
robots_txt | object | 2 of 2 | |
robots_txt. | string | 2 of 2 | "# AI training crawlers — explicit a…" |
robots_txt. | string | 2 of 2 | "8a12283023f4ae29941dca8de0633eb44d9…" |
robots_txt. | array< | 2 of 2 | "/404.html" |
version | number | 2 of 2 | 1 |
check_date | string | 2 of 2 | "2026-09-23T14:45:19.751Z" |
connection_status | string | 2 of 2 | "success" |
technology | object | 2 of 2 | |
technology. | array< | 2 of 2 | |
technology. | string | 2 of 2 | "cloudflare" |
technology. | string | 2 of 2 | "Cloudflare" |
technology. | number | 2 of 2 | 100 |
technology. | string | 2 of 2 | "CloudFlare.svg" |
technology. | string | 2 of 2 | "http://www.cloudflare.com" |
technology. | string | 2 of 2 | "cpe:/a:deepvendor:cloudflare" |
technology. | string | 2 of 2 | "" |
technology. | array< | 2 of 2 | "CDN" |
technology. | string | 2 of 2 | "Cloudflare is a web-infrastructure…" |
technology. | array | 2 of 2 | |
technology. | null | 2 of 2 | |
screenshot | object | 1 of 2 | |
screenshot. | string | 1 of 2 | "https://diss-999.storage.googleapis…" |
screenshot. | string | 1 of 2 | "2026-09-23T14:45:36.767Z" |
screenshot. | string | 1 of 2 | "success" |
screenshot. | string | 1 of 2 | "https://www.deepinfo.com/" |
screenshot. | string | 1 of 2 | "https://www.deepinfo.com" |
Errors
400 (10400) if url is missing or invalid, or a parameter is out of range. The validation error currently names the parameter domain, although the request parameter is url (see the example). A 400 without parameters (only details, e.g. "Target is not allowed.") means the target cannot be analyzed. 500 for an unexpected error. 503 means the service or its browser is busy: retry after the number of seconds in the Retry-After header. See Getting Started → Errors.
Examples
Saved examples from the Deepinfo API. Selecting one loads it into the request and response panels.
Worked Examples
Worked examples of this endpoint, each on its own page with the exact request and the response it returns.
- Everything About a Web PageEverything about developer.mozilla.org in one call: technologies, meta tags, HTML, visible text, links, scripts, favicons, robots.txt and headers.
- With a ScreenshotWeb data with screenshot=true: the same data plus a screenshot object with a signed link to a JPEG of the page.
- Tracking IDs on a PageWeb data of www.mozilla.org: html.trackers finds a Google Tag Manager container and two Google Analytics IDs.