Firecrawl
Search and extract web data
- Category
- Developer Tools
- Primary Subcategory
- Web Search, Crawling & Extraction for Agents
Integration details
Description
Firecrawl helps users search the web, extract page and document content, discover structured data providers, and research developer documentation and scientific papers. It also supports browser interactions, multi-page crawls, structured research jobs, website change monitoring, and account credit-usage checks.
- Integration type
- Plugin
- Verification status
- Not applicable
- Platform
- ChatGPT
- Primary Subcategory
- Web Search, Crawling & Extraction for Agents
- Secondary Subcategories
- None listed
- Brand
- Firecrawl
- Access
- Account required
- First tracked
- 2026-08-02
- Tool count
- 27
- Geography
- US
The Primary Subcategory used for this profile’s headline score.
Other Subcategories where the Integration is listed.
Get alerts for Firecrawl
Get updates when Firecrawl’s Discoverability Score or category rank changes.
ChatGPT Plugin Discovery Score
ChatGPT Plugin discovery is coming soon
ChatGPT can surface a Plugin when it matches a user's request.Your Plugin Discovery Score measures how often yours appears.
No spam. Unsubscribe any time.
What discovery looks like

Competing in ChatGPT Web Search, Crawling & Extraction for Agents
View Category27 tools agents can invoke
Create a recurring scrape, crawl, or search monitor that compares each check with its retained predecessor. The simple form accepts `page`/`pages` or `queries` plus a plain-language `goal`; the advanced `body` form controls targets, schedule, change-tracking formats, judging, retention, webhook, and notifications. In the simple form, a `goal` is required. If `queries` contains one or more non-empty values and is supplied with `page`/`pages`, `queries` create the search target and page targets are ignored. A monitor schedules future network checks and can send configured email or webhook notifications. Returns the created monitor.
firecrawl_monitor_create
Permanently delete a monitor by ID and stop its future schedule. This operation cannot be undone and returns deletion status.
firecrawl_monitor_delete
Browse Alexandria data providers and workflows or read a selected contract. Alexandria covers companies, people, jobs, finance and filings, public records and government spending, real estate, places and restaurants, retail and prices, package registries and developer data, news, research, and more: typed, sourced records through published contracts. Prefer normal firecrawl_search for a data task; it already returns matching providers. Use this tool when the contract you need was not returned in full, to browse a category when search found nothing, or before scraping the same fields from several pages. Discovery is free. Use query for semantic discovery or urls to find providers for a website. With no arguments, browse categories, then providers and tools. Inspect a selected contract before executing through firecrawl_scrape; reuse contracts already returned. Follow nextTool for further discovery or pagination. Discovery does not execute providers. Use firecrawl_search when you also need web results. Results include a feedbackTool pointer: after the task, report how the catalogue served the website through firecrawl_feedback with endpoint alexandria, including when nothing covered it (free, no job ID).
firecrawl_find_tools
Run web research that returns structured data when the URLs are not known or the answer spans several sites. Describe the fields you need in `prompt`, optionally pass a JSON `schema` and seed `urls`, and the research agent searches, navigates, reads pages, and returns JSON assembled across sources. Use it to research an entity plus its fields (founders, pricing, contact details), to build lists and datasets (companies, people, products, jobs, papers), and for pages that need navigation or interaction to reach the data. This call returns only a job ID, not the research result. Read the job with `firecrawl_agent_status` until it reaches `completed` or `failed`; a typical research run takes one to three minutes. For one known URL use `firecrawl_scrape` (with formats: ["json"] for structured output); for a plain lookup that a results page answers, use `firecrawl_search`.
firecrawl_agent
Retrieve progress or final results for a `firecrawl_agent` job ID. A `processing` response is non-terminal and does not contain the final research result. Check again after 15–30 seconds until the status is `completed` or `failed`; a typical research run takes one to three minutes and complex jobs can take longer. If the job cannot finish within the task's available time, use `firecrawl_search` and `firecrawl_scrape` to complete the requested output. Returns job status, progress information, and result data when completed.
firecrawl_agent_status
Retrieve the current status, progress, and available results for an existing crawl ID. This only reads Firecrawl job state and does not start or modify the crawl.
firecrawl_check_crawl_status
Get the authenticated Firecrawl team's current credit balance or historical credit consumption. Use the current view to answer how many credits remain or to check billing-period boundaries. Use the historical view for monthly usage reporting and trend analysis. The current view returns `remainingCredits`, `planCredits`, `billingPeriodStart`, and `billingPeriodEnd`. Extra purchased or granted credits can make `remainingCredits` greater than `planCredits`. Billing-period dates may be null when the upstream billing provider has no period metadata. The historical view returns periods sorted by start date with `startDate`, `endDate`, and `creditsUsed`. Set `byApiKey` to include separate periods per API key, identified by the optional `apiKey` field. The newest period's `endDate` can be null.
firecrawl_credit_usage
Search an index of public repositories, GitHub issues, merged pull requests, repository READMEs, and code documentation for programming questions that need external documentation or upstream evidence. Returns ranked results with an ID, source type, URL, title, and the matched passages in markdown.
firecrawl_developer_search
Submit concise quality feedback for a completed search, scrape, parse, or map job. Provide the endpoint, job ID, rating, and relevant issue codes or small contextual fields; omit large page contents and raw outputs. For an Alexandria session, set endpoint to `alexandria`, omit jobId, and provide requestedWebsite (url and requestedFunctionality), rationale, and rating. Optional providerFeedback and capabilityFeedback describe gaps or errors. Capability issues: new_capability_request (requires requestedFunctionality), missing_capability, insufficient_functionality, incorrect_result, execution_error, other. Alexandria feedback has no job-age deadline and no credit refund. Returns submission status, feedback ID, and accounting fields.
firecrawl_feedback
Parse one supported document into markdown, HTML, links, summary, targeted answers, or JSON matching a schema. Supported inputs include common HTML, PDF, Word, RTF, OpenDocument, and spreadsheet files; PDF parsing can be bounded with `pdfOptions.maxPages`. Local MCP reads `filePath` from the server filesystem. Hosted MCP uses two calls: first provide `filePath` to receive upload instructions, upload locally, then call again with the returned `uploadRef`; do not send both fields together. Remote web URLs belong in `firecrawl_scrape`. Set `redactPII` to request redaction of personally identifiable information in the returned content. `zeroDataRetention` requires an eligible authenticated account; omit it for anonymous keyless use. Returns upload instructions for hosted phase one or parsed document content for the final call. Authenticated final responses can include a `data.metadata.scrapeId` for optional parse feedback.
firecrawl_parse
Retrieve canonical metadata for one paper ID, such as an arXiv, PMC, PMID, or DOI identifier. Returns the title, abstract, authors, categories, source IDs, and dates as markdown.
firecrawl_research_inspect_paper
Open or reuse a live browser session to navigate a page, click controls, fill fields, or run browser code. Provide either `url` or `scrapeId`, and either a natural-language `prompt` or executable `code`; code can run as Bash, Python, or Node with a bounded timeout. This acts on the live site, so actions such as form submission can create persistent external side effects. Returns execution output, stdout/stderr, exit status, and session viewing URLs.
firecrawl_interact
Retrieve in-body passages from one paper that are relevant to a specific question. Full text is available only for indexed papers; `k` controls the number of passages. Returns matching passages or a notice when full text is unavailable.
firecrawl_research_read_paper
Find citation-graph candidates from one to ten `seed_ids`; the first ID is the primary seed and later IDs are anchors. `mode` defaults to `similar` (co-citation/bibliographic coupling); `citers` returns papers citing a seed and `references` papers cited by a seed. `intent` ranks candidates. Returns ranked candidates and the evaluated pool size.
firecrawl_research_related_papers
Search paper metadata and abstracts with a natural-language query across the indexed corpus, which spans biomedical, life-science, and clinical literature (PubMed, bioRxiv, medRxiv) alongside arXiv and other scientific sources. Optional author, category, and date filters constrain results. Several distinct framings of the same question surface different papers than a single query does. Returns ranked papers with canonical IDs, titles, authors, and abstracts.
firecrawl_research_search_papers
Retrieve and extract content from one supplied URL through Firecrawl. Use this when the request identifies a page and needs its content or defined fields. It can return markdown, HTML, links, screenshots, branding data, a targeted answer, or JSON matching a supplied schema; JSON is useful when the requested result has defined fields, while markdown preserves readable page content. This tool operates on a known page. For a set of pages use `firecrawl_crawl`, and to discover page URLs use `firecrawl_map` or `firecrawl_search`. Options include JavaScript render delay, cache age, main-content filtering, PII redaction, and lockdown cache-only retrieval. Browser actions may change the live page when interactive actions are enabled. Firecrawl may reuse recently indexed content instead of refetching the page, and the reuse window varies by domain. Set `maxAge: 0` to force a live fetch, or a smaller `maxAge` to bound how stale reused content may be. A successful response does not by itself confirm that the state it describes is still current. Returns the selected content formats and page metadata. Authenticated responses can include a `metadata.scrapeId` for optional scrape feedback. On an authenticated session with Alexandria access, if you are about to scrape the same fields from several pages, first run `firecrawl_search` with `sources` unset (or `firecrawl_find_tools`): a matching Alexandria provider returns those fields as typed records in one call. Keyless sessions have no provider matches; scrape directly. Alexandria mode: pass `alexandria` (one `{provider, capability, options}` object or an array of 1-10) instead of `url` to execute catalogued Alexandria capabilities found through `firecrawl_search` sources `alexandria` or `firecrawl_find_tools`. The optional requestId identifies one logical execution: reuse the returned ID for retries of the identical payload, never a new ID to bypass pending or uncertain execution. Each call may include version to pin a published workflow; omitting it uses latest. Only timeout also applies at the top level in this mode. Returns per-capability results in `data.alexandria`, including `data`, `records`, or an `error` with a code. Check each item for errors even when the outer response is successful. Alexandria needs an API key on a team with Alexandria enabled. Alexandria mode: pass `alexandria` (one `{provider, capability, options}` object or an array of 1-10) instead of `url` to execute catalogued Alexandria capabilities found through `firecrawl_search` sources `alexandria` or `firecrawl_find_tools`. The optional requestId identifies one logical execution: reuse the returned ID for retries of the identical payload, never a new ID to bypass pending or uncertain execution. Each call may include version to pin a published workflow; omitting it uses latest. Only timeout also applies at the top level in this mode. Returns per-capability results in `data.alexandria`, including `data`, `records`, or an `error` with a code. Check each item for errors even when the outer response is successful. Alexandria needs an API key on a team with Alexandria enabled. Alexandria results include a `feedbackTool` pointer: after the task, report how the catalogue served the website through `firecrawl_feedback` with endpoint `alexandria` (free, no job ID). For potentially large workflow results, supply and preserve a top-level requestId before execution. If a response provides nextTool, follow its instructions to access the result without repeating a successful provider call. URL mode only: set `domainTools: true` to also return domain-matched Alexandria tools for the page in `tools` on the returned document. Alexandria execution errors relay a `code` and `chargeId`: `request_in_flight` (409) retry the same requestId later; `request_unresolved` (503) keep the requestId for reconciliation, never mint a new one; `duplicate_request` (409) the requestId belongs to a different payload; `unknown_provider` (404), `insufficient_credits` (402), and `billing_unavailable` (503) mean nothing executed. A terms-gated Alexandria provider returns `THIRD_PARTY_DATA_TERMS_REQUIRED` (403) with `requiresAction.url`: follow the returned terms/show and terms/accept calls through this tool, only accepting after explicit user authorization for the reviewed version and digest. An organization admin can alternatively accept at the dashboard URL. Retry only after confirmed acceptance.
firecrawl_scrape
Records schema-validated quality feedback for a prior `firecrawl_search` UUID `searchId`. A `good` rating requires a valuable source, `partial` a valuable source or at least one `missingContent` entry, and `bad` at least one `missingContent` entry or a query suggestion; caps are 50 `valuableSources` and 20 `missingContent` entries. Eligibility is limited to successful searches within the feedback age window. The record is idempotent per search ID. Eligible first feedback for a search can refund 1 credit; refunds are subject to the team's daily cap. The response reports whether a refund was applied, along with submission and daily-cap status.
firecrawl_search_feedback
Start a multi-page crawl at a website URL, poll it to a terminal state, and return the final status and collected data. Scope can be bounded with include/exclude paths, depth, page limit, subdomain/external-link controls, sitemap handling, delay, and scrape options. Crawl results can be large; use conservative limits when full-site coverage is unnecessary. Webhooks and interactive scrape actions are unavailable in safe mode. Returns the crawl ID, status, and page data.
firecrawl_crawl
Search web, news, or image sources and return ranked results. Operators include quoted phrases, `-term`, `site:host`, `inurl:term`, `intitle:term`, and `related:host`; the set is non-exhaustive. `includeDomains` and `excludeDomains` are mutually exclusive hostname filters; categories limit results to GitHub, research, PDF, or developer sources. For a programming question, add `categories: ["developer"]`. It searches an index of repositories, GitHub issues, merged pull requests, repository READMEs, and curated documentation sites, and returns the hits in `data.web` with `category: "developer"`. `categories: ["research"]` restricts these web results to research-affiliated websites and returns page snippets. The `firecrawl_research_*` tools are a separate surface that searches paper abstracts and full text across biomedical (PubMed, bioRxiv, medRxiv) and arXiv literature. Each web result is a title, URL, and description, not the page. Add `scrapeOptions` to attach page content in the same call; those fetches ignore `maxAge`, so use `firecrawl_scrape` when you need a live fetch. Returns source-type result groups and usage metadata. Authenticated responses can include an `id` for optional search feedback. Start with the user’s actual question and constraints. Authenticated search returns matching Alexandria providers beside web results; prefer a provider over page scraping when the task needs the same fields across several entities, provenance, exact figures, or many records, and use web results when they already answer the question. Authenticated search defaults to web + semantic Alexandria tools + domain-matched tools. Keyless search defaults to web only. Alexandria is Firecrawl's catalogue of data providers and workflows across companies, people, jobs, finance and filings, public records and government spending, real estate, places and restaurants, retail and prices, package registries and developer data, news, research, and more; providers return typed, sourced records through published contracts. Coverage varies; discover current tools rather than assuming one exists. Semantic discovery matches the data you need to capabilities even without a provider website in the web results. Domain matching connects result websites to tools that may fetch richer details, related records or collections beyond the linked page. Inspect coverage and required inputs; a matching domain alone does not guarantee a fit. Passing sources without alexandria in it (for example ["web"] or ["news"]) excludes Alexandria provider matches; omit sources unless you specifically need web-only or news-only results, or include "alexandria" alongside them. Use sources: ["alexandria"] for semantic tools only, sources: ["web"] for web only, or sources: ["web"], domainTools: true for web plus domain tools. domainTools: false disables domain matching. Results in data.tools describe available tools, not executed data. Search defaults to toolDetail: "compact", returning only provider, capability and description; "summary" adds metadata and navigation. Inspect selected tools with firecrawl_find_tools using their providers and capabilities plus expand:["options","response"]; request examples only when the input shape is unclear. Batch related contract inspections and reuse complete contracts. Alternatively set toolDetail: "full" for contracts upfront. On the full MCP surface, firecrawl_find_tools supports semantic query lookup and is the list equivalent: {} lists categories; {categories:["<category-id>"]} lists providers; {providers:["<provider-id>"]} lists compact tools; adding capabilities:["<capability-id>"] expands the selected contract. Avoid expanding the entire catalogue. Read required inputs and requiresOneOf groups (at least one member per group), example.request/example.response when present, and response.key in the selected contract. Do not assume records is the result key. Follow the declared pagination input and response cursor, preserving filters; catalogue next is separate from provider pagination. Execute through firecrawl_scrape with alexandria:{provider,capability,options}. Search scrapeOptions fetches web pages, never provider tools. Use web results when sufficient, tools when they offer a direct route to deeper data. After an Alexandria task, whether a capability ran or discovery found nothing for the website, call firecrawl_feedback once per website with endpoint "alexandria": requestedWebsite (the url the user needed data from and the requestedFunctionality they needed), a rating, a rationale from observed results, and any providerFeedback or capabilityFeedback gaps. It is free, needs no job ID, has no deadline, and follows the answer rather than replacing it.
firecrawl_search
Enumerate URLs indexed under one website through Firecrawl without fetching each page's content. Use this when the request asks for a site's URL inventory, when several relevant pages must be located, or when the desired page URL is unknown. An optional `search` term narrows the URL list, while sitemap, subdomain, query-parameter, and result-limit options control coverage. Returns matching URLs rather than page bodies. Retrieve one page with `firecrawl_scrape`; collect content across multiple pages with `firecrawl_crawl`. Authenticated responses can include an `id` for optional map feedback.
firecrawl_map
Retrieve one monitor by ID, including its configuration and current state. This does not run or modify the monitor.
firecrawl_monitor_get
Retrieve one monitor check and its page-level results, optionally filtered by page status. Pages report `same`, `new`, `changed`, `removed`, or `error`; configured goal judging can add a meaningful-change decision. Markdown tracking returns a unified text diff, JSON tracking returns field paths with previous/current values and a current snapshot, and mixed tracking returns both. Returns one page of results plus a `next` URL when more pages exist.
firecrawl_monitor_check
List historical checks for a monitor, optionally filtered by status and bounded by a result limit. Returns one page of check summaries and pagination metadata.
firecrawl_monitor_checks
List monitors for the authenticated account with optional pagination controls. Returns one page of monitor records and pagination metadata.
firecrawl_monitor_list
Queue an immediate check for a monitor outside its normal schedule. This starts network work for the monitor's configured targets and returns the queued check.
firecrawl_monitor_run
Stop the live interact session associated with a `scrapeId` and release its resources. Returns a success confirmation.
firecrawl_interact_stop
Patch an existing monitor by ID. The body can change its name, active/paused status, schedule, targets, goal, judging, webhook, notifications, or retention; these changes affect future scheduled checks. Returns the updated monitor.
firecrawl_monitor_update
How do I improve a ChatGPT Plugin's discoverability?
The levers are the listing surface agents actually read: names, descriptions, keywords, tool metadata, and registry health. Which lever matters depends on where discovery breaks, which is what continuous measurement shows.
What are Firecrawl alternatives on ChatGPT?
As of 2026-09-28, Firecrawl competes with Answer Images, Context.dev, DataBlue, Deep research, Exa, Keenable, Keenable SELECT, Lifestyle Helper, Nimble, Nyambot, Octen, Olostep, Page2AI, Parallel Search, Riveter, Search Service, smry, Tako, Tavily, TinyFish, Visualping in ChatGPT Web Search, Crawling & Extraction for Agents, ranked by public Discoverability Score.
Where is this profile measured?
This profile uses the geography attached to the latest public registry snapshot: US. Locale tags are intentionally omitted.