Block 1 - Load the xml file as JSON
- Type / Role
- n8n-nodes-base.extractFromFile - extractFromFile
- Config choices
- Version 1
This workflow is provided as-is. Please review and test before using in production.
This workflow contains community nodes that are only compatible with the self hosted version of n8n. This workflow allows users to extract site...
n8n-nodes-base.extractfromfile, n8n-nodes-base.if, n8n-nodes-base.webhook, n8n-nodes-scrapingbee.scrapingbee, n8n-nodes-base.compression, n8n-nodes-base.code, n8n-nodes-base.googlesheets, n8n-nodes-base.stickynote
This workflow is cataloged by N8N Workflows and links back to its original n8n.io source page by Sahil Sunny.
Original n8n.io sourceThis workflow contains community nodes that are only compatible with the self-hosted version of n8n.
This workflow allows users to extract sitemap links using ScrapingBee API. It only needs the domain name www.example.com and it automatically checks robots.txt and sitemap.xml to find the links. It is also designed to recursively run the workflow when new .xml links are found while scraping the sitemap.
domain=www.example.comrobots.txt file, if not found it checks sitemap.xmlWhen the workflow is finished, you will see the output in the links column of the Google Sheet that we added to the workflow.
links. Connect to the sheet by signing in using your Google Credential and add the link to your sheet.domain as query parameter. Example:curl "https://webhook_link?domain=scrapingbee.com"
Scrape robots.txt file, Scrape sitemap.xml file, and Scrape xml file nodes.Append links to sheet node with a relevant node.If you wish to scrape the pages using the extracted links, then you can implement a new workflow that reads the sheet or file (output generated by this workflow) for links and for each link send a request to ScrapingBee's HTML API and save the returned data.
NOTE: Some heavy sitemaps could result in a crash if the workflow consumes more memory than what is available in your n8n plan or self-hosted system. If this happens, we would recommend you to either upgrade your plan or use a self-hosted solution with a higher memory.
This catalog entry is organized from the workflow JSON. The node-level section below shows the executable blocks available for review before importing the template.
| Workflow | ScrapingBee and Google Sheets integration template |
|---|---|
| Complexity | advanced |
| Nodes | 22 |
| Categories | Market Research |
| Author | Sahil Sunny |
| Published | 27 Aug 2025 |
Use the JSON export at /data/workflows/7927/7927.json as the source template for this automation.
Open n8n, import the downloaded JSON, and review each node before activating the workflow.
Replace placeholder credentials, API keys, webhook URLs, account IDs, and environment-specific values with your own settings.
Run the workflow manually or in a staging workspace, inspect node output, and confirm downstream systems receive the expected data.
Enable the workflow only after testing, then monitor executions, errors, and rate limits during the first production runs.
Review imported nodes carefully before activation. This catalog entry is intended to help you inspect the workflow structure, understand required services, and find related templates faster.
Node names, credentials, schedules, webhook paths, and external service limits may need adjustment for your workspace.
This workflow contains community nodes that are only compatible with the self hosted version of n8n. This workflow allows users to extract site...
Review the workflow JSON, configure any required credentials in n8n, and test the automation in a safe workspace before using it in production.
Yes. Use the block-by-block analysis and the downloadable JSON to inspect each node, then adjust credentials, prompts, schedules, filters, or destinations for your Market Research use case.