Add website sources
Crawl a website, select discovered pages, add a single URL, and diagnose pages that cannot be indexed.
Crawl a website
Open Website sources
Go to Training Data → Website and enter the full public origin, including https://.
Discover and filter pages
Start the crawl, review discovered links, and filter the list. Select only pages that contain reliable customer-facing information.
Add selected pages
Submit the selection. Each URL is processed separately and receives its own status.
Verify the result
Wait for Indexed, preview extracted text, and ask a question that requires the page.
Use the individual URL option for a single page that is not linked from the main site.
Manage crawled pages
Below the crawl form, the Websites table shows one row per crawled site with its page counts; removing a website deletes all of its pages from the agent's knowledge. The Crawled Pages table lists every page with its status. Search by title or address, filter by status (for example only failed pages), retry a failed page, open the original page, or delete a page you no longer want the agent to use.
Pages with very little readable text can fail. Fetchply rejects extracted pages with fewer than about ten words.
Login walls, robots rules, bot protection, client-only rendering, redirects, and network errors can prevent extraction. Use a file or text source when the public page cannot be read safely.