Attributions
Last updated: 2026-02-01
Portions of our domain discovery and crawl scheduling inputs are derived from publicly available datasets and resources. We provide attribution to the creators and maintainers below. These references help explain where public discovery signals come from; use of these resources does not imply endorsement.
Tranco Top Sites List
Used as a seed/discovery signal for crawling and verification.
- Project: https://tranco-list.eu/
- License: CC BY 4.0
Common Crawl
Used as a discovery signal (e.g., extracting hostnames/domains observed in public crawl indexes).
- Project: https://commoncrawl.org/
- License: CC BY 4.0
Cisco Umbrella Popularity List
Used as a seed/discovery signal for crawling and verification.
- Info page: https://s3-us-west-1.amazonaws.com/umbrella-static/index.html
- Terms/license: See the provider's page above.
Notes
- This page covers attribution for dataset inputs used in domain discovery and crawl scheduling.
- We honor website instructions (e.g., robots.txt and relevant HTML directives) during crawling where applicable.
- No endorsement by data providers is implied.
Attribution questions: info@kommento.app
Crawler opt-out or abuse reports: abuse@kommento.app
Phone: 1-866-KMT-SCAN