Core Differentiators
Three capabilities buyers value most
Built-In Governance Signals — Quality, Safety, And Provenance In Every Response
Every API result includes five pre-computed quality and ranking signals, two NSFW moderation scores, and full crawl provenance — no separate pipeline required.
Why this matters for you
- Product teams: Ship search and retrieval features with quality signals and safety controls already built in. No post-processing layer to build.
- Legal and compliance: Governance data is baked into every record — provenance, safety score, and source are auditable at the row level.
- Enterprise procurement: Answer “how do you handle content safety and data origin?” with a concrete, per-record answer rather than a policy statement.
Most search APIs return title, URL, and a snippet. Quality scoring, NSFW moderation, and crawl provenance are left to the customer to build. We include all three by default.
Built-In NSFW Moderation On Every Record
Every document carries two content safety scores: a raw NSFW signal and a post-processed model score. Both are returned in every API response.
Why this matters for you
- Consumer AI products: Filter unsafe content at query time using a threshold — no separate classification pass required.
- Enterprise and regulated teams: Document your moderation layer as part of data provenance. Auditors and legal teams can verify it.
- Dataset curators: Build clean retrieval sets by excluding records above an NSFW threshold. One field, zero extra tooling.
NSFW moderation at the raw index level is rare. Most providers expect you to run your own classification pipeline after retrieval. We do it upstream.
Crawl Provenance On Every Document
Every record includes the source URL, parent domain, and the exact timestamp when the page was crawled and indexed.
Why this matters for you
- AI model teams: Document retrieval data origin for audit, licensing disclosure, and legal review without guesswork.
- Enterprise procurement: Answer "where did this data come from?" with a precise, auditable answer per record — not a general policy statement.
- Compliance teams: Propagate takedown and deletion requests accurately because you know exactly which records came from which source.
Aggregators that blend data from multiple providers often cannot give you per-record origin. We can, because the data comes from our own crawler only.
Supporting Data
Quality signals that save you weeks of engineering
Importance Score
A measure of the page's structural and content-level authority within the index. Use it to surface high-value documents and filter out thin or low-signal pages.
Quality Score
Content quality rating derived from page-level signals. Lets you exclude low-quality content from retrieval sets or rank results by document quality rather than just keyword match.
Freshness Score
Temporal relevance signal based on crawl recency and content change patterns. Critical for news monitoring, trend detection, and time-sensitive retrieval applications.
Popularity Score
Observed engagement and reach signal. Useful for weighting results toward pages that real users visit rather than technically indexed but rarely seen content.
PageRank
Link-graph derived authority signal. Standard for filtering spam and thin domains and for ranking documents by the authority of their source within the web graph.
Why pre-computed matters
Building your own scoring pipeline over raw web data takes weeks. We run these signals at index time so you receive a production-ready, multi-signal ranked dataset on day one.
Scale And Collection
An independent index at a scale that makes it useful
2.3B+
Documents indexed
One of the largest independently maintained web indexes available via API. Deep enough to support benchmark construction, evaluation datasets, and broad retrieval tasks.
First-party
Crawler only
Data comes exclusively from our own search crawler — not resold from Google, Bing, or any third-party provider. Independent collection means independent provenance.
English
Initial language coverage
Current index focuses on English-language content with language detection on every record. Multilingual expansion is on the roadmap.
Comparison
How we compare
| Feature | timpi.data | Brave / Exa / Tavily | SerpAPI | Common Crawl |
|---|---|---|---|---|
| Pre-computed quality signals | 5 signals | Partial | No | No |
| Built-in NSFW moderation | Raw + post-processed | No | No | No |
| Per-record crawl provenance | URL + domain + timestamp | URL only | URL only | Partial |
| Independent crawl (not resold) | Yes | Yes (some) | No — aggregates engines | Yes |
| Publisher takedown workflow | Documented spec | Ad hoc | Ad hoc | Partial |
Competitor data is based on publicly documented features. Verify against each provider's current terms before making procurement decisions.
Ready to see the data?
Run a live query against our 2.3 billion document index — no signup required.