📊
Product Database Directory

AI Data Directory

Tools for web scraping, data extraction, and AI data pipelines. Browse all 8 cataloged software options, compare specifications, and find the right tool for your workflow.

8 Total Tools
🏆 EDITORIAL RANKINGS

Looking for our top recommendations?

View our strictly curated Top 9 Best AI Data Tools ranking for 2026, reviewed by lead editor Bill.

View Top 9 Ranking →
Sorted by: Editorial Rating
AI Data

Crawl4AI

4.3

Open-source web crawler designed for LLMs and AI agents with structured extraction and browser automation.

#LLM-first output format #Built-in browser automation with anti-bot support #Structured data extraction via LLM-guided parsing
AI Data

Atlas

4.3

Open-source local-first cognitive memory system implementing AGM-compatible belief revision that automatically re-evaluates downstream beliefs when facts change, with SHA-256 hash chain for data integrity.

#Highly secure & local-first #Boosts workflow efficiency #User-friendly interface
AI Data

sqlsure

4.3

A deterministic SQL semantic inspector that catches silently-wrong AI-generated queries — double-counting, bad joins, exposed PII — in about 0.1 ms before they run. Works as a CI gate, an MCP server, or a library.

#Deterministic semantic checks — catches double-counting, wrong joins, and exposed PII #Three doors: CI gate, MCP server, and embeddable library #Judges SQL against facts from dbt tests, PK/FK declarations, or live DB introspection
AI Data

Adaptive Recall

4.0

Adaptive Recall is a hosted memory system for AI applications that goes far beyond simple vector search. It stores, recalls, and manages long-term memory for agents and apps over MCP or a plain REST API, and — unlike a static embeddings store — it actively learns. Four retrieval strategies run in parallel (vector similarity, temporal recency, full-text keyword, and knowledge-graph traversal), and the system learns which to prioritize for each query type. Results are ranked with ACT-R cognitive scoring from 30 years of cognitive-science research, factoring in recency, access frequency, entity connections, and validated confidence. A knowledge graph is built automatically from stored memories, memories move through a confidence-based lifecycle and fade when unused, and an ML pipeline trains on your usage patterns — validating every parameter change against real query history before adopting it. A simple eight-tool API (store, recall, update, forget, graph, status, snapshot, feedback) covers everything, with Bearer-token auth and JSON in/out. Free, Starter, Pro, and Business plans are available.

#Four retrieval strategies learned per query #ACT-R cognitive scoring surfaces the right memory #Automatic knowledge graph from stored memories
AI Data

Polygres

4.0

Postgres extensions for retrieval-augmented generation: graph search, multi-path hybrid retrieval, and token-budgeted context assembly, self-hosted or managed.

#pgGraph 图检索引擎 #pgContext 十路融合检索 #混合检索 + 上下文组装
AI Data

TamedTable

3.9

TamedTable is an AI ETL tool you drive with natural language. Load a CSV, JSONL, Parquet, or Arrow file, type 'normalize phone numbers' or 'drop duplicate emails', and the LLM writes a JSON spec that transforms the data — with a 96.7% label-match benchmark at about $0.15 per 1,000 rows. It cleans, enriches, classifies, validates, and translates; every change saves as a replayable recipe or exportable Python script. Source-available, runs on your own API keys (BYOK).

#no-code data prep #replayable and exportable #multi-format
AI Data

MothRAG

3.8

An open-source RAG framework (Apache 2.0) that hits research-SOTA parity on multi-hop QA benchmarks using only commodity LLM APIs — no GPU, no training, no graph rebuild.

#SOTA parity on multi-hop benchmarks, no GPU/training #Deterministic orchestration, zero run variance #Graph-free: no expensive rebuild on corpus change
AI Data

ParseHawk

3.5

ParseHawk is a fully local document AI processing toolkit — no data leaves your machine. It ships with an API server, CLI, and Web UI, making it easy to integrate into existing workflows or use standalone for document parsing, chunking, OCR, and Q&A over documents.

#100% Local Processing #Multi-Interface Support #Document Format Support

❓ Frequently Asked Questions

How many ai data tools are listed in this directory?

+

Our database currently tracks 8 ai data tools and platforms, covering free, freemium, open-source, and commercial solutions.

How do I find the best ai data tools?

+

For quick decision-making, see our curated Top 9 rankings at /best/ai-data, where lead editor Bill selects and ranks the 9 best-performing options based on real-world reliability and value.