Where data agents
are built and scaled.
Factory is the platform underneath Radar and Benchmark. Describe the data you need in plain English — an agent builds the pipeline, runs it on a schedule, and rebuilds it when the web changes.
Just want the data, not the machinery? See Radar and Benchmark.
Home / Build
What do you want to build?
Pipelines
5Datasets
5Recent runs
Home / Understand
What do you want to know?
Sources
12Datasets
5Recent alerts
3Clear allFrom a sentence to a running pipeline.
No scraping scripts, no selectors, no maintenance rota. You describe the outcome — Factory's agents do the engineering.
You
Hi Jason! I want to collect the prices of all my products from all around the web, every day. I guess probably 100k rows across a dozen sites. I want to collect that, turn it into a dashboard and newsletter for me. Keep it up to date automatically.
Can you help?
Preparing the smallest useful catalog handoff.
Yes — that volume and the twelve-site shape are feasible. I can build the daily collection, canonical product matching, price history, dashboard, alerts, and newsletter.
Took 48s
You
Here's our current product catalog. It has the SKUs, UPCs, priority retailers, and a handful of known product URLs.
Please go ahead, schedule it daily, and fix routine site changes automatically when you can.
Attached: product_catalog.csv
Running final checks and organizing the workspace links.
Task completed
Market Radar is live — 12 retailers, 5 pipelines, 5 datasets, dashboard, newsletter, and alerting are ready.
Open dashboardDone — I built the Market Radar system from your catalog and verified the first daily run. Here's everything:
- Live product offers — the canonical 8.4M-row product and offer view
- Price observations and inventory signals
- Competitive Market Pulse — the dashboard for coverage and price movement
- Daily Market Brief — scheduled every morning at 07:00
- Five daily pipelines with routine site-change recovery and independent validation
Let me know if you need anything else, or if you’d like me to help analyse the data.
Took 1h 3m 42s
Ask Jason to build, inspect, or explain…
Describe the goal
"Track competitor pricing across 40 grocery sites." Plain English is the whole spec — sources, schema, and cadence are inferred or asked for.
An agent team builds it
In an isolated sandbox, the agent probes sources, picks fetch strategies, locks a schema, and assembles a multi-step pipeline — then validates it end to end.
It runs on schedule
Daily, weekly, or on demand. Datasets, dashboards, and alerts are delivered to Slack, Snowflake, PowerBI, Excel, or your API — and shared via live links.
The machinery under every dataset.
Everything Radar monitors and Benchmark simulates is produced by this stack — built once, then run and repaired automatically.
Agent-built pipelines
An LLM agent plans the extraction, writes the pipeline, and tests it against real pages — the same work a data engineer would do, in minutes.
A fetch chain that learns
10+ backends, from plain HTTP to full browser automation. Factory remembers what works per domain and starts there next time.
Sandboxed execution
Every agent and every pipeline runs in an isolated sandbox — no shared state, no access beyond the job it was given.
Cell-level provenance
Every value in every dataset traces back to the page it came from. Click a cell, see the source — no "trust us" data.
Self-healing runs
When a site redesign breaks a step, a healing agent rebuilds just that step and retries the run — before anyone notices.
Scheduled delivery
Dashboards, alerts, newsletters, and exports — Slack, Snowflake, PowerBI, Excel, CSV, or API — on the cadence you choose.
The web changes. Factory notices first.
Scrapers rot — that's why most teams give up on them. Factory treats a broken step as a job for another agent, not a ticket for your engineers.
Every factory needs a foreman.
Jason is the copilot inside every Factory workspace. Ask him anything about your data and he answers from the live datasets — then acts on what you decide, right in the chat.
-
Answers from your live data
"Why did prices drop yesterday?" gets a real answer, queried from the dataset — not a canned reply.
-
Builds and edits pipelines
"Track two more retailers" or "add a promo column" — Jason changes the pipeline, validates it with a sample build, and versions every edit so you can revert.
-
Runs the delivery side too
Dashboard widgets, alerts, newsletters, scheduled summaries — and when a build or rerun finishes, he reports back in the same thread, in the app or in Slack.
Jason
watching your workspace
Ask Jason about your data…
answers · builds · delivers
Inspectable by design.
You don't have to look inside — but you always can. Every pipeline Factory builds is real, versioned, testable software.
Code is the source of truth
Factory generates Python pipelines — not opaque configuration. Every pipeline is readable, diffable, and reviewable, like any other code your team ships.
Skills: extraction knowledge, versioned
Extraction logic is packaged as skills — versioned, installable units with typed input/output schemas, tests, and fixtures, stored in a registry. Skills are domain-aware and reusable, so every pipeline in a domain gets smarter than the last.
Every pipeline is a DAG
Pipelines render as inspectable DAGs — sources, extraction steps, transforms, and outputs, visible end to end. When something breaks, you can see exactly where.
Typed schemas, validated every run
Outputs are typed and validated against their schema on every run. Anomalies get flagged, accuracy gets tracked — 98.7% extraction accuracy across complex web & app sources.
Sandboxed execution, escalating fetch
Pipelines execute in isolated sandboxes. The fetch layer escalates automatically — plain HTML, headless browser, and beyond, across 10+ backends — until the data flows.
One factory, two products
Every Radar dataset and every Benchmark table is a Factory pipeline underneath. Most teams meet Factory through one of them.
Jsonify Radar
Radar continuously scans menus, retailers, and e-commerce websites and apps to collect product, price and promotion data at scale.
great for: e-commerce, f&b, retail
Learn more →Jsonify Benchmark
Benchmark simulates real customer journeys by automatically filling live quote flows on competitor websites.
great for: insurance, ISP, mobile
Learn more →What data do you need?
Competitor Pricing ›
retailTrack product prices across retailers. Compare discounts and promotions.
Insurance Quotes ›
insuranceCompare insurance quotes across providers. Track premiums by persona.
Restaurant Menus ›
foodMonitor menu prices across restaurant chains. Track items and placement.
ISP & Telecom Plans ›
telecomCompare internet service plans across providers. Track speeds and pricing.
Property Listings ›
real estateCompare property listings across cities. Track prices and availability.
Custom ›
from scratchDescribe what you need and we'll build it.
Who Factory is built for.
Built for people who read the code before they trust the output.
Engineers who've maintained scrapers
You've written the XPath, watched it break, and rewritten it. Factory generates the pipeline, watches the source, and repairs itself — you review DAGs and diffs instead of firefighting.
Technical evaluators
Your strategy team is already talking to us. This page is your diligence stop: generated code, typed schemas, versioned skills, sandboxed execution, DAG-level observability. Ask us anything — or book a call and we'll walk the architecture.
Platform teams building on structured web data
You need external data as a dependable input, not a side project. Typed, validated, scheduled outputs delivered into your warehouse — with the pipeline itself inspectable when you need it.
Run it yourself, or let us run it for you.
Demo access
Free
Build pipelines and run them on synthetic demo data. Free to explore — no card required.
Self-serve
Usage-based
Pay per pipeline run and rows extracted. You build, you operate, you export. Scales with what you actually use.
Managed
Custom
We build, run, monitor, and guarantee your pipelines. SLAs, quality validation on every run, dedicated infrastructure, delivery straight into your stack.
Book a call →Self-serve tiers cover pipelines you operate yourself. Managed datasets — built, monitored, and quality-guaranteed by Jsonify — are scoped per engagement.
SOC 2 Type 2 · GDPR compliant · Independent controller model — see legal FAQ
FAQ
What about personal data and GDPR?
Can I see what a pipeline actually does?
Watch Factory build your first pipeline.
Describe the data you need and see a pipeline assembled live — runs instantly with demo data, no setup required.