Privacy Policy
Version 2.0 · effective 2026-05-23. GDPR-aligned. Written to be understood without a lawyer; if anything is unclear, email support@probily.tech — we will explain in normal language.
1. Who we are
Probily («we», «us», «the Operator») — the brand and website at probily.tech — is the data controller for personal data processed by the Service. Contact: support@probily.tech.
We have not appointed a separate Data Protection Officer because we do not meet the GDPR Article 37 thresholds. The Operator handles all DPO functions personally. If you would prefer to escalate a privacy concern, you may also lodge a complaint with your national data-protection authority.
2. What we collect
The Service collects data in the categories below. We collect only what we need to run the Service; we do not collect anything for marketing or for sale.
2.1. Account data
- Your email address (required for sign-up).
- A salted hash of your password, managed by SuperTokens (we never see your plaintext password).
- Your account creation timestamp + last sign-in timestamp.
- A session cookie (essential, set on sign-in, expires on sign-out or after 30 days of inactivity).
2.2. Inputs & reports you generate
- The article URL or query you submit for a report.
- The structured data and rendered HTML of every report.
- The case ID and signed share-token of each report.
- Timestamps of generation and last access.
- The Wayback Machine snapshot URL + a SHA-256 hash of the HTML the report was built from (the captured-source archive that lets readers verify what the article said at report-time).
2.3. Project data
- Project name, description, colour swatch you choose.
- Entities you add, with the role classification (main / supporting) and any manual notes you attach.
- Manual relations you draw between entities.
- Markdown notes attached to a project or to an individual entity within a project.
2.4. Contribution data
- The text of any correction or addition you submit to the global entity archive (with your handle attached for public attribution if approved).
- The citation URL you provide for each contribution.
- The moderation decision (approve / reject / revise) and the reviewer's notes.
2.5. Public profile data (opt-in)
- If you turn on the public profile: your username, display name, bio, and the public links you list (Twitter, LinkedIn, website — all optional).
- By default profiles are NOT public.
2.6. Operational telemetry
- HTTP access logs (request path, response code, IP, duration). Retained for 30 days for security; aggregated for longer.
- Pipeline structured logs (which adapter returned what, which fetch failed). No personal data beyond what the request contains.
- Container health metrics. No personal data.
2.7. Billing data (when paid tiers launch)
- Paddle Customer ID + subscription state. We do NOT store your card number or any other payment-method detail. Paddle is our merchant of record and payment processor; see paddle.com/legal/privacy.
- Itemised billing-ledger entries (top-ups, debits, adjustments) for accounting.
We do not collect: location, contacts, device fingerprints, browsing history outside the Service, any analytics from third-party trackers (we do not use Google Analytics, Facebook Pixel, or any similar service).
2.5. Personal data about people named in reports (third parties)
A report may contain personal data about individuals and organisations named in the article you submit. This data is not obtained from those individuals — it comes from the news article you provide and from public, open-source records (Wikipedia / Wikidata, official sanctions and watch-lists, court and regulatory dockets, corporate registries). Where such an individual is an EU/UK data subject, this is personal data obtained from a source other than the data subject within the meaning of GDPR Article 14.
We process it under our legitimate interests (Art. 6(1)(f)) in supporting journalism, research, fact-checking and due diligence; and, for data revealing criminal offences or special categories, under the substantial-public-interest condition (Art. 9(2)(g) / Art. 10 GDPR and the corresponding national provisions). We surface only information already in the public record, always linked to its original public source. Because reports are generated on demand from public sources and may concern a large, unbounded number of people, individually notifying every data subject would involve disproportionate effort; the Article 14 exemption for disproportionate effort therefore applies, and this Policy serves as the public notice of that processing. A named person may still exercise the rights in §6 (including objection and erasure) by contacting us.
3. Legal basis for processing
Under GDPR Article 6 we rely on the following lawful bases:
- Contract (Art. 6(1)(b)) — for account data, your inputs, reports, projects, contributions. Without this data we cannot provide the Service you requested.
- Legitimate interest (Art. 6(1)(f)) — for operational telemetry (security, fraud prevention, rate-limit enforcement, debugging).
- Consent (Art. 6(1)(a)) — for the public-profile opt-in and for product-update emails (off by default).
- Legal obligation (Art. 6(1)(c)) — where we must retain billing records for tax/audit, or respond to a lawful court order.
4. Who we share data with
Probily does NOT sell your data, ever. We share with:
- Internet Archive (Wayback Machine). When you submit a URL for a report we ask Wayback to save a snapshot of it. The URL is sent to Wayback; nothing about you is. The snapshot becomes part of Internet Archive's permanent record.
- Public databases queried by L0. Wikipedia, Wikidata, OpenSanctions, CourtListener, OpenAlex, Crossref, PubMed, GDELT, ICIJ Offshore Leaks, ECHR HUDOC, NACP, GLEIF, SEC EDGAR, Companies House, and a few more. We send these APIs the query terms only (typically a name or a QID); nothing about you is sent.
- Anthropic API. Article text is sent to Anthropic to drive the entity-extraction pass. Anthropic states it does not retain inputs sent via the API for training purposes.
- Paddle. Billing operations route through Paddle, our merchant of record.
- Authorities. Only where required by a lawful court order or a binding legal obligation (e.g. a mutual legal-assistance request).
We do not share with marketing partners, analytics vendors, advertising networks, or data brokers.
5. Cookies
Probily uses one essential cookie set by SuperTokens to keep you signed in. We do not use tracking, advertising, or analytics cookies. We do not display a cookie banner because GDPR does not require consent for strictly-necessary cookies.
6. Your rights
Under GDPR you have the right to:
- Access — ask for a copy of the personal data we hold about you.
- Rectify — correct inaccurate data.
You can edit your profile in
/cabinet/settings; contact us for anything else. - Erase (right to be forgotten) — delete your account and private data. See §7 below for what we keep as historical record vs delete in full.
- Restrict — pause processing while a complaint is investigated.
- Object — to processing based on legitimate interest; we will weigh your rights against our interest.
- Portability — receive your data in a structured, commonly-used format (JSON), available on request by email.
- Withdraw consent — for any processing based on consent. Withdrawing does not affect the lawfulness of prior processing.
- Lodge a complaint — with your national data-protection authority.
To exercise any right, email support@probily.tech. We respond within 30 days (Art. 12(3)).
7. Data retention
| Data | Retention |
|---|---|
| Account email + auth hash | Kept while your account is active; deleted on request |
| Private reports + projects + notes | Kept while your account is active; projects deletable in your workspace, otherwise deleted on request |
| Public reports + approved contributions | Indefinite (historical record). You may request retraction; we evaluate per case |
| Captured-source snapshots | Wayback's retention (independent of us) |
| HTTP access logs | 30 days; aggregated metrics longer |
| Billing records (when launched) | 10 years (tax/audit requirement) |
| Backup snapshots | 90 days rolling; deleted data clears on next rotation |
8. Children
The Service is not directed at children under 16 and we do not knowingly collect personal data from anyone under 16. If you become aware that a child has provided us with personal data, please contact support@probily.tech and we will delete it.
9. International transfers
Probily's servers are hosted in the EU. The processors we use (Anthropic, Paddle) are based in the United States / United Kingdom and rely on the EU Standard Contractual Clauses or equivalent safeguards for any EU-to-US data transfer. Internet Archive is based in the US and operates as a non-profit information service; the URL you submit is the only transferred information.
10. Security
We protect data in transit with TLS (Let's Encrypt) and at rest with database encryption (PostgreSQL with disk-level encryption). Passwords are bcrypt-hashed by SuperTokens. Session cookies are HTTP-only, Secure, SameSite=Lax. Backups are encrypted. We log security events and review them weekly.
If we become aware of a personal-data breach affecting your data, we will notify you within 72 hours (GDPR Art. 34) and describe what happened, what data was affected, and what steps we are taking.
11. Changes to this Policy
Material changes will be posted here with a new effective date and a changelog. We will email active accounts at least 30 days before changes take effect. Continued use of the Service after a change constitutes acceptance.
12. Contact
Data controller: Probily. Privacy questions: support@probily.tech. Postal address available on request.
Changelog
- v2.0 · 2026-05-23 — Full GDPR-aligned rewrite for the cabinet era. New sections: legal basis (Art. 6), explicit data-rights chapter, retention table, breach notification commitment, processors list.
- v1.0 · 2026-05-18 — Initial release for the Telegram bot.