Ryota Labs

AnkiAI Capture: from any page to a deck you can trust.

Photograph a textbook page, drop in a lecture PDF, or paste notes. Capture writes flashcards, ties every card to the sentence it came from, and hands them to AnkiAI's FSRS scheduler only after you approve them.

What learners get

Four card types chosen per concept: question and answer, cloze deletion for terms and formulas, reversible cards for vocabulary, and reading cards that drill kanji readings with furigana.

Every card has a source button. Tapping it opens the original page with the supporting sentence highlighted, so a learner can settle any doubt in one second instead of trusting the AI blindly.

A review tray before saving: accept, edit, or reject each card. Rejections ask for a one-tap reason (too easy, ambiguous, wrong, duplicate), and those reasons shape the next deck.

During study, a premium "Why was I wrong?" button explains the answer using only the cited source, with the page reference attached.

Large textbooks can be queued as an overnight build that is ready the next morning at lower cost.

Pipeline

  1. IngestCamera pages come in through VisionKit's document scanner, PDFs through the share sheet or Files app. Each source is uploaded once to the Claude Files API and referenced by file_id for every later pass.
  2. Transcribe imagesPhotos and handwritten notes go through Claude's vision input and become a clean plain-text document, page by page. PDFs skip this step because Claude reads their text and page images natively.
  3. GroundThe source is sent as a document block with citations enabled. Claude lists the facts worth memorizing, and the API attaches the exact supporting spans: page ranges for PDFs, character ranges for transcribed text.
  4. GenerateA second call turns the grounded facts into cards using structured outputs, so every response matches the card schema below and parses without a retry loop.
  5. VerifyA low-effort pass checks each card against its cited span and drops anything not answerable from the source, any duplicate of an existing card, and any front that admits two answers.
  6. Review and scheduleApproved cards enter FSRS. Cards from the same fact are linked as siblings so they are never shown on the same day.

Why Claude

Generation quality is the product. A card that is wrong, vague, or untraceable costs a learner more than no card at all. These are the specific Claude capabilities the design depends on.

Native PDF understandingClaude processes both the extracted text and the rendered page images of a PDF. Lecture slides put meaning in diagrams, tables, and layout, so a text-only OCR pipeline loses it. One API call replaces an OCR stage, a layout parser, and a captioning model.
Citations API for groundingWith citations enabled, PDFs are chunked into sentences and responses carry citation blocks with page ranges. The source anchor on each card comes from the API, not from text the model writes, so a card can never point to a quote that does not exist. This is the core trust feature of Capture.
Structured outputs for the card schemaJSON outputs constrained by a schema guarantee every card has a valid type, front, back, and source reference. Citations and structured outputs cannot share one request, which is why grounding and generation are separate passes joined by span IDs.
Vision on mixed Japanese and EnglishJapanese study material mixes kanji, kana, Latin script, formulas, and handwriting on the same page. Claude transcribes these in a single pass, and the transcript becomes a citable text document so photographed pages get source anchors too (images embedded in PDFs cannot be cited directly).
Long context across a whole documentAn entire lecture PDF fits in one request, so a term defined on page 3 and applied on page 40 becomes one concept with linked cards instead of two disconnected ones.
Prompt cachingThe source document and the system prompt with the learner's style preferences are cached once and reused across the ground, generate, verify, and "make more cards" calls, cutting both latency and input cost on every follow-up.
Message Batches APIOvernight builds for full textbooks run through the Batches API at half the price of standard calls, which keeps a large deck inside the subscription's margin.
Model and effort routingHaiku handles transcription and simple vocabulary decks; Sonnet handles concept extraction and card writing. The effort parameter is set low for verification and higher for concept-heavy material, so cost follows difficulty.

Card schema

{
  "cards": [{
    "type": "basic | cloze | reverse | reading",
    "front": "string",
    "back": "string",
    "concept_id": "string",        // siblings share this; FSRS buries them
    "source_span_ids": ["s12"],    // joins to citations from the Ground pass
    "difficulty_hint": 1,          // 1-5, initial ordering only; FSRS owns scheduling
    "tags": ["string"]
  }]
}

Card-writing rules in the system prompt

One fact per card. Fronts must have exactly one correct answer. Prefer cloze for terms and formulas, reverse for vocabulary pairs, reading cards for kanji. Never introduce facts absent from the cited span. Match the learner's language and target exam level.

Security, privacy, and cost control

ConcernDesign
API key exposureNo key in the app. Calls go through a Cloudflare Workers proxy that holds the key.
AbuseThe proxy accepts only requests from genuine app installs (App Attest) and checks the StoreKit 2 signed transaction for premium features.
Runaway costPer-user daily page quotas, plus caching and batch builds to keep each deck's cost predictable.
User documentsUploaded files are deleted from the Files API once the deck is built. Decks themselves stay on device and in the user's own sync account.

Milestones

  1. PrototypeGround and generate passes working on PDFs, measured on my own study material.
  2. TestFlightCamera ingest, review tray, source jump, and the proxy with App Attest.
  3. LaunchCapture in the premium plan, followed by overnight batch builds and "Why was I wrong?".