Khalil Hebachi
Email meGitHubLinkedIn
Writing

· Updated

Building RavenClip: recovery, AI checks, and costs

Why I chose persistent stages, what automated checks can prove, and what a hidden provider failure taught me.

RavenClip turns news into short videos and publishes them on a schedule. I built and operate it on Laravel, React, a Node/Skia renderer and FFmpeg. The case study covers the product; this post covers the technical choices and their limits.

The short version

  1. 01Recovery9 s → 2 sMedian wait before rendering, once each stage starts the moment the last one finishes.
  2. 02Prompts15-18% fewerInput tokens per call after replacing a banned-phrase list with three principles.
  3. 03ChecksShape, not truthGates catch drift, bad hooks and length. They cannot confirm that a claim is true.
  4. 04Fallback25 daysA provider sat out of credit while fallback hid it. Monitors now catch this within hours.
  5. 05Costs$0.36Recorded API spend per rendered video, 1-25 September 2026.
  6. 06AI agents18 hoursScene-asset jobs failed after an agent’s bulk edit. Three safeguards followed.

Persist progress where recovery matters

Every stage saves its status in the database, so a failure resumes where it stopped instead of paying for the same generation twice. Voice and visuals run in parallel, then meet again at rendering.

From a news story to a published video
News feedsContextScript

voice branch

VoiceTranscript

visual branch

ScenesAssets
RenderReviewPublish
voice branchvisual branchNews feedsContextScriptVoiceTranscriptScenesAssetsRenderReviewPublish

Voice and visual work run in parallel before rendering. Persisted stage results let retries reuse completed work. Publishing runs automatically or after customer review, according to the channel settings.

One worker per stage

A runner moves forward whatever is ready. A row lock stops two workers from claiming the same stage.

Ordered timeouts

Job timeout, then supervisor timeout, then queue redelivery. Recovery retries a stuck render with the assets it already has.

Publish once

State checks plus a unique database constraint: the same video cannot post twice.

One measurement corrected me. Starting each stage the moment the previous one finished fixed the median, not the whole delay.

9 s → 2 s
Median wait before rendering
~1 min
Slowest waits, unchanged

When I skip it

This adds state and recovery paths to maintain. For the short, fixed sequence after a user edits a scene, a simple job chain is still the better choice. The full machinery is for generation that runs unattended and publishes in public.

Give models a structure, then check the result

Outputs are structured and checked against schemas. Each stage gets its own thinking budget: simple sorting does not need the budget of scriptwriting.

// One narration beat, as the writer model returns it
type Beat = {
  adds: string;        // what new information this line brings; checked, then thrown away
  contentText: string; // the narration itself
};

Naming the new information before writing the line pushes the model to say something new every time.

RavenClip scene editor: each scene lists its narration line, its photos and its timing, next to a phone preview of the video
Where the beats end up: each scene keeps its narration line, photos and timing, and a user can edit any of it.

Before

A long list of banned phrases. Fixing one phrase just produced a new version of the same problem.

After: three principles

True
Use facts from the source.
Dense
Add information in each sentence.
Spoken
Write words that work read aloud.

Plus one example for each recurring failure.15-18% fewer input tokens per call

Simpler instructions helped, but prompts alone still let some bad output through. That is what the gates are for.

Be precise about what a check proves

The system ranks drafts by their defects and tells a failed attempt exactly what to fix. A critic’s rewrite wins only if it also passes the gates and scores higher.

Checked or enforced

  • The script stays on its subject
  • The hook has the right shape
  • Length and endings fit
  • Later stages keep the approved narration, or fall back to the checked lines
  • On-screen scores, standings and prices come from data sources, not from the model

Not proven

  • That a claim is true. A false sentence built from words in the source can still pass.
  • Spoken figures. They are not checked.
  • Non-Latin scripts. They skip the subject check.
RavenClip publish tab with three options for one video: Auto, Hold and Schedule
So customers keep review controls, and flagged videos wait for approval.

Monitor the fallback path

External calls share one mechanism for retries, key rotation and switching providers. Each task also has its own rule for what a fallback must keep.

When a call fails

  • Temporary errorA few retries with growing delays
  • 2 timeouts in a rowSwitch provider
  • Empty balanceShared 30-minute cooldown

What a fallback must keep

  • VoiceThe same speaker
  • ImagesThe same quality level
  • EmbeddingsNever mixed: their outputs do not compare

Checking only whether videos finished hid one failure for weeks.

Record costs and understand the omissions

Billable calls are recorded before the output is validated, because an unusable answer can still cost money. Each entry keeps the rate that applied at the time, and replaced attempts keep their cost rows.

~$0.45per rendered video, with a share of infrastructure added

$0.36 recorded API spend

~$0.09 infrastructure share

Period
1-25 September 2026
Recorded API spend
$66.59
Rendered videos
183

Not includedPayment feesMy timeSame-key retries, not fully priced apart

Code-scanning test

Finds paid integrations with no cost recording. It proves the recording exists, not that the price is right.

Daily audit

Checks for missing prices and for differences from what providers report.

Caching also needs enough traffic to pay off.

Caching, by call type

  • EnrichmentFrequent: provider cache hits
  • Scene directionToo rare to keep a cache warm, so I cut prompt size instead and left caching until volume justifies it

Building with AI and owning the mistakes

Coding agents

BuildInvestigateTest

Under my direction

Product scopePricingArchitecture

My session records include corrections to both code and business proposals, down to which features belong in paid plans. Agents still make mistakes that reach production.

23
stale or leaking tests, collected after I removed CI to speed up releases, cleaned on 25 September
~1,300
test definitions: a measure of coverage, not the absence of bugs

Next time I would keep a small syntax-and-smoke check from day one. The useful question is which failure paths are checked, and what happens when a check misses.

Keep the next improvement grounded

  1. 01

    Keep the raw data

    Store the data behind each measurement, not just the result.

  2. 02

    Keep incident evidence

    Hold it past short log windows.

  3. 03

    Close the cost gaps

    Finish the cost accounting, including retries on the same key.

RavenClip is an early product with three paying customers as of 26 September 2026. Running it shows me which abstractions help and which assumptions need another look. The case study tells the product side.

Ask about my work

AI answers from my CV and case studies. It can make mistakes.

Ask anything about my projects, results, or how I work.