Jev launched on 15 September and hasn't left my feed since. It's a model from TypeSafe that only makes decisions: you give it some text and a question, and it gives you back a number between 0 and 1. No paragraph, no chat. Jev pricing is simple: $0.042 per million input tokens, output free.
Everything I saw about it was speed and price. Nevo, who runs Postiz, put it better than I could: nobody is talking about the quality. So I tested it on the only data I really know, my own posts.
I run seenpaid, a social media scheduler that AI agents can drive. Before an agent schedules anything it calls a tool called validate_post, which checks the caption against each platform's rules. I wanted Jev in that check, but not before I knew whether it was accurate.
The setup
Every caption goes to Jev with four yes/no questions. I wrote them the way a normal person would judge a post:
- Would a typical person scrolling social media see this post as spam, engagement bait, or a hard sell?
- Does the first line of this post give a reader a specific reason to keep reading?
- Does this post make sense to someone who has never heard of the author, without extra context?
- Does this post promote a brand, product, service or offer?
I ran it over 522 captions from my own X account. Most are drafts I scheduled and then cancelled, so the pile includes plenty of posts I decided weren't good enough. 520 came back, 2 failed.
Then I picked 100 of them at random and had Fable answer the same four questions, same wording, without seeing Jev's answers. Fable isn't ground truth. It's the model I'd have used if Jev didn't exist, so it's the honest baseline.
Jev vs Fable: how often they agreed
The promotion question is where Jev is solid. 94 out of 100, with a correlation of 0.90. If a post is selling something, Jev sees it.
The other three are shakier. The spam question agrees 74 times at 0.5, but seenpaid doesn't warn at 0.5. It warns at 0.65, because an agent that gets warned about half its captions learns to ignore the warnings. At 0.65 it looks better:
85 of 100. And look at where the misses go: Jev flagged 12 posts Fable was fine with, and Fable flagged only 3 that Jev let through. Jev is the stricter one.
Where Jev gets it wrong
The 12 are mostly short opinion posts. This one got 0.69 from Jev and 0.20 from Fable: "every platform rewards the same thing: people staying on your post. write for that." This one got 0.76 from Jev and 0.20 from Fable: "a stranger signed up and then went quiet for twelve days. what did you do that finally made them pay?"
Neither is spam. They're short, lowercase and punchy, and Jev reads that style as bait. It's the same thing Nevo ran into when Jev-based tools labelled his handwritten tweets as AI.
It misses in the other direction too, just less often. "Who's still awake building? Reply." is bait by any definition. Fable gave it 0.80. Jev gave it 0.51 and let it through.
What it said about my own posts
I have a lot of "drop your startup below" style posts. Jev flagged 62 of 63 of them, average 0.82. It flagged 124 of the 160 posts that mention seenpaid, average 0.74. It doesn't care how nicely you mention your product.
Two things surprised me. I assumed my "founders be honest:" question posts were my cheapest ones. Jev scored them 0.60 on average, the same as everything else, and flagged only 12 of 43. And the posts I actually published averaged 0.58 while the ones I cancelled averaged 0.60. So Jev didn't agree with my own taste either.
Can you rewrite your way out of it?
I took my worst ad and rewrote it, then scored both live. Before: "post everywhere. connect stripe. see the real number per post. 7 days free at seenpaid.com". After: "Likes tell you which posts people enjoyed. They don't tell you which ones sold anything. That's the number I built seenpaid to show: every Stripe payment, matched to the post that sent the buyer."
The opening score more than doubled. The spam score dropped from 0.94 to 0.65 and stopped right at the line. You can fix your first line. You can't write your way out of being promotional.
What it cost
- 520 posts, four questions each: 265,488 input tokens, about 511 per caption.
- Total: $0.011. About one cent.
- Speed, round trip from my laptop: 341 ms median, 601 ms at p95, 1.9 seconds at worst.
- TypeSafe's own signup said it was full when I signed up on 22 September, so I went through Vercel's AI Gateway, which speaks the same API. Their $5 free credit covered all of it.
- A fresh key got rate limited at 8 requests in parallel. 2 at a time with retries was fine.
How I use it now
It's live in validate_post, with three rules. Jev never blocks or publishes a post; it adds a signals field to the response and, past 0.65, one plain sentence of advice. It runs in parallel with the rest of the check and has a 2.5 second timeout, so if it's slow or down nothing changes. And it never gets a veto, because captions are user text and VentureBeat already showed a fake "pre-approved" field talking a Jev verdict down from 0.76 to 0.48.
If you're thinking about adding Jev to your own product: use it as a cheap first pass, not a judge. Don't point it at someone's writing style. Pick your own threshold instead of 0.5. And before you trust any number, run 100 of your own rows through it next to a model you already trust. That took me an afternoon and it's the only reason I know where the 12 misses are.
If you want to see the scores on your own captions, connect an agent to seenpaid and call validate_post. The Jev signals come back next to the platform checks.