For RL task data vendors · Korea’s sovereign AI labs
To data vendors: there is a real opportunity selling into Korea’s sovereign AI models, but it will not be easy.
Since Artificial Analysis’ post on Korea’s Sovereign AI program (and AfterQuery’s on Motif), my inbox has picked up noticeably. I’ve been doing RL task QA and research with US data vendors. Over the past weeks, most of my time has shifted to helping Korean labs with acceptance testing on the data they buy.
I grew up and did my graduate work in Korea, and worked at one of its stronger AI labs (Krafton Deep Learning Center). The Korean AI world is small; I know people, researchers and decision-makers, across most of these teams. What I see is a market with a problem on both sides.
02 / Buyer side
These teams are small relative to the volume of vendor outreach hitting them, and most of that outreach goes unanswered. The reason is cost. Evaluating one vendor properly takes an NDA, a sample, a scoping call, and an internal eval. Doing that dozens of times is not something a team this size can absorb while training a frontier model.
Their need is no secret. They want to improve on the capabilities the Artificial Analysis Intelligence Index (AAII) measures, and they want data they can trust at a price they can afford. AAII is assembled from well-known public benchmarks, all of them public knowledge, so a great deal of relevant data already exists off the shelf. What’s scarce is a way to compare it. That is the gap I work in.
03 / Seller side
I’m currently working with eight vendors building Terminal-Bench-style tasks, and in practice the only thing separating them is price. Samples are cherry-picked, so they stop differentiating. Catalogues have the same problem. The QA criteria vendors cite largely overlap from one deck to the next. I run my own checks, and where I find something worth fixing, I tell the vendor.
The most convincing evidence a vendor can offer is a measured lift from training an open model on their data, on the target benchmark and on adjacent ones (tb2.1 → deepswe). Benchmaxxing has already been a story in the Korean press, and these buyers are scored in public, so nobody wants to be the next one. A vendor needs that evidence ready, and most arrive without it.
Approach a lab without this context and you bounce off: they cannot tell how you compare, and they do not have time to find out. The feedback here is grounded in the pool I already hold:
20+
vendors working with me
10+
under NDA · samples & pricing in hand
8
building Terminal-Bench-style tasks
Free · Quality feedback
How a lab actually reads your catalogue, and what to fix before one sees it.
Free · Price feedback
Your position in the current price pool, and a recommended price where your strengths show best. Vendors pay me nothing, so the advice serves nothing but your sale.
Position once · Every lab
The positioning we set together carries to every lab I serve. When one is interested, I introduce you directly.
None of this costs you anything, and none of it touches the measurement: the draw below is random, and neither of us controls it. What you get is a sharper catalogue, a defensible price, and more shots at a real sale.
05 / The protocol
Ten tasks are drawn from your dataset.toml at dataset_draw by a public randomness round that does not exist until your catalogue is committed, so no one can aim the draw at their best tasks. Catalogues need at least 100 tasks, so the draw stays a sample, not the product. You send me exactly those tasks, I verify they match the committed digests, and I run them against the candidate models repeatedly: a usable difficulty band (roughly 0 < p < 0.8), my task-implementation and trial-analysis rubrics (largely borrowed from Harbor check / analysis), and how closely the failure modes match the target benchmark (I have already run the candidate models on it). The lab gets every drawn task as measured, raw trajectories and judgments included, and the receipt lets anyone recompute the draw.
Scoping
Send a short blurb and your catalogue.
If there's a fit, we take a short call: pitch your data's strengths to me the way you would to a lab.
We sign an NDA. You send samples and price.
Feedback · on the catalogue
I review the samples and send you feedback: how your catalogue reads from the buyer's side, and what to fix.
You check the rest of your dataset against that feedback and fix what needs fixing.
Measurement · on the random draw
You run the draw yourself at dataset_draw: paste your dataset.toml (minimum 100 tasks), and the first public drand round published after your commit picks ten. One draw per catalogue.
You send me the draw receipt and exactly those ten tasks. I verify they match the committed digests, run them against the Korean models, and analyze the results.
Listing
You see all of it: the analysis, the price range across my current pool, and where you sit in it. On your confirmation you go into my listings.
I let the labs know a new dataset is available.
When a lab is interested, I introduce you directly.
Vendors pay nothing. The buyer pays, at acceptance.
No vendor money touches this service: nothing to be tested, nothing to be listed, nothing on a closed deal. The fee sits on the buyer side, paid when a lab accepts a dataset and the deal closes, and it scales with what the channel is worth in that deal: lower for large, well-known vendors a lab could have found on its own, higher for smaller vendors that would otherwise have had a low chance of being evaluated at all. The lab pays me, not the vendor being measured, so a good score is not something a vendor can buy. That is what keeps the number worth something. And this is early: right now I care less about the fee than about finding the shape that serves more labs and vendors well.
08 / Get listed
Agentic coding and terminal work, math, reasoning, long-horizon tool use. Send a short blurb and your catalogue.
jongwon.park@posttrain.dev