The Salesforce Org Argues Back: Building a Data Seeder That Reads Your Validation Rules

I filed the AI integration as an epic on day two of this project. It took a year to land — and the part I’d actually been dreading arrived in the twenty-three minutes between two commits.

Every Salesforce project starts in the same room: a fresh sandbox with nobody in it. No accounts, no contacts, no opportunities, no assets. You cannot demo a report with no rows in it. You cannot test a flow that has nothing to fire on. You cannot show a customer their own process working when the screen is a polite empty state saying No records to display.

So the first day of every engagement is spent making up people. A spreadsheet of invented companies, Data Loader, an insert that fails, a column deleted, another insert, a different failure, another column deleted, and eventually a few hundred rows of something that technically exists. Then the next project starts and you do it again from scratch, because the data you made was shaped for the last org’s configuration and this one has different ideas.

I have been doing that for years. It is not hard. It is not interesting. It is the definition of a someday problem: obviously worth automating, never quite worth the afternoon.

I’ve written up projects like this before — a Grafana data platform, the homelab dashboard I’d abandoned three times, an alert that paid for the whole lab, and a Proxmox cluster with a MAGI panel bolted to it. This one isn’t the house. It’s the day job, which makes the someday excuse considerably harder to justify — I have been doing something the slow way and invoicing for it.

Same principle either way: I supply direction, the agent supplies execution.

The tool is on GitHub: github.com/angusmaul/salesforce-sandbox-data-seeder.

Why this is harder than “generate some fake companies”

Generating a thousand plausible businesses is trivial. Faker does it in about four milliseconds. The difficulty is that a Salesforce org is not a schema — it is a schema that somebody has been configuring for six years, and every one of those configuration decisions is a landmine your generated record has to walk past.

A modest developer org hands back 802 creatable objects. Twelve of them are custom — the things somebody in this business actually built. The other 790 are Salesforce’s own plumbing, and none of them are labelled as such.

Underneath that:

  • Picklist dependencies are real and enforced. Country controls State, and the mapping is not published as a list — it’s a base64 validFor bitmap you have to decode per value to find out which states are legal under which country.
  • Restricted picklists reject anything outside the active set, so a perfectly sensible value that used to be valid is now an error.
  • Precision and scale are per field. 9477.9 is a fine number and it will not go into a Number(4,1), whose ceiling is 999.9.
  • Lookups need real IDs, which means the objects have to be loaded in dependency order, and each child needs the parent IDs the previous step actually produced.
  • Compound address fields are quietly derived — send both BillingState and BillingStateCode and the org tells you off for a mismatched integration value.

All of that is discoverable. Tedious, but discoverable: it’s in the metadata, and a sufficiently patient program can read it.

And then there are validation rules, which are a different kind of problem entirely.

A validation rule is a formula somebody wrote, in this org, for a business reason that isn’t recorded anywhere in the schema. The metadata will tell you that Active__c is a picklist of Yes and No. Nothing in the metadata tells you that in this org, an Account whose Active__c is No cannot be created at all. That fact lives in a formula and a sentence of English, and the only way you learn it is by having your record thrown back at you.

Everything else about seeding a Salesforce org is a data problem. Validation rules are a reading-comprehension problem.

Two false starts, six months apart

The repository is honest about how this went. The first commit is August 2025: a CLI that discovers the data model and fills it with Faker output. It worked, in the sense that it ran.

The very next commit, the same day, files an epic called AI-Integration. I knew from day two what the tool needed. Then nothing happened for six months.

In February 2026 the epic shipped: field classification via Claude Haiku, analysing the org’s field schemas once and mapping them into a semantic library of 8 categories and 66 subcategories, so the generated data is correlated rather than merely plausible. A contact in Engineering gets an engineering job title. An Australian address gets an Australian phone format. A name and an email address belong to the same person. Two and a half thousand lines, 141 tests, one commit.

That was a genuine improvement and it fixed the wrong problem. The data got much more realistic and it still didn’t go in. Because “realistic” and “acceptable to this org” are unrelated properties, and I had spent the whole effort on the first one.

Then nothing happened for six more months.

The 24% run

What broke the deadlock in August wasn’t an idea. It was running the thing against a real sandbox and reading what came back.

A 275-record load succeeded at 24%. 67 records in, 208 rejected. Account limped in at 55 of 100 and Lead at 12 of 75, and Contact came in at 0 out of 100 — not degraded, not partial, zero. Every single one refused.

The logs sorted the wreckage into two deterministic classes almost immediately.

166 of roughly 208 failures were state and country integrity. And the reason is the detail I keep turning over, because it’s not a bug so much as an indictment. The app already decoded the org’s country-to-state dependency map — pulled the validFor bitmaps during field analysis, decoded them, and stashed the result on the session. That work was being done on every run. It was then never read. Generation was reaching for a hardcoded four-country dictionary that had been written months earlier as a stopgap, and the elaborate, correct, org-specific answer was sitting untouched in memory the whole time.

Sitting under that was the reason Contact scored zero: a function called extractAddressPrefix, which is meant to tell MailingCity from ShippingCity, returned Shipping for Mailing* fields. A copy-paste error, one word wrong, which guaranteed that every address lookup on Contact missed. Accounts have Billing and Shipping addresses and limped in at some percentage. Contacts have Mailing addresses and could never have worked at all.

I would never have found that by reading the code. It looks right. You find it by asking why one object scored zero when a structurally identical one didn’t, and being unwilling to accept “probably something to do with addresses” as an answer.

46 further failures were numeric range, which is the boring one and was fixed the boring way: field analysis now stores precision, scale and digit count, and values get clamped to the field’s maximum representable number before they’re sent.

Lesson: Before you add capability, check what the program already knows and isn’t asking itself. The most expensive bug in that run was a correct answer computed on every single run and read by nothing.

Remediation you can write down

That fixed the deterministic classes. The next increment dealt with the fact that a failed insert was, until that afternoon, simply logged and abandoned.

Salesforce error codes are unusually well-behaved: most of them tell you precisely what to do. So each failure now gets up to two remediated resubmissions, with the remediation chosen by code:

STRING_TOO_LONG truncates to the field’s real length. INVALID_OR_NULL_FOR_RESTRICTED_PICKLIST re-picks from the active values. NUMBER_OUTSIDE_VALID_RANGE clamps, then drops the field if clamping didn’t satisfy it. REQUIRED_FIELD_MISSING regenerates the named fields — unless it’s an unfillable required lookup, in which case the record is dropped rather than retried into the same wall. Duplicate-rule violations uniquify the identity fields, and do it email-aware, so a mangled address is still an address. Row locks get one unchanged resubmit, because a row lock is a timing complaint and not a data complaint.

The piece I like most is the learned blocklist. If removing a particular field is what rescued at least three records, and those recoveries account for at least 80% of everything recovered this pass, that field stops being generated for the rest of the session. The tool notices what it keeps getting wrong and stops doing it, which means object number seven doesn’t repeat the mistake object number two already paid for.

One error code was deliberately left terminal: FIELD_CUSTOM_VALIDATION_EXCEPTION.

Every other code is a mechanical instruction. That one is a human being telling you no. There is no generic remediation for it, because the correct action depends entirely on what the formula says — and the formula is not in the error. All you get back is somebody’s error message.

The hundred percent that proved nothing

With those two increments in, the same 275-record load went in clean. 100%. Account 100, Contact 100, Lead 75, in thirteen seconds.

Which proved almost nothing, and I knew it while I was looking at it. That org’s Accounts had no awkward rules on them. A green run against a permissive org is not evidence that the tool handles a strict one; it’s evidence that you picked an easy org. Every consultant reading this has inherited the other kind — the org with eleven rules nobody documented, written by someone who left in 2019.

So I went into Setup and wrote a validation rule specifically to break my own tool. Accounts must have Active__c set to Yes at creation, or they don’t get created.

The next run came back at 72%. Contact still 100, Lead still 75, and Account collapsed to 23 of 100**–77 records**, failed terminally, every one of them with the same sentence attached:

Active is required to be Yes when creating accounts

A rule a competent human satisfies in about four seconds of reading, and the tool could not touch it.

Lesson: A passing test against convenient conditions is a measurement of the conditions, not the tool. If you can’t find something that breaks it, go and build the thing that breaks it.

The rule that says no and won’t say why

This was the blocker, and it had been the blocker long before I wrote that rule down. Not “difficult” — blocked. Every other failure class had a mechanical answer and this one required understanding a formula written by a stranger. It’s why the project had sat for six months twice: I could see the shape of the work and could not see the shape of the solution.

The solution turned out to hinge on one inversion, and once you’ve seen it the whole thing collapses into something small.

A validation rule’s errorConditionFormula evaluates to TRUE when the record should be rejected. It isn’t a description of a valid record. It’s a description of an invalid one. So you don’t ask a model “what does this rule want?” — a vague question with an essay for an answer. You tell it: this formula rejects when true; produce constraints that make it false.

That’s the whole trick, and it changes the task from open-ended interpretation into a translation with a right answer.

The second decision matters just as much, and it’s the one I’d defend hardest: the model is not allowed to answer freely. It gets five constraint types and nothing else.

fixedValue      — field must equal value
picklistSubset  — field must be one of these values
range           — numeric bounds
notNull         — field must be populated
maxLength       — string length cap

Anything the model cannot express in that vocabulary — cross-field logic, record types, user context, anything requiring the model to be clever — it is instructed to name under unsupported rather than improvise. Those rules are then reported to the user, alongside a suggestion to use the disable-and-restore fallback for those specific rules if they want them out of the way.

So the model does comprehension, in a domain where comprehension is genuinely required, and hands back a data structure. Ordinary deterministic code does the enforcement: applyConstraints() walks the record, applies each constraint, skips anything Salesforce won’t accept as a write anyway — formula fields, auto-numbers, non-createable fields — and returns human-readable notes about every change it made. You can hit /api/validation-rules/list/:sessionId and see every rule the org has, next to exactly what was derived from it.

The retry loop then gets one carefully hedged upgrade. FIELD_CUSTOM_VALIDATION_EXCEPTION is promoted from terminal to adjust-and-resubmit — but only when constraints exist for that object and aren’t already satisfied. If the rule was one the model couldn’t express, or the record already complies and failed anyway, it stays terminal. Without that guard, a rule nobody can interpret would burn every retry pass resubmitting identical records to an org that has already made its position clear.

The test for all this is the failure I’d manufactured: Active__c=No fails, constraints are applied, the resubmit succeeds. 169 of 169 server tests green.

Then the same 275-record load, against the same org, with my hostile rule still active and untouched: 275 of 275. The two runs are twelve minutes apart in the logs — 72% either side of a change that reads, in the load log, as nothing at all. Same records requested, same rule enforcing, same org. The tool just read the rule this time.

The whole interpreter is 127 lines. The constraint applier is 93. That is the entirety of the thing that had blocked the project for a year.

Lesson: When you put a model in a pipeline, shrink its job until the output is a data structure you can validate, and give it a way to say “I can’t express this.” A model with five allowed answers and an escape hatch is inspectable. A model asked to be helpful is a second bug you can’t step through in a debugger.

The commits either side of that work are twenty-three minutes apart. I’ve stared at that gap a fair bit since. It is not that the code was hard — 127 lines is not hard. It’s that arriving at “invert the formula and constrain the vocabulary” required knowing that validation failures were a distinct, terminal, 77-record failure class, and knowing that required the retry loop from an hour earlier, which required the failure taxonomy from an hour before that. The blocker was never the typing. It was two hundred failures nobody had read.

The slow model that found three bugs

One more thread worth pulling, because it’s the best bug of the day and it isn’t really an AI bug at all.

Field classification had been hardcoded to Anthropic, via a server environment variable. That’s fine for me and useless for the actual use case, because this tool reads your org’s field names, picklist values, validation-rule formulas and error messages — which is to say, a fairly complete description of a client’s business logic. Plenty of engagements would never permit that leaving the building, and a tool that requires it is a tool that doesn’t get used.

So it became bring-your-own: Anthropic, any OpenAI-compatible endpoint — OpenAI, Groq, OpenRouter, LM Studio, vLLM — or a native Ollama running on your own hardware, all configured in the UI rather than in a server env var. Failures return null and generation falls back to pattern-based rules, so a missing or broken provider degrades the output instead of breaking the run.

The generation plan: 275 records across three objects in dependency order, and the banner stating exactly what goes to the model.
The generation plan: 275 records across three objects in dependency order, and the banner stating exactly what goes to the model.

The banner in that screenshot is the whole argument in one sentence: field metadata (not your data) is sent to your local Ollama model for classification. The org’s schema is the thing being analysed, the records never are, and on that setup neither leaves the building. Note the amber Production badge too — a Developer Edition org is, in Salesforce’s own taxonomy, a production org, which is exactly why the tool flags it rather than quietly proceeding.

And then a local gpt-oss:20b took 77 seconds to classify a set of fields, and three separate bugs fell out of a latency regime nothing had ever been in before.

  1. Next.js’s rewrite proxy severs connections at 30 seconds by default. The browser reported socket hang up. The server, meanwhile, was completely fine and finished the job.
  2. A session lost-update race. The long-running analysis endpoints hold a reference to the session object across a multi-minute await, while the wizard’s ordinary PUTs replace that object in the store. Last writer wins. In the observed run, the finished AI plan was overwritten by a stale snapshot seconds after being successfully cached — so the work completed, was saved, and then vanished.
  3. The UI gave up when the request died, despite the server completing the analysis regardless.

All three diagnosed by reading a real run’s logs alongside the session store, rather than by reasoning about what ought to have happened. The fixes are unglamorous: raise the proxy timeout, have PUT mutate the stored object in place instead of replacing it, have long-running endpoints re-fetch the session before writing their results back, and have the UI poll the cached-plan endpoint instead of assuming a dead socket means dead work.

Not one of those three is a bug about AI. They’re a read-modify-write race and two timeout assumptions, all of which had been sitting there since the beginning, all of them invisible while every call returned in two seconds.

Lesson: Supporting slow inference is a load test. A model that takes 77 seconds instead of 2 doesn’t just make things slower — it drags your code into a latency regime where every implicit timeout and every race you’d been getting away with becomes reproducible.

Smaller traps, for the record

  • jsforce’s sObject Collections call fails wholesale past 200 records unless you pass allowRecursive: true, at which point it chunks for you. Discovered the direct way.
  • Result-to-record pairing was using indexOf, which quietly does the wrong thing the moment two generated records are identical. The API returns results in input order; positional mapping is both correct and a prerequisite for retrying anything.
  • Never write GeocodeAccuracy, and suppress the text State/Country twins whenever the *Code fields exist — Salesforce derives the text itself and objects loudly if you send both.
  • A StateCode must never be written without its CountryCode. There’s now a backstop for this in both record loops, because the rule was getting violated from two different code paths that each thought the other was handling it.
  • Credentials are stored server-side at 0600 and never returned to the browser — but they’re plaintext JSON on disk, which is written down in the README rather than glossed over. It’s a sandbox tool. It should still say what it is.

What it looks like now

  • A seven-step wizard — connect, discover, select, configure, preview, execute, results — with live progress over WebSocket and a results dashboard that will hand you the whole run as a ZIP.
  • Saved org connections, so authenticating once via an External Client App means one click on every subsequent session.
  • Any AI provider you like, including none. Hosted, local, or pattern-based fallback. The org’s metadata never has to leave your network.
  • Validation-rule awareness, with the rules it couldn’t interpret named explicitly rather than silently skipped.
  • 169 passing server tests, including the exact 77-record failure that started the last increment, frozen as a regression case.
  • Deployable via Docker Compose, LXC with systemd units, or bare Node.

The honest part

Every one of these write-ups ends with me working out what actually changed, and it’s never the thing I expected going in. With the dashboard it was that the tedious middle could be delegated. With the energy alert it was that the previous projects existing made the next one trivial. With the MAGI panel it was that an aesthetic constraint had done the engineering work.

This one lands somewhere I didn’t like at first.

The AI’s job got smaller, and that’s what made it work.

The instinct — mine, certainly, and I think most people’s — is to hand the model more. When the generated data kept getting rejected, my first thought was that the model should generate better data. February’s version is that thought, fully committed: an AI plan for every field, semantic categories, correlation maps, 2,456 lines. It made the data lovely and it did not get one extra record into the org.

What actually shipped does the opposite. The model never sees a record. It never generates a value. It reads a set of formulas and a field list and answers in five permitted shapes, with an explicit option to decline. Everything downstream — picking, clamping, truncating, uniquifying, retrying, blocklisting — is deterministic code you can step through at three in the morning.

And the reason that design was even available is that the logs existed. You can only shrink a model’s job to “translate this formula into a constraint” once you know that formula-shaped rejections are a distinct failure class of exactly 77 records, and that everything else in the pile is mechanical. In February I didn’t know that, so the AI had to be the whole feature. In August the AI is 127 lines in the middle of a loop whose actual engine is a log file.

So when people say agentic engineering is moving quickly — and it is; three increments with tests in ninety-four minutes would not have happened in February — I don’t think the interesting part is that the models write better code. That’s true and it’s the least of it. The interesting part is that the cost of going and looking collapsed. Running the thing against a real org, pulling two hundred failures apart, sorting them into classes, and letting each class dictate its own fix used to be an afternoon I would never spend. It’s now the cheapest step in the process, which means evidence has become cheaper than theorising.

That’s the actual shift. Not that the agent is better at answering. That it’s finally cheap enough to stop guessing.

Six years of hand-rolled CSVs, two abandoned attempts, and the thing that broke it open was a 24% success rate and somebody willing to read all of it.

Point it at a real org. Read every failure. Then ask the model one small question.