HRIS Migrations Don't Fail on the API. They Fail on the Data Model.
Every HRIS implementation plan has a data migration phase that reads like a file transfer: export from the old system, map the columns, import into the new one, reconcile the headcount, go live. The integration work gets estimated in story points. The data model gets a spreadsheet.
That ordering is backwards, and it is why go-lives slip by a quarter. The endpoints are rarely the hard part — rate limits, pagination and webhook retries are solved problems with published patterns. What breaks migrations is the quiet assumption that an employee is a row with columns.
HR data is temporal, and most schemas are not
In a typical application table, a row holds current state. In an HR system, nearly every meaningful field is effective-dated: compensation, job title, manager, department, location, cost center, FTE percentage. Each carries a validity window, and the history is the point. You cannot answer "what was this person's salary in March" or "who approved that backdated promotion" from current state.
Load current-state columns and you get a system that is accurate on day one and wrong forever after. Retroactive pay, accrual recalculation and any audit request all fail against it. The load has to target effective-dated records, which usually means one source row fans out into several target rows that must be inserted in chronological order — and most vendor APIs reject an out-of-sequence effective date outright.
That single constraint, ordering, is what breaks most first-pass migration scripts.
One human is not one record
The second mismatch is identity. In the source system a person is usually one row. In the target, the same human may legitimately be several: an employee record, a separate contractor record, a rehire with a new hire date, or a worker holding two concurrent assignments in different legal entities.
The join keys are worse than they look. Email addresses change. Some vendors reuse employee IDs after termination. National identifiers are jurisdiction-specific and frequently absent for non-payrolled workers entirely.
The fix is to stop treating any vendor-supplied field as a primary key. Generate a durable surrogate key you control, persist it on both sides, and join on that. It is one extra column, and it is the difference between a clean load and re-importing everyone because somebody got married.
Stored values versus derived values
This is the one that produces silent corruption instead of loud failures.
PTO balances, tenure, accrued liability, FTE-normalized headcount, compa-ratio — these are computed from events, not stored facts. Import a balance snapshot and the new system holds a number with no ledger underneath it. The first time someone books a day off, the balance moves from a starting point nobody can explain, and HR and payroll stop agreeing.
Import the events instead: hires, terminations, accrual rules, grants, leave taken. Let the target compute the balances. Then compare the computed balance against what the old system reported — any gap is a genuine configuration difference, and it is far better to find it during migration than during the first payroll run.
"The counts match" is not reconciliation
A matching headcount is the most common false green light in a migration. Row counts say nothing about whether the values are correct.
Reconcile on aggregates that a finance or HR stakeholder can independently verify:
- active headcount by department and by legal entity
- year-to-date gross pay by legal entity
- total PTO liability, in hours and in currency
- workers with no manager assigned
- workers with a null cost center
The last two matter more than they look. Nulls in structural fields are what actually break downstream reporting, and they are invisible in a row count.
Make the load idempotent before you make it fast
Migration runs fail partway. A vendor gateway times out at record 600. Someone re-runs the job. If the load is not idempotent, you now have 600 duplicated employees and a cleanup project considerably nastier than the migration itself.
Every upsert should carry an external ID you own, and the job should be safe to run twice against the same dataset. Test that explicitly: run a full load into a sandbox, run it again, assert the record count is unchanged. It is a short test that prevents the most common migration disaster there is.
HR data is the most sensitive data you will touch
Salaries, health information, leave reasons, disciplinary records, immigration status, bank details. Handle it accordingly:
- no production HR data in staging or local environments
- no employee records in logs, error messages or support tickets — log the surrogate key, not the payload
- redaction before anything reaches third-party observability tooling
- a retention and deletion path defined before you load, not after someone files a request
If a script dumps a full employee payload to stdout on failure, that payload is now sitting in CI logs with a longer retention period than the HR system has.
The order worth running it in
- Freeze the source data, or snapshot it with an explicit as-of date.
- Model the target's effective-dated fields first, before writing any load code.
- Generate and persist your own surrogate keys.
- Build an idempotent load, and prove it by running it twice.
- Reconcile on finance-verifiable aggregates, not row counts.
- Run a parallel or dual-write window before the hard cutover.
- Only then automate webhooks and ongoing sync.
Steps 1 through 6 are data engineering. Step 7 is the part every team staffs for.
The API is not where HRIS migrations go wrong. The data model is: temporal records, ambiguous identity, and derived values that should never have been imported in the first place. Get those three right and the integration really is the easy part.
If you are scoping the other half of this — vendor selection, change management, training and the go-live sequence — the full walkthrough lives here: HR software implementation guide.