Why should long-running exports, batch jobs and AI analysis run in the background?
An export, bulk update, file-generation process or AI analysis should become a background job when it cannot reliably finish within one normal interactive wait. The product should validate the request, create a durable job record and let the user leave. That record must preserve the parameters, owner, state, progress, result, error and audit history. Moving a spinner elsewhere is not background processing; the deliverable is a recoverable business operation.
Where this decision matters
This guidance is for leaders specifying management reporting, data import and export, bulk record changes, document conversion, batch notifications, search indexing, OCR or AI analysis. A single-record lookup and a quick validation usually belong in the immediate interaction. Work that depends on data volume, queue capacity, several external services or a generated result may take tens of seconds or minutes and varies too much to hold a page open safely.
Microsoft's Asynchronous Request-Reply pattern separates accepting a request from finishing the work and gives the caller a status location. Google's approved long-running operations guidance independently models an operation identifier, progress metadata, final response and error. They describe a broadly useful contract rather than requiring Azure or Google Cloud.
Do not use one time threshold as the whole rule
Google's guidance and NN/g's research both use roughly ten seconds as a practical signal, not a universal service level. Ten seconds is excessive for an authentication check. Several minutes may be reasonable for a monthly analysis that a user explicitly started.
Choose the interaction from four properties:
| Decision | Immediate response is more suitable | A background job is more suitable |
|---|---|---|
| Duration | Short and predictable | Varies with volume, capacity or suppliers |
| Interruption | Closing the page can safely abandon it | Work should continue or be resumed |
| Result | Needed immediately on the current screen | Creates a file, batch outcome or later record |
| Failure | The user can correct input now | Partial success, retry or recovery is possible |
Streaming may be more appropriate when the value is the output as it arrives, such as incremental generation. A workflow state is needed when the operation waits for approval or an external callback. The product need not use one technical pattern for every slow action, but every path must tell the user whether the request was accepted, what is happening and where the outcome will appear.
Acceptance must be a real product event
Before creating a job, perform the checks that can complete quickly: permission, valid parameters, permitted file type, safe data range, current capacity and an existing equivalent request. Reject an invalid request immediately. A system must not say “submitted” and reveal several minutes later that a required parameter was missing.
After acceptance, return and display at least:
- a unique job identifier and creation time;
- the initiating user, organisation and authorised viewers;
- the job type and a readable parameter snapshot;
- an initial state such as queued or waiting for capacity;
- a persistent route to status and the eventual result; and
- an estimate only when it has evidence, otherwise meaningful stages.
Parameters should be immutable for that run. If a user starts “Export September orders” and later changes the screen filters, the job must still explain which dates, stores, columns and version it used. For changing data, define whether the result reflects request time, execution time or completion time. Otherwise two identically named reports can disagree without an accountable reason.
Give the job centre an explicit state model
A useful first state model can include queued, running, succeeded, partially succeeded, failed, cancelling, cancelled and expired. These must correspond to reality. Placing a message on a queue is not export success; producing a file that cannot be stored or retrieved is not completion.
Microsoft recommends status data such as creation time, last update, current state, optional progress and a structured error. A business product should add total, successful and failed item counts, the result asset, a failure report, expiry and the next available action.
Progress must be honest:
- show “3,240 of 10,000 processed” when the denominator is reliable;
- show “validate data — create file — store result” when stages are known;
- show current stage and last update when an external AI or supplier cannot provide a credible percentage; and
- never treat an indefinite spinner as sufficient progress.
NN/g's research on long waits and interruptions recommends letting lengthy processes continue in the background and restoring context with completion time, summary and result links. That supports the user experience; it does not replace reliable execution, authorisation or reconciliation.
Define what cancellation and retry mean
Cancellation is not always immediate or reversible. A queued job may stop cleanly. A running job may stop at a checkpoint. Work that has already sent notifications, created financial instructions or changed official records can often stop only future steps; completed effects must be reconciled or compensated. The interface should state this consequence rather than imply that a button erased every action.
Retry must also be a business rule, not “submit the same thing again”. Microsoft's background-job practices explain that queues, overlapping schedules and infrastructure restarts can execute the same logical item more than once. Jobs therefore need an idempotency design:
- a repeated export may create a new file, but each version and timestamp remains visible;
- a bulk update skips items already completed with the intended result and never charges, rewards or notifies twice;
- an AI analysis may be regenerated, but the earlier output is not silently overwritten; and
- a timeout or double click finds the existing job before creating another equivalent operation.
Partial failure must be an explicit outcome. If 970 of 1,000 customer records update and 30 fail validation, the system can retain the 970 results and provide a correction file for the remainder. If the business requires atomic completion, it must instead roll back the batch. This is a scope and acceptance decision, not an implementation detail to discover after launch.
AI jobs need provenance and review
An AI job should retain more than its final prose. Record the input scope, instruction or rule version, model or service version, initiator, timestamps, cited sources, exceptions and human amendments. Customer, employee, order or contract data must keep the same access, logging, retention and deletion rules while it sits in a queue, log, notification or result file.
Batch classification, summarisation and risk suggestions can run in the background, but high-impact decisions need an authorised review path. A reviewer should be able to return from a result to the supporting records, mark an error and choose whether to rerun. Retries against a model supplier must be bounded and must not duplicate official state changes.
A proportionate first release
The first release does not need a general workflow platform. Start with one frequent long-running operation, such as an order export or one batch AI analysis, and deliver the complete loop:
- validate access, parameters and data range quickly;
- create an identifier and immutable parameter summary;
- separate execution from the current browser or app session;
- provide a job list with filters, state, last update and result access;
- retain structured errors and a defined retry scope;
- survive duplicate requests, redelivery and service restart consistently;
- enforce result access, download audit, expiry and deletion; and
- notify in product on completion or failure, adding email, SMS or enterprise messaging only where the business warrants it.
Do not promise that background work will “always finish within one minute” without a measured basis. Define priority, concurrency and service expectations instead: whether daily small exports share capacity with month-end reports, how many jobs one organisation may run, whether an executive can displace operational work, and what degrades during a peak. Background execution adds queues, state, storage and operations. A consistently quick action may remain simpler and clearer as a synchronous request.
Accept with interruption and duplication scenarios
Use the site's software acceptance-criteria method to put state, evidence, ownership and remediation into signed scenarios. Acceptance must include leaving immediately after submission, opening the same job on another device, resubmitting after a network timeout, restarting a worker mid-job, a transient supplier failure, partial batch failure, two users submitting equivalent work, a stalled job, an unauthorised user guessing an identifier, downloading after expiry, and cancelling while queued, running and nearly complete.
Business evidence should confirm that:
- displayed totals reconcile to processed items, failures and result assets;
- notifications open the correct job without exposing another organisation's data;
- a stale last-update time and monitoring detect a stalled operation;
- retry cannot create duplicate orders, messages or charges;
- the audit trail explains who ran which parameters and received which result version; and
- expiry, deletion and evidence retention match the contract and data policy.
The value of a long-running feature is not a progress bar. It is the user's ability to leave safely, return with context, trust the outcome and recover from failure. When those states and evidence become signed acceptance criteria, a custom system turns a fragile click into an operable business capability.
References
- Microsoft Azure Architecture Center: Asynchronous Request-Reply pattern, accessed 28 September 2026; request acceptance, status, completion, cancellation and retention.
- Google AIP-151: Long-running operations, accessed 28 September 2026; operation identity, progress metadata, result, error, parallel work and expiry.
- Microsoft Azure Architecture Center: Best practices for background jobs, accessed 28 September 2026; result communication, idempotency, checkpoints, redelivery and recovery.
- Nielsen Norman Group: Designing for Long Waits and Interruptions, accessed 28 September 2026; progress, context recovery and background interaction.