How should software maintenance and operations be priced?
Do not price maintenance and operations as a fixed percentage of development. Separate defect warranty, production operations, on-call coverage, reliability engineering, and feature iteration. Then price the responsibility from service hours, system scope, SLO and recovery targets, ticket and release capacity, tools, and known operational risk.
“Maintenance” is often used for fixing defects, watching servers, responding overnight, and adding features, leaving the parties with different expectations of one annual fee. The development price says little about runtime responsibility: a low-frequency internal tool and a 24-hour transaction service require very different staffing, redundancy, and incident communication.
| Service | Work included | Typical pricing basis | Excluded | Evidence |
|---|---|---|---|---|
| Defect warranty | Baseline function fails signed acceptance | Included for a defined scope and period | New rules, platform changes, UX improvement | Reproduction, issue, fix release |
| Core operations | Monitoring, backup, patches, deployment, capacity, in-hours incidents | Base fee plus ticket/change allowance | New business features | Reports, alerts, restore tests, releases |
| SLA/on-call | Out-of-hours response, incident command, recovery, review | Coverage readiness, incident effort, redundant resources | Unlimited work of every kind | SLO and response/recovery records |
| Reliability engineering | Automation, performance, resilience, toil removal | Work package or team capacity | Routine feature iteration | Before/after SLI, exercises, time saved |
| New iteration | Pages, processes, APIs, and rules | Effort, staged price, or dedicated capacity | Original-scope warranty | Requirement, acceptance, release |
A fixed annual fee can purchase readiness even in a month without incidents. Per-ticket support fits a non-critical, low-frequency service but provides less predictable recovery. A dedicated team supports continuing change but does not automatically include 24×7 on-call. State service hours, severity, first response, recovery objective, communication, exclusions, and excess rates in the agreement.
When defining budget, scope, and cost assumptions, also compare Why are infrastructure, third-party, and operations costs separate from software development? and How are custom software projects usually priced?; the linked guidance adds context that should be considered in the same decision.
Price the operating responsibility
A useful model is service readiness and monitoring baseline + coverage/on-call staffing + included ticket and release capacity + tools and exercises + excess work. Cloud, messaging, AI, and monitoring SaaS remain separate supplier bills. Any reseller arrangement should disclose original price, markup, ownership, and exit.
More services, databases, regions, and environments create more monitoring, patching, backup, and release objects. Extended or 24×7 coverage needs a rota, not one permanently reachable phone. Tighter availability, latency, RTO, and RPO require redundancy, automation, and exercises. Frequent releases and supplier upgrades consume testing and rollback effort. A noisy, debt-heavy system may need stabilization before a responsible SLA is possible. Security and sector requirements add evidence, scanning, remediation, and exercise work.
Google's SRE Workbook chapters on monitoring and on-call explain user-impact SLI/SLO monitoring, actionable alerts, incident tickets, escalation, and review. They do not set a Chinese provider's price, but they show why 24×7 operations is a people, process, and tooling capability rather than “contact us anytime.”
Accept operations through results and records
A monthly service record should show SLO performance, major incidents, detection and recovery time, backup success and restore exercises, vulnerabilities and patches, capacity, release and rollback outcomes, supplier-cost anomalies, and unresolved risks. “The server was normal” does not prove that restoration works or an alert receives action.
First response, service recovery, permanent correction, and closure of review actions are different timings and should be defined separately. For a cloud or carrier outage, clarify whether the operations provider coordinates and degrades service or is being asked to guarantee a third party's availability, which it cannot control.
Before launch, Wavesteam establishes the service inventory, dashboards, alert severity, contacts, backup restoration, and release rollback, then recommends business-hours support or a higher SLA from actual criticality. The client keeps production accounts and bills. Warranty defects, operations tickets, and new requirements remain distinct records so an annual fee does not hide a scope dispute.