Today, service managers in the mechanical engineering industry almost always track the same key performance indicators—first-time fix rate, mean time to repair, and net promoter score. But the numbers often seem strangely stable: somewhere between “acceptable” and “actually too low,” without anyone being able to explain why. The problem rarely lies with the KPIs themselves, but rather with the underlying data.
Service metrics are only as meaningful as the data from which they are derived. If FTFR and MTTR are calculated from manually compiled Excel spreadsheets—using different definitions of “resolved” and without the ability to drill down by equipment type or technician—they merely reflect the current state of affairs—but offer no leverage for improvement. If you really want to make use of KPIs, you need an integrated database in which every number can be traced back to a specific machine, a ticket, and a contractual relationship.
- Problem: Fragmented service data – KPIs show symptoms but do not provide a drill-down option. Improvement initiatives fail due to lack of root cause analysis.
- Solution: KPI architecture along the 4-step model – digitize data, network systems, make decisions with source verification.
- Result: FTFR and MTTR can be filtered by machine type, region and technician. Economic KPIs such as service margin and aftermarket turnover per machine become controllable.
This article describes the complete KPI set for service in mechanical engineering – operational core KPIs (FTFR, MTTR, MTBF), customer and contract-related KPIs (NPS, SLA compliance), economic KPIs (cost-per-ticket, margin, aftermarket turnover). And above all: how these figures become management tools instead of retrospective reports.
Why service KPIs lose their impact without an integrated database
In most service organizations today, KPIs are created in exactly the same way: an employee collects ticket data from the helpdesk at the end of the quarter, compares it with the ERP system for the contract data, adds service report PDFs for the deployment times, puts everything into Excel and calculates FTFR and MTTR from it. The result is a number – but not a story.
Three structural problems are blocking real control:
Missing reference objects. An FTFR of 65% says little as long as it is unclear whether the value applies to system type A or B, whether it occurs in the North or South region, which service technician was deployed or which error class was present. Without a machine file as a structured point of reference, every KPI remains an isolated figure without a basis for action.
Unclear definitions. What counts as “successfully resolved on the first attempt”? Resolving the symptom or analyzing the root cause? Is it enough to simply restart the system, or must the root cause of the problem be addressed? As long as this definition isn’t codified in the system, every service organization measures something different—and comparisons between locations or quarters lose their meaning.
Manual time stamps. MTTR calculations fail at exactly one point: When exactly did the problem occur? When was it reported? When could the technician start? When was it fixed? As long as these time stamps are compiled manually from reports, they are inaccurate – and the KPIs based on them are correspondingly unreliable.
Where your service KPIs stand today – clarified in 30 minutes.
We take a specific type of system and show which drill-down options arise with a structured machine file – and which stage your service organization should tackle next. Not a slide presentation, but a data view of your specific situation.
→ Arrange an initial consultation
The 4-step model as a KPI foundation
logicline works according to a four-stage model that also applies to KPI architecture: each stage creates the prerequisites for the next. Anyone who wants to invest in level 3 (Decide) without having level 1 (Digitize) is building on an unstable foundation. You can find a detailed treatment of the four stages as a strategy framework in the Pillar Service Strategies in Mechanical Engineering.
For the KPI question, this means specifically:
| Stage | What happens | KPI effect |
|---|---|---|
| 1. digitize | Digital machine file on Salesforce – machines, components, configurations, service history structured | Reference objects for each KPI: per system, per component, per machine type |
| 2. networking | IoT telemetry, tickets, contract data and service report in the same platform | Automated timestamps replace manual entries; KPIs become available in real time |
| 3. decision making | Service Decision Intelligence (SDI) uses networked data for triage and diagnosis | KPIs become a lever: more precise triage increases FTFR, networked diagnosis shortens MTTR |
| 4. automation | Workflow automation, KPI-driven service control | KPIs control routines autonomously – escalation only in case of deviation |
The order is not arbitrary. If you set up a KPI dashboard without a clean machine file, you get nice charts with questionable content. If you place an SDI layer over fragmented data, you get AI answers without a reliable basis. The clean sequence saves a lot of time along the way.
Operational core KPIs – FTFR, MTTR, MTBF
First-time fix rate (FTFR) – how often the problem is solved on first use
Formula: FTFR = Stakes with first solution / total number of stakes × 100
In the mechanical engineering service industry, FTFR rates typically range from two-thirds to three-quarters—the exact figure depends heavily on machine complexity, service structure, and definition. The definition itself is the first critical question: Does “successful” mean that the symptom has disappeared, or does the root cause have to be resolved? A technician who restarts a system without fixing the underlying malfunction can statistically improve the FTFR—even though the problem will resurface in two weeks.
Without a structured machine file, the FTFR is a total figure that contributes nothing to the action. With a machine file, it can be filtered according to machine type, region, technician and fault class. Only then does it become clear whether a certain machine type is systematically difficult to triage, whether a service partner performs significantly worse than the average, or whether certain fault classes recur.
Service Decision Intelligence (SDI) improves FTFR through more precise triage before deployment. Based on service history and telemetry, the intelligence layer identifies typical error patterns and provides recommendations on spare parts, technician qualifications and diagnostic steps – with proof of source. The Pillar When Agentforce hallucinates describes how this layer works and why it does not hallucinate: Why AI in service needs its own knowledge base.
Mean Time to Repair (MTTR) – how long the solution takes
Formula: MTTR = total unplanned downtime / number of faults
The MTTR measures the time between ticket opening and closing. The biggest problem is the distortion caused by waiting times: If a technician has to wait three hours for a spare part, this time is included in the MTTR even though the actual problem lies in the logistics, not the repair. A clean KPI architecture therefore divides MTTR into sub-KPIs – active repair time, waiting time for spare part, waiting time for specialist escalation, waiting time for customer feedback.
A second typical problem is incomplete time stamps. If the exact start of a fault is not recorded – whether due to manual input or missing IoT signals – the ticket cannot be reliably included in the MTTR calculation. Accordingly, the DIN EN 15341 standard recommends only including faults with fully documented start and end times in the calculation. In practice, this means: automated time stamps via PLC or IoT signals instead of relying on service report entries.
How SDI reduces MTTR through networked diagnostic recommendations – and how Empolis Service Express plays a role in this – is described in the Knowledge Management in Service pillar.
Mean Time Between Failures (MTBF) – on the system side instead of the process side
Formula: MTBF = sum of operating times between failures / number of failures
MTBF is shifting from a service-oriented perspective to an engineering one: it measures the reliability of the system itself. The main challenge is establishing a uniform definition of “failure”—is it defined by an error code, a production shutdown, or a safety shutdown? Without a clear, system-level definition, comparisons between plant types or locations become meaningless.
What really matters: Deriving MTBF from telemetry rather than from service reports. Telemetry provides objective timestamps; service reports are manual and reactive. A digital machine record that feeds IoT data directly into the asset structure enables MTBF analyses by machine type, year of manufacture, or operating environment—and often reveals that not all machines are equally “reliable,” but rather that there are clusters based on maturity levels.
Customer and contract KPIs – NPS and SLA compliance
Net Promoter Score (NPS) – when it becomes meaningful
Formula: NPS = % promoters – % detractors, scale -100 to +100
The NPS is a simple metric with a complex interpretation problem. The standard question is “How likely are you to recommend our company?” on a scale of 0 to 10. However, annual global surveys often provide only a general sense of sentiment without any reference to specific service interactions. If a customer had a frustrating service experience in February and gives a general response in July—what does the score actually measure?
NPS only becomes meaningful if it is collected immediately after each service call, on a ticket-specific and machine-related basis. The evaluation can then be correlated with FTFR, MTTR, technician, system type and contract class. This makes it possible to recognize: On which machine types are customers systematically dissatisfied, even though the operational KPI is average? Which service partners generate NPS detractors?
SDI takes NPS analysis to the next level by linking operational data with NPS feedback. Patterns become visible that would disappear in the noise without a structured database.
SLA compliance – ticket level instead of contract average
SLA compliance measures whether agreed response and resolution times are adhered to. Two central challenges:
Reference level. An SLA compliance of 96% at contract level can hide individual customers with chronic violations. Only a drill-down to ticket, asset and contract level makes systematic weaknesses visible. Structurally, this means that SLA compliance must be evaluated in at least two views – per contract and per asset class.
Data situation. Without automated time recording, the exact time of the fault report is often missing. Manual time stamps distort the response time measurement. DIN EN 15341 therefore recommends only including tickets with complete time stamps in the SLA calculation – which in practice means that the recording must be done via IoT and ticket system integration, not via manual reports.
An integrated database of machine files (reference object), tickets in Salesforce (workflow) and telemetry (time source) turns SLA compliance into a controllable KPI. Chronic underperformance – for example in a specific contract model or region – can be identified and addressed before it leads to contract termination.
Economic KPIs – cost-per-ticket, service margin, aftermarket turnover
Operational KPIs show service quality. Economic KPIs show whether the service is also financially viable.
Cost-per-ticket – the complete cost overview
Formula: Cost-per-ticket = total costs of all service assignments / number of completed tickets
The most common mistake: only including personnel costs in the calculation. If you want to calculate cost-per-ticket in a way that is relevant to management, you need a complete cost model – personnel, material, travel costs, overheads and, if applicable, opportunity costs for the tied-up technician capacity. Only then can comparisons between contract classes, machine types or regions show where cost drivers lie.
A higher FTFR directly lowers the number of repeat operations – and therefore the cost-per-ticket. If you look at FTFR and cost-per-ticket together, you can see whether first-time fix improvements actually have an economic impact or whether other cost drivers dominate.
Service margin – profitability of the service business
The service margin measures the ratio of service revenue to service costs. Successful mechanical engineering service organizations achieve margins that are significantly higher than those of the new machine business – service business is structurally higher-margin because it is less material-intensive. Low margins are generally not a market issue, but have internal causes: inefficient deployment planning, manual spare parts scheduling, a reactive service organization that carries out emergency assignments instead of planned maintenance.
Structured service based on networked data shifts the margin in two directions: fewer emergency call-outs (cost reduction) and more maintenance contracts with predictable sales (revenue increase). The pillar Service sales campaigns from the installed base describes how structured service sales campaigns from the installed base work.
Aftermarket sales per machine – the life cycle view
The aftermarket turnover per machine shows the financial contribution of each system over its entire life cycle – spare parts, maintenance contracts, modernizations, digital services. The KPI turns the installed base into a controllable business: Which machine clusters contribute the most? Which ones will soon fall out of the service window? Where are modernization or replacement offers worthwhile?
Without a structured machine file, this view remains in the fog. With a machine file, aftermarket sales per system, per generation, per contract model can be evaluated – and thus become triggers for targeted sales campaigns. If you want to use the end-of-life business strategically, you will find an in-depth look in Pillar End-of-Life Management in Mechanical Engineering.
Integrated KPI architecture – from measurement to control
Individual KPIs are not enough. The effect comes from the link: FTFR with cost-per-ticket, MTTR with NPS, MTBF with aftermarket turnover per machine. Only these correlations show where improvements actually have an impact on business success.
A typical scenario from mechanical engineering
A manufacturer with approximately 4,500 installed systems worldwide currently manages its master data in Excel and an ERP system. The FTFR is calculated manually on a quarterly basis and stands at “somewhere between 60 and 65%.” There is no drill-down capability by system type, region, technician, or fault class. No one can provide a clear explanation for why the FTFR has been stagnant for the past three years—and so attempts at improvement remain a matter of trial and error.
Step 1: Structured machine file on Salesforce. The installed systems are recorded in a structured manner with components, configurations and service history – the Installed Base Assessment provides the inventory in four to six weeks.
Step 2: Tickets, telemetry, and contract information are integrated into the same platform. FTFR and MTTR can be analyzed in real time—by system type, region, and technician. It turns out that the 62% average masks an FTFR of 78% for the main product line and 41% for a smaller specialized system. The leverage lies not in “the overall FTFR,” but in the specialized system.
Step 3: SDI analyzes the networked data and identifies the pattern – the special system requires a specific tool configuration and technician qualification, which is missing in 60% of initial deployments. The triage is adjusted. Within two quarters, the FTFR of the special plant increases from 41% to 70%.
This is the difference between KPI as a report and KPI as a management tool.
SDI as a lever – not just measurement, but improvement
Service Decision Intelligence is more than just a dashboard. The intelligence layer works on the networked service data and actively improves KPIs:
- FTFR increases through more precise triage before deployment – based on patterns from the entire service history
- MTTR decreases through context-related diagnostic recommendations with source reference
- NPS becomes controllable because detractor patterns become visible across plant types and regions
- SLA compliance can be proactively monitored because risk tickets can be identified at an early stage
Three properties make SDI particularly relevant for KPI applications:
Data sovereignty. SDI runs on the machine manufacturer’s infrastructure – Azure or AWS in the EU region – not in shared LLM environments. Service data does not flow into external training models, the AI Act is fulfilled.
Source reference. Every recommendation comes with an audit trail. When SDI says, “The likely cause of the error is Sensor X, based on comparable cases Y and Z”—those sources are available for review. This is crucial for regulatory-sensitive service cases (claims, warranty).
GRAX as Lakehouse. The complete Salesforce service history is accessible without API limits. Service history, telemetry and contract data are stored in the same context and are available for KPI analyses.
The Pillar IoT data in service deployment describes how this changes the service console on site – and how the technician actually benefits from it.
Next step
Service KPIs in mechanical engineering are not a reporting problem, but a data structure problem. The tools (Salesforce Service Cloud, machine file, IoT platforms, SDI) are mature. What makes the difference is the sequence: structuring data, networking systems, making decisions with proof of source.
The first step is therefore not the next BI selection, but an honest inventory: Where does the service data stand today? Which KPIs can really be filtered? Where does the drill-down capability break down?
Two pragmatic approaches:
- Installed Base Assessment – when the machine data is scattered today and KPI reports are compiled manually from Excel. In four to six weeks, the database is structured and a concrete KPI leverage report is delivered.
- Initial consultation – if the database is already in place and you want to specifically assess which level will bring the greatest KPI leverage next – and whether SDI makes sense on your service history as a lever on FTFR and MTTR.
FAQs
What data is required to correctly measure FTFR and MTTR?
At least three types of data must be available in a structured form: the number of assignments with first solution versus total number of assignments (for FTFR), the exact time stamps of ticket opening and closing (for MTTR) and the reference objects for drill-down – machine type, component, region, technician. Without these reference objects, FTFR and MTTR remain calculable but not controllable.
How do I define "resolved" in FTFR so that the KPI remains comparable?
A service case is considered to have been resolved on first use if the underlying cause has been eliminated – not just the symptom. A system restart that temporarily suppresses the fault does not count as an initial solution. This definition must be defined in the system (e.g. via a mandatory field “Cause resolved” in the service report), otherwise every technician and every region will interpret the FTFR differently.
How do I link NPS with FTFR and MTTR after a service call?
Collect the NPS immediately after each service call, on a ticket-specific and machine-related basis. This allows the NPS value to be correlated with FTFR, MTTR, technician, system type and contract class. This link makes patterns visible – for example, systematically low NPS values for a certain type of system despite an average MTTR. Such patterns can only be recognized if the KPIs are based on the same reference object.
What role does IoT data play in KPI recording?
IoT data provides objective time stamps and supplements or replaces manual entries. The start of a fault can be recorded automatically via PLC or sensor signals instead of waiting for a call from the customer. This makes MTTR calculations more precise and MTBF can be reliably evaluated in the first place. The Pillar [IoT data in the service portal](https://www.logicline.de/iot-daten-im-serviceportal-nutzbar-machen) describes how these data flows run from the sensor to the service console.