A Finnish manufacturer buying an AI fault-diagnosis assistant should require every useful suggestion to connect to relevant machine evidence and a clear operating context. The pilot should show what the system observed, which sources support its explanation and when it cannot reach a dependable conclusion. Start with a bounded advisory workflow that maintenance staff can evaluate.
VTT’s 22 June 2026 announcement of the CLEAR project describes work on trustworthy industrial AI, including the challenge of combining diverse data and explaining behaviour. It is a research programme, not proof that a particular commercial assistant is ready for a factory. The practical buying opportunity is to turn those concerns into explicit engineering deliverables.
Choose a diagnosis task that people can review
Imagine a fictional Finnish equipment manufacturer supporting a production cell. Operators record alarms, technicians write maintenance notes, and sensor histories live in a separate system. An assistant could assemble a short evidence pack for a technician investigating a recurring stoppage.
Define the initial task narrowly: retrieve the relevant history, identify comparable recorded incidents and suggest questions for investigation. The technician remains responsible for the maintenance decision. Changes to machine settings, safety functions or production control are outside this proposed first release.
Ask technicians which missing information slows diagnosis today. Their answer may reveal that inconsistent asset identifiers or unavailable manuals cause more delay than the absence of a language model. The discovery phase should be allowed to reach that conclusion.
Align the records before combining them
Create a shared asset reference linking the sensor source, alarm log, maintenance record and applicable documentation. Record when equipment was replaced or reconfigured. A familiar asset name can conceal a different component generation with different operating behaviour.
Then inspect time. Source systems may record local time, UTC, collection time or upload time. The proposed data contract should distinguish those meanings and preserve the original value. Test the ordering of a known incident across all sources before using it as an evaluation case.
Make missing measurements visible. A gap in a sensor stream should not be presented as a stable reading, and a late-uploaded note should not silently appear to have been available when the original decision was made.
Build an evidence pack with an answer boundary
For the example workflow, a useful output could contain the affected asset, relevant alarm sequence, selected measurements, linked maintenance observations and applicable document sections. It should separate those observations from the assistant’s proposed explanation.
Ask the vendor to show exactly how the output changes when a source is unavailable. A confident paragraph with a missing manual reference is difficult for a technician to assess. The interface should identify the gap and offer a useful next step, such as checking the asset version or retrieving an approved document.
Language needs should come from the actual team. Finnish and English maintenance notes may use different abbreviations for the same component. Build an approved vocabulary with technicians, and preserve equipment codes and measurement units during any translation or summarisation.
Evaluate difficult cases, including ordinary operation
Agree a versioned evaluation set using approved historical or synthetic material. Include more than memorable failures. A proposed set should cover:
- ▸A well-documented fault with a known resolution.
- ▸Similar symptoms produced by different underlying causes.
- ▸Normal operation that should not trigger an alarming diagnosis.
- ▸Missing sensor data during an otherwise familiar event.
- ▸An outdated manual linked to a newer component.
- ▸A maintenance note that contradicts an earlier working hypothesis.
- ▸An unfamiliar operating condition outside the pilot’s evidence base.
Keep review cases separate from examples repeatedly used during development. Ask domain reviewers to assess evidence relevance, factual consistency and whether the output supports a sensible next investigation. A fluent answer alone is an inadequate acceptance criterion.
Measure assistance at the maintenance desk
Choose measures tied to the task. For an evidence-gathering assistant, useful observations might include time to locate the required records, omitted relevant facts and unnecessary investigation prompts. Record how often the technician accepts, changes or rejects the suggested explanation, along with the reason.
Separate performance across operating conditions. A model that works well on the most common product run may struggle after a tooling change. Review results by the contexts that maintenance considers meaningful rather than relying only on an overall average.
VTT’s industrial AI service description highlights combining sensor streams, textual information and domain expertise. For this proposed pilot, domain expertise belongs in the review process and acceptance decision, not merely in an initial interview.
Run beside the existing process first
A shadow-mode trial lets the team compare the assistant with the current workflow without making it an operational authority. Define who may see outputs, how feedback is captured and what causes the trial to pause. Preserve the existing maintenance process throughout the agreed evaluation period.
Include the application’s ordinary failure modes: an unavailable data source, slow response, expired access and a model update that changes answers. The technician should still be able to reach the underlying records and continue the established investigation route.
The handover needs a versioned evaluation set, source inventory, access model and named owner for future changes. If a sensor mapping or maintenance taxonomy changes, someone must know which tests to rerun and who approves the result.
Commission a pilot with an honest decision at the end
The outcome may justify broader use, a narrower feature or more investment in data quality. Write those possible decisions into the proposal so the team is paid to produce useful evidence rather than defend a predetermined launch.
TuniCyberLabs can help build industrial data connections and advisory applications through custom software development. Describe your Finnish maintenance workflow, including the asset type, data sources and decision the technician must make. Our AI vendor-evaluation guide can help compare the proposed evidence, operating boundaries and handover across suppliers.
