Industry

Robotics Software in Waterloo: Can You Reproduce What the Robot Did?

TuniCyberLabs Team
Archive date:
Published
7 min read

A robot incident needs more than a video. Connect recorded messages, configuration and software versions so a team can investigate without guessing.

The mobile robot stopped beside a clear aisle. By the time an engineer arrived, the pallet had moved, the operator had restarted the system and a newer software build was running. There was a video of the stop, but nobody could establish which map, calibration or configuration had produced it.

This illustrative incident captures a useful question for buyers of robotics software in Kitchener-Waterloo: can the team reconstruct a decision after the physical scene has disappeared? Better dashboards help, but the deeper requirement is a coherent record of the robot's software, inputs and operating context.

From a working demonstration to an explainable fleet

The University of Waterloo's RoboHub fleet includes different robot capabilities coordinated through a custom ROS-based communications network. It is a concrete local example of robotics as connected systems, without implying that every regional manufacturer uses the same architecture.

Once several components exchange messages, an incident can cross boundaries. A perception component sees an obstacle, a planner changes a route and a controller executes the resulting command. A fleet interface may show only the final status. Each component can be functioning as designed while their combined behaviour surprises the operator.

The practical trend is toward treating diagnostic evidence as part of the product. Instead of asking an engineer to assemble fragments after a failure, the system prepares enough context to investigate ordinary operational incidents consistently.

What a recording contains, and what it cannot promise

The ROS 2 project's rosbag2 documentation describes recording and playback of timestamped communications. That makes it useful for examining message sequences and feeding recorded inputs into software. The project also documents recording choices and message-loss considerations. Check capabilities against the distribution and package version you deploy; the repository's development branch is not a promise about every installed release.

A recording is not a complete copy of the physical world. It may omit a topic, start after the relevant event or contain gaps. Replaying messages does not recreate wheel friction, lighting, a moving person or the exact scheduling of every process. Treat replay as an investigation tool with known coverage, rather than proof that every physical outcome will repeat.

That distinction changes the project scope. The developer needs to define what evidence is captured, how completeness is assessed and which questions still require a controlled physical test.

Build an incident bundle around one question

Return to the robot that stopped beside the aisle. The first question might be whether a recent perception change classified a reflection as an obstacle. A useful incident bundle links the relevant recording to the precise software build and configuration used at the time.

It also identifies the robot, sensor calibration, map revision, selected operating mode and the operator action that ended the event. Where clocks differ, the bundle records how timestamps should be interpreted instead of silently presenting every event as perfectly synchronized.

The software team can then compare the old and proposed perception versions against the same recorded inputs in an isolated environment. The comparison should show what changed and what did not. A result such as fewer obstacle detections is an observation to investigate, not automatic permission to deploy a physical robot.

This narrower experiment often provides more value than collecting every stream indefinitely. It connects recording volume to a specific engineering decision.

Keep replay away from unintended physical actions

An investigation environment needs explicit boundaries. Recorded command messages should not reach live actuators through an accidentally shared network or namespace. The team responsible for the physical system must define the test environment and the process for returning changes to equipment.

The application work can support those responsibilities through separate credentials, clear environment indicators, controlled exports and recorded approvals. It should not imply that a software replay substitutes for machinery safety assessment or site commissioning.

Operational staff also need to know whether a diagnostic session is observational or can change configuration. A remote engineer viewing an incident and a technician authorizing a machine change should not receive indistinguishable screens or permissions.

Buy the viewer, build the missing connection

Before commissioning a custom platform, evaluate the recording, visualization and device-management tools already available in your stack. A product may already display sensor streams and messages well. Rebuilding that interface can consume effort while leaving the real gap untouched.

The missing connection may be a reliable link between a recording and the configuration deployed on a particular robot. It may be a support workflow that gathers a bounded incident window or a permissions model that shares diagnostic data with a supplier without exposing unrelated operations.

Compare tools using your own representative incident bundle. Ask whether another engineer can open it, identify missing evidence and understand the versions involved. Our guide to testing strategies that survive refactoring explains why useful tests should preserve observable behaviour rather than merely mirror implementation.

Price the evidence pipeline, not just the screen

Cost depends on sensor volume, recording duration, device connectivity, retention, access controls and the number of hardware and software variants. Transferring a large recording from a constrained site may be a bigger challenge than rendering a timeline.

Start with a representative device and one recurring incident category. Define the capture trigger, export path, retention rule and diagnostic question. Test incomplete uploads and missing metadata before adding more robots. A remote software team can work on these application layers while the customer's authorized specialists retain control of physical equipment and site procedures.

The related device identity and firmware guide covers the identity side of connected equipment. For help connecting diagnostics, configuration history and support workflows, explore custom software development or describe your robotics software project. Include the robot software stack and the incident your team currently struggles to explain.

TAGS
WaterlooCanadaKitchenerRobotics softwareROS 2Diagnostics

Frequently Asked Questions

Can ROS 2 recordings reproduce a robot incident exactly?

+

They can replay captured communications for investigation, but they do not recreate every physical condition or execution detail. Record the relevant configuration and versions, check coverage and distinguish software replay from controlled physical validation.

Should we build a custom robotics visualization tool?

+

Evaluate existing viewers first. Custom development may be better spent connecting recordings to deployed configurations, device identities and support workflows when standard viewers already handle the underlying data well.

Can a remote software team help a Waterloo robotics project?

+

A remote team can develop diagnostic applications, data pipelines and support integrations using agreed access and test environments. Physical equipment changes, machinery safety and site commissioning require clearly assigned customer and specialist responsibilities.

Need help with
this topic
?

Our team specializes in the technologies and strategies discussed in this article. Let’s talk about how we can help your business.

Get in Touch