Wheel type inspection robot and camera motion understanding method and system thereof

By using a wheeled inspection robot's data acquisition, camera motion classification, SfM-VLM fusion inference, and multimodal decision-making modules, the problems of decoupling camera motion from scene dynamics, target tracking, and multimodal information fusion in complex environments are solved, improving pose estimation accuracy and inspection report accuracy.

CN120766376BActive Publication Date: 2025-11-04ANHUI SYMMETRY AXIS INTELLIGENT SECURITY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511250330.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-11-04
Estimated Expiration
2045-09-03

AI Technical Summary

Technical Problem

Existing wheeled inspection robots face challenges in complex environments, including decoupling camera motion from scene dynamics, target tracking failure, and multimodal information fusion defects. These issues result in large pose estimation errors, decreased target detection accuracy, and high false alarm rates.

Method used

The system employs a data acquisition module to acquire displacement, video stream, and pose data. A camera motion classification module performs video frame processing and visual inspection. A natural language description is generated by combining the SfM-VLM fusion inference module. Inspection decisions are generated through a structured annotation module and a multimodal decision module, thereby achieving the fusion and decision-making of multimodal information.

Benefits of technology

It improves the accuracy of pose estimation in dynamic environments, reduces the target loss rate and false alarm rate, and enhances the consistency of inspection dataset annotation and the accuracy of report generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766376B_ABST
    Figure CN120766376B_ABST
Patent Text Reader

Abstract

The application discloses a wheeled inspection robot and a camera motion understanding method and system thereof, wherein displacement, video stream data and pose data are acquired in real time through a data acquisition module; a camera motion classification module classifies motion and detects vision for a video frame, outputs geometric primitives and semantic primitive labels and device state detection results, and maps to a reference frame; an SfM-VLM fusion inference module extracts geometric parameters of camera motion through an SfM algorithm, inputs a VLM model after fusing visual features to generate natural language semantic description; a structured annotation module generates double-track annotations of a label layer and a subtitle layer based on geometric primitives, semantic primitive labels and natural language semantic description; and a multi-modal decision module fuses geometric primitives and semantic primitive labels, device state detection results and data collected by the data acquisition module to generate an inspection decision; the problems of motion decoupling, target tracking, data annotation and multi-modal decision in a dynamic environment are solved, and the inspection reliability and intelligent level are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent inspection robot technology, and specifically discloses a wheeled inspection robot and its camera motion understanding method and system. Background Technology

[0002] In the field of industrial automation and intelligent monitoring, wheeled inspection robots have been widely used for equipment condition monitoring and fault diagnosis, but existing technologies face the following technical bottlenecks in complex environments:

[0003] 1. The challenge of decoupling camera motion from scene dynamics: Traditional SLAM algorithms have difficulty distinguishing between the robot's own motion and the motion signals of dynamic objects in the scene. For example, during factory assembly line inspection, the rotation of the conveyor belt captured by the camera during the robot's forward movement is easily misjudged as camera translation, leading to the accumulation of errors in 3D map construction and a decrease in target detection accuracy. Experimental data shows that the pose estimation error of traditional SLAM in dynamic scenes can reach more than ±5cm.

[0004] 2. Target tracking failure under complex motion modes: When a robot performs complex motions such as translation and rotation or tracks a moving target, traditional visual tracking algorithms lack the ability to understand motion semantics and have difficulty keeping the target in the center of the field of view. For example, when an inspection drone tracks a power transmission line, the camera performs horizontal and vertical rotation at the same time, and the target loss rate can reach more than 40%.

[0005] 3. Deficiencies in multimodal information fusion and decision-making: Robots have difficulty associating camera motion semantics with data from sensors such as infrared and lidar, and cannot form structured inspection reports. For example, when the camera detects a hot spot on the equipment, it cannot combine the motion trajectory to determine whether the hot spot is a real fault or a motion artifact, resulting in a false alarm rate as high as 15%.

[0006] Therefore, it is necessary to invent a wheeled inspection robot and its camera motion understanding method and system to solve the above problems. Summary of the Invention

[0007] To overcome the aforementioned deficiencies in the prior art, this invention provides a wheeled inspection robot and its camera motion understanding method and system. The system acquires displacement, video stream data, and pose data in real time through a data acquisition module; a camera motion classification module performs motion classification and visual detection on video frames, outputting geometric primitives and semantic primitives labels, along with equipment status detection results, which are then mapped to a reference frame; an SfM-VLM fusion inference module extracts the geometric parameters of camera motion using the SfM algorithm, fuses visual features, and inputs them into a VLM model to generate a natural language semantic description; a structured annotation module generates dual-track annotations (label layer and subtitle layer) based on geometric primitives, semantic primitives labels, and natural language semantic descriptions; and a multimodal decision module integrates geometric primitives and semantic primitives labels, equipment status detection results, and data acquired by the data acquisition module to generate inspection decisions, effectively solving the problems mentioned in the background art.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a wheeled inspection robot and its camera motion understanding method and system, specifically including: a wheeled inspection robot, a data acquisition module, a camera motion classification module, an SfM-VLM fusion inference module, a structured annotation module, a multimodal decision module, and a computer-readable storage medium. The wheeled inspection robot is equipped with a high-precision encoder. The data acquisition module includes a vision module and a multi-sensor fusion unit. The vision module includes an RGB camera and an infrared camera. The multi-sensor fusion unit includes an inertial measurement unit (IMU) and a lidar.

[0009] Data acquisition module: Real-time acquisition of displacement data, video stream data, and pose assistance data;

[0010] Camera motion classification module: Based on the ResNet+Transformer model, it preprocesses video frames, extracts visual features and performs temporal modeling, outputs geometric primitives and semantic primitives labels and device status detection results, and maps them to the reference frame;

[0011] SfM-VLM Fusion Inference Module: Extracts the geometric parameters of camera motion using the Structure for Motion Recovery (SfM) algorithm, and uses the Visual Language Model (VLM) to convert the geometric parameters into semantic intents described in natural language.

[0012] Structured annotation module: Generates dual-track annotations for both tag and subtitle layers;

[0013] Multimodal decision module: integrates geometric primitives and semantic primitives, equipment status detection results and data collected by the data acquisition module to generate inspection control instructions and structured reports;

[0014] Computer-readable storage medium: storing programs that implement the functions of the data acquisition module, camera motion classification module, SfM-VLM fusion inference module, structured annotation module, and multimodal decision module described above.

[0015] Preferably, the displacement data includes the translation distance and heading angle output by the chassis encoder; the video stream data includes RGB video and infrared thermal imaging video; the pose assistance data includes the angular velocity, acceleration and lidar point cloud of the IMU; and the geometric parameters include the translation vector, rotation matrix and focal length variation.

[0016] Preferably, the geometric primitives include translation, rotation, and scaling; the semantic primitives include tracking, arc motion, and scene revealing; and the reference frame includes the camera center, the object center, and the ground center.

[0017] Preferably, the specific execution steps of the camera motion classification module are as follows:

[0018] Preprocessing of input video frames, including image scaling and normalization;

[0019] Visual features are extracted using the ResNet model;

[0020] Temporal modeling of continuous frame features is performed using the Transformer model;

[0021] Using a pre-trained camera motion classification model, motion primitive labels and confidence scores are output based on visual features, and device visual detection is performed simultaneously. The motion primitive labels include geometric and semantic primitive categories.

[0022] Map motion primitive labels to a camera center, object center, or ground center reference frame.

[0023] Preferably, the specific execution steps of the SfM-VLM fusion inference module are as follows:

[0024] Based on video frame sequences and IMU data, the camera pose trajectory is calculated using an improved version of the COLMAP algorithm;

[0025] Extract geometric parameters;

[0026] Input is constructed by fusing geometric parameters and visual features;

[0027] The input is transformed into a natural language semantic description by fine-tuning the generative visual language model Qwen2.5VL.

[0028] Preferably, the specific execution steps of the structured annotation module are as follows:

[0029] Tag layer generation, labeling geometric motion type, semantic motion type, stability status and equipment detection results;

[0030] Subtitle layer generation combines motion intent and scene information to generate natural language descriptions using VLM;

[0031] Integrate the tag layer and subtitle layer to output structured annotations with timestamps.

[0032] Preferably, the specific execution steps of the multimodal decision module are as follows:

[0033] Aligning camera motion semantics with multi-sensor data based on timestamps;

[0034] Determine the authenticity of equipment anomalies based on motion stability status;

[0035] If an anomaly is found, an alarm command and a fault report will be generated; otherwise, a normal inspection report will be generated.

[0036] Output robot control commands based on motion semantic intent.

[0037] Preferably, the wheeled inspection robot and its camera motion understanding method specifically include the following steps: S1, acquiring displacement data, video stream data and pose data in real time through the data acquisition module;

[0038] S2, the camera motion classification module performs motion classification and visual detection on video frames, outputs geometric primitive and semantic primitive labels and device status detection results, and maps them to the reference frame;

[0039] The S3 and SfM-VLM fusion inference modules extract the geometric parameters of camera motion using the SfM algorithm, fuse visual features, and then input them into the VLM model to generate a natural language semantic description.

[0040] S4. The structured annotation module generates dual-track annotations for the tag layer and the subtitle layer based on geometric primitives, semantic primitive tags, and natural language semantic descriptions.

[0041] S5, the multimodal decision module integrates geometric primitives and semantic primitives labels, equipment status detection results and data collected by the data acquisition module to generate inspection decisions.

[0042] The technical effects and advantages of this invention are as follows:

[0043] 1. Enhanced adaptability to dynamic environments: By decoupling from motion primitives, the pose estimation error of SLAM in dynamic scenes is reduced by more than 40%. For example, in the inspection of a conveyor belt during operation, the 3D reconstruction accuracy is improved to ±2cm;

[0044] 2. Enhanced target tracking reliability: The tracking strategy based on semantic primitives reduces the target loss rate by 60% in complex motion scenarios. For example, when simultaneous translation and rotation occur, the valve detection accuracy increases from 72% to 91%.

[0045] 3. Standardized data annotation: The structured annotation framework improves the annotation consistency of the inspection dataset by 85% and enhances the model's generalization ability across scenarios by 20-30%;

[0046] 4. Intelligent multimodal decision-making: Motion-enhanced report generation reduces human error rate by 50%, and the false alarm rate of equipment overheating faults is reduced from 15% to below 5%. Attached Figure Description

[0047] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.

[0048] Figure 1 This is a schematic diagram of the overall structure of the present invention.

[0049] Figure 2 This is a flowchart illustrating the overall steps of the present invention. Detailed Implementation

[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] This invention provides a wheeled inspection robot and its camera motion understanding system, the structure of which is as follows: Figure 1 As shown, it specifically includes: a wheeled inspection robot, a data acquisition module, a camera motion classification module, an SfM-VLM fusion inference module, a structured annotation module, a multimodal decision module, and a computer-readable storage medium. The wheeled inspection robot is equipped with a high-precision encoder. The data acquisition module includes a vision module and a multi-sensor fusion unit. The vision module includes an RGB camera and an infrared camera. The multi-sensor fusion unit includes an inertial measurement unit (IMU) and a lidar.

[0052] Furthermore, in the above technical solution, the data acquisition module is used to acquire displacement data, video stream data, and pose assistance data in real time;

[0053] The camera motion classification module is used to preprocess video frames, extract visual features and perform temporal modeling based on the ResNet+Transformer model, output geometric primitives and semantic primitives labels and device status detection results, and map them to the reference frame.

[0054] The SfM-VLM fusion inference module is used to extract the geometric parameters of camera motion through the Structure for Motion Recovery (SfM) algorithm, and to convert the geometric parameters into semantic intents described in natural language using the Visual Language Model (VLM).

[0055] The structured annotation module is used to generate dual-track annotations for the tag layer and the subtitle layer;

[0056] The multimodal decision module is used to integrate geometric primitives and semantic primitives, equipment status detection results and data collected by the data acquisition module to generate inspection control instructions and structured reports;

[0057] The program is a computer-readable storage that implements the functions of the data acquisition module, camera motion classification module, SfM-VLM fusion inference module, structured annotation module, and multimodal decision module described above.

[0058] like Figure 2 As shown, the wheeled inspection robot and its camera motion understanding method specifically include the following steps:

[0059] S1. Real-time acquisition of displacement data, video stream data, and pose data via the data acquisition module;

[0060] Furthermore, in the above technical solution, the displacement data includes the translation distance and heading angle output by the chassis encoder; the video stream data includes RGB video and infrared thermal imaging video; the pose assistance data includes the angular velocity, acceleration and lidar point cloud of the IMU; and the geometric parameters include translation vector, rotation matrix and focal length change.

[0061] S2, the camera motion classification module performs motion classification and visual detection on video frames, outputs geometric primitive and semantic primitive labels and device status detection results, and maps them to the reference frame;

[0062] Furthermore, in the above technical solution, the geometric primitives include translation, rotation, and scaling; the semantic primitives include tracking, arc motion, and scene revealing; and the reference frame includes the camera center, the object center, and the ground center.

[0063] Furthermore, in the above technical solution, the specific execution steps of the camera motion classification module are as follows:

[0064] Preprocessing of input video frames, including image scaling and normalization;

[0065] Visual features are extracted using the ResNet model;

[0066] Temporal modeling of continuous frame features is performed using the Transformer model;

[0067] Using a pre-trained camera motion classification model, motion primitive labels and confidence scores are output based on visual features, and device visual detection is performed simultaneously. The motion primitive labels include geometric and semantic primitive categories.

[0068] Map motion primitive labels to camera center, object center, or ground center references.

[0069] It should be further explained that the camera motion classification model is trained based on a classification system containing more than 50 motion primitives, taking the video stream as input and outputting motion primitive labels.

[0070] It should be further explained that after outputting motion primitive labels, the camera motion classification module needs to select a reference frame based on scene dynamics and motion stability. The mapping logic is as follows:

[0071] When the target object dominates the field of view, that is, the target detection box covers more than 30% of the image area, the motion semantics are based on the target, that is, the motion primitive label is mapped to the central frame of the object.

[0072] When the robot's motion is stable, i.e., the IMU acceleration variance is <0.1 m / s 2 Motion semantics are based on the ground coordinate system, that is, the motion primitive labels are mapped to the ground center frame;

[0073] When the above two rules are not met, i.e. the target is small or the robot's movement is violent, the default is to use the camera's own coordinate system as the reference, that is, to map the motion primitive label to the camera's central frame.

[0074] The S3 and SfM-VLM fusion inference modules extract the geometric parameters of camera motion using the SfM algorithm, fuse visual features, and then input them into the VLM model to generate a natural language semantic description.

[0075] Furthermore, in the above technical solution, the specific execution steps of the SfM-VLM fusion inference module are as follows:

[0076] Based on video frame sequences and IMU data, the camera pose trajectory is calculated using an improved version of the COLMAP algorithm;

[0077] Extract geometric parameters;

[0078] Input is constructed by fusing geometric parameters and visual features;

[0079] The input is transformed into a natural language semantic description by fine-tuning the generative visual language model Qwen2.5VL.

[0080] It should be further explained that the improved COLMAP algorithm initializes the camera pose using IMU data and uses dynamic object masks to filter point clouds, reducing dynamic scene errors; the Qwen2.5VL model fine-tuning method uses the COCO-Motion dataset, with geometric parameters + ResNet features as input and semantic descriptions as labels for supervised fine-tuning.

[0081] S4. The structured annotation module generates dual-track annotations for the tag layer and the subtitle layer based on geometric primitives, semantic primitive tags, and natural language semantic descriptions.

[0082] Furthermore, in the above technical solution, the specific execution steps of the structured annotation module are as follows:

[0083] Tag layer generation, labeling geometric motion type, semantic motion type, stability status and equipment detection results;

[0084] Subtitle layer generation combines motion intent and scene information to generate natural language descriptions using VLM;

[0085] Integrate the tag layer and subtitle layer to output structured annotations with timestamps.

[0086] S5, the multimodal decision module integrates geometric primitives and semantic primitives labels, equipment status detection results and data collected by the data acquisition module to generate inspection decisions.

[0087] Furthermore, in the above technical solution, the specific execution steps of the multimodal decision module are as follows:

[0088] Aligning camera motion semantics with multi-sensor data based on timestamps;

[0089] Determine the authenticity of equipment anomalies based on motion stability status;

[0090] If an anomaly is found, an alarm command and a fault report will be generated; otherwise, a normal inspection report will be generated.

[0091] Output robot control commands based on motion semantic intent.

[0092] Taking power line inspection as an example, a wheeled robot starts from the starting point and inspects high-voltage equipment along a preset path:

[0093] 1. Data Acquisition: The chassis encoder outputs displacement data in real time: translational distance and heading angle; the RGB camera acquires 1080p video stream; the infrared camera outputs 640×480 thermal imaging video; the IMU provides angular velocity and acceleration; and the lidar generates point cloud.

[0094] 2. Camera motion classification and device detection: RGB frames are scaled to 256×256 and pixel values ​​are normalized to [0,1]; visual features are extracted using ResNet-50 and YOLOv7 model is run simultaneously to detect device status; Transformer encoder processes features from 10 consecutive frames and outputs motion primitive labels; motion primitive labels are mapped to reference frames based on the position of the target detection box.

[0095] 3. SfM-VLM Fusion Inference: Based on the SfM algorithm, the camera trajectory is calculated to obtain the translation vector and rotation matrix. The geometric parameters and extracted visual features are fused to construct the input. The input is converted into a natural language semantic description through a fine-tuned generative visual language model, Qwen2.5VL. For example, in a certain data acquisition, the translation vector is 0.3 / s and the rotation matrix is ​​an upward rotation of 20°. Visual inspection reveals a 5mm crack in the top insulator of the transformer. The Qwen2.5VL model converts this into "The camera moves forward at a speed of 0.3m / s and rotates upward by 20°, intending to check the top insulator of the transformer". It also associates this with the visually detected insulator crack to generate the annotation "The crack is about 5mm long and is located on the third insulator from the left".

[0096] 4. Multimodal decision generation: If an infrared camera detects a temperature of 65℃ in a crack area, which is 20℃ higher than the ambient temperature, and the camera is stationary, eliminating motion artifacts, the multimodal decision module determines that this anomaly is a real fault, triggers an alarm, and generates a report that includes motion trajectory, visual image, temperature heat map, and semantic description.

[0097] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A wheeled inspection robot and its camera motion understanding system, characterized in that, include: The system includes a wheeled inspection robot, a data acquisition module, a camera motion classification module, an SfM-VLM fusion inference module, a structured annotation module, a multimodal decision module, and a computer-readable storage medium. The wheeled inspection robot is equipped with a high-precision encoder. The data acquisition module integrates a vision module and a multi-sensor fusion unit. The vision module includes an RGB camera and an infrared camera. The multi-sensor fusion unit includes an inertial measurement unit (IMU) and a lidar. Data acquisition module: Real-time acquisition of displacement data, video stream data, and pose assistance data; Camera motion classification module: Based on the ResNet+Transformer model, it preprocesses video frames, extracts visual features and performs temporal modeling, outputs geometric primitives and semantic primitives labels and device status detection results, and maps them to the reference frame; SfM-VLM Fusion Inference Module: Extracts the geometric parameters of camera motion using the Structure for Motion Recovery (SfM) algorithm, and uses the Visual Language Model (VLM) to convert the geometric parameters into semantic intents described in natural language. Structured annotation module: Generates dual-track annotations for both tag and subtitle layers; Multimodal decision module: integrates geometric primitives and semantic primitives, equipment status detection results and data collected by the data acquisition module to generate inspection control instructions and structured reports; Computer-readable storage medium: storing programs that implement the functions of the data acquisition module, camera motion classification module, SfM-VLM fusion inference module, structured annotation module, and multimodal decision module described above.

2. The wheeled inspection robot and its camera motion understanding system as described in claim 1, characterized in that: The displacement data includes the translation distance and heading angle output by the chassis encoder; the video stream data includes RGB video and infrared thermal imaging video; the pose assistance data includes the angular velocity and acceleration of the IMU and the point cloud of the lidar; and the geometric parameters include the translation vector, rotation matrix and focal length variation.

3. The wheeled inspection robot and its camera motion understanding system as described in claim 1, characterized in that: The geometric primitives include translation, rotation, and scaling; the semantic primitives include tracking, arc motion, and scene revealing; and the reference frame includes the camera center, the object center, and the ground center.

4. The wheeled inspection robot and its camera motion understanding system as described in claim 1, characterized in that: The specific execution steps of the camera motion classification module are as follows: Preprocessing of input video frames, including image scaling and normalization; Visual features are extracted using the ResNet model; Temporal modeling of continuous frame features is performed using the Transformer model; Using a pre-trained camera motion classification model, motion primitive labels and confidence scores are output based on visual features, and device visual detection is performed simultaneously. The motion primitive labels include geometric and semantic primitive categories. Map motion primitive labels to a camera center, object center, or ground center reference frame.

5. The wheeled inspection robot and its camera motion understanding system as described in claim 1, characterized in that: The specific execution steps of the SfM-VLM fusion inference module are as follows: Based on video frame sequences and IMU data, the camera pose trajectory is calculated using an improved version of the COLMAP algorithm; Extract geometric parameters; Input is constructed by fusing geometric parameters and visual features; The input is transformed into a natural language semantic description by fine-tuning the generative visual language model Qwen2.5VL.

6. The wheeled inspection robot and its camera motion understanding system as described in claim 1, characterized in that: The specific execution steps of the structured annotation module are as follows: Tag layer generation, labeling geometric motion type, semantic motion type, stability status and equipment detection results; Subtitle layer generation combines motion intent and scene information to generate natural language descriptions using VLM; Integrate the tag layer and subtitle layer to output structured annotations with timestamps.

7. The wheeled inspection robot and its camera motion understanding system as described in claim 1, characterized in that: The specific execution steps of the multimodal decision module are as follows: Aligning camera motion semantics with multi-sensor data based on timestamps; Determine the authenticity of equipment anomalies based on motion stability status; If an anomaly is found, an alarm command and a fault report will be generated; otherwise, a normal inspection report will be generated. Output robot control commands based on motion semantic intent.

8. A wheeled inspection robot and its camera motion understanding method, characterized in that: Specifically, the following steps are included: S1. Real-time acquisition of displacement data, video stream data, and pose data via the data acquisition module; S2, the camera motion classification module performs motion classification and visual detection on video frames, outputs geometric primitive and semantic primitive labels and device status detection results, and maps them to the reference frame; The S3 and SfM-VLM fusion inference modules extract the geometric parameters of camera motion using the SfM algorithm, fuse visual features, and then input them into the VLM model to generate a natural language semantic description. S4. The structured annotation module generates dual-track annotations for the tag layer and the subtitle layer based on geometric primitives, semantic primitive tags, and natural language semantic descriptions. S5, the multimodal decision module integrates geometric primitives and semantic primitives labels, equipment status detection results and data collected by the data acquisition module to generate inspection decisions.

Citation Information

Patent Citations

  • Semantic intelligent substation inspection operation robot navigation system and method

    CN111968262A

  • Automatic inspection method and system for orchard and electronic equipment

    CN118658056A