Unmanned vehicle and robot dog cooperative inspection system and method based on confidence

By using unmanned vehicles and robot dogs for collaborative inspections and calculating confidence correction coefficients using multimodal data, the problem of insufficient confidence in assessment based on single-modal data from unmanned vehicles is solved, resulting in more accurate and reliable facility inspection.

CN121661613BActive Publication Date: 2026-05-01BEIJING YIZHUANG DIGITAL INFRASTRUCTURE TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING YIZHUANG DIGITAL INFRASTRUCTURE TECH DEV CO LTD
Filing Date
2026-02-06
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing unmanned vehicle inspection systems rely on single-modal data for confidence assessment, which cannot fully reflect the multi-dimensional attributes of the target object. This results in low confidence of the initial detection results, making it easy to make misjudgments or omissions. Furthermore, they lack a collaborative processing mechanism for multi-source heterogeneous data.

Method used

The collaborative inspection method of unmanned vehicles and robot dogs involves the robot dog collecting multimodal data (linear acceleration, infrared images, and high-definition images) when the confidence level of the initial detection results is lower than a threshold. Based on this data, multimodal features are calculated and weighted fusion is performed to generate confidence correction coefficients to correct the initial detection results.

Benefits of technology

It improves the accuracy and reliability of target detection results, reduces the risk of misjudgment and omission during inspections, and enhances the overall efficiency and safety of facility inspections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661613B_ABST
    Figure CN121661613B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of robots, in particular to a confidence-based unmanned vehicle and robot dog cooperative inspection system and method, multi-modal data is collected; based on linear acceleration data, the structural response of a target detection object after being subjected to force is analyzed to calculate first modal characteristics; based on first image information, the thermal distribution anomaly of the target detection object is analyzed to calculate second modal characteristics; based on the second image information, the apparent defects of the target detection object are analyzed through a visual detection model to calculate third modal characteristics; the first modal characteristics, the second modal characteristics and the third modal characteristics are weighted and fused to generate a confidence correction coefficient; the preliminary detection result is corrected by using the confidence correction coefficient to obtain a corrected target detection result. In this way, the multi-modal data collected by the robot dog cooperating with the unmanned vehicle is fused and corrected, the accuracy and reliability of the target detection result are improved, and the misjudgment and omission risk in the inspection is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

A confidence-based collaborative inspection system and method for unmanned vehicles and robot dogs Technical Field

[0001] This invention relates to the field of robotics technology, specifically to a confidence-based collaborative inspection system and method for unmanned vehicles and robotic dogs. Background Technology

[0002] Facility inspection is a crucial link in ensuring the safe operation of urban infrastructure. While unmanned vehicles (RVs) are the mainstream inspection vehicle, their limited operation by road conditions restricts them to non-motorized vehicle lanes, resulting in excessively long detection distances for objects such as manhole covers and power poles. This physical distance limitation makes it difficult for the visual sensors on RVs to acquire clear images of targets, especially in complex lighting or occlusion environments, leading to blurry, distorted, or incomplete initial detection results.

[0003] More critically, existing inspection systems rely excessively on single-modal data for confidence assessment, such as analyzing target conditions solely through visible light images. Because single-modal data has limited information dimensions, it cannot comprehensively reflect the target's structural integrity, thermal distribution characteristics, and surface defects, among other multi-dimensional attributes. When the target contains internal voids, microcracks, or hidden corrosion, the system struggles to effectively identify them. This lack of information directly leads to persistently low initial detection confidence, thereby increasing the risk of misjudgments or missed detections.

[0004] Existing technical solutions lack a collaborative processing mechanism for multi-source heterogeneous data, and cannot dynamically supplement key information when the initial confidence level is insufficient, resulting in low reliability of inspection results and difficulty in meeting practical application needs. Summary of the Invention

[0005] To address the technical problem of low reliability in inspection results due to the lack of a collaborative processing mechanism for multi-source heterogeneous data and the inability to dynamically supplement key information when the initial confidence level is insufficient, this invention provides a confidence-based collaborative inspection system and method for unmanned vehicles and robot dogs. This improves the accuracy and reliability of target detection results and reduces the risk of misjudgment and omission during inspection.

[0006] To achieve the above objectives, in a first aspect, this application proposes a confidence-based collaborative inspection method for unmanned vehicles and robotic dogs. This confidence-based collaborative inspection method is applied to a confidence-based collaborative inspection system for unmanned vehicles and robotic dogs. The confidence-based collaborative inspection method for unmanned vehicles and robotic dogs includes:

[0007] When the confidence level of the autonomous vehicle's preliminary detection result of the target object is lower than a preset threshold, a collaborative inspection command containing the target location information is sent to the robot dog.

[0008] In response to the collaborative inspection command, the robot dog is controlled to move to the target location and collect multimodal data, including linear acceleration data, first image information and second image information;

[0009] Based on linear acceleration data, the structural response of the target object after being subjected to force is analyzed to calculate the first mode feature;

[0010] Based on the first image information, the thermal distribution anomalies of the target object are analyzed to calculate the second modal features;

[0011] Based on the second image information, the apparent defects of the target detection object are analyzed by a visual detection model to calculate the third modality features;

[0012] The first modal feature, the second modal feature, and the third modal feature are weighted and fused to generate confidence correction coefficients;

[0013] The preliminary detection results are corrected using confidence correction coefficients to obtain the corrected target detection results.

[0014] In one embodiment, the linear acceleration data includes triaxial linear acceleration data. The step of analyzing the structural response of the target object after being subjected to force based on the linear acceleration data to calculate the first modal characteristics includes:

[0015] Multiple first and second triaxial linear acceleration data are acquired; wherein, the first triaxial linear acceleration data is the data after the robot dog has processed the target processing position, and the second triaxial linear acceleration data is the data before processing.

[0016] Based on multiple first and second triaxial linear acceleration data, the defect tilt characteristics reflecting the tilt degree of the target object are calculated;

[0017] Calculate the position correction coefficient for distance compensation of defect tilt features based on the distance between the target processing position and the center position of the target detection object;

[0018] Based on the defect tilt characteristics, select valid target processing locations that are non-zero.

[0019] The first modal feature is obtained by weighted averaging based on the effective target processing position and its corresponding position correction coefficient.

[0020] In one embodiment, the step of calculating the defect tilt feature reflecting the tilt degree of the target object to be detected based on multiple first triaxial linear acceleration data and second triaxial linear acceleration data includes:

[0021] Calculate the directional similarity between the first triaxial linear acceleration data and the second triaxial linear acceleration data;

[0022] The directional similarity is transformed to obtain the feature value that is positively correlated with the degree of tilt, thus obtaining the defect tilt feature; where the directional similarity is measured by the cosine value of the angle between the vectors.

[0023] In one embodiment, the step of calculating the position correction coefficient for distance compensation of defect tilt features based on the distance between the target processing position and the center position of the target detection object includes:

[0024] Obtain the first coordinate of the target processing position and the second coordinate of the center position of the target detection object;

[0025] Based on the first and second coordinates, calculate the first straight-line distance between the target processing position and the center position of the target detection object;

[0026] Obtain the second straight-line distance between the edge of the target object and the center position of the target object;

[0027] Calculate the ratio between the first straight-line distance and the second straight-line distance to obtain the position correction coefficient.

[0028] In one embodiment, the first image information includes an infrared image. The step of analyzing the thermal distribution anomalies of the target object based on the first image information to calculate the second modal features includes:

[0029] Segment the target detection area containing the target object from the infrared image;

[0030] A threshold is set based on the overall grayscale distribution of the target detection area to divide the pixels within the area into defect area pixels and non-defect area pixels.

[0031] Calculate the first ratio reflecting the proportion of the defect area based on the ratio of the number of pixels in the defect area to the total number of pixels.

[0032] A second ratio reflecting the severity of the defect is calculated based on the degree of difference between the gray values ​​of pixels in the defective area and the average gray values ​​of pixels in the non-defective area.

[0033] Based on multiple infrared images acquired, several first ratios and second ratios are calculated, and the average of the products of these multiple first ratios and second ratios is calculated to obtain the second modal features.

[0034] In one embodiment, the step of dividing pixels within the target detection area into defect region pixels and non-defect region pixels by setting a threshold based on the overall grayscale distribution of the target detection area includes:

[0035] Calculate the average grayscale value of all pixels within the target detection area, and set a grayscale threshold based on the average grayscale value;

[0036] Pixels with gray values ​​greater than the gray threshold are classified as defect area pixels;

[0037] Pixels with gray values ​​less than or equal to the gray value threshold are classified as non-defect region pixels.

[0038] In one embodiment, the step of analyzing the apparent defects of the target object using a visual detection model based on the second image information to calculate the third modality features includes:

[0039] Multiple high-definition images collected by the robot dog are acquired, and each high-definition image is input into a pre-trained visual defect detection model.

[0040] Based on the visual defect detection model, the defect detection confidence level corresponding to each high-definition image is obtained;

[0041] The average confidence level of defect detection for each high-resolution image is calculated to obtain the third modality feature.

[0042] In one embodiment, weighted fusion of the first modal feature, the second modal feature, and the third modal feature to generate confidence correction coefficients includes:

[0043] The first modality features and the second modality features are normalized respectively;

[0044] Preset weighting coefficients are assigned to the first modal feature, the second modal feature, and the third modal feature; wherein the sum of the weighting coefficients of the first modal feature, the second modal feature, and the third modal feature is one.

[0045] The confidence correction coefficients are obtained by multiplying the normalized first, second, and third modal features by their corresponding weighting coefficients and then summing the results.

[0046] In one embodiment, the weighted fusion of the first modal feature, the second modal feature, and the third modal feature to generate the confidence correction coefficient further includes:

[0047] The first modality features and the second modality features are normalized respectively;

[0048] Preset error parameters are added to the first modal feature, the second modal feature, and the third modal feature respectively; wherein, the error parameters are used to determine that the first modal feature, the second modal feature, and the third modal feature are non-zero values;

[0049] The confidence correction coefficients are obtained by multiplying the first, second, and third modal features after adding the error parameters.

[0050] Secondly, this application proposes a confidence-based collaborative inspection system between an unmanned vehicle and a robot dog, which includes:

[0051] The communication module is used to send a collaborative inspection command containing target location information to the robot dog when the confidence level of the unmanned vehicle's preliminary detection result of the target detection object is lower than a preset threshold.

[0052] The data acquisition module is used to respond to the collaborative inspection command, control the robot dog to the target position, and acquire multimodal data, including linear acceleration data, first image information, and second image information.

[0053] The first calculation module is used to analyze the structural response of the target object after being subjected to force based on linear acceleration data, so as to calculate the first modal characteristics;

[0054] The second calculation module is used to analyze the thermal distribution anomalies of the target detection object based on the first image information, so as to calculate the second modal features;

[0055] The third calculation module is used to analyze the apparent defects of the target detection object based on the second image information using a visual detection model, so as to calculate the third modality features;

[0056] The weighted fusion module is used to perform weighted fusion of the first modality features, the second modality features, and the third modality features to generate confidence correction coefficients;

[0057] The correction module is used to correct the preliminary detection results using confidence correction coefficients to obtain the corrected target detection results.

[0058] One or more technical solutions proposed in this application have, but are not limited to, the following technical effects:

[0059] This application provides a confidence-based collaborative inspection system and method for unmanned vehicles and robot dogs based on multi-principle sensors. When the confidence level of the unmanned vehicle's preliminary detection result of a target object is lower than a preset threshold, a collaborative inspection command containing target location information is sent to the robot dog. In response to the collaborative inspection command, the robot dog is controlled to move to the target location and collect multimodal data, including linear acceleration data, first image information, and second image information. Based on the linear acceleration data, the structural response of the target object under force is analyzed to calculate the first modal feature. Based on the first image information, thermal distribution anomalies of the target object are analyzed to calculate the second modal feature. Based on the second image information, apparent defects of the target object are analyzed using a visual inspection model to calculate the third modal feature. The first, second, and third modal features are weighted and fused to generate a confidence correction coefficient. The confidence correction coefficient is used to correct the preliminary detection result to obtain the corrected target detection result. By sending instructions to the robot dog to collect multimodal data and fusing and correcting the confidence level when the confidence level is low, the accuracy and reliability of target detection results are improved, and the risk of misjudgment and missed judgment during inspection is reduced. Attached Figure Description

[0060] Figure 1 is a flowchart of the first embodiment of the confidence-based collaborative inspection method of unmanned vehicle and robot dog of this application;

[0061] Figure 2 is a flowchart of the second embodiment of the confidence-based collaborative inspection method of unmanned vehicles and robot dogs in this application.

[0062] Figure 3 is a flowchart illustrating the third embodiment of the confidence-based collaborative inspection method between unmanned vehicles and robot dogs in this application.

[0063] Figure 4 is a flowchart illustrating the fourth embodiment of the confidence-based collaborative inspection method of unmanned vehicles and robot dogs in this application.

[0064] Figure 5 is a flowchart illustrating the fifth embodiment of the confidence-based collaborative inspection method of unmanned vehicles and robot dogs in this application;

[0065] Figure 6 is a schematic diagram of the structure of the confidence-based unmanned vehicle and robot dog collaborative inspection system of this application. Detailed Implementation

[0066] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0067] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0068] To make the technical solutions and advantages of the embodiments of this application clearer, the exemplary embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not an exhaustive list of all embodiments.

[0069] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.

[0070] During facility inspections, unmanned vehicles are limited to non-motorized vehicle lanes, preventing them from conducting close-up observations of facilities at greater distances. Furthermore, the confidence level of the inspection results is primarily calculated based on single-modal data. Due to insufficient information in this modality, it cannot fully reflect the true state of the target object, resulting in a low confidence level and consequently affecting the accuracy and reliability of the inspection results.

[0071] Specifically, in the scenario of manhole cover inspection, when an unmanned vehicle is traveling in the non-motorized vehicle lane, its onboard high-definition camera performs visual inspection of manhole covers in the center of the motorized vehicle lane. Due to the long shooting distance and limited angle, the edge areas of the manhole cover in the image are often obscured or blurred, causing the confidence level output by the visual inspection model to be lower than the preset threshold. Furthermore, single-modal data cannot confirm whether the manhole cover is damaged, and the unmanned vehicle cannot enter the motorized vehicle lane for close-up inspection, resulting in unreliable inspection results.

[0072] If the above problems are not addressed, the inspection system will frequently interrupt operations due to insufficient confidence levels or output unreliable inspection results. Potential structural defects may be overlooked, increasing public safety hazards. Furthermore, low-confidence results require manual verification, reducing overall inspection efficiency and impacting the timeliness of municipal management.

[0073] Based on this, this application provides a confidence-based collaborative inspection method for unmanned vehicles and robotic dogs, which is applied to a confidence-based collaborative inspection system for unmanned vehicles and robotic dogs. Please refer to Figure 1, which is a flowchart illustrating the first embodiment of the confidence-based collaborative inspection method for unmanned vehicles and robotic dogs of this application.

[0074] In this embodiment, the above-mentioned confidence-based unmanned vehicle and robot dog collaborative inspection method includes steps S10 to S70:

[0075] Step S10: When the confidence level of the unmanned vehicle's preliminary detection result of the target object is lower than the preset threshold, a collaborative inspection command containing the target location information is sent to the robot dog.

[0076] When the autonomous vehicle performs a preliminary detection of the target object, it can obtain a preliminary detection result and its confidence level. If the confidence level is lower than a preset threshold, it indicates that the autonomous vehicle has doubts about the reliability of the detection result and further verification is required. At this time, the autonomous vehicle will send a collaborative inspection command to the robot dog, which may include the location information of the target object so that the robot dog can accurately go to the target area.

[0077] For example, an autonomous vehicle can use its onboard visual sensors to make a preliminary identification of distant municipal facilities and calculate a confidence level. If the confidence level is lower than the system's set standard, the autonomous vehicle will send a command to the robot dog via its wireless communication module, informing it of the geographical coordinates of the target facility.

[0078] Step S20: In response to the collaborative inspection command, control the robot dog to the target location and collect multimodal data.

[0079] The multimodal data includes linear acceleration data, first image information, and second image information.

[0080] In response to the received collaborative inspection command, the robot dog is controlled to move to the designated target location. Upon arrival, the robot dog uses its onboard sensors to collect multimodal data. This multimodal data can be used to comprehensively reflect the state of the target object from different dimensions.

[0081] Specifically, the multimodal data includes linear acceleration data, first image information, and second image information. For example, the robot dog can navigate to the target manhole cover based on the GPS coordinates in the command, and use its built-in inertial measurement unit (IMU) to collect linear acceleration data when pressing the manhole cover. Simultaneously, it uses an infrared camera to acquire a thermal image of the manhole cover and a high-definition camera to capture an image of the manhole cover's surface. It should be noted that the change in linear acceleration data primarily reflects the change in the projection of the gravity vector in the robot's coordinate system, and does not refer to the acceleration due to motion.

[0082] In one specific implementation, the steps of controlling the robot dog to the target location and collecting multimodal data in response to a collaborative inspection command include:

[0083] (1) Use GPS positioning to control the robot dog to move to the target location, and use an obstacle avoidance control framework to dynamically avoid obstacles during the movement.

[0084] Among them, GPS positioning control refers to obtaining the robot dog's real-time location information through the Global Positioning System and combining it with a preset path or target location to accurately guide and control the robot dog's movement, so as to ensure that the robot dog can accurately reach the designated inspection area.

[0085] For example, differential GPS or real-time dynamic GPS technology can be used in conjunction with inertial measurement units for data fusion to achieve centimeter-level positioning accuracy, thereby improving the accuracy and stability of robot dog navigation. An obstacle avoidance control framework refers to a set of algorithms and strategies used to detect obstacles in the environment and plan safe paths.

[0086] During the robot dog's movement, this obstacle avoidance control framework can perceive the surrounding environment in real time, identify potential obstacles, and dynamically adjust the robot dog's trajectory to avoid collisions. For example, environmental information can be acquired using lidar, ultrasonic sensors, or visual sensors, and an environmental map can be constructed through simultaneous localization and mapping (SMR) technology. This is then combined with path planning algorithms such as A* algorithm, fast random tree algorithm, or dynamic window method to achieve autonomous obstacle avoidance and safe movement for the robot dog.

[0087] Dynamic obstacle avoidance refers to the ability of the obstacle avoidance control framework to make real-time and flexible path adjustments as the robot dog moves, based on real-time changes in obstacle information. This is particularly important for inspections in complex and dynamically changing industrial or field environments, as it can effectively prevent the robot dog from colliding with moving or sudden obstacles, ensuring the smooth progress of inspection tasks.

[0088] (2) At the target location, at different positions and angles of the target detection object, perform a preset number of multimodal data acquisition cycles to obtain the first multimodal data.

[0089] The first multimodal data includes linear acceleration data corresponding to multiple time sequences before the robot dog processes the target detection object, first image information, and second image information.

[0090] By collecting data from different positions and angles of the target object, we can obtain comprehensive spatial information about the target object, capture potential local defects or anomalies, and avoid missing key information due to data collection from a single viewpoint or position.

[0091] For example, a robot dog can execute a pre-programmed scanning path around a target object, or use its robotic arm to adjust the orientation of sensors to collect data from multiple sides, top, or bottom.

[0092] The preset number of multimodal data acquisition cycles refers to the robot dog repeatedly performing data acquisition actions a certain number of times at each specified acquisition point or area. This repeated acquisition mechanism can improve the reliability and robustness of the data. By averaging or statistically analyzing the data collected multiple times, the impact of random noise and measurement errors can be effectively reduced, ensuring the accuracy of subsequent feature calculations.

[0093] For example, you can set each data collection cycle to repeat 3-5 times to obtain more stable data samples.

[0094] The first multimodal data refers to the raw multimodal data collected by the robot dog before applying pressure to the target object. This first multimodal data can serve as a baseline or initial state information for subsequent comparative analysis with processed data. It includes linear acceleration data, first image information (such as infrared images), and second image information (such as high-resolution images). These data together constitute a multidimensional feature description of the target object when it is not subjected to external disturbances. Multiple time-series corresponding linear acceleration data refers to linear acceleration sensor data continuously collected over a period of time.

[0095] It should be noted that the first multimodal data can be the recording of minute vibrations or structural responses of the target object under natural conditions when the robot dog is not applying pressure to it. By analyzing the time-series first multimodal data, the inherent characteristics of the target object under static or quasi-static conditions can be understood, providing a comparative basis for subsequent structural analysis under stress.

[0096] (3) In each acquisition cycle, the robot dog is controlled to apply pressure to the target detection object with a preset pressure value and perform a preset number of multimodal data acquisition cycles to obtain the second multimodal data.

[0097] The second multimodal data includes linear acceleration data corresponding to multiple time sequences after the robot dog processes the target detection object, first image information, and second image information.

[0098] Controlling the robot dog to apply pressure to the target object for processing with a preset pressure value can mean that the robot dog, through its onboard actuator, applies a precisely controlled force to the target object for processing. The purpose of actively applying pressure to the target object is to induce a structural response in the target object, so that its internal defects or weak points exhibit more obvious characteristics under stress, thereby facilitating detection and analysis.

[0099] For example, a robot dog can be equipped with a force sensor feedback-controlled robotic arm that can apply a constant or variable pressure, such as a vertical pressure of 50 Newtons, to the target object by adjusting the position and force of the end effector of the robotic arm.

[0100] Here, the second multimodal data refers to the multimodal data collected by the robot dog after applying pressure to the target object. This second multimodal data reflects the response state of the target object after being subjected to external stimuli, and, in contrast to the first multimodal data, serves as a key basis for analyzing the structural response and defect characteristics of the target object. It also includes linear acceleration data, first image information, and second image information, but these data are collected during the process of the target object being subjected to force or recovering from force.

[0101] It should be noted that the linear acceleration data corresponding to multiple time series in the second multimodal data can be used to record the dynamic response process of the target object under preset pressure and afterwards. By analyzing these time series data, the vibration mode, deformation characteristics, or abnormal response of the target object after being subjected to force can be identified, and then the first modal characteristics reflecting its structural integrity can be calculated. For example, the acceleration change curves during and within a few seconds after the pressure is applied can be recorded to capture the transient and steady-state structural responses.

[0102] In this way, through precise navigation and safe obstacle avoidance mechanisms, the robot dog can accurately reach the location of the target object. Once there, the robot dog first performs multiple loops of data acquisition from various spatial positions and angles of the target object without applying external force, acquiring the first multimodal data. This data provides comprehensive baseline information for the initial state of the target object, especially the linear acceleration data, which records the subtle vibration characteristics of the target object under natural conditions. Subsequently, the robot dog actively applies a force to the target object with a preset pressure value, simulating external stress or load, and performs multiple loops of data acquisition during and after the application of pressure, acquiring the second multimodal data. The linear acceleration data acquired at this time can capture the dynamic structural response of the target object after being subjected to force. By comparing and analyzing the linear acceleration data before and after processing, the structural changes or defect responses of the target object after being subjected to force can be identified more accurately, thus providing a more reliable basis for calculating the first modal features. At the same time, the multiple image acquisitions at different positions and angles also provide richer and more comprehensive visual and thermodynamic information for the subsequent calculation of the second and third modal features. This strategy, which combines proactive pressure with multi-dimensional and multi-time-series data acquisition, enables the acquired multimodal data to more comprehensively and deeply reflect the physical characteristics and potential defects of the target object, significantly improving the accuracy and reliability of subsequent confidence correction.

[0103] Step S30: Based on linear acceleration data, analyze the structural response of the target object after being subjected to force to calculate the first mode feature.

[0104] The first modal feature can refer to a quantitative indicator that reflects the structural response or physical deformation characteristics of the target object, calculated based on linear acceleration data.

[0105] First-modal features can be used to quantify the physical deformation or structural stability of a target object. For example, when a robot dog presses a target facility, its built-in accelerometer records minute vibrations or tilt data of the facility. By processing this raw acceleration data, features reflecting the structural integrity of the facility can be extracted; for example, by calculating the simple difference between the acceleration vectors before and after pressing, the degree of deformation of the facility (target object) can be assessed.

[0106] Step S40: Based on the first image information, analyze the thermal distribution anomaly of the target detection object to calculate the second modal features.

[0107] The second modal feature can refer to a quantitative indicator calculated based on the first image information, which reflects the abnormal thermal distribution or thermal defect characteristics of the target object.

[0108] Based on the acquired first image information, the thermal distribution anomalies of the target object can be analyzed to calculate the second modal features. These second modal features can be used to reveal whether the target object exhibits temperature anomalies, which are typically associated with internal defects or material degradation.

[0109] For example, by performing simple grayscale analysis on infrared images, regions with significantly higher or lower temperatures can be identified. These regions are then designated as anomalous areas, and their area or average temperature value is extracted as a second modal feature reflecting abnormal thermal distribution within a preset range of the target object.

[0110] Step S50: Based on the second image information, analyze the appearance defects of the target detection object through a visual detection model to calculate the third modality features.

[0111] The third modality feature can refer to a quantitative indicator that reflects the appearance defect characteristics of the target object, calculated based on the second image information through a visual detection model.

[0112] Based on the acquired second image information, a pre-trained visual detection model can be used to analyze the apparent defects of the target object to calculate the third modality features. The third modality features can be used to assess whether there is visible damage or abnormality on the surface of the target object.

[0113] For example, a high-resolution image can be input into a basic image classification model, which can identify whether there are apparent defects such as cracks, damage, or corrosion in the image. After identification, it can output a numerical value indicating the probability of the defect's existence, which serves as a third modality feature.

[0114] Step S60: Perform weighted fusion of the first modal feature, the second modal feature and the third modal feature to generate confidence correction coefficients.

[0115] After determining the first modal feature, the second modal feature, and the third modal feature, the calculated first modal feature, the second modal feature, and the third modal feature can be weighted and fused to generate a comprehensive confidence correction coefficient.

[0116] It should be noted that fusing the first modality features, the second modality features, and the third modality features can be used to comprehensively consider information from different modalities, thereby more comprehensively and accurately assessing the true state of the target object.

[0117] For example, the first modal feature, the second modal feature, and the third modal feature can be added together and then divided by a constant to obtain a preliminary correction coefficient. This correction coefficient can be a comprehensive correction value calculated by integrating data from multiple different sources.

[0118] Step S70: Correct the preliminary detection results using the confidence correction coefficient to obtain the corrected target detection results.

[0119] The target detection result can refer to the final detection result obtained after the confidence level of the preliminary detection result is adjusted by the confidence correction coefficient. The target detection result has higher accuracy and reliability.

[0120] Specifically, the generated confidence correction coefficients can be used to correct the initial detection results of the unmanned vehicle, thereby obtaining the corrected target detection results. The target detection results can be used to improve the confidence of the initial detection results, making them closer to the true state of the target object.

[0121] The following example illustrates this embodiment. Suppose that during urban road inspection, an unmanned vehicle (UAV) is responsible for conducting preliminary inspections of municipal manhole covers along its route. Since the UAV can only travel in non-motorized vehicle lanes, its onboard visual sensors may have limited viewing angles or insufficient lighting for manhole covers at greater distances, resulting in low confidence levels in the preliminary detection of manhole cover damage. For example, the UAV might capture an image of a distant manhole cover using its high-definition camera and input it into a pre-trained visual detection model for analysis.

[0122] At this point, the autonomous vehicle immediately sends a collaborative inspection command to its onboard robot dog. This command, in a structured data format (e.g., JSON), contains the precise geographical location of the manhole cover and the type of inspection task the robot dog needs to perform. Upon receiving the command, the robot dog uses its built-in navigation system to plan a path and autonomously moves to the target manhole cover. During the movement, the robot dog uses its obstacle avoidance system to perceive the surrounding environment and dynamically adjusts its route to avoid collisions.

[0123] Upon reaching the manhole cover, the robot dog will initiate a multimodal data acquisition process. It will collect data multiple times from different areas and angles of the manhole cover to obtain comprehensive information. Specifically, the robot dog will apply slight pressure to the manhole cover using its mechanical legs, while simultaneously acquiring triaxial acceleration data of the manhole cover after being subjected to force via its built-in inertial measurement unit (IMU). At the same time, the robot dog will use its onboard infrared camera to acquire infrared images of the manhole cover surface to detect abnormal thermal distribution; and use a high-definition camera to capture high-definition images of the manhole cover surface to capture surface defects.

[0124] After acquiring multimodal data, this data will be analyzed to calculate the modal characteristics. First, based on the acquired linear acceleration data, the structural response of the manhole cover under stress will be analyzed. For example, by comparing the simple changes in the acceleration vector before and after pressing, it can be preliminarily determined whether the manhole cover has macroscopic tilting or deformation, thus calculating the first modal characteristic, which reflects the structural stability of the manhole cover. Second, based on the acquired infrared image information, abnormal thermal distribution on the surface of the manhole cover will be analyzed. For example, by performing simple grayscale statistics on the infrared image, areas with abnormal temperatures can be identified, and the area ratio or average temperature difference of these areas can be calculated to obtain the second modal characteristic. The second modal characteristic indicates whether the manhole cover has thermal anomalies such as internal voids, leakage, or material degradation. Third, based on the acquired high-resolution image information, the apparent defects of the manhole cover will be analyzed using a pre-trained visual detection model. For example, multiple high-resolution images can be input into a basic image recognition model, which can identify whether there are visible defects such as cracks, wear, corrosion or missing parts on the surface of the manhole cover, and calculate the average confidence score of these defects as a third modal feature, which reflects the surface integrity of the manhole cover.

[0125] After calculating the first, second, and third modal features, the system performs a weighted fusion of these three features to generate a comprehensive confidence correction coefficient. For example, a preset weight can be assigned to each modal feature (e.g., visual features have a higher weight, thermal features a lower weight, and acceleration features a lower weight), and then the weighted feature values ​​are summed to obtain a confidence correction coefficient between 0 and 1. This confidence correction coefficient comprehensively reflects the significance of defects in the manhole cover across the structural, thermal, and visual dimensions.

[0126] The confidence correction coefficient will be used to correct the initial detection results of the autonomous vehicle. This raises the initial detection results, which were originally below the threshold, to a higher confidence level, enabling the system to more accurately determine the true health condition of the manhole cover and avoid misjudgments due to insufficient information from a single modality. If the corrected confidence level reaches or exceeds the preset threshold, the robot dog will return to the autonomous vehicle, which will continue its inspection; otherwise, further collaborative inspections or human intervention may be required.

[0127] This application significantly improves the accuracy and reliability of target detection object inspection results through the collaborative operation of unmanned vehicles and robot dogs.

[0128] Based on the first embodiment of this application, a second embodiment of this application is proposed. Please refer to Figure 2, which is a flowchart of the second embodiment of the unmanned vehicle and robot dog collaborative inspection method based on confidence level.

[0129] As a refinement of step S30 in the first embodiment, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment can be referred to the above description, and will not be repeated hereafter. Based on this, the confidence-based unmanned vehicle and robot dog collaborative inspection method of this application also includes steps S31 to S35:

[0130] Step S31: Obtain multiple first triaxial linear acceleration data and second triaxial linear acceleration data.

[0131] Among them, the first three-axis linear acceleration data is the data after the robot dog processes the target processing position, and the second three-axis linear acceleration data is the data before processing.

[0132] It should be noted that linear acceleration data includes triaxial linear acceleration data. Triaxial linear acceleration data refers to acceleration components measured in the three orthogonal directions of X, Y, and Z. Triaxial linear acceleration data can be used to comprehensively reflect the linear motion state and force conditions of an object in space. For example, it shows the changes in acceleration in different directions when an object is subjected to impact, vibration, or tilting. By collecting triaxial data, the dynamic response of the target object can be captured more accurately, rather than just changes in a single direction.

[0133] Here, the first and second sets of triaxial linear acceleration data represent the data of the robot dog before and after processing the target object at the target location, respectively. Processing the target object can refer to the robot dog applying a certain physical action to the target object, such as applying a preset pressure or performing a slight tap. By collecting multiple sets of triaxial linear acceleration data before and after processing, a comparison can be made, thereby more clearly identifying the structural response caused by the action applied by the robot dog and effectively distinguishing between the inherent vibration of the target object itself or environmental noise.

[0134] Step S32: Calculate the defect tilt characteristics that reflect the tilt degree of the target object based on multiple first and second triaxial linear acceleration data.

[0135] Among them, the defect tilt feature can measure the degree to which the attitude or structure of the target object deviates from its normal state after being subjected to external forces. The defect tilt feature can be obtained by comparing the triaxial linear acceleration data before and after processing, for example, by calculating the change in the direction of the acceleration vector, the change in the attitude angle, or the change in the gravitational component. For example, the cosine value of the angle between the acceleration vectors before and after processing can be calculated, or the attitude angle before and after processing can be calculated by the inertial measurement unit (IMU) fusion algorithm, thereby obtaining the degree of tilt.

[0136] In one specific implementation, the step of calculating the defect tilt characteristics reflecting the tilt degree of the target object to be detected, based on multiple first triaxial linear acceleration data and second triaxial linear acceleration data, includes:

[0137] (1) Calculate the directional similarity between the first triaxial linear acceleration data and the second triaxial linear acceleration data.

[0138] Among them, directional similarity can effectively filter out differences in acceleration amplitude and focus on the directional changes in the posture or deformation of the structure after it is subjected to force.

[0139] It should be noted that by calculating the directional similarity between the first and second triaxial linear acceleration data, the structural response or tilt change of the target object can be evaluated by comparing the directions of the triaxial linear acceleration data collected before and after processing.

[0140] One approach is to treat the three-axis linear acceleration data before and after processing as three-dimensional vectors, and then calculate the angle between these two vectors. The smaller the angle, the more similar the directions; the larger the angle, the greater the difference in directions, which may indicate that the structure has undergone significant tilting or deformation.

[0141] As another approach, the acceleration data can be preprocessed, for example, by performing low-pass filtering to remove high-frequency noise, and then calculating the directional similarity between the filtered triaxial linear acceleration data vectors to improve the robustness of the calculation.

[0142] (2) The direction similarity is converted to obtain the feature value that is positively correlated with the degree of tilt, and the defect tilt feature is obtained.

[0143] In this context, directional similarity is measured using the cosine of the angle between the vectors.

[0144] After determining the directional similarity, the original directional similarity value can be transformed into a quantitative indicator that intuitively reflects the degree of tilt. Through this transformation, the numerical value of the defect tilt feature can be directly proportional to the severity of the tilt of the target object, facilitating subsequent defect assessment and fusion.

[0145] As one implementation, if directional similarity is measured using the cosine of the vector angle within the range [-1, 1], it can be converted to 1 minus the absolute value of the cosine, or the angle itself can be used as a measure of tilt using the arcsine function. For example, when the cosine value is 1, the directions of the first and second triaxial linear acceleration data are completely consistent, with a tilt of 0; when the cosine value is -1, the directions of the first and second triaxial linear acceleration data are completely opposite, with the maximum tilt.

[0146] As another approach, a non-linear mapping function can be designed based on actual application scenarios and experience to map the orientation similarity value to a tilt score of 0 to 100, where 0 represents no tilt and 100 represents maximum tilt, in order to better match human perception of tilt.

[0147] Step S33: Calculate the position correction coefficient for distance compensation of defect tilt features based on the distance between the target processing position and the center position of the target detection object.

[0148] It should be noted that when the robot dog processes the target object, there may be a distance between the point of action of the robot dog and the geometric center of the target object. This distance difference will cause different attenuation or amplification effects in the structural response data collected at different locations. The position correction coefficient is used to quantify this distance effect and compensate for defect tilt characteristics, so as to eliminate the measurement bias introduced by different processing positions and ensure the comparability of data collected at different locations.

[0149] In one specific implementation, the step of calculating the position correction coefficient for distance compensation of defect tilt features based on the distance between the target processing position and the center position of the target detection object includes:

[0150] (1) Obtain the first coordinate of the target processing position and the second coordinate of the center position of the target detection object.

[0151] It should be noted that the actual location of the robot dog and the specific spatial location of the core area of ​​the detected object can be determined in real time by the positioning module on the robot dog itself (e.g., a SLAM system combining GPS, inertial measurement unit (IMU) and visual odometry). At the same time, the second coordinate of the center position of the target object can be obtained from pre-stored device model data, CAD drawings, or by real-time scanning and geometric center calculation by the sensors on the robot dog (such as depth camera and LiDAR).

[0152] In addition, external high-precision positioning systems, such as differential GPS (RTK-GPS) or ultra-wideband (UWB) positioning systems, can be used to perform high-precision positioning of the robot dog and the target detection object, thereby obtaining the coordinate information of the target detection object and the target processing position.

[0153] (2) Calculate the first straight-line distance between the target processing position and the center position of the target detection object based on the first coordinate and the second coordinate.

[0154] The first linear distance can be used to quantify the actual spatial separation between the pressure point applied by the robot dog and the core area of ​​the object being detected.

[0155] Specifically, the first straight-line distance can be calculated using the three-dimensional Euclidean distance formula, which calculates the spatial distance between two three-dimensional coordinate points (x1, y1, z1) and (x2, y2, z2). In some simplified scenarios, the distance between planar coordinates can also be calculated on a two-dimensional plane.

[0156] (3) Obtain the second straight-line distance between the edge of the target object and the center position of the target object.

[0157] The second straight-line distance can be used to provide information about the size or range of the target object itself, serving as a reference for subsequent distance compensation.

[0158] It should be noted that the second straight-line distance can be directly extracted from the preset 3D model or CAD drawing of the target object, for example, obtaining the distance from its geometric center to the farthest edge point, or the average radius. Alternatively, the target object can be scanned by the depth camera or LiDAR mounted on the robot dog to construct its point cloud model in real time, and the distance from the center to the edge can be calculated based on the point cloud data.

[0159] (4) Calculate the ratio between the first straight-line distance and the second straight-line distance to obtain the position correction coefficient.

[0160] The position correction coefficient can be obtained by calculating the ratio between the first straight-line distance and the second straight-line distance, or by normalizing the actual processing position with the object's own size to obtain a dimensionless correction coefficient. The position correction coefficient can be used for subsequent defect tilt feature compensation.

[0161] The position correction coefficient can be obtained directly by dividing the first straight-line distance by the second straight-line distance. In some applications, to adjust the sensitivity or range of the correction coefficient, this ratio can also be transformed by a function, such as a logarithmic function, an exponential function, or a sigmoid function.

[0162] Step S34: Based on the defect tilt characteristics, select non-zero valid target processing positions.

[0163] Here, it is understandable that not all processing locations will produce effective structural responses or tilt characteristics. For example, in some robust or defect-free areas, even if an action is applied, there may be no significant tilt change, or the tilt characteristic value may be close to zero. Filtering out non-zero effective target processing locations aims to exclude acquisition points that are insensitive to defects or have poor data quality, thereby focusing on key areas that truly reflect the structural state of the target object being detected.

[0164] Step S35: Based on the effective target processing position and its corresponding position correction coefficient, the first modal feature is obtained by weighted average calculation.

[0165] The defect tilt features corresponding to the selected valid target processing locations are weighted and averaged using their respective position correction coefficients. The weights can be determined based on the magnitude of the position correction coefficient or other preset rules; for example, locations closer to the center or with smaller correction coefficients may be assigned higher weights. In this way, data from multiple valid processing locations can be combined to obtain a more comprehensive and accurate first modal feature reflecting the overall structural response of the target object.

[0166] As an example, taking the i-th pressing operation as an example, the robot dog records its triaxial linear acceleration moment before applying pressure, which serves as the reference acceleration for that pressing operation. During the pressing process, the triaxial linear acceleration data collected in real time by the IMU sensor is the pressing response data. To analyze the tilt of the manhole cover, a cosine similarity algorithm is used to calculate the similarity s (directional similarity) between the pressing response data vector and the reference acceleration vector. According to physical principles, if the manhole cover has defects (such as surrounding collapse or partial breakage), pressing will cause the manhole cover to tilt, thereby changing the robot dog's force-bearing posture and causing the response acceleration direction to deviate from the reference direction, resulting in a decrease in similarity s. In particular, under normal operating conditions, the deviation of the robot dog's body tilt direction after pressing from the reference direction is designed to be no greater than 90 degrees to ensure that the similarity s is always greater than or equal to 0. To convert the similarity (negatively correlated with the degree of tilt) into an intuitive positive correlation metric, the defect tilt characteristic of that pressing operation can be obtained by calculating (1-s). The larger this characteristic value, the more significant the tilt of the manhole cover at that pressing point caused by the defect.

[0167] Secondly, to correct the impact of the pressing position on the assessment of defect tilt characteristics, a position correction coefficient needs to be calculated. According to the lever principle, under the same applied force and defect severity, the closer the pressing point is to the edge of the manhole cover (away from the center), the more likely the manhole cover is to tilt. To avoid the calculated defect tilt characteristic value being artificially underestimated due to the pressing point being close to the center, the system introduces a position correction mechanism. Specifically, the position coordinates of the pressing center (first coordinates) can be obtained, and the current center position coordinates of the manhole cover can be obtained from historical sewer manhole cover construction data (second coordinates). The Euclidean distance between the current pressing point position coordinates and the manhole cover center position coordinates is calculated and denoted as the current pressing distance (first straight-line distance). Simultaneously, the maximum Euclidean distance from the edge of the manhole cover to its center is defined as the reference distance (second straight-line distance). The position correction coefficient is the ratio of this maximum distance to the current pressing distance. Therefore, when the pressing point is located at the edge of the manhole cover, the correction coefficient is 1; when the pressing point moves towards the center, the distance decreases, and the correction coefficient is greater than 1, thus amplifying the defect tilt characteristic value at the corresponding point in subsequent calculations to compensate for signal attenuation caused by the shortening of the lever arm.

[0168] Subsequently, a crucial data filtering step is performed. Since pressing on a solid, defect-free area of ​​the manhole cover should theoretically not cause tilting, its defect tilt characteristic value should be zero or close to zero. Including this data in the overall calculation would dilute the defect signal, resulting in an underestimation of the first modal characteristic value, failing to accurately reflect the salience of the defect area. Therefore, before calculating the first modal characteristic, the system automatically filters out all pressing point data with a defect tilt characteristic value of zero, retaining only valid pressing points (valid target processing locations) with a defect tilt characteristic value greater than zero for analysis.

[0169] Finally, based on the selected valid pressure point data, the first modal feature A is calculated. Specifically, the formula for calculating the first modal feature A is as follows:

[0170]

[0171] The number of effective pressure points is: For the j-th effective pressing point, its defect tilt characteristic is denoted as... The position correction factor is denoted as .

[0172] It should be noted that the first modal feature A is the weighted average of the defect tilt characteristic values ​​of all effective pressing points after position correction. Through this calculation, the first modal feature A comprehensively reflects the overall defect offset trend and degree exhibited by the manhole cover after being pressed at different locations. The larger its characteristic value, the higher the probability and severity of structural defects in the manhole cover. When the number of selected effective points... When the value is zero, it indicates that no pressing has caused a measurable tilt. At this time, the first modal feature A is set to zero, and it is determined that there is no obvious defect offset.

[0173] This application uses a robotic dog to apply force to a target object at a target processing location and acquires multiple triaxial linear acceleration data before and after processing. By comparing these triaxial linear acceleration data before and after processing, defect tilt features reflecting the tilt degree of the target object can be calculated, thereby quantifying its structural response. Considering that different processing positions of the robotic dog on the target object may lead to differences in structural response, this application further calculates the distance between the target processing location and the center position of the target object, and generates a position correction coefficient based on this distance to compensate for the defect tilt features and eliminate measurement errors caused by position differences. Subsequently, to ensure the validity of the data used, non-zero valid target processing locations can be screened based on the defect tilt features, excluding those regions that fail to produce significant structural responses. Finally, the defect tilt features corresponding to these valid target processing locations are combined with their respective position correction coefficients and fused using a weighted average method to obtain a comprehensive and accurate first modal feature. In this way, by collecting data from multiple points, comparing data before and after, compensating for location, and weighting fusion, the limitations that may exist in single or raw acceleration data are effectively overcome, making the calculation of the first mode feature more accurate and robust, and providing a more reliable basis for subsequent confidence correction.

[0174] Based on the above embodiments, a third embodiment of this application is further proposed. Please refer to Figure 3, which is a flowchart of the third embodiment of the confidence-based unmanned vehicle and robot dog collaborative inspection method of this application.

[0175] As a refinement of step S40 in the first embodiment, in the first and second embodiments of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description and will not be repeated hereafter. Based on this, the confidence-based unmanned vehicle and robot dog collaborative inspection method of this application also includes steps S41 to S45:

[0176] Step S41: Segment the target detection area where the target object is located from the infrared image.

[0177] In particular, segmenting the target detection area from the infrared image can accurately identify and isolate the pixel areas related to the target detection object in the infrared image, eliminate background interference, and ensure that subsequent analysis focuses only on the target itself.

[0178] Specifically, the target detection region can be determined in various ways. For example, a deep learning-based semantic segmentation model can be used to identify target objects in infrared images through pre-training and output their pixel-level masks. Alternatively, traditional image processing methods can be employed. For instance, image enhancement (e.g., histogram equalization) can be performed first, followed by applying edge detection algorithms combined with morphological operations (such as dilation and erosion) to outline the target contour, and finally, the target detection region can be extracted through region growing or connected component analysis.

[0179] Step S42: Set a threshold based on the overall grayscale distribution of the target detection area to divide the pixels in the area into defect area pixels and non-defect area pixels.

[0180] Here, by analyzing the infrared thermal image characteristics of the target area, pixels that may have thermal anomalies (defects) are distinguished from normal pixels. Here, grayscale values ​​in infrared images typically represent temperature or thermal radiation intensity.

[0181] Specifically, the threshold setting can automatically determine an optimal threshold based on the grayscale histogram of the target area, so that the inter-class variance or information entropy between the foreground (defect area) and the background (non-defect area) is maximized.

[0182] As another example, adaptive thresholding can also be used, such as local thresholding, which dynamically sets the threshold based on the grayscale distribution of the local area around the pixel to adapt to uneven lighting or heat distribution in the image.

[0183] In one specific implementation, the step of dividing pixels within the target detection area into defect region pixels and non-defect region pixels by setting a threshold based on the overall grayscale distribution of the target detection area includes:

[0184] (1) Calculate the average gray value of all pixels in the target detection area, and set the gray value threshold based on the average gray value.

[0185] Calculating the average grayscale value of all pixels within the target area can refer to summing the grayscale values ​​of all pixels within the target area segmented from the infrared image and then dividing by the total number of pixels in that area to obtain a value representing the overall thermal level of the target area.

[0186] It should be noted that this grayscale average value can be used as a dynamic and adaptive benchmark for subsequent threshold setting, thereby better reflecting the actual thermal distribution of the target object under the current environment and working conditions.

[0187] One approach is to iterate through each pixel within the target area, accumulating its grayscale value to calculate the average. Alternatively, one can use functions provided by an image processing library to directly obtain the average grayscale value of a specified area. Setting a grayscale threshold based on the average grayscale value means using the calculated average grayscale value as a reference point to determine the boundary used to distinguish defective and non-defective pixels. This grayscale threshold can be set equal to the average grayscale value, or appropriately offset or scaled based on the average grayscale value to adapt to different detection needs and defect types.

[0188] For example, the average grayscale value can be directly used as the grayscale threshold; or, based on experience or experimental results, the grayscale threshold can be set to the average grayscale value plus a preset constant to improve the sensitivity to abnormally high temperature areas.

[0189] (2) Pixels with gray values ​​greater than the gray threshold are classified as defect area pixels.

[0190] After setting the grayscale threshold, each pixel in the target area can be judged. If its grayscale value is higher than the set grayscale threshold, the pixel is considered to belong to the thermal anomaly area, that is, the defect area.

[0191] Defective region pixels can be used to indicate abnormal phenomena such as temperature rise or heat accumulation at the physical location corresponding to these pixels.

[0192] (3) Pixels with gray values ​​less than or equal to the gray threshold are classified as non-defect area pixels.

[0193] After setting the grayscale threshold, each pixel in the target area can be judged. If the grayscale value of a pixel is lower than or equal to the set grayscale threshold, the pixel is considered to belong to the normal thermal distribution area, that is, the non-defect area.

[0194] Non-defect areas can be used to indicate that these pixels reflect the absence of abnormal phenomena such as temperature rise or heat accumulation at the corresponding physical location, or they can represent the thermal performance of the target object under normal working conditions.

[0195] Here, by introducing the average grayscale value of all pixels within the target area as a benchmark for setting the grayscale threshold, adaptive identification of thermal anomaly regions in infrared images is achieved. This adaptive threshold setting method based on the thermal characteristics of the target itself makes the identification of defect areas no longer limited by the influence of the external environment or the inherent thermal differences of the target, thus enabling more accurate capture of true thermal anomalies. This provides more reliable basic data for subsequent calculation of the first ratio reflecting the proportion of defect area and the second ratio reflecting the severity of the defect.

[0196] Step S43: Calculate the first ratio reflecting the proportion of the defect area based on the ratio of the number of pixels in the defect area to the total number of pixels.

[0197] The first ratio is used to quantify the relative size of the thermal anomaly area on the surface of the target object, which can intuitively reflect the extent of the defect.

[0198] Specifically, after pixel division is completed, the total number of pixels identified as defective areas is counted, and then divided by the total number of pixels in the target area to obtain the ratio.

[0199] Step S44: Calculate a second ratio reflecting the severity of the defect based on the difference between the gray values ​​of pixels in the defective area and the average gray values ​​of pixels in the non-defective area.

[0200] The second ratio is used to quantify the intensity of thermal anomalies in the defect area, reflecting the depth or severity of the defect.

[0201] Specifically, the average gray value of pixels in the non-defective region can be calculated first. Then, the difference between the average gray value of pixels in the defective region and the average gray value of pixels in the non-defective region can be calculated and normalized with the average gray value of the non-defective region to obtain the second ratio.

[0202] Step S45: Based on the multiple infrared images acquired, calculate multiple first ratios and second ratios respectively, and calculate the product average of the multiple first ratios and second ratios to obtain the second modal features.

[0203] Here, by collecting data multiple times and averaging the products, the robustness and accuracy of the second modality features can be improved, the random errors caused by a single measurement can be reduced, and the thermal distribution anomalies of the target can be reflected more stably.

[0204] It should be noted that when the robot dog is inspecting the target object, it can acquire multiple infrared images at different times or from different angles. For each infrared image, a first ratio and a second ratio are calculated according to the steps described above. Then, all the acquired first ratios are multiplied together, and the result is taken to the power of N (where N is the number of acquisitions) to obtain the geometric mean of the first ratios. Similarly, the product average of all the second ratios is calculated. Finally, the average of the first ratios and the average of the second ratios can be used as components of the second modal feature, and multiplication or weighted summation calculations are performed to obtain the final second modal feature.

[0205] As an example, the manhole cover region can be accurately segmented from the first image information (infrared image). This embodiment accomplishes this task by constructing a dedicated image segmentation model. Specifically, 2000 infrared images containing both normal and various defective manhole covers can be collected as training samples, and the outline of the manhole cover region can be labeled for each image to form a "manhole cover label" dataset. Using this dataset, a fully convolutional neural network is trained for 100 rounds with cross-entropy as the loss function and Adam as the optimizer to obtain a pre-trained segmentation model. By inputting the i-th infrared image to be analyzed into this model, the accurate pixel-level mask of the manhole cover region can be obtained.

[0206] After obtaining the manhole cover area, it is necessary to further identify the defective parts within this area. Based on the physical principle that the thermal radiation characteristics of defective areas change due to thinning, thus presenting different grayscale values ​​in infrared images, an automatic segmentation method based on grayscale thresholds is adopted. Specifically, the average grayscale value of all pixels within the manhole cover area can be calculated and used as the grayscale threshold for segmentation. Pixels with grayscale values ​​greater than this threshold are classified as defective area pixels, while pixels with grayscale values ​​less than or equal to this threshold are classified as non-defective area pixels. In this way, it is possible to adaptively separate suspected defective areas that are abnormally bright (usually corresponding to thinner, potentially hotter defective areas) or abnormally dark (potentially corresponding to water accumulation, different materials, etc.) based on the overall brightness characteristics of each image.

[0207] Based on the above pixel classification, calculate the first ratio (proportion of defective areas) and the second ratio (severity of defects).

[0208] The first ratio is calculated as the ratio of the total number of pixels classified as defective areas to the total number of pixels in the manhole cover area. First ratio It can directly reflect the proportion of the defective area in the visible area of ​​the manhole cover. The larger the value, the wider the spatial range of the defect.

[0209] The second ratio (defect severity) measures the degree of thermal anomaly of the defect. The average grayscale value of all pixels in the non-defective area is calculated, denoted as p, to represent the baseline thermal state of the normal area of ​​the manhole cover. For each pixel in the defective area (a total of K), the absolute value of the difference between its grayscale value pk and the baseline value p is calculated. This absolute value is then divided by the baseline value p to obtain the relative difference of that pixel. Finally, the arithmetic mean of the relative differences of all K defective pixels is taken; this average value is the second ratio. This ratio quantifies the relative extent to which the defective region deviates from its normal thermal state. The larger the value, the more significant the difference in thermal properties between the defect area and the normal area, indirectly reflecting the severity of the defect (such as the degree of material loss or thickness reduction).

[0210] To obtain a robust and comprehensive evaluation, results from multiple data collections can be combined. Assume the robot dog collected a total of... Zhang (specifically, in this embodiment) =20) Infrared image. The first ratio is calculated for the i-th image. Second ratio The second modal feature B can be calculated. Specifically, the formula for calculating the second modal feature B is as follows:

[0211]

[0212] in, This indicates the number of infrared images, and its value is equal to the number of acquisitions (20). This represents the first ratio of the infrared image acquired in the i-th acquisition. This represents the second ratio of the infrared image acquired in the i-th acquisition.

[0213] It should be noted that the calculation of the second modality feature B essentially involves multiplying the two ratios of each image (to obtain the overall defect significance score for that image), and then averaging the overall scores of all images. Therefore, the second modality feature B is a dimensionless index that integrates information on the "area size" and "abnormality" of the defect. The larger its value, the more significant the defect feature of the manhole cover is when viewed from the perspective of infrared thermal imaging.

[0214] This application embodiment accurately segments the target area from infrared images, ensuring that subsequent analysis focuses on the target itself and effectively eliminating background interference. Next, by analyzing the overall grayscale distribution of the target area and setting a threshold, it can intelligently distinguish potential thermal anomaly areas (defective region pixels) and normal areas (non-defective region pixels), laying the foundation for defect identification. Based on this, the scheme further quantifies the first and second ratios of the defects. These two ratios comprehensively characterize the thermal distribution anomalies of the target object from different perspectives. To further improve the stability and reliability of the features, the scheme calculates the product average of multiple first and second ratios based on multiple acquired infrared images, effectively smoothing out any instantaneous fluctuations or noise that may exist in a single measurement, thereby obtaining a more robust and accurate second modal feature. This refined and multi-verified second modal feature can more accurately reflect the thermal distribution anomalies of the target, providing high-quality input for subsequent confidence correction and significantly improving the correction accuracy and reliability of the target detection results by the autonomous vehicle and robot dog collaborative inspection method.

[0215] Based on the above implementation methods, a fourth embodiment of this application is proposed. Please refer to Figure 4, which is a flowchart of the fourth embodiment of the confidence-based unmanned vehicle and robot dog collaborative inspection method of this application.

[0216] As a refinement of step S50 in the first embodiment, the same or similar content in the above embodiments of this application can be referred to the above description, and will not be repeated hereafter. Based on this, the confidence-based unmanned vehicle and robot dog collaborative inspection method of this application also includes steps S51~S53:

[0217] Step S51: Acquire multiple high-definition images collected by the robot dog, and input the multiple high-definition images into the pre-trained visual defect detection model respectively.

[0218] High-definition images refer to digital images with higher resolution and clarity, which can display more detailed information.

[0219] For example, a robot dog could be equipped with a high-resolution CMOS or CCD sensor to capture images with a resolution of 1080p (1920x1080 pixels) or 4K (3840x2160 pixels). Alternatively, the robot dog could employ image enhancement technology to perform super-resolution processing on the captured raw images to improve their clarity and detail, bringing them up to high-definition standards.

[0220] Here, acquiring multiple high-resolution images allows for the capture of the appearance information of the target object from different angles, under different lighting conditions, or at different times, thereby providing a more comprehensive and richer dataset, which helps to improve the accuracy and robustness of defect detection.

[0221] For example, a robot dog can move in a circular motion around a target object or adjust the camera angle at a fixed position to acquire multiple high-definition images covering different sides of the target object. Alternatively, the robot dog can use its multiple high-definition cameras to simultaneously acquire images from different perspectives or capture multiple frames of images in a short period of time to capture the dynamic or static appearance features of the target object.

[0222] It should also be noted that pre-trained visual defect detection models refer to deep learning models that have been trained on large-scale datasets and have learned general image features and defect patterns. These models can identify various apparent defects in images, such as cracks, scratches, corrosion, and deformation. For example, the model could be based on a convolutional neural network (CNN) architecture. These models perform exceptionally well in industrial defect detection, enabling end-to-end defect localization and classification.

[0223] Step S52: Based on the visual defect detection model, obtain the defect detection confidence level corresponding to each high-definition image.

[0224] Among them, the defect detection confidence score can be a quantitative evaluation of the accuracy of the visual defect detection model in identifying potential defects in an image. It is usually expressed as a value between 0 and 1, with a higher value indicating that the model is more confident in the detected defect.

[0225] It should be noted that confidence level can be used as an indicator to measure the likelihood and severity of defects, which can help with subsequent decision-making and analysis.

[0226] For example, the confidence score output by the model can be the output probability of the classification layer (such as the Softmax layer), representing the probability that a defect exists in the image. Alternatively, for object detection models, the confidence score can be the confidence score of the bounding box prediction, combined with the classification confidence score to comprehensively evaluate the reliability of defect detection.

[0227] Step S53: Calculate the average value of the defect detection confidence corresponding to each high-definition image to obtain the third modality feature.

[0228] Understandably, averaging can effectively reduce the randomness or error of single image detection results, improve the stability and reliability of third-modal features, and thus more accurately reflect the overall appearance defects of the target object.

[0229] For example, all the acquired high-definition images can be input into the model to obtain their respective defect detection confidence scores, and then these confidence scores can be simply averaged arithmetically. Alternatively, a weighted average method can be used, for example, by assigning different weights to the confidence scores of different images based on factors such as image quality, shooting angle, or distance from the target object, and then averaging them to obtain the third modality feature.

[0230] As an example, this embodiment analyzes the apparent visual defects on the surface of manhole covers using high-resolution visible light images to supplement the shortcomings of infrared images in texture detail recognition. High-resolution images provide rich information on edges, cracks, and colors, but their imaging quality is easily affected by ambient lighting conditions. To overcome this limitation and achieve accurate recognition, this method employs a deep learning-based visual detection model. Specifically, multiple high-resolution images collected by the robot dog around the manhole cover during its inspection are input one by one into a pre-trained convolutional neural network model for defect detection. This model has been trained on a large dataset of images labeled "manhole cover defective" and "manhole cover normal," and is capable of recognizing various types of apparent damage. For each input high-resolution image, the model outputs a defect detection confidence score between 0 and 1, reflecting the model's degree of certainty that a defect exists in the image.

[0231] To comprehensively assess the overall visual condition of manhole covers and reduce the risk of misjudgment caused by factors such as lighting, angle, or momentary occlusion in a single image, the system integrates the detection confidence scores of all acquired high-definition images. Specifically, the arithmetic mean of these image confidence scores is calculated, and this average is used as the final third modality feature C. The third modality feature C characterizes the overall quantitative assessment of the salience of manhole cover defects from a high-definition visual perspective; a higher value indicates that the defect signs detected from the visible light image are more explicit and consistent. By analyzing multiple images and taking the average, the stability and reliability of the visual assessment are effectively improved, enabling it to effectively complement and fuse with the first and second modality features.

[0232] This application embodiment sets the second image information to high-definition images, enabling the robot dog to acquire image data with rich details. These multiple high-definition images are input one by one into a pre-trained visual defect detection model, which can identify various appearance defects in the images and output corresponding defect detection confidence scores. To eliminate the randomness or error that may exist in the detection results of a single image and to obtain a more stable and representative appearance defect assessment, this application further calculates the average value of the defect detection confidence scores corresponding to all high-definition images. In this way, multi-view and multi-time-series image information can be integrated, so that the final third modality feature can more accurately and robustly reflect the overall appearance defect situation of the target object. This processing method significantly improves the reliability of the third modality feature, thereby providing a more solid foundation for the subsequent generation of confidence correction coefficients, effectively solving the uncertainty problem caused by relying on a single or low-quality image for defect assessment, and thus improving the accuracy of the final corrected target detection result.

[0233] Based on any of the above embodiments of this application, a fifth embodiment of this application is proposed. Please refer to Figure 5, which is a flowchart illustrating the fifth embodiment of the confidence-based collaborative inspection method of unmanned vehicles and robot dogs in this application.

[0234] As a refinement of step S60 in the first embodiment, the same or similar content in the above embodiments of this application can be referred to the above description, and will not be repeated hereafter. Based on this, the confidence-based unmanned vehicle and robot dog collaborative inspection method of this application also includes steps S61~S63:

[0235] Step S61: Normalize the first modal features and the second modal features respectively.

[0236] Here, normalizing the first and second modal features can scale their values ​​to a predetermined standard range, thus eliminating differences in dimensions and numerical ranges between different features. For example, the min-max normalization method can be used to linearly transform the feature values ​​to the range of [0, 1]; or the Z-score normalization method can be used to transform the feature values ​​into a distribution with a mean of 0 and a standard deviation of 1.

[0237] Normalization ensures that features of different modalities are comparable in the subsequent fusion process, preventing features with larger values ​​from dominating the fusion result.

[0238] Step S62: Assign preset weighting coefficients to the first modal feature, the second modal feature, and the third modal feature.

[0239] The sum of the weighting coefficients of the first modal feature, the second modal feature, and the third modal feature is one.

[0240] Here, a weight value can be assigned to each modal feature based on its relative importance or reliability in reflecting the defects of the target object.

[0241] It should be noted that the weighting coefficients can be set based on the experience of domain experts, learned from training data through machine learning algorithms, or adjusted through multiple experiments and performance evaluations. The constraint that the sum of the weighting coefficients is one ensures that these weights represent the proportion of each feature in the total contribution, making the fusion result highly interpretable.

[0242] Step S63: Multiply the normalized first modal features, second modal features, and third modal features by their corresponding weighting coefficients and sum them to obtain the confidence correction coefficients.

[0243] Here, a linear weighted summation method can be used to integrate the contributions of each modality feature. Specifically, the normalized value (or original value, for unnormalized third modality features) of each feature is multiplied by its corresponding weighting coefficient, and then all products are summed to obtain a comprehensive value, namely the confidence correction coefficient. In this way, information from different modalities can be effectively combined and adjusted according to their importance to generate a correction factor that can more comprehensively and accurately reflect the state of the target detection object.

[0244] As an example, confidence correction coefficient The specific formula is as follows:

[0245]

[0246] in, Indicates the weighting coefficient. A, B, and C represent the first modal feature, the second modal feature, and the third modal feature, respectively.

[0247] It should be noted that since driverless vehicles primarily rely on vision technology for inspection, therefore Generally greater than or equal to and However, the first modal feature is difficult to identify subtle features of the manhole cover, therefore Generally less than and , This represents the Sigmoid normalization function.

[0248] For example, as a concrete example, the confidence correction coefficient The calculation process is explained below. First, the first and second modal features can be normalized. For example, if the current value of the first modal feature is 50, after min-max normalization to the range [0, 1], the normalized value of the first modal feature can be 0.5. If the current value of the second modal feature is 0.8, since it is already within the range [0, 1], the normalized value of the second modal feature can be 0.8. Assume the current value of the third modal feature is 0.7. Next, preset weighting coefficients are assigned to these three features. For example, the weighting coefficient of the first modal feature can be set to 0.3, the weighting coefficient of the second modal feature can be set to 0.4, and the weighting coefficient of the third modal feature can be set to 0.3. The sum of these weighting coefficients is 1. Finally, the normalized first modal feature (0.5), the normalized second modal feature (0.8), and the third modal feature (0.7) are multiplied by their corresponding weighted coefficients, and then summed. Specifically, the confidence correction coefficient = (0.5 * 0.3) + (0.8 * 0.4) + (0.7 * 0.3) = 0.15 + 0.32 + 0.21 = 0.68. Thus, the confidence correction coefficient, which comprehensively considers the information from each modality and has been weighted, can be obtained.

[0249] Here, multiple complementary or different analytical features are usually integrated to comprehensively characterize the saliency of a feature. The defect features of manhole covers in three complementary modes are analyzed separately, and the defect features of manhole covers are characterized by weighted fusion.

[0250] This application first normalizes the first and second modal features, eliminating differences in their numerical ranges and ensuring they contribute information fairly during fusion. Then, by assigning preset weighting coefficients to each modal feature and ensuring the sum of these coefficients is one, this application can flexibly adjust the weighting based on the importance or reliability of each modal feature in practical applications, thus enabling the confidence correction coefficient to more accurately reflect the comprehensive information of the multimodal data. Finally, by multiplying the normalized features by their corresponding weighting coefficients and summing the results, this application generates a comprehensive confidence correction coefficient. This coefficient not only considers the independent information of each modal feature but also effectively balances their influence through normalization and weighting mechanisms, overcoming the biases that may result from direct fusion and making the correction of the initial detection results more accurate and reliable.

[0251] In another embodiment, as another method for fusing the first modal feature, the second modal feature, and the third modal feature, the weighted fusion of the first modal feature, the second modal feature, and the third modal feature to generate the confidence correction coefficient further includes:

[0252] (1) Normalize the first modal features and the second modal features respectively.

[0253] Similarly, normalizing the first and second modal features respectively can transform feature data with different dimensions or numerical ranges into a unified scale space, thereby eliminating the impact of dimensional differences on subsequent fusion calculations and ensuring that each feature has equal importance in the fusion process.

[0254] (2) Add preset error parameters to the first modal feature, the second modal feature and the third modal feature respectively.

[0255] The error parameter is used to determine that the first modal feature, the second modal feature, and the third modal feature are non-zero values.

[0256] Adding error parameters to the first, second, and third modal features can prevent the final confidence correction coefficient from becoming zero due to any modal feature value being zero during subsequent multiplication fusion, thus losing the effective information provided by other modal features.

[0257] It should be noted that by adding a preset, extremely small positive value as an error parameter, even if the original feature value is zero, it can still maintain a small non-zero contribution in the fusion calculation. For example, the error parameter can be a preset, fixed, small constant; or it can be implemented through a function that replaces the feature value with a certain minimum threshold when the feature value is less than that threshold, otherwise keeping the original value.

[0258] (3) Multiply the first modal feature, the second modal feature and the third modal feature after adding the error parameter to obtain the confidence correction coefficient.

[0259] When all modal features indicate high confidence, the product result will be high; however, when any modal feature indicates low confidence (even with the addition of an error parameter), the product result will be significantly lower, thus providing a more conservative assessment of the overall confidence.

[0260] In another implementation, the confidence correction coefficient is calculated. You can also refer to the following formula:

[0261]

[0262] in, All represent error parameters, which are added to the first, second, and third modal features, respectively, and their values ​​range from 0.01 to 0.1. Let represent the Sigmoid normalization function, and A, B, and C represent the first modal feature, the second modal feature, and the third modal feature, respectively.

[0263] This embodiment first normalizes the first and second modal features to ensure numerical comparability of feature data from different sources and dimensions. Based on this, preset error parameters are added to all modal features, effectively avoiding the complete loss of overall confidence information due to a single feature value of zero during multiplicative fusion. This ensures that all modal features influence the final confidence correction coefficient. Subsequently, confidence correction coefficients are generated by multiplying these processed modal features. This multiplicative fusion mechanism can more sensitively capture the combined effect between modal features; that is, the final confidence correction coefficient will be high only when all modal features exhibit relatively high confidence. Conversely, if the confidence of any modal feature is low, even if other features are high, the final confidence correction coefficient will decrease accordingly, thus providing a more rigorous and comprehensive defect assessment. Compared to simple weighted summation, this method can better reflect the effect of multimodal data in jointly indicating defects, and avoids the weak signal of a single modality being masked by other strong signals, thereby improving the accuracy and reliability of confidence correction.

[0264] In another embodiment, when determining the confidence correction coefficient Then, the confidence correction coefficient can be used to correct the preliminary detection results to obtain the corrected target detection results. Specifically, the correction process can be to correct the confidence level corresponding to the step detection results using the confidence correction coefficient. If the corrected confidence level is greater than a preset confidence level threshold, then the corresponding detection result is taken as the target detection result.

[0265] It should be noted that the confidence threshold can be a dynamic threshold constant determined based on the environmental information of the target object, or a fixed global constant determined based on the accuracy required for the current inspection task. As an example, in key areas with high safety requirements, such as main roads in city centers and around schools, a lower confidence threshold (e.g., 80%) can be used to more comprehensively capture potential risks. Conversely, in suburban auxiliary roads or internal parks with lower traffic volume, the system can use a relatively low threshold (e.g., 60%), and the system will only issue an alarm when the evidence is very sufficient and conclusive, aiming to reduce false alarms. As another example, based on historical data, model performance testing, or domain expert experience, when conducting collaborative inspections of facilities on a main road, the detection result can only be considered a valid target detection result and trigger subsequent inspection actions if its corrected final confidence level reaches or exceeds 70%.

[0266] For example, the corrected confidence level This can be expressed by the following formula, specifically:

[0267]

[0268] in, The confidence level of the autonomous vehicle's preliminary detection results of the target object. This is the confidence correction coefficient.

[0269] For example, if The confidence level is greater than or equal to a preset confidence threshold, where the confidence threshold can be the maximum set confidence level of 100%. When the corrected confidence level... Once the confidence level is greater than or equal to the threshold, the robot dog returns to the unmanned vehicle cabin along the original route, closes the cabin door, and continues to proceed according to the pre-set inspection route.

[0270] If the corrected confidence level If the confidence level is less than a preset confidence threshold, the robot dog can be controlled to re-execute the confidence-based autonomous vehicle and robot dog collaborative inspection method proposed in the above embodiments within the area corresponding to the target detection object, until the confidence level of the target detection result after a preset number of inspections is greater than or equal to the preset confidence level. If the confidence level of the target detection result is still less than the confidence threshold after a preset number of inspections, the target detection object is marked, and corresponding prompt information is output to the server to prompt manual inspection of the target detection object. Here, when marking the target detection object, the location information of the target detection object and the detection information corresponding to the preset number of inspections of the target detection object (including detailed detection content and detection time, etc.) can be marked.

[0271] Please refer to Figure 6, which is a schematic diagram of a confidence-based autonomous vehicle and robot dog collaborative inspection system 60 provided in this application. The confidence-based autonomous vehicle and robot dog collaborative inspection system 60 includes:

[0272] The communication module 61 is used to send a collaborative inspection command containing target location information to the robot dog when the confidence level of the unmanned vehicle's preliminary detection result of the target detection object is lower than a preset threshold.

[0273] The acquisition module 62 is used to respond to the collaborative inspection command, control the robot dog to the target position, and acquire multimodal data, including linear acceleration data, first image information and second image information;

[0274] The first calculation module 63 is used to analyze the structural response of the target object after being subjected to force based on linear acceleration data, so as to calculate the first modal characteristics;

[0275] The second calculation module 64 is used to analyze the thermal distribution anomaly of the target detection object based on the first image information in order to calculate the second modal features;

[0276] The third calculation module 65 is used to analyze the appearance defects of the target detection object based on the second image information through a visual detection model in order to calculate the third modality features;

[0277] The weighted fusion module 66 is used to perform weighted fusion of the first modal feature, the second modal feature and the third modal feature to generate confidence correction coefficients;

[0278] The correction module 67 is used to correct the preliminary detection results using confidence correction coefficients to obtain the corrected target detection results.

[0279] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A confidence-based collaborative inspection method between unmanned vehicles and robot dogs, characterized in that, The method includes: when the confidence level of the preliminary detection result of the unmanned vehicle on the target detection object is lower than a preset threshold, sending a collaborative inspection command containing target location information to the robot dog; responding to the collaborative inspection command, controlling the robot dog to the target location and collecting multimodal data, wherein the multimodal data includes linear acceleration data, first image information, and second image information; based on the linear acceleration data, analyzing the structural response of the target detection object after being subjected to force to calculate a first modal feature; the force is generated after the robot dog applies pressure to the target detection object; based on the first image information, analyzing the thermal distribution anomaly of the target detection object to calculate a second modal feature; based on the second image information, analyzing the appearance defects of the target detection object through a visual inspection model to calculate a third modal feature; performing weighted fusion of the first modal feature, the second modal feature, and the third modal feature to generate a confidence correction coefficient; and using the confidence correction coefficient to correct the preliminary detection result to obtain a corrected target detection result.

2. The confidence-based collaborative inspection method between unmanned vehicles and robot dogs according to claim 1, characterized in that, in, The linear acceleration data includes triaxial linear acceleration data. The step of analyzing the structural response of the target detection object after being subjected to force based on the linear acceleration data to calculate the first modal feature includes: acquiring multiple first triaxial linear acceleration data and second triaxial linear acceleration data; wherein, the first triaxial linear acceleration data is data collected after the robot dog applies pressure to the target processing position of the target detection object, and the second triaxial linear acceleration data is data collected before the robot dog applies pressure to the target processing position of the target detection object; calculating a defect tilt feature reflecting the tilt degree of the target detection object based on the multiple first triaxial linear acceleration data and the second triaxial linear acceleration data; calculating a position correction coefficient for distance compensation of the defect tilt feature based on the distance between the target processing position and the center position of the target detection object; filtering non-zero valid target processing positions based on the defect tilt feature; and calculating the first modal feature by weighted average based on the valid target processing positions and their corresponding position correction coefficients.

3. The confidence-based collaborative inspection method between unmanned vehicles and robot dogs according to claim 2, characterized in that, The step of calculating the defect tilt feature reflecting the tilt degree of the target detection object based on multiple first triaxial linear acceleration data and second triaxial linear acceleration data includes: calculating the directional similarity between the first triaxial linear acceleration data and the second triaxial linear acceleration data; converting the directional similarity to obtain a feature value positively correlated with the tilt degree, thereby obtaining the defect tilt feature; wherein the directional similarity is measured by the cosine value of the vector angle.

4. The confidence-based collaborative inspection method between unmanned vehicles and robot dogs according to claim 2, characterized in that, The step of calculating the position correction coefficient for distance compensation of the defect tilt feature based on the distance between the target processing position and the center position of the target detection object includes: obtaining a first coordinate of the target processing position and a second coordinate of the center position of the target detection object; calculating a first straight-line distance between the target processing position and the center position of the target detection object based on the first coordinate and the second coordinate; obtaining a second straight-line distance between the edge of the target detection object and the center position of the target detection object; and calculating the ratio between the first straight-line distance and the second straight-line distance to obtain the position correction coefficient.

5. The confidence-based collaborative inspection method between unmanned vehicles and robot dogs according to claim 1, characterized in that, The first image information includes an infrared image. The step of analyzing the thermal distribution anomaly of the target detection object based on the first image information to calculate the second modal feature includes: segmenting the target detection area where the target detection object is located from the infrared image; setting a threshold according to the overall grayscale distribution of the target detection area to divide the pixels in the area into defect area pixels and non-defect area pixels; calculating a first ratio reflecting the proportion of defect area based on the ratio of the number of defect area pixels to the total number of pixels; calculating a second ratio reflecting the severity of defect based on the difference between the average grayscale value of the defect area pixels and the average grayscale value of the non-defect area pixels; calculating multiple first ratios and second ratios based on multiple acquired infrared images, and multiplying and averaging the multiple first ratios and second ratios to obtain the second modal feature.

6. The confidence-based collaborative inspection method between unmanned vehicles and robot dogs according to claim 5, characterized in that, The step of setting a threshold based on the overall grayscale distribution of the target detection area to divide the pixels in the area into defect area pixels and non-defect area pixels includes: calculating the average grayscale value of all pixels in the target detection area, and setting a grayscale threshold based on the average grayscale value; classifying pixels with grayscale values ​​greater than the grayscale threshold as defect area pixels; and classifying pixels with grayscale values ​​less than or equal to the grayscale threshold as non-defect area pixels.

7. The confidence-based collaborative inspection method between unmanned vehicles and robot dogs according to claim 1, characterized in that, The second image information is a high-definition image. The step of analyzing the apparent defects of the target detection object based on the second image information and calculating the third modality feature by means of a visual detection model includes: acquiring multiple high-definition images collected by the robot dog, and inputting the multiple high-definition images into a pre-trained visual defect detection model respectively; based on the visual defect detection model, obtaining the defect detection confidence score corresponding to each high-definition image; calculating the average value of the defect detection confidence score corresponding to each high-definition image to obtain the third modality feature.

8. The confidence-based collaborative inspection method between unmanned vehicles and robot dogs according to claim 1, characterized in that, The step of weightedly fusing the first modal feature, the second modal feature, and the third modal feature to generate a confidence correction coefficient includes: normalizing the first modal feature and the second modal feature respectively; assigning preset weighting coefficients to the first modal feature, the second modal feature, and the third modal feature; wherein the sum of the weighting coefficients of the first modal feature, the second modal feature, and the third modal feature is one; multiplying the normalized first modal feature, the second modal feature, and the third modal feature by the corresponding weighting coefficients and summing the results to obtain the confidence correction coefficient.

9. The confidence-based collaborative inspection method between unmanned vehicles and robot dogs according to claim 1, characterized in that, The step of weightedly fusing the first modal feature, the second modal feature, and the third modal feature to generate a confidence correction coefficient further includes: normalizing the first modal feature and the second modal feature respectively; adding a preset error parameter to the first modal feature, the second modal feature, and the third modal feature respectively; wherein the error parameter is used to make the first modal feature, the second modal feature, and the third modal feature non-zero; and multiplying the first modal feature, the second modal feature, and the third modal feature after adding the error parameter to obtain the confidence correction coefficient.

10. A confidence-based collaborative inspection system for unmanned vehicles and robot dogs, characterized in that, The system includes: a communication module, used to send a collaborative inspection command containing target location information to the robot dog when the confidence level of the preliminary detection result of the unmanned vehicle on the target detection object is lower than a preset threshold; a data acquisition module, used to control the robot dog to the target location in response to the collaborative inspection command and acquire multimodal data, wherein the multimodal data includes linear acceleration data, first image information, and second image information; a first calculation module, used to analyze the structural response of the target detection object after being subjected to force based on the linear acceleration data, so as to calculate the first modal feature; the force is generated after the robot dog applies pressure to the target detection object; a second calculation module, used to analyze the thermal distribution anomaly of the target detection object based on the first image information, so as to calculate the second modal feature; a third calculation module, used to analyze the appearance defects of the target detection object through a visual inspection model based on the second image information, so as to calculate the third modal feature; a weighted fusion module, used to perform weighted fusion of the first modal feature, the second modal feature, and the third modal feature to generate a confidence correction coefficient; and a correction module, used to correct the preliminary detection result using the confidence correction coefficient to obtain the corrected target detection result.

Citation Information

Patent Citations

  • Air-ground cooperative autonomous inspection system and method for unmanned aerial vehicle and robot dog

    CN119200657A

  • Unmanned aerial vehicle and robot dog collaborative inspection system

    CN120973035A