Inspection path generation method and device, storage medium and computer equipment
By integrating data from multiple sensors in the same time and space and using a spatiotemporal-semantic alignment model to correct positioning errors, a more accurate inspection path is generated, solving the problem of inaccurate paths caused by multimodal data matching errors and improving positioning accuracy and system robustness.
Patent Information
- Application Number
- CN202511827178.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-01-09
AI Technical Summary
In existing technologies, the problem of inaccurate inspection path generation arises due to matching errors between multimodal data.
By acquiring target point cloud data and target image data collected by multiple pre-integrated sensors at the same time and in the same space, the positioning error is corrected using a spatiotemporal-semantic alignment model, and inspection paths and control commands are generated.
It improves dynamic positioning accuracy and system robustness, reduces data processing latency, adapts to complex environments, and enhances the accuracy and autonomy of inspection paths.
Smart Images

Figure CN121300380A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of industrial automation technology, and in particular to a method, apparatus, storage medium and computer equipment for generating inspection paths. Background Technology
[0002] Currently, in the fields of industrial automation and intelligent operation and maintenance, robots and drones are gradually replacing manual labor in performing high-risk, high-frequency, or high-precision tasks. For example, in substation inspection scenarios, quadruped robots equipped with infrared cameras and LiDAR autonomously navigate and detect equipment temperature and abnormal noises; in warehousing and logistics scenarios, drone swarms collaboratively inventory high-bay racks, and AGV (Automated Guided Vehicle) robots sort goods; in disaster relief scenarios, drones search for survivors, and ground robots enter collapsed buildings to deliver supplies. These scenarios rely on the autonomous perception, real-time decision-making, and precise control of robots or drones.
[0003] In existing technologies, when inspecting a target environment, current environmental data can be acquired, and a target map can be constructed based on this data. A corresponding inspection path can then be generated based on the environmental data and the target map. However, when the acquired environmental data is multimodal data, such as point cloud data or image data, the matching errors between different modalities can lead to inaccurate inspection paths when using this type of data to generate inspection paths. Summary of the Invention
[0004] The purpose of this application is to at least solve one of the above-mentioned technical defects, especially the technical defect in the prior art where, when the acquired environmental data is multimodal data, there is a certain matching error between different modal data, and the generated inspection path is easily inaccurate when using such data to generate an inspection path.
[0005] This application provides a method for generating inspection paths, the method comprising:
[0006] Acquire target point cloud data and target image data collected by multiple pre-integrated sensors at the same time and in the same space, as well as a target map corresponding to the target point cloud data and the target image data;
[0007] The current positioning error is corrected based on the target point cloud data and the target image data. An inspection path and control instructions are generated based on the corrected positioning results and the target map. The inspection operation is then executed according to the inspection path and the control instructions.
[0008] Optionally, the step of correcting the current positioning error based on the target point cloud data and the target image data to obtain the corrected positioning result includes:
[0009] The target point cloud data and the target image data are input into a pre-constructed spatiotemporal-semantic alignment model to obtain the alignment result output by the spatiotemporal-semantic alignment model;
[0010] The current positioning error is corrected based on the alignment result until the corrected positioning error is minimized, thus obtaining the corrected positioning result.
[0011] Optionally, the step of inputting the target point cloud data and the target image data into a pre-constructed spatiotemporal-semantic alignment model to obtain the alignment result output by the spatiotemporal-semantic alignment model includes:
[0012] Determine the state variables corresponding to the target point cloud data and the target image data, and construct a spatiotemporal-semantic joint element;
[0013] The target point cloud data, the target image data, the state variables, and the spatiotemporal-semantic joint voxels are input into a pre-constructed spatiotemporal-semantic alignment model to obtain the alignment result output by the spatiotemporal-semantic alignment model.
[0014] Optionally, the step of correcting the current positioning error based on the alignment result until the corrected positioning error is minimized, to obtain the corrected positioning result, includes:
[0015] Based on the state variable and the alignment result, calculate the compensation amount of the state variable. After updating the state variable according to the compensation amount, return to execute the input of the target point cloud data, the target image data, the state variable and the spatiotemporal-semantic joint voxel into the pre-constructed spatiotemporal-semantic alignment model and its subsequent steps until the calculated compensation amount is optimal.
[0016] The state variable at which the compensation amount is optimal is used as the corrected positioning result.
[0017] Optionally, the alignment result includes residual values of visual factors and residual values of semantic constraint factors;
[0018] The calculation of the compensation amount for the state variables based on the state variables and the alignment result includes:
[0019] The residual value of the laser factor is determined based on the target point cloud data;
[0020] The cross-modal residual is obtained by weighted summing of the residual values of the visual factor, the semantic constraint factor, and the laser factor.
[0021] The compensation amount of the state variable is calculated based on the state variable and the cross-modal residual.
[0022] Optionally, the step of weighted summing of the residual values of the visual factor, the semantic constraint factor, and the laser factor to obtain the cross-modal residual includes:
[0023] The scene entropy and environmental change sensitivity of the current environment are determined based on the target point cloud data and the target image data.
[0024] Based on the scene entropy and the sensitivity to environmental changes, determine the visual weight corresponding to the residual value of the visual factor, the semantic weight corresponding to the residual value of the semantic constraint factor, and the laser weight corresponding to the residual value of the laser factor.
[0025] The cross-modal residual is obtained by weighted summation of the residual values and corresponding visual weights of the visual factors, the residual values and corresponding semantic weights of the semantic constraint factors, and the residual values and corresponding laser weights of the laser factors.
[0026] Optionally, the corrected localization result includes semantic feature variables, semantic topological constraints, and dynamic obstacle probability distribution;
[0027] The generation of inspection paths and control commands based on the corrected positioning results and the target map includes:
[0028] A semantic state set is constructed based on the semantic feature variables and the target map, the semantic state set is mapped onto a semantic manifold, and an optimization target is generated based on the mapping result;
[0029] A stochastic chance constraint is constructed based on the probability distribution of the dynamic obstacle;
[0030] The optimization objective is optimized hierarchically based on the semantic topology constraints and the random chance constraints, and the inspection path and control instructions are determined based on the optimization results.
[0031] This application also provides an apparatus for generating inspection paths, including:
[0032] The data acquisition module is used to acquire target point cloud data and target image data collected by multiple pre-integrated sensors at the same time and in the same space, as well as the target map corresponding to the target point cloud data and the target image data;
[0033] The path generation module is used to correct the current positioning error based on the target point cloud data and the target image data, and generate an inspection path and control instructions based on the corrected positioning results and the target map, and execute the inspection operation according to the inspection path and the control instructions.
[0034] This application also provides a computer-readable storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the inspection path generation method as described in any of the above embodiments.
[0035] This application also provides a computer device, including: one or more processors, and memory;
[0036] The memory stores computer-readable instructions, which, when executed by the one or more processors, perform the steps of the inspection path generation method as described in any of the above embodiments.
[0037] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0038] The inspection path generation method, apparatus, storage medium, and computer equipment provided in this application can, when generating an inspection path, first acquire target point cloud data and target image data collected simultaneously by multiple pre-integrated sensors in the same time and space, as well as a target map corresponding to the target point cloud data and target image data. This can not only synchronize the data collected by different sensors in time and space, ensuring the geometric and physical consistency of the motion trajectory, thereby improving the dynamic positioning accuracy, but also maintain the function through the synchronized redundant data when some sensors fail, effectively enhancing the robustness of the system. It can also reduce data processing latency through time and space synchronization, thereby meeting the real-time control requirements of dynamic scenarios, and can update the map in real time according to the current environment to improve the accuracy of navigation and positioning and adapt to complex environments. Next, this application can correct the current positioning error based on the target point cloud data and target image data to further improve the positioning accuracy. Finally, this application can generate a more accurate inspection path and control commands based on the corrected positioning results and target map, and execute the inspection operation according to the inspection path and control commands. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 A flowchart illustrating an inspection path generation method provided in this application embodiment;
[0041] Figure 2A flowchart illustrating the generation of inspection paths and control commands provided in this application embodiment;
[0042] Figure 3 This is a schematic diagram of the structure of an inspection path generation device provided in an embodiment of this application;
[0043] Figure 4 This is a schematic diagram of the internal structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0044] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0045] In one embodiment, such as Figure 1 As shown, Figure 1 This is a flowchart illustrating an inspection path generation method provided in an embodiment of this application; this application provides an inspection path generation method, which may include:
[0046] S110: Acquire target point cloud data and target image data collected by multiple pre-integrated sensors at the same time and in the same space, as well as the target map corresponding to the target point cloud data and target image data.
[0047] In this step, when generating the inspection path, multiple pre-integrated sensors can be used to acquire target point cloud data and target image data collected at the same time and in the same space. This allows the current location, equipment distribution, and surrounding environment information to be determined based on the target point cloud data and target image data, and a target map to be constructed.
[0048] Specifically, when this application performs high-precision tasks, it can acquire target point cloud data and target image data collected simultaneously by multiple pre-integrated sensors in the same time and space. Among them, the sensors pre-integrated in this application include, but are not limited to, LiDAR, vision sensors, inertial measurement units (IMU), millimeter-wave radar, BeiDou navigation system, etc., and can be configured according to the actual application environment, without limitation.
[0049] Among the aforementioned sensors, lidar can provide high-precision 3D point cloud data for environmental modeling and localization; visual sensors can include monocular, binocular, or multi-view cameras for image feature extraction, target recognition, and depth estimation; inertial measurement units can measure acceleration and angular velocity and provide high-frequency motion state information; millimeter-wave radar is used for long-distance obstacle detection and velocity measurement, and performs stably, especially in adverse weather conditions; and the BeiDou Navigation Satellite System can provide absolute position information to assist in global positioning.
[0050] Understandably, in multi-sensor systems, due to the inconsistent data acquisition times of different sensors, if each sensor uses its own local clock, factors such as crystal oscillator errors and temperature variations can lead to different clock frequencies and gradual time desynchronization, resulting in asynchronous initial times between sensors. This application addresses this by synchronizing the data acquisition times of each sensor to ensure timestamp alignment across multiple data sources. Furthermore, to eliminate spatial positional biases, this application can align the coordinate systems of each sensor to the same reference frame. This improves data fusion accuracy, avoids feature misalignment caused by spatiotemporal deviations (such as matching errors between laser point clouds and image pixels), enhances system robustness (allowing for continued functionality even when some sensors fail, relying on synchronized redundant data for short-term localization using only IMU and vision), and optimizes real-time performance by reducing data processing latency through synchronization, meeting the real-time control requirements of dynamic scenarios (such as robot obstacle avoidance and path planning).
[0051] For example, when a quadruped robot needs to traverse equipment areas and avoid energized structures and oil pipelines, the lidar on the quadruped robot can scan the three-dimensional outline of the equipment (such as scanning to generate a point cloud skeleton of substation equipment); the binocular camera can identify equipment nameplates (such as "#3 main transformer 101 switch"), instrument readings, insulator cracks and other defects; after collecting these raw data, they can be spatiotemporally synchronized to obtain target point cloud data and target image data.
[0052] Furthermore, after acquiring target point cloud data and target image data collected by various sensors at the same time and in the same space, the feature points in the target point cloud data and target image data can be decoupled from static and dynamic states. A static map can be constructed based on the static features obtained after decoupling, and then a corresponding inspection path can be generated based on the static map.
[0053] Specifically, after acquiring target point cloud data and target image data, this application utilizes multimodal features, such as geometric features in the target point cloud data and semantic features in the target image data, to determine the target through feature extraction and target recognition. This allows for the capture of kinematic anomalies and the quantification of feature stability through multimodal features, suppressing false detections in textured regions and improving dynamic detection accuracy.
[0054] Furthermore, after decoupling the feature points in the target point cloud data and target image data, dynamic and static feature points can be obtained. At this point, the application can filter out the dynamic feature points and use the static feature points to construct a static map. For example, the application can automatically filter out moving maintenance personnel, vehicle point clouds, etc. through dynamic-static decoupling. When the A-phase bushing thermometer is found to be obscured by a shadow, the shadow can be eliminated. Equipment status changes (such as the trip indicator light illuminating) can also be marked as semantic features. Then, the static features are used to construct the corresponding static map, thereby improving the accuracy of the map.
[0055] To improve real-time response performance and reduce computational complexity when constructing a static map, this application can also locally update the latest acquired global map based on the decoupled target point cloud data and target image data, thereby obtaining the target map. For example, this application can add static equipment (current transformers, insulators) from static features to the permanent map, and automatically associate visually recognized equipment nameplate information (such as "main transformer B phase") with point cloud clusters and substation drawing coordinates; when a newly installed surge arrester is detected, a local area in the latest acquired global map can be updated, thereby further improving the accuracy of the target map.
[0056] S120: Correct the current positioning error based on the target point cloud data and target image data, and generate inspection path and control instructions based on the corrected positioning results and target map, and execute inspection operations according to the inspection path and control instructions.
[0057] In this step, after acquiring target point cloud data and target image data collected by multiple pre-integrated sensors in the same time and space through S110, as well as the target map corresponding to the target point cloud data and target image data, the current positioning error can be corrected based on the target point cloud data and target image data. Based on the corrected positioning results and target map, inspection paths and control commands are generated. In this way, when performing inspection operations according to the inspection paths and control commands, the inspection efficiency and accuracy can be further improved, and the operational risks can be reduced.
[0058] Specifically, when correcting the current positioning error based on target point cloud data and target image data, this application can determine the pose of the LiDAR, the pose of the vision sensor, the laser-vision extrinsic parameters, and other auxiliary parameters using the target point cloud data and target image data. This allows for joint optimization, thereby reducing positioning errors. Furthermore, this application, through multimodal fusion positioning, can not only solve the problem of LiDAR noise caused by reflective surfaces such as ceramic insulators, but also address the visual degradation problem under insufficient nighttime lighting.
[0059] For example, this application can locate the device outline and the position of the main transformer oil level gauge using lidar, verify the oil level scale reading and identify the nameplate fine-tuning position using a visual sensor. When the reflective surfaces such as porcelain insulators fail, visual recognition is the primary method. When visual recognition degrades due to insufficient nighttime lighting, laser positioning is the primary method. Furthermore, laser outline matching is preferred under strong light, and the IMU magnetometer is turned off in areas with electromagnetic interference to improve positioning accuracy.
[0060] After correcting the current positioning error, this application can generate inspection paths and control commands based on the corrected positioning results and target map, and execute inspection operations according to the inspection paths and control commands. For example, the corrected positioning results of this application may include the current global pose, laser-visual extrinsic parameters, environmental features, etc. Therefore, this application can generate obstacle avoidance paths (such as bypassing temporary fences) based on relevant information in the positioning results and target map. It can also optimize gait parameters for anti-slip grid ground, adjust the center of gravity height when crossing cable trenches, and stop and replan when encountering sudden obstacles (falling tools), thereby reducing operational risks and improving autonomy and intelligence.
[0061] In the above embodiments, when generating the inspection path, target point cloud data and target image data collected by multiple pre-integrated sensors at the same time and in the same space, as well as the target map corresponding to the target point cloud data and target image data, can be acquired first. This can not only synchronize the data collected by different sensors in time and space, ensuring the geometric and physical consistency of the motion trajectory, thereby improving the dynamic positioning accuracy, but also maintain the function through the synchronized redundant data when some sensors fail, effectively enhancing the robustness of the system. It can also reduce data processing latency through time and space synchronization, thereby meeting the real-time control requirements of dynamic scenes. Furthermore, it can update the map in real time according to the current environment to improve the accuracy of navigation and positioning and adapt to complex environments. Next, this application can correct the current positioning error according to the target point cloud data and target image data to further improve the positioning accuracy. Finally, this application can generate a more accurate inspection path and control commands based on the corrected positioning results and target map, and execute the inspection operation according to the inspection path and control commands.
[0062] In one embodiment, S120, correcting the current positioning error based on the target point cloud data and the target image data to obtain a corrected positioning result, may include:
[0063] S121: Input the target point cloud data and the target image data into a pre-constructed spatiotemporal-semantic alignment model to obtain the alignment result output by the spatiotemporal-semantic alignment model.
[0064] S122: Correct the current positioning error based on the alignment result until the corrected positioning error is minimized, and obtain the corrected positioning result.
[0065] In this embodiment, when correcting the current positioning error based on target point cloud data and target image data, the pose of the lidar, the pose of the vision sensor, the laser-visual extrinsic parameters, and other auxiliary parameters can be determined using the target point cloud data and target image data. This allows for joint optimization, thereby reducing positioning errors. Furthermore, this application, through multimodal fusion positioning, can not only solve the lidar noise caused by reflective surfaces such as ceramic insulators, but also address the visual degradation problem under insufficient nighttime lighting.
[0066] For example, this application can locate the device outline and the position of the main transformer oil level gauge using lidar, verify the oil level scale reading and identify the nameplate fine-tuning position using a visual sensor. When the reflective surfaces such as porcelain insulators fail, visual recognition is the primary method. When visual recognition degrades due to insufficient nighttime lighting, laser positioning is the primary method. Furthermore, laser outline matching is preferred under strong light, and the IMU magnetometer is turned off in areas with electromagnetic interference to improve positioning accuracy.
[0067] Specifically, this application can pre-construct a spatiotemporal-semantic alignment model, which can be a multi-objective optimization function used in the vision-laser fusion system of robots or UAVs to jointly optimize feature matching, semantic consistency, and projection geometry error, thereby improving the localization and mapping accuracy in complex environments. Once the spatiotemporal-semantic alignment model is constructed, target point cloud data and target image data can be input into it to obtain the alignment result output by the model. Then, this application can correct the current localization error based on the alignment result until the corrected localization error is minimized, thus obtaining the final localization result. After generating the corresponding inspection map using this localization result, accurate navigation can be achieved.
[0068] In one embodiment, inputting the target point cloud data and the target image data into a pre-constructed spatiotemporal-semantic alignment model in S121 to obtain the alignment result output by the spatiotemporal-semantic alignment model may include:
[0069] S1211: Determine the state variables corresponding to the target point cloud data and the target image data, and construct a spatiotemporal-semantic union element.
[0070] S1212: Input the target point cloud data, the target image data, the state variables, and the spatiotemporal-semantic joint voxels into a pre-constructed spatiotemporal-semantic alignment model to obtain the alignment result output by the spatiotemporal-semantic alignment model.
[0071] In this embodiment, when using a pre-built spatiotemporal-semantic alignment model to optimize the localization error, the state variables corresponding to the target point cloud data and the target image data can be determined first, as follows:
[0072] The laser feature set corresponding to the target point cloud data in this application is:
[0073]
[0074] in, Let J be the laser feature set, which contains j laser features, each of which is a triplet. For the j-th laser feature Point cloud coordinates in three-dimensional space, For the j-th laser feature The laser descriptor is a 128-dimensional real vector. For the j-th laser feature The local surface covariance matrix is usually a 3×3 symmetric positive definite matrix, which can be obtained by calculating the distribution of the neighborhood point cloud using PCA.
[0075] The visual feature set corresponding to the target image data in this application is:
[0076]
[0077] in, Let be a visual feature set, containing i visual features, each of which is a triplet. For the i-th visual feature pixel coordinates, For the i-th visual feature The visual descriptor is a 256-dimensional real vector. For the i-th visual feature The semantic categories consist of C classes, where C is the total number of semantic categories. These categories can be generated by a pre-trained semantic segmentation network and aligned with the COCO categories.
[0078] Next, this application can determine the state variables corresponding to the target point cloud data and the target image data, these state variables... It is expressed as follows:
[0079]
[0080] in, For global pose, For the Lie algebraic form of laser-vision extrinsics, For other auxiliary parameters, such as environmental feature descriptors, sensor internal states (e.g., IMU bias), or deep learning encoded features, It is a special European style group. For Lie algebra. When this application... When SGID (Semantic-Geometric Integrated Descriptor) is used as a network parameter, its network structure is as follows:
[0081]
[0082] in, For the i-th pair of image features and laser features, This provides spatial consistency information for geometric difference vectors. This is the inverse of the camera intrinsic matrix, used to back-project pixel coordinates onto the normalized camera coordinate system. It is a multilayer perceptron used to fuse visual, laser, and geometric difference information.
[0083] Next, this application can construct spatiotemporal-semantic union elements, as follows:
[0084]
[0085] in, A dynamic set of voxels, used to represent dynamic environments or dynamic objects. The k-th voxel, where k is the voxel index. Let the coordinates be the center coordinates of the voxels. For timestamp windows (ΔT=0.1s), The semantic probability distribution can be updated using Bayesian methods. The length of the timestamp window. The current time is as follows:
[0086]
[0087] in, This represents the semantic probability distribution after the nth update. This represents the semantic probability distribution (prior) for the (n-1)th update. For the observation term (likelihood), i.e., at location ,time Semantic categories observed The probability of the observation item The dynamic feature set obtained after the above decoupling of static and dynamic features can be used for initialization.
[0088] Furthermore, the spatiotemporal-semantic alignment model of this application can be:
[0089]
[0090] in, For the residuals of the visual factors, For describing sub-matching items, For semantic consistency items, For the projection geometry error term, This is a cross-modal descriptor mapping network (two-stream CNN + PointNet structure), with visual descriptors as input. Output and laser descriptor By mapping visual descriptors to the same feature space as laser descriptors, direct comparison can be facilitated. The matrix is a Mahalanobis distance matrix, which can be updated through online learning. This represents a matching pair between the i-th visual feature and the j-th laser feature. For camera projection model, This is the intrinsic parameter matrix. Is with The semantic probability distribution of the corresponding modality (or the distribution from the prior / map). For laser-visual extrinsics, KL divergence measures the difference between two distributions. These are semantic weight coefficients, used to balance the importance of semantic terms. These are geometric weighting coefficients, which control the contribution of geometric errors.
[0091] When the target point cloud data, target image data, state variables, and spatiotemporal-semantic joint voxels are input into the above spatiotemporal-semantic alignment model, the alignment result output by the spatiotemporal-semantic alignment model can be obtained. By optimizing the alignment result, the present application achieves accurate alignment of multimodal data (visual, laser, semantic), thereby solving the limitation problem of a single sensor in complex environments.
[0092] In one embodiment, step S122, correcting the current positioning error based on the alignment result until the corrected positioning error is minimized, to obtain the corrected positioning result, may include:
[0093] S1221: Calculate the compensation amount of the state variable based on the state variable and the alignment result. After updating the state variable according to the compensation amount, return to execute the input of the target point cloud data, the target image data, the state variable and the spatiotemporal-semantic joint voxel into the pre-constructed spatiotemporal-semantic alignment model and its subsequent steps until the calculated compensation amount is optimal.
[0094] S1222: Use the state variable when the compensation amount is optimal as the corrected positioning result.
[0095] In this embodiment, after obtaining the alignment result, the joint optimization engine can be used to optimize the alignment result to achieve accurate alignment of multimodal data (visual, laser, semantic).
[0096] Specifically, this application can first calculate the compensation amount of the state variables based on the state variables and the alignment results, and then update the state variables based on the compensation amount. Next, this application can continue to input the target point cloud data, target image data, updated state variables, and spatiotemporal-semantic joint voxels into the pre-constructed spatiotemporal-semantic alignment model, and obtain the alignment result output by the spatiotemporal-semantic alignment model. Then, based on the updated state variables and the alignment result, the compensation amount of the updated state variables is calculated, and it is determined whether the compensation amount is the optimal compensation amount. If the compensation amount is the optimal compensation amount, the updated state variables are updated again using the compensation amount to obtain the corrected positioning result. If the compensation amount is not the optimal compensation amount, the state variables can be updated again based on the compensation amount, and the alignment result can be calculated again until the calculated compensation amount is optimal. In this way, accurate alignment of multimodal data (visual, laser, semantic) can be achieved, thereby reducing positioning errors.
[0097] In one embodiment, the alignment result may include residual values of visual factors and residual values of semantic constraint factors.
[0098] S1221, calculating the compensation amount of the state variable based on the state variable and the alignment result, may include:
[0099] S12211: Determine the residual value of the laser factor based on the target point cloud data.
[0100] S12212: The cross-modal residual is obtained by weighted summing of the residual values of the visual factor, the semantic constraint factor, and the laser factor.
[0101] S12213: Calculate the compensation amount of the state variable based on the state variable and the cross-modal residual.
[0102] In this embodiment, when calculating the compensation amount of the state variable, the residual value of the laser factor can be determined first based on the target point cloud data. Then, the residual values of the visual factor and the semantic constraint factor can be determined based on the alignment results. Next, the residual values of the visual factor, the semantic constraint factor, and the laser factor are weighted and summed to obtain the cross-modal residual. Finally, this application can calculate the compensation amount of the state variable based on the state variable and the cross-modal residual.
[0103] Specifically, the residual value of the laser factor in this application can be obtained by the following formula:
[0104]
[0105] in, Is with The corresponding 3D coordinates of the global map point in the global coordinate system. Let be the covariance matrix (3×3) of the j-th laser point, describing the uncertainty of the point's location.
[0106] In this application, when determining the residual values of the visual factors and the semantic constraint factors based on the alignment results, the residual values of the visual factors can be those in the aforementioned spatiotemporal-semantic alignment model. The residual value of the semantic constraint factor can be from the spatiotemporal-semantic alignment model described above. Once the individual residual values are obtained, they can be weighted and summed to obtain the cross-modal residual. Then, this application can calculate the compensation amount of the state variable based on the state variable and the cross-modal residual.
[0107] Furthermore, when calculating the compensation amount of the state variables based on the state variables and cross-modal residuals, this application can use a joint optimization method of multi-sensor fusion and semantic guidance. This method achieves high-precision incremental updates of pose and extrinsic parameters by decomposing the Jacobian matrix and dynamically adjusting the regularization parameters. Its core objective is to calculate the optimal increment of the state variables by minimizing the cross-modal residuals to compensate for the state variables.
[0108] For example, this application can first construct a joint Jacobian matrix:
[0109]
[0110] in, For the joint Jacobian matrix, For cross-modal residuals, For state variables, The linearization point of the current state (the state value used when differentiating).
[0111] Next, this application can decompose the joint Jacobian matrix into the following using the implicit differentiation chain rule:
[0112]
[0113] in, To align the derivatives of the residuals with respect to the state variables, Output the derivative of the SGID network with respect to the network parameter θ. To match the derivative of the residual with respect to the laser-visual extrinsic parameters in the laser point cloud, For external parameters to state variables The derivative of (usually the identity mapping or a simple transformation). The derivative of semantic constraints with respect to the descriptor. semantic variables to state variables The derivative of .
[0114] Furthermore, the compensation amount can be calculated using the following formula:
[0115]
[0116] in, This represents the compensation amount (increment) for the state variables, i.e., the update amount of the state variables in this iteration. For the sensor noise matrix, The values represent the measurement noise variance for vision and laser, respectively, and the weighted reliability of different sensors. Let be the regularization matrix, where the extrinsic regularization term is . Based on semantic confidence Dynamic adjustment The coefficients of the pose regularization term are: For SGID network parameters, the regularization coefficients are... This is the inverse noise matrix, used to weight the residuals based on sensor reliability.
[0117] Parameter update after compensation:
[0118]
[0119] in, Let be the extrinsic parameter matrix of the laser radar to the camera at the t-th iteration. Let be the extrinsic parameter matrix after the (t+1)th iteration. For Lie algebra increments, The global pose is updated after the (t+1)th iteration. Let be the global pose at the t-th iteration. For the increment of global pose, This refers to incremental update operations on a Lie group.
[0120] This application achieves breakthroughs in positioning accuracy, dynamic adaptability, and cross-modal consistency through a triple innovative mechanism of spatiotemporal-semantic joint voxels, multimodal factor graph optimization, and semantic-guided error compensation. Deep coupling with the dynamic environment SLAM module constructs a complete dynamic scene cognition system from low-level perception to high-level understanding, providing a new paradigm for multimodal fusion positioning in fields such as autonomous driving and embodied intelligence.
[0121] In one embodiment, the weighted summation of the residual values of the visual factor, the semantic constraint factor, and the laser factor in step S12212 to obtain the cross-modal residual may include:
[0122] S22121: Determine the scene entropy and environmental change sensitivity of the current environment based on the target point cloud data and the target image data.
[0123] S22122: Determine the visual weight corresponding to the residual value of the visual factor, the semantic weight corresponding to the residual value of the semantic constraint factor, and the laser weight corresponding to the residual value of the laser factor based on the scene entropy and the environmental change sensitivity.
[0124] S22123: The cross-modal residual is obtained by weighted summation of the residual values and corresponding visual weights of the visual factors, the residual values and corresponding semantic weights of the semantic constraint factors, and the residual values and corresponding laser weights of the laser factors.
[0125] In this embodiment, when performing weighted summation, the scene entropy and environmental change sensitivity of the current environment can be determined first based on the target point cloud data and target image data. Then, the visual weights corresponding to the residual values of the visual factors, the semantic weights corresponding to the residual values of the semantic constraint factors, and the laser weights corresponding to the residual values of the laser factors can be determined based on the scene entropy and environmental change sensitivity. In this way, the cross-modal residuals can be obtained by performing weighted summation based on the residual values of the visual factors and their corresponding visual weights, the residual values of the semantic constraint factors and their corresponding semantic weights, and the residual values of the laser factors and their corresponding laser weights.
[0126] Specifically, when determining the scene entropy and environmental change sensitivity of the current environment based on target point cloud data and target image data, this application can calculate the entropy of grayscale or gradient histograms for target image data by dividing the image into blocks, and then calculate the scene entropy after global averaging. For target point cloud data, this application can calculate the scene entropy based on the entropy values of point density, normal changes, or semantic distribution. The environmental change sensitivity of this application can be determined through the calculation process of the above-mentioned environmental change quantification indicators.
[0127] Once the scene entropy and environmental change sensitivity of the current environment are determined, this application can calculate the visual weight corresponding to the residual value of the visual factor, the semantic weight corresponding to the residual value of the semantic constraint factor, and the laser weight corresponding to the residual value of the laser factor according to the following formulas:
[0128]
[0129]
[0130] in, For scene entropy, Let be the scene entropy at the current time t. Let be the scene entropy at the previous time t-1. The change in scene entropy is calculated as the difference between the Euclidean norm (2-norm) of the scene entropy at the current moment and the scene entropy at the previous moment. Sensitivity to environmental changes As weights, different weights can be obtained when different scene entropy and environmental change sensitivity are input into this application.
[0131] In one embodiment, such as Figure 2 As shown, Figure 2 This is a flowchart illustrating the generation of inspection paths and control commands provided in an embodiment of this application; the corrected positioning results may include semantic feature variables, semantic topological constraints, and dynamic obstacle probability distributions.
[0132] S120, which generates inspection paths and control commands based on the corrected positioning results and the target map, may include:
[0133] S123: Construct a semantic state set based on semantic feature variables and target map, map the semantic state set onto semantic manifold, and generate optimization target based on the mapping result.
[0134] S124: Constructing stochastic chance constraints based on the probability distribution of dynamic obstacles.
[0135] S125: Perform hierarchical optimization of the optimization objective based on semantic topological constraints and stochastic chance constraints, and determine the inspection path and control instructions based on the optimization results.
[0136] In this embodiment, after the current positioning error is corrected, an inspection path and control commands can be generated based on the corrected positioning result and the target map, and the inspection operation can be executed according to the inspection path and control commands. For example, the corrected positioning result may include the current global pose, laser-visual extrinsic parameters, environmental features, etc. Therefore, the application can generate obstacle avoidance paths (such as bypassing temporary fences) based on the relevant information in the positioning result and the target map. It can also optimize gait parameters for anti-slip grid ground, adjust the center of gravity height when crossing cable trenches, and stop and replan when encountering sudden obstacles (fallen tools), thereby reducing operational risks and improving autonomy and intelligence.
[0137] In one specific implementation, this application can first construct a semantic state set based on semantic feature variables and a target map, then map the semantic state set onto a semantic manifold, and generate an optimization objective based on the mapping result. Next, this application can also construct stochastic chance constraints based on the probability distribution of dynamic obstacles. The specific construction process can be set according to existing technologies and will not be elaborated here. Once the constraints and optimization objective are determined, this application can perform hierarchical optimization of the optimization objective based on semantic topological constraints and stochastic chance constraints, and determine the inspection path and control commands based on the optimization results.
[0138] Wherein, the semantic state set at time k in this application It can be defined as follows:
[0139]
[0140] in, Let be the position variable of the robot in three-dimensional space. For the robot's pose variables, The roll angle is the angle of rotation about the x-axis. The pitch angle is rotated about the y-axis. The yaw angle is the angle of rotation about the z-axis. Let the robot's velocity variables be the three coordinate axes. Let there be m semantic feature variables (such as predicted speed of dynamic obstacles, friction coefficient of ground material, etc.), where m is the total number of semantic variables. It has 9+m dimensions (the first 9 dimensions are geometric states, and the last m dimensions are semantic variables). The position, attitude, and velocity variables mentioned above can be predicted by combining the target map and semantic feature variables.
[0141] Since the aforementioned set of semantic states is a semantic representation of high-level objectives (such as "safe driving" and "optimal energy consumption"), this application can associate it with underlying physical quantities through manifold projection to determine the optimization objective. For example, if the automated inspection robot in this application wants to achieve the semantic objective of "emergency obstacle avoidance," it can map the corresponding set of semantic states "obstacle distance" and "vehicle speed" onto the semantic manifold to form a risk level, and take the minimum risk level as the optimization objective.
[0142] Furthermore, this application can also embed the set of semantic states into the dynamic model, as shown in the following formula:
[0143]
[0144] in, This is the state transition matrix, which contains physical parameters such as robot mass and moment of inertia. This is the system state vector at time k, which contains the robot's dynamic physical quantities (such as position, velocity, attitude angle, etc.). Let k+1 be the system state vector. To control the input matrix (corresponding to the joint torque of 6 degrees of freedom). This is the semantic coupling matrix, which describes the influence of semantic features on motion. The process noise covariance matrix is... The control input vector corresponds to the control commands of the actuator (such as joint torque and motor voltage). It is a set of semantic states.
[0145] Next, this application can perform hierarchical optimization of the optimization objective based on semantic topological constraints and stochastic chance constraints, specifically the hierarchical objective function. as follows:
[0146]
[0147] in, To track performance items, To control the cost item, For semantic penalty terms, To predict the actual system output at step k in the time domain, The reference trajectory is the target trajectory or state sequence that the system is expected to track in the future time domain. The output weight matrix is denoted as , and the control input is denoted as at step k. To control the input weight matrix, For semantic trade-off coefficients, For semantic penalty function, For semantic manifold mapping matrix, Let the semantic constraint space be the semantically valid region (e.g., "collision risk < threshold" or "energy efficiency > minimum standard"). The typical form of the semantic penalty function is as follows:
[0148]
[0149] in, It is a distance metric (such as Riemannian geometric distance) on a semantic manifold, used to evaluate the actual semantic state and the legal region. The degree of deviation, This refers to mapping the system output at step k to the semantic feature space. As a distance metric on the semantic manifold, it computes the mapped semantic state and the legal region. The degree of deviation, To predict the length of the time domain (optimize the number of steps).
[0150] In the above embodiments, this application reduces computational complexity by separating high-level global planning from low-level local control, and improves robustness by explicitly modeling environmental dynamics (such as moving obstacles) and sensor noise. Thus, during substation inspections, it can generate an initial path based on a semantic map (equipment distribution, safety level) and mark high-risk areas. It can also replan local trajectories after detecting dynamic obstacles (such as birds), and adjust foot force in real time to avoid instability on slippery insulators, while ensuring the robotic arm is aligned with the detection point.
[0151] The inspection path generation device provided in the embodiments of this application is described below. The inspection path generation device described below and the inspection path generation method described above can be referred to in correspondence.
[0152] In one embodiment, such as Figure 3 As shown, Figure 3 This application provides a schematic diagram of an inspection path generation device according to an embodiment of the present application; the present application also provides an inspection path generation method device, which may include a data acquisition module 210 and a path generation module 220, specifically including the following:
[0153] The data acquisition module 210 is used to acquire target point cloud data and target image data collected by multiple pre-integrated sensors at the same time and in the same space, as well as the target map corresponding to the target point cloud data and the target image data.
[0154] The path generation module 220 is used to correct the current positioning error based on the target point cloud data and the target image data, and generate an inspection path and control instructions based on the corrected positioning results and the target map, and execute the inspection operation according to the inspection path and the control instructions.
[0155] In the above embodiments, when generating the inspection path, target point cloud data and target image data collected by multiple pre-integrated sensors at the same time and in the same space, as well as the target map corresponding to the target point cloud data and target image data, can be acquired first. This can not only synchronize the data collected by different sensors in time and space, ensuring the geometric and physical consistency of the motion trajectory, thereby improving the dynamic positioning accuracy, but also maintain the function through the synchronized redundant data when some sensors fail, effectively enhancing the robustness of the system. It can also reduce data processing latency through time and space synchronization, thereby meeting the real-time control requirements of dynamic scenes. Furthermore, it can update the map in real time according to the current environment to improve the accuracy of navigation and positioning and adapt to complex environments. Next, this application can correct the current positioning error according to the target point cloud data and target image data to further improve the positioning accuracy. Finally, this application can generate a more accurate inspection path and control commands based on the corrected positioning results and target map, and execute the inspection operation according to the inspection path and control commands.
[0156] In one embodiment, this application also provides a storage medium storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of the inspection path generation method as described in any of the above embodiments.
[0157] In one embodiment, this application also provides a computer device, including: one or more processors, and memory.
[0158] The memory stores computer-readable instructions, which, when executed by the one or more processors, perform the steps of the inspection path generation method as described in any of the above embodiments.
[0159] Indicatively, such as Figure 4 As shown, Figure 4 This is a schematic diagram of the internal structure of a computer device 300 provided in an embodiment of this application. The computer device 300 can be provided as a server. (Refer to...) Figure 4 The computer device 300 includes a processing component 302, which further includes one or more processors, and memory resources represented by memory 301 for storing instructions, such as application programs, that can be executed by the processing component 302. The application programs stored in memory 301 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 302 is configured to execute instructions to perform the inspection path generation method of any of the above embodiments.
[0160] The computer device 300 may also include a power supply component 303 configured to perform power management of the computer device 300, a wired or wireless network interface 304 configured to connect the computer device 300 to a network, and an input / output (I / O) interface 305. The computer device 300 may operate on an operating system stored in memory 301, such as Windows Server™, Mac OS X™, Unix™, Linux™, Free BSD™, or similar.
[0161] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0162] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0163] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.
[0164] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating inspection paths, characterized in that, The method includes: Acquire target point cloud data and target image data collected by multiple pre-integrated sensors at the same time and in the same space, as well as a target map corresponding to the target point cloud data and the target image data; The current positioning error is corrected based on the target point cloud data and the target image data. An inspection path and control instructions are generated based on the corrected positioning results and the target map. The inspection operation is then executed according to the inspection path and the control instructions.
2. The inspection path generation method according to claim 1, characterized in that, The step of correcting the current positioning error based on the target point cloud data and the target image data to obtain the corrected positioning result includes: The target point cloud data and the target image data are input into a pre-constructed spatiotemporal-semantic alignment model to obtain the alignment result output by the spatiotemporal-semantic alignment model; The current positioning error is corrected based on the alignment result until the corrected positioning error is minimized, thus obtaining the corrected positioning result.
3. The inspection path generation method according to claim 2, characterized in that, The step of inputting the target point cloud data and the target image data into a pre-constructed spatiotemporal-semantic alignment model to obtain the alignment result output by the spatiotemporal-semantic alignment model includes: Determine the state variables corresponding to the target point cloud data and the target image data, and construct a spatiotemporal-semantic joint element; The target point cloud data, the target image data, the state variables, and the spatiotemporal-semantic joint voxels are input into a pre-constructed spatiotemporal-semantic alignment model to obtain the alignment result output by the spatiotemporal-semantic alignment model.
4. The inspection path generation method according to claim 3, characterized in that, The step of correcting the current positioning error based on the alignment result until the corrected positioning error is minimized, to obtain the corrected positioning result, includes: Based on the state variable and the alignment result, calculate the compensation amount of the state variable. After updating the state variable according to the compensation amount, return to execute the input of the target point cloud data, the target image data, the state variable and the spatiotemporal-semantic joint voxel into the pre-constructed spatiotemporal-semantic alignment model and its subsequent steps until the calculated compensation amount is optimal. The state variable at which the compensation amount is optimal is used as the corrected positioning result.
5. The inspection path generation method according to claim 4, characterized in that, The alignment result includes the residual values of the visual factors and the residual values of the semantic constraint factors; The calculation of the compensation amount for the state variables based on the state variables and the alignment result includes: The residual value of the laser factor is determined based on the target point cloud data; The cross-modal residual is obtained by weighted summing of the residual values of the visual factor, the semantic constraint factor, and the laser factor. The compensation amount of the state variable is calculated based on the state variable and the cross-modal residual.
6. The inspection path generation method according to claim 5, characterized in that, The step of weighted summing of the residual values of the visual factor, the semantic constraint factor, and the laser factor to obtain the cross-modal residual includes: The scene entropy and environmental change sensitivity of the current environment are determined based on the target point cloud data and the target image data. Based on the scene entropy and the sensitivity to environmental changes, determine the visual weight corresponding to the residual value of the visual factor, the semantic weight corresponding to the residual value of the semantic constraint factor, and the laser weight corresponding to the residual value of the laser factor. The cross-modal residual is obtained by weighted summation of the residual values and corresponding visual weights of the visual factors, the residual values and corresponding semantic weights of the semantic constraint factors, and the residual values and corresponding laser weights of the laser factors.
7. The inspection path generation method according to any one of claims 1-6, characterized in that, The corrected localization results include semantic feature variables, semantic topological constraints, and dynamic obstacle probability distribution; The generation of inspection paths and control commands based on the corrected positioning results and the target map includes: A semantic state set is constructed based on the semantic feature variables and the target map, the semantic state set is mapped onto a semantic manifold, and an optimization target is generated based on the mapping result; A stochastic chance constraint is constructed based on the probability distribution of the dynamic obstacle; The optimization objective is optimized hierarchically based on the semantic topology constraints and the random chance constraints, and the inspection path and control instructions are determined based on the optimization results.
8. An apparatus for generating inspection paths, characterized in that, include: The data acquisition module is used to acquire target point cloud data and target image data collected by multiple pre-integrated sensors at the same time and in the same space, as well as the target map corresponding to the target point cloud data and the target image data; The path generation module is used to correct the current positioning error based on the target point cloud data and the target image data, and generate an inspection path and control instructions based on the corrected positioning results and the target map, and execute the inspection operation according to the inspection path and the control instructions.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of the inspection path generation method as described in any one of claims 1 to 7.
10. A computer device, characterized in that, include: One or more processors, and memory; The memory stores computer-readable instructions, which, when executed by the one or more processors, perform the steps of the inspection path generation method as described in any one of claims 1 to 7.
Citation Information
Cited By
Chemical safety inspection robot control system and method based on sensor fusion
CN121492063A