Switch cabinet detection robot navigation method based on laser radar and vision fusion

By fusing LiDAR and vision to generate an enhanced environmental map, the problem of inaccurate navigation and positioning of robots in substations has been solved, enabling precise identification and navigation of switchgear and improving the accuracy and robustness of navigation.

CN121764080APending Publication Date: 2026-03-31STATE GRID HEBEI ELECTRIC POWER CO LTD XIONGAN NEW DISTRICT POWER SUPPLY CO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing robot navigation solutions struggle to accurately distinguish and locate specific target switchgear in the dense, unstructured environment of highly similar equipment in substations, leading to inaccurate navigation and positioning.

Method used

By employing a fusion approach of LiDAR and vision, an enhanced environmental map that combines geometric and semantic information is generated by acquiring LiDAR point cloud data and visual image data. This map is then combined with real-time detection commands to perform global path planning, thereby enabling robot navigation.

Benefits of technology

It improves the accuracy and intelligence of robot navigation, ensures the accuracy of navigation targets, and enhances robustness and reliability in complex lighting environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121764080A_ABST
    Figure CN121764080A_ABST
Patent Text Reader

Abstract

The invention provides a switch cabinet detection robot navigation method based on laser radar and vision fusion, and relates to the technical field of power grids. According to the method, the visual semantic label is introduced, so that the robot has the capability of understanding a scene function, different target switch cabinets with similar appearances can be uniquely identified, the spanning from geometric positioning to semantic identification is realized, and the accuracy of a navigation target is ensured. The generated enhanced environment map provides a unified information source with both geometric constraints and semantic constraints for path planning, so that the planned path not only can safely avoid obstacles, but also can actively guide an operable area of a target cabinet body, and the intelligence and task fitting degree of the navigation path are improved. According to the invention, through tightly coupled multi-sensor information fusion, the overall robustness and reliability of the transformer substation in a complex illumination environment are effectively improved, the problem of inaccurate navigation and positioning in a current robot navigation scheme is solved, and the accuracy of robot navigation and positioning is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power grid technology, and in particular to a navigation method for a switchgear inspection robot based on the fusion of lidar and vision. Background Technology

[0002] In the field of substation indoor equipment inspection, the use of robots for automated inspection of switchgear has become a development trend. Switchgear rooms typically contain dozens or even hundreds of switchgear units with highly similar appearances and uniform models, forming a dense, unstructured environment of highly similar equipment. In this specific environment, autonomous navigation of robots faces significant challenges, the core pain point being the difficulty of existing technologies in accurately distinguishing and locating specific target switchgear units.

[0003] Currently, the mainstream robot navigation solution is based on LiDAR (Light Detection and Ranging) technology. This technology utilizes point cloud data to construct a high-precision geometric map of the environment and achieves simultaneous localization and map building. However, LiDAR perceives the environment at the geometric level and cannot understand the functional semantics of the environment. When faced with rows of nearly identical switch cabinets, the robot can only perceive a series of similar cubic obstacles and cannot uniquely distinguish "switch cabinet 1" from "switch cabinet 2" based on geometric features. This leads to frequent localization confusion for the robot; although it knows its approximate position on the global map, it cannot accurately match its position with the target switch cabinet identifier in the task instructions, thus failing to reliably navigate to the correct target cabinet to perform the inspection task. Furthermore, switch cabinet inspection often requires the robotic arm or inspection instrument to be precisely aligned with specific components (such as meters and indicator lights) on the cabinet door, which LiDAR-based navigation technology cannot accurately and reliably identify.

[0004] Therefore, robot navigation solutions suffer from inaccurate navigation and positioning. Summary of the Invention

[0005] This invention provides a navigation method for switch cabinet inspection robots based on the fusion of lidar and vision, which solves the problem of inaccurate navigation and positioning in current robot navigation solutions.

[0006] In a first aspect, the present invention provides a navigation method for a switchgear inspection robot based on the fusion of lidar and vision. The method includes: acquiring lidar point cloud data collected by a lidar sensor mounted on the robot, and visual image data collected by a vision sensor; performing synchronous localization and map construction based on the lidar point cloud data to generate a geometric map; performing visual feature recognition and semantic segmentation based on the visual image data to determine pixel-level semantic labels; fusing the pixel-level semantic labels and the geometric map at the feature layer, and generating an enhanced environment map that combines geometric and semantic information through coordinate transformation and data association; and performing global path planning based on the enhanced environment map and the detection instructions of the target switchgear received in real time to obtain the optimal path to the target switchgear, thereby realizing robot navigation.

[0007] Secondly, embodiments of the present invention provide a navigation device for a switchgear inspection robot based on the fusion of LiDAR and vision. The device includes a communication module and a processing module. The communication module is used to acquire laser point cloud data collected by the LiDAR sensor mounted on the robot, and visual image data collected by the vision sensor. The processing module is used to perform synchronous localization and map construction based on the laser point cloud data to generate a geometric map; perform visual feature recognition and semantic segmentation based on the visual image data to determine pixel-level semantic labels; fuse the pixel-level semantic labels and the geometric map at the feature layer, and generate an enhanced environment map that combines geometric and semantic information through coordinate transformation and data association; and perform global path planning based on the enhanced environment map and the real-time received detection instructions for the target switchgear to obtain the optimal path to the target switchgear, thereby achieving robot navigation.

[0008] Thirdly, embodiments of the present invention provide a switch cabinet inspection robot navigation system based on lidar and vision fusion. The system includes an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor is used to call and run the computer program stored in the memory to perform the steps of the method as described in the first aspect and any possible implementation thereof.

[0009] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of the method as described in the first aspect and any possible implementation thereof.

[0010] This invention provides a robot navigation method for switchgear inspection based on the fusion of LiDAR and vision. By introducing visual semantic tags, this invention enables the robot to understand the scene and uniquely identify different target switchgear with similar appearances, achieving a leap from geometric localization to semantic recognition and ensuring the accuracy of navigation targets. The generated enhanced environment map provides a unified information source with both geometric and semantic constraints for path planning, enabling the planned path to not only safely avoid obstacles but also actively guide to the operable area of ​​the target switchgear, improving the intelligence and task relevance of the navigation path. Through tightly coupled multi-sensor information fusion, this invention effectively improves the overall robustness and reliability of the robot in complex lighting environments in substations, solving the problem of inaccurate navigation and positioning in current robot navigation solutions and improving the accuracy of robot navigation and positioning. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart illustrating a navigation method for a switch cabinet inspection robot based on the fusion of lidar and vision, provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of a switch cabinet inspection robot navigation device based on the fusion of lidar and vision provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.

[0014] In the description of this invention, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" and "more than one" refer to two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.

[0015] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.

[0016] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the steps or modules listed, but may optionally include other steps or modules not listed, or may optionally include other steps or modules inherent to such process, method, product, or device.

[0017] To make the objectives, technical solutions, and advantages of the present invention clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0018] like Figure 1 As shown, this embodiment of the invention provides a navigation method for a switch cabinet inspection robot based on the fusion of lidar and vision. The method includes steps S101-S104.

[0019] S101. Acquire laser point cloud data collected by the lidar sensor on the robot and visual image data collected by the vision sensor.

[0020] In some embodiments, laser point cloud data is a three-dimensional spatial data format that precisely records the geometric contours of an environment in the form of discrete points. Visual image data is a two-dimensional pixel array containing color and texture information.

[0021] For example, embodiments of the present invention can activate the robot's lidar and camera. The lidar measures the distance to the surrounding environment by emitting a laser beam and receiving its return signal, forming laser point cloud data (i.e., a set of a large number of three-dimensional coordinate points, each point representing a spatial position on the surface of an object). The vision sensor simultaneously captures color images, generating visual image data.

[0022] S102. Based on laser point cloud data, perform synchronous positioning and map construction to generate a geometric map.

[0023] In some embodiments, simultaneous localization and mapping (SMR) refers to a technique where a robot simultaneously estimates its own position and builds a map of the environment in an unknown environment. A geometric map is a spatially structured map that only describes the shape, size, and location of objects in the environment, without containing any semantic information about what those objects are.

[0024] For example, embodiments of the present invention can employ a graph-optimized SLAM algorithm to process point cloud data. This algorithm estimates the robot's own motion (localization) by matching point clouds of consecutive frames (inter-frame matching) and corrects accumulated errors by identifying revisited locations (loop closure detection), ultimately constructing a globally consistent geometric map containing only the outlines of objects such as walls and switch cabinets. This map is typically represented as a two-dimensional grid map (dividing the environment into grids, with each grid marked as occupied, free, or unknown) or a three-dimensional point cloud map.

[0025] As one possible implementation, step S102 can be specifically implemented as steps S1021-S1024.

[0026] S1021. Perform cluster analysis on the laser point cloud data to identify and filter out dynamic point clouds generated by moving personnel or equipment in the environment, and generate static environmental point cloud data.

[0027] In some embodiments, clustering analysis is a technique for dividing objects in a dataset into multiple groups (clusters) such that objects within the same group are similar to each other, while objects in different groups are significantly different. Static environment point cloud data is 3D point cloud data containing only the fixed and unchanging environmental structure after filtering out all temporary and movable objects.

[0028] For example, embodiments of the present invention can use the Euclidean clustering algorithm to process each frame of laser point cloud. This algorithm groups points according to the spatial distance between them, classifying points that are close to each other as the same object. By comparing the positional changes of each cluster in multiple consecutive frames of point cloud, moving clusters (such as walking workers or moving vehicles) can be identified. These point clouds identified as dynamic objects will be directly discarded, retaining only the point clouds of static objects such as walls and switch cabinets, thereby generating clean static environmental point cloud data.

[0029] S1022. Extract multi-scale geometric features from static environmental point cloud data.

[0030] In some embodiments, the multi-scale geometric features include global point cloud features for loop closure detection and local line and surface features for inter-frame matching.

[0031] In some embodiments, multi-scale geometric features include both features describing local fine structures (lines, surfaces) and features describing the overall macroscopic layout (global descriptors). Loop closure detection refers to the robot's ability to recognize when it has revisited a previously visited location, which is crucial for eliminating accumulated drift errors in the SLAM process.

[0032] For example, embodiments of the present invention can calculate the normal vector and curvature of the neighborhood surrounding each point in a single frame of the point cloud. Regions with drastic changes in normal vector direction are extracted as line features (such as wall edges), while regions with gentle changes are extracted as surface features (such as wall surfaces, cabinet surfaces). These features are computationally efficient and used for rapid matching between adjacent frames to estimate robot motion. Global point cloud feature extraction: For the point cloud of an entire frame or a keyframe, its global descriptor is calculated. This feature can highly summarize the global geometric distribution characteristics of the point cloud at that location, and is used for fast and accurate loop closure detection when the robot revisits a previous location.

[0033] S1023. Tightly couple and optimize the local line and surface features in the multi-scale geometric features with the pre-integrated data of the inertial measurement unit to output the robot pose estimate.

[0034] In some embodiments, tightly coupled optimization: a multi-sensor fusion architecture that jointly optimizes raw observation data or low-level features from different sensors in the same state estimator, resulting in higher accuracy. Robot pose estimation: estimates of the robot's position (x, y, z coordinates) and attitude (pitch, yaw, roll angle) in three-dimensional space.

[0035] For example, the inertial measurement unit (IMU) provides high-frequency angular velocity and acceleration measurements. During the interval between two LiDAR scans, the IMU data is pre-integrated to obtain a prediction of relative motion. Subsequently, an optimization problem is constructed to jointly optimize the results obtained from the LiDAR observation model (pose transformation calculated based on line-surface feature matching) and the IMU pre-integration model. Through this tightly coupled approach, the IMU provides motion priors, compensating for the limitations of LiDAR in fast-moving or feature-deficient scenarios, ultimately outputting a higher-frequency, smoother, and more accurate robot pose estimate.

[0036] S1024. Based on robot pose estimation and global point cloud features, a pose graph optimization method is used to construct a globally consistent 3D point cloud semantic map skeleton to obtain a geometric map.

[0037] In some embodiments, pose graph optimization: a backend optimization technique that represents the robot's pose and the spatial constraints between them as a graph model, and adjusts the pose through optimization algorithms so that the entire graph satisfies these constraints. Geometric map: a point cloud map that accurately describes the three-dimensional geometric structure of the environment after global optimization, serving as a precise spatial skeleton for subsequent semantic fusion.

[0038] For example, in this embodiment of the invention, a pose graph can be constructed using robot poses as nodes and the transformation relationships between poses (given by inter-frame matching or IMU) as edges. When the loop closure detection module detects that the robot has returned to its previous position by comparing global point cloud features, a new constraint edge is added to the graph. Finally, graph optimization techniques are used to adjust and optimize the entire graph, minimizing the error of all constraint edges. This process can evenly distribute the accumulated error across the entire trajectory, thereby obtaining a robot motion trajectory with extremely high global consistency. By stitching together the point clouds corresponding to all optimized poses, a final accurate and non-overlapping 3D point cloud map, i.e., a geometric map, is obtained.

[0039] S103. Based on visual image data, perform visual feature recognition and semantic segmentation to determine pixel-level semantic labels.

[0040] In some embodiments, semantic segmentation is a computer vision task aimed at classifying each pixel in an image into a predefined semantic category. Pixel-level semantic labeling refers to the precise annotation of the object category to which each pixel in an image belongs.

[0041] For example, in embodiments of the present invention, the acquired image can be input into a pre-trained deep learning semantic segmentation model. This model performs pixel-by-pixel classification of the image, labeling each pixel. For instance, the model can identify which pixels belong to cabinet doors, which to indicator lights, and which to signs, and output a pixel-level semantic label map of the same size as the original image, where each pixel value represents its semantic category.

[0042] As one possible implementation, step S103 can be specifically implemented as steps S1031-S1034.

[0043] S1031. Adaptive histogram equalization and homomorphic filtering are applied to the visual image data to suppress the interference of uneven lighting and cabinet door reflections on the recognition effect, resulting in a processed visual image.

[0044] In some embodiments, adaptive histogram equalization is a technique that enhances image contrast through local optimization, effectively improving image quality under different lighting conditions. Homomorphic filtering is a technique that simultaneously compresses the image brightness range and enhances contrast in the frequency domain, particularly suitable for handling uneven illumination and eliminating specular reflections.

[0045] For example, embodiments of the present invention can employ two image processing techniques working in tandem to improve image quality. Adaptive histogram equalization: This method divides the image into multiple small blocks and performs histogram equalization independently within each block. Compared to global equalization, it better enhances the contrast of each local area, allowing details of the switch cabinet (such as signage text) hidden in dark or bright areas to emerge, without excessively amplifying noise in the overall image. Homomorphic filtering: This technique treats the image as a product of illumination components (low-frequency, uneven lighting) and reflection components (high-frequency, texture details of the object itself). By designing a specific filter in the frequency domain, it can specifically suppress low-frequency uneven lighting and enhance high-frequency object edges and textures, thereby effectively compressing areas of concentrated highlight reflection and restoring details on the cabinet door surface obscured by reflections. The combination of these two methods outputs a processed visual image with clear details and suppressed lighting effects.

[0046] S1032. Input the processed visual image into the deep learning model and perform pixel-level segmentation to generate a segmentation result that simultaneously includes semantic category and instance identifier.

[0047] In some embodiments, semantic categories include switch cabinet doors, indicator lights, meters, and status signs.

[0048] In some embodiments, instance segmentation is a computer vision task whose goal is not only to perform pixel-level classification (semantic segmentation) but also to distinguish different individuals within the same category. Segmentation result: The output of the instance segmentation model contains the category, location (bounding box), and precise pixel contour (mask) of each individual object in the image.

[0049] For example, embodiments of the present invention may use a deep learning model based on an instance segmentation network. This network first uses its region proposal network to find candidate boxes in the image that may contain objects, and then performs two tasks in parallel: Classification: determining which semantic category (e.g., cabinet door, indicator light, etc.) the object within each candidate box belongs to; Mask generation: at the pixel level, generating a precise binary mask for each identified individual object, clearly outlining the object's contour.

[0050] S1033. Based on the predefined spatial topology rule library of switch cabinet components, perform context logic verification and repair on the segmentation results, and determine the repaired segmentation results.

[0051] In some embodiments, the spatial topology rule base is a knowledge set describing the relative positions, inclusion, adjacency, and other relationships that different objects in the environment should follow in space. The corrected segmentation result is a segmentation result whose completeness and accuracy are improved after contextual logical reasoning and error correction.

[0052] For example, embodiments of the present invention may include a built-in spatial topology rule base for switch cabinet components, which stores prior knowledge, such as: indicator lights should be located on the surface of the cabinet door, and instruments on the same cabinet door should not have significant spatial overlap, etc. The initial segmentation results are compared with these rules. If a rule violation is found (e.g., a floating indicator light is identified but its attached cabinet door is not detected, or the masks of two instrument instances overlap significantly), a correction mechanism is activated. This mechanism may include: inferring and filling in the missing cabinet door area based on the indicator light's position; or analyzing overlapping instruments and merging them into a single correct instance.

[0053] S1034. Based on the repaired segmentation results from different perspectives, the segments are fused using three-dimensional spatial geometric consistency to obtain pixel-level semantic labels.

[0054] In some embodiments, 3D spatial geometric consistency refers to the principle that the 3D reconstruction results of the same object should remain consistent, continuous, and without contradictions in spatial location when observed from different perspectives. Pixel-level semantic labels are the final, high-quality semantic segmentation output obtained after multi-view fusion optimization, providing reliable input for subsequent fusion with geometric maps.

[0055] For example, in this embodiment of the invention, a robot is controlled to photograph the same target switch cabinet at multiple adjacent points to obtain repaired segmentation results from different perspectives. Using the robot's own positioning information and camera parameters, all two-dimensional segmentation results are projected into the same three-dimensional space. In three-dimensional space, for the same physical component (such as a specific cabinet door), recognition results from different perspectives will superimpose and corroborate each other. Through three-dimensional spatial geometric consistency checks, for example, isolated pixels that cannot form a continuous surface in three-dimensional space or are contradictory from multiple perspectives are removed, and semantic information that is stable in all perspectives is fused. Finally, a unified, complete set of pixel-level semantic labels is generated for the entire scene, eliminating single-view occlusion and ambiguity.

[0056] S104. The pixel-level semantic labels and geometric maps are fused at the feature layer, and an enhanced environmental map with both geometric and semantic information is generated through coordinate transformation and data association.

[0057] In some embodiments, feature layer fusion integrates information from different sources at the data level (e.g., points, pixels), rather than through simple voting at the decision-making level. Augmented environment maps are a new type of map that simultaneously incorporates precise environmental geometry and object semantic information.

[0058] For example, embodiments of the present invention can utilize the pre-calibrated relative position and attitude relationship (extrinsic parameter matrix) between the LiDAR and the camera to accurately project pixel-level semantic labels from a two-dimensional image onto a three-dimensional LiDAR point cloud through perspective transformation. In this way, each three-dimensional point in the point cloud not only possesses spatial coordinates but is also assigned a semantic label (e.g., a point belongs to the door of switch cabinet number 1). Ultimately, all these labeled point clouds are integrated into a map to form an augmented environment map.

[0059] As one possible implementation, step S104 can be specifically implemented as steps S1041-S1045.

[0060] S1041. Based on pixel-level semantic tags, perform precise alignment in timestamps and space, and project them onto the 3D point cloud semantic map skeleton of the geometric map to obtain a geometric map with semantic tags.

[0061] In some embodiments, timestamp alignment ensures that data collected by different sensors are synchronized in time to avoid data misalignment caused by robot movement. A geometrically labeled map: the initial fusion product is a 3D point cloud where each point is associated with a semantic tag.

[0062] For example, in this embodiment of the invention, timestamp alignment is first performed to ensure that the visual image and laser point cloud data used for fusion were acquired at the same time or within a very short time interval. Next, using a pre-calibrated camera-LiDAR extrinsic parameter matrix (describing the relative position and orientation between the two sensors), each pixel-level semantic label on the two-dimensional image is precisely projected onto the corresponding three-dimensional laser point cloud using the principle of perspective projection. After projection, each three-dimensional point in the point cloud not only possesses spatial coordinate information but is also assigned a semantic label from the image (such as cabinet door number 1, indicator light), thus forming a geometric map with semantic labels.

[0063] S1042. Divide the three-dimensional space of the geometric map into multiple voxel grids.

[0064] In some embodiments, a voxel mesh is the smallest volumetric unit obtained by discretizing a three-dimensional space, analogous to a pixel in a two-dimensional image. Voxelization is the process of converting a continuous three-dimensional space into a set of discrete voxel meshes.

[0065] For example, in this embodiment of the invention, the entire space containing the 3D point cloud map is considered as a large cuboid bounding box. This bounding box is then uniformly divided along the X, Y, and Z coordinate axes at a fixed resolution (e.g., 5 cm x 5 cm x 5 cm), generating a large number of tiny, regularly arranged cubic units. This process is called voxelization, and each cubic unit is a voxel grid. All subsequent semantic information processing is performed on these voxel units, thereby transforming the irregular, sparse point cloud data into regular, dense voxel grid data, facilitating probabilistic statistics and spatial reasoning.

[0066] S1043. Based on the geometric map with semantic labels, perform Bayesian probability fusion on the semantic labels that fall within the same voxel, and calculate the probability distribution of each voxel belonging to different semantic categories.

[0067] In some embodiments, Bayesian probabilistic fusion is a probabilistic reasoning method based on Bayes' theorem that iteratively updates the belief (probability) of an event based on new evidence (observational data). Probability distribution describes the likelihood of a voxel belonging to various predefined semantic categories (such as cabinet door, indicator light, background, etc.).

[0068] For example, for each voxel grid, the semantic labels carried by all 3D points falling within it are statistically analyzed. A probability distribution is maintained for each voxel, with all categories having equal probabilities initially. When new semantic observation data (i.e., point clouds carrying labels) falls in, the probability distribution of the voxel is iteratively updated using a Bayesian update rule. For instance, if a voxel is observed multiple times to belong to the cabinet door category, its probability of belonging to the cabinet door category will increase with the number of observations. Conversely, occasional incorrect or contradictory labels (noise) will be overwhelmed by a large number of correct observations. Through this fusion, a stable and reliable semantic category probability distribution is calculated for each voxel.

[0069] S1044. Based on the probability distribution of each voxel belonging to different semantic categories, cross-validate the visual semantic information and the geometric features of the laser point cloud to determine the validated geometric map.

[0070] For example, embodiments of the present invention utilize geometric features in laser point cloud data, excluding spatial coordinates, primarily reflection intensity, to verify visual semantic probabilities. Objects of different materials exhibit varying laser reflection characteristics. For instance, metal cabinet doors typically have high reflection intensity, while insulating materials or dark-colored signs have low reflection intensity. The process checks whether the dominant semantic category (the category with the highest probability) of each voxel matches the average reflection intensity characteristics of its internal point cloud. If a significant contradiction occurs (e.g., a voxel is semantically classified as a metal cabinet door, but its average reflection intensity is very low), the semantic information at that location is deemed unreliable, and the probability value of its corresponding semantic category is lowered, or it is marked as unknown. Through this cross-validation step, the reliability of the semantic information is further improved, forming a validated geometric map.

[0071] In some embodiments, reflection intensity: the echo intensity signal received by the lidar, which is related to the material, color, and roughness of the object's surface, and can serve as an auxiliary feature to distinguish different objects. Cross-validation: using the characteristics of one data source (laser geometric features) to test and correct the results of another data source (visual semantic information).

[0072] S1045. Based on the verified geometric map, extract the isosurface of the voxel probability distribution to generate a hierarchical 3D semantic map containing occupancy probability, semantic category probability, and component instances, as an enhanced environment map.

[0073] In some embodiments, an isosurface is a surface formed by all points with the same scalar value (e.g., an occupancy probability of 0.5) in a three-dimensional scalar field. A hierarchical three-dimensional semantic map is a high-level map representation that integrates geometric information, semantic probabilistic information, and object instance information of the environment at different levels.

[0074] For example, embodiments of the present invention reconstruct a continuous, smooth object surface from a discrete voxel mesh. A Marching Cubes isosurface extraction algorithm is employed, which traverses all voxels and constructs an isosurface through a three-dimensional scalar field based on the occupancy probability (representing the likelihood that a voxel is occupied by an object) of each voxel. This isosurface represents the surface model of the object in the environment. In the final generated hierarchical three-dimensional semantic map, each surface region not only contains its geometry but also is associated with the probability of the most likely semantic category at that location and the component instance ID (obtained through associated instance segmentation results).

[0075] S105. Based on the enhanced environment map and the real-time received detection instructions of the target switch cabinet, global path planning is performed to obtain the optimal path to the target switch cabinet, thereby realizing robot navigation.

[0076] In some embodiments, global path planning involves planning a complete path from the starting point to the destination based on known global environmental information. Optimal path refers to the best solution obtained after comprehensively considering path length, safety, and task reachability.

[0077] For example, an embodiment of the present invention can receive instructions such as detecting switch cabinet number 101. In the augmented environment map, by querying semantic information, the switch cabinet entity labeled 101 is accurately located. Subsequently, a search is performed in the free area of ​​the map to calculate a collision-free optimal path from the robot's starting point to the target switch cabinet's operating position (e.g., directly in front of the cabinet door). This path consists of a series of path points, and the robot chassis control system achieves autonomous movement by tracking these path points.

[0078] As one possible implementation, step S105 can be specifically implemented as steps S1051-S1053.

[0079] S1051. Parse the target switchgear number in the detection command; and query the geometric location and operable area of ​​the target switchgear in the layered three-dimensional semantic map.

[0080] In some embodiments, a function number is a functional string that uniquely identifies the switch cabinet and is typically marked on a label on the cabinet. An operable area is the optimal working space that a robot end effector or sensor needs to reach to complete a detection task.

[0081] For example, in this embodiment of the invention, a detection command is received from an upper-level scheduling system. This command typically includes the function number of the target switchgear (e.g., KYN28-101). A semantic query is performed in a hierarchical 3D semantic map. By matching the text information associated with the component instance ID stored in the map (e.g., nameplate text obtained from visual recognition) or according to a preset installation layout order, the switchgear entity corresponding to the target number is uniquely identified. Once the target is found, its operable area is further determined based on the semantic model of the switchgear. This is typically defined as a virtual workspace frame on the front of the cabinet door, suitable for the robot to perform detection tasks.

[0082] S1052. Based on the geometric position and operable area of ​​the target switch cabinet, the robot's position data, and the multi-objective cost function, path optimization is performed to obtain the optimal movement path.

[0083] In some embodiments, the objectives of the multi-objective cost function include path length, smoothness, safe distance from energized equipment, and accessibility toward the operating surface.

[0084] In some embodiments, the multi-objective cost function is a mathematical function that quantifies and weights multiple optimization objectives (sometimes conflicting) to comprehensively evaluate the quality of a path. The optimal movement path is the path that best meets all constraints and comprehensive evaluation metrics; it is typically represented as a series of ordered spatial coordinates.

[0085] For example, this embodiment of the invention employs a sampling-based planning algorithm for path search. This algorithm intelligently and randomly samples within the free space of the enhanced environment map and progressively constructs a path tree growing from the starting point to the destination. When evaluating the merits of each potential path, the algorithm calculates based on a carefully designed multi-objective cost function. This function comprehensively considers path length (striving for the shortest), smoothness (reducing sharp turns), safe distance from electrical equipment (using semantic information to avoid high-risk areas), and accessibility to the operating surface (ensuring the robot ultimately faces the switch cabinet in the correct posture).

[0086] S1053. Control the robot to travel along the optimal movement path to the operable area of ​​the target switch cabinet.

[0087] In some embodiments, model predictive control is an advanced control strategy that solves for the optimal control input at the current moment by predicting the robot's motion state over a future period of time, thereby achieving accurate trajectory tracking. The motion chassis is the robot's mobile platform, responsible for executing movement commands, and typically includes wheels, motors, drives, and associated control units.

[0088] For example, in this embodiment of the invention, the planned path point sequence is input into the robot's underlying motion controller. The controller uses advanced tracking algorithms such as model predictive control to calculate the linear velocity and angular velocity commands sent to the robot's motion chassis, driving the wheels to move, enabling the robot to move smoothly and accurately along the preset path trajectory, and finally stop stably within the operable area of ​​the target switch cabinet, completing the global navigation task.

[0089] For example, step S1053 can be specifically implemented as steps A1-A5.

[0090] A1. During the robot's movement along the optimal path, the robot's lidar and vision sensors perceive the local environment in front of it in real time and obtain local environment data.

[0091] In some embodiments, while the robot moves, the LiDAR and vision sensors continue to operate, but the focus shifts to the local environment of a fan-shaped area in front of the robot. The LiDAR provides precise distances and outlines of obstacles ahead, while the vision sensors provide color and texture information, together forming local environmental data used to determine the passage conditions ahead. This local environmental data—the real-time spatial information perceived by the robot along its direction of movement—is primarily used for dynamic obstacle avoidance.

[0092] A2. Based on local environmental data, dynamic and static obstacles are identified in real time through clustering algorithms, and temporary dynamic obstacles are identified by combining with an enhanced environmental map.

[0093] For example, embodiments of the present invention can perform cluster analysis on real-time acquired local laser point clouds to identify individual objects. Subsequently, the real-time locations of these objects are compared with prior static objects (such as switch cabinets and walls) in the augmented environment map. Any object that does not exist in the map or whose location does not match the map will be identified as a temporary dynamic obstacle (such as a worker temporarily passing by or a moving tool vehicle).

[0094] In some embodiments, temporary dynamic obstacles are objects whose positions change over time and are not pre-labeled on a global map. Identification is the process of distinguishing and recognizing specific targets by comparing real-time observations with prior knowledge.

[0095] A3. Determine obstacle avoidance strategies based on temporary dynamic obstacles and a pre-defined multi-obstacle semantic classifier.

[0096] In some embodiments, the obstacle avoidance strategy includes reducing the safety distance weight in the multi-objective cost function if the obstacle is classified as a flexible obstacle that can be temporarily traversed, and triggering local path replanning while maintaining the safety distance weight if the obstacle is classified as a rigid obstacle that must be avoided.

[0097] In some embodiments, a multi-obstacle semantic classifier is an intelligent model capable of identifying the specific type of obstacle (such as people, vehicles, cables) based on perception data. An obstacle avoidance strategy is a set of differentiated navigation behavior rules adopted for obstacles of different natures.

[0098] For example, in this embodiment of the invention, the identified obstacle information (such as point cloud shape and visual appearance) is input into a multi-obstacle semantic classifier (a pre-trained lightweight deep learning model). This classifier categorizes obstacles based on their physical characteristics, for example: flexible obstacles that can be temporarily traversed, such as temporary cables hanging on the ground. The obstacle avoidance strategy adopted is tolerant traversal, that is, temporarily reducing the safety distance weight during planning, allowing the robot to cautiously approach or cross it. Rigid obstacles that must be strictly avoided, such as people and precision instruments, are addressed by triggering local path replanning and maintaining or increasing the safety distance weight in the new plan to ensure absolute safety.

[0099] A4. If the obstacle avoidance strategy is local path replanning, then based on the updated multi-objective cost function, multiple local trajectories are generated, and the trajectory with the lowest cost value among the multiple local trajectories is selected as the detour path.

[0100] For example, when replanning is required, this embodiment of the invention employs a dynamic window method between the robot's current position and a path point ahead. This method, considering the robot's current speed and dynamic constraints, samples and generates multiple local trajectories (i.e., feasible future motion paths within a short time) in the velocity space. Each trajectory undergoes cost calculation based on an updated multi-objective cost function (e.g., maintaining a high safety weight for rigid obstacles). Finally, the trajectory with the lowest cost value is selected as the currently executed detour path.

[0101] In some embodiments, the dynamic window method is a local path planner that considers robot dynamics constraints and selects the optimal short-term motion command by sampling in the velocity space and simulating the trajectory. A detour path is a locally alternative path replanned to avoid temporary dynamic obstacles.

[0102] A5. Control the robot to move along the bypass path and continuously evaluate the robot's relative position to the optimal movement path; when it is confirmed that the temporary dynamic obstacle has been removed or the robot has bypassed it, and the safety conditions are met, smoothly guide the robot back to the optimal movement path until it reaches the operable area of ​​the target switch cabinet.

[0103] For example, the robot executes a detour path. Simultaneously, the lateral and longitudinal distances between the robot and the original optimal path are continuously evaluated. Once the temporary obstacle has been detected as having disappeared (confirmed by sensors) and the robot is in a position where it can safely return to the original path, the motion controller generates a guide trajectory that smoothly transitions from the current position to the original path. The robot moves along this transition trajectory, eventually seamlessly returning to the initial globally optimal path and continuing towards the target until the task is completed.

[0104] In some embodiments, evaluation refers to the process of continuously monitoring and determining whether the system state meets specific conditions (such as obstacles being cleared or safe entry conditions being met). Recovery refers to the robot's behavior of returning to and tracking the global reference path after completing local obstacle avoidance.

[0105] This invention provides a robot navigation method for switchgear inspection based on the fusion of LiDAR and vision. By introducing visual semantic tags, the robot gains the ability to understand scene functions and uniquely identify different target switchgear with similar appearances. This achieves a leap from geometric localization to semantic recognition, ensuring the accuracy of navigation targets. The generated enhanced environment map provides a unified information source with both geometric and semantic constraints for path planning, enabling the planned path to not only safely avoid obstacles but also actively guide to the operable area of ​​the target switchgear, improving the intelligence and task relevance of the navigation path. Through tightly coupled multi-sensor information fusion, this invention effectively improves the overall robustness and reliability of the robot in complex lighting environments in substations, solving the problem of inaccurate navigation and positioning in current robot navigation solutions and improving the accuracy of robot navigation and positioning.

[0106] Optionally, the switch cabinet detection robot navigation method based on lidar and vision fusion provided in this embodiment of the invention further includes steps S201-S205 after step S105.

[0107] S201. When the robot moves to the preset adjacent area of ​​the target switch cabinet according to the optimal path, the vision servo controller is activated, and the vision sensor continuously collects real-time images of the target switch cabinet.

[0108] For example, when the robot determines through laser positioning that it has entered a preset adjacent area within a certain range (e.g., within 2 meters) in front of the target switch cabinet, it automatically switches from global navigation mode to visual servo fine-tuning mode. The vision sensor begins to continuously capture images of the target switch cabinet at a higher frequency.

[0109] The preset proximity area is a transitional zone defined around the target switchgear, used to trigger a mode switch from coarse positioning to fine-tuning positioning. The visual servo controller is a control system that uses visual feedback information to control the robot's movement in real time.

[0110] S202. Extract at least one visual feature from the real-time image.

[0111] In some embodiments, stable and easily trackable visual features are extracted from each frame of real-time image. These features may be inherent, high-contrast parts of the switch cabinet, such as the four corners of a sign, the edges of a cabinet door handle, or specific feature points learned through training. Visual features: Local structures in an image that have distinct characteristics (such as corners, edges, spots) and can be continuously tracked by a computer.

[0112] S203. Compare the visual features with the expected state of the visual features under the preset accurate detection pose to determine the six-degree-of-freedom pose error between the robot's current pose and the accurate detection pose.

[0113] In some embodiments, during system initialization, the ideal position (desired state) of the target visual features in the image is predefined and stored when the robot is in a precise detection pose (i.e., the optimal position and posture where the robotic arm or sensor can perfectly perform the detection task). The feature positions extracted in real time are compared with this template, and their pixel deviations in the image coordinate system are calculated. Then, the six-degree-of-freedom pose error of the robot in three-dimensional space relative to the ideal pose (i.e., translation errors in the forward, backward, left, right, and up / down directions, and rotation errors in the pitch, yaw, and roll directions) is calculated using the camera model.

[0114] For example, precise detection pose: the optimal position and orientation of the robot preset to complete the detection task. Six-DOF pose error: a quantity describing the difference between the robot's current pose and the target pose in six directions in three-dimensional space.

[0115] S204. Based on the six-degree-of-freedom pose error and combined with the image Jacobian matrix, the linear velocity and angular velocity control commands of the robot are determined through the algorithm rules of visual servo control.

[0116] In some embodiments, the image Jacobian matrix describes the mathematical relationship between changes in image features and robot motion. The visual servo control law (typically a proportional-derivative controller) calculates in real-time the linear velocity (forward / backward speed) and angular velocity (rotational speed) commands that the robot chassis should generate to eliminate the pose error, based on the calculated pose error and the Jacobian matrix. Image Jacobian matrix: A mathematical matrix that maps robot motion speed to the rate of change of image features; it is the core of visual servo control. Visual servo control law: An algorithmic rule for calculating robot motion control quantities based on visual errors.

[0117] In some embodiments, the algorithm rules for visual servo control are the core control strategy for driving robot motion based on image feature errors. The core of this strategy is to establish a mathematical relationship between image pixel changes and robot motion speed using the image Jacobian matrix. The basic rule employs a proportional control method, directly calculating the linear and angular velocity commands required to eliminate the error by mapping the deviation (feature error) between the real-time extracted feature positions and the desired positions in the image through the inverse of the Jacobian matrix. For switchgear fine-tuning scenarios, this rule incorporates depth prior information to ensure computational stability and performs weighted processing on feature points to improve anti-interference capabilities. Finally, through closed-loop iteration, the image error converges to zero, thereby achieving millimeter-level precise pose alignment.

[0118] S205. Send control commands to the robot's motion chassis and / or actuator to drive the robot to move. During the movement, repeatedly collect real-time images of the target switch cabinet and generate control commands until the magnitude of the six-degree-of-freedom pose error is less than a preset threshold, so that the robot can achieve accurate pose detection.

[0119] In some embodiments, the calculated speed command can be sent to the robot's underlying motion controller. The robot begins micro-movements, while the vision system continuously acquires new images, extracts features, calculates new errors, and generates new control commands, forming a closed-loop feedback control. This process is repeated continuously, and the robot's pose is continuously corrected until the overall error between it and the target pose (i.e., the magnitude of the error) is less than a preset small threshold (e.g., a few millimeters or degrees). At this point, it is determined that the robot has reached the accurate detection pose, and the fine-tuning process ends. Closed-loop feedback control: A control method that adjusts the system behavior (motion) by continuously measuring the output (current pose) and comparing it with the expected value (target pose). Magnitude of error: A scalar measure of the overall magnitude of the six-degree-of-freedom pose error.

[0120] Thus, by introducing closed-loop feedback control based on image features, this invention achieves millimeter-level real-time fine-tuning of the robot's pose, ensuring that its detection unit can ultimately and accurately align with the operating part of the target switch cabinet. This improves navigation accuracy from the macroscopic path level to the operational level that meets the requirements of the detection task, greatly enhancing the success rate and reliability of automated detection.

[0121] Optionally, the switch cabinet inspection robot navigation method based on lidar and vision fusion provided in this embodiment of the invention further includes steps S301-S305.

[0122] S301. During robot navigation, calculate the Euclidean distance from the robot's current position to the target operable area in real time.

[0123] In some embodiments, Euclidean distance, the straight-line distance between two points in three-dimensional space, is the most intuitive way to measure distance.

[0124] For example, in this embodiment of the invention, during the entire process of the robot's autonomous movement, the navigation system performs the following calculations cyclically at a high frequency (e.g., 10Hz): It obtains the robot's current precise three-dimensional coordinates through its own positioning module, and simultaneously obtains the three-dimensional coordinates of the center point of the operable area of ​​the target switch cabinet from the task information. Then, based on the Euclidean distance formula between the two points, it directly calculates the straight-line distance between the current location and the target point. This calculated Euclidean distance is the core basis for subsequent adaptive decision-making.

[0125] S302. Based on the Euclidean distance and the piecewise linear function between the weights of each objective and the Euclidean distance in the multi-objective cost function, the dynamic weights of each objective are obtained.

[0126] In some embodiments, a piecewise linear function is a function composed of multiple line segments, each of which is linear within its domain. It is often used to describe relationships where different patterns of change exist across different intervals. Dynamic weights are weight coefficients that change in real-time according to specific rules during operation, used to adjust the importance of different optimization objectives.

[0127] For example, this embodiment of the invention pre-defines an adaptive weight model, which defines a weight-distance piecewise linear function for each sub-objective (path length, smoothness, safe distance, reachability) in the cost function. These functions describe how the weight of the sub-objective should change with the Euclidean distance. The real-time calculated distance values ​​are input into these functions, and through simple table lookup and interpolation calculations, the dynamic weight value corresponding to each sub-objective at the current time is dynamically calculated. For example, when the distance is far, the output value of the path length weight function is higher; when the distance is short, the output value of the reachability weight function increases sharply.

[0128] S303. Update the multi-objective cost function based on the dynamic weights of each objective.

[0129] For example, in this embodiment of the invention, dynamic weight values ​​are assigned to the corresponding variables of the multi-objective cost function in the path planner, overriding the previous weight settings. This means that the inherent evaluation criteria of the cost function have been updated, and its evaluation criteria for paths will immediately reflect the focus of the current navigation stage (determined by distance). For example, after the update, the cost function will become more "preferential" to choose the path that allows the robot to better align with the cabinet door, rather than the absolute shortest path.

[0130] S304. Based on the updated multi-objective cost function, the robot performs local replanning of the path ahead during its movement and controls the robot to move along the replanned path.

[0131] In some embodiments, exemplary, after the weights are updated, the present invention immediately triggers a local replanning. The planner (typically using the same algorithm as the global planner, but within a limited local window) utilizes the updated multi-objective cost function to replan a new segment path from the robot's current position to a look-ahead point on the original path. This new path incorporates the latest navigation strategy for the current stage. Subsequently, the robot's motion controller immediately begins tracking this replanned path, thereby adaptively matching its movement behavior to the task stage (distance from the target). The entire process is repeated cyclically until the robot reaches the target, ensuring that its behavior is always optimal.

[0132] Thus, this invention enables the robot to prioritize movement efficiency when far from the target, and automatically switch to a navigation strategy that emphasizes safety and operational precision when approaching the target. This achieves a smooth transition from extensive movement to refined approach, effectively balancing the contradiction between efficiency and precision, and improving overall navigation performance.

[0133] Optionally, the switch cabinet inspection robot navigation method based on lidar and vision fusion provided in this embodiment of the invention further includes steps S401-S404.

[0134] S401. Based on the detection instructions of the target switchgear, analyze the task objective and decide on the passage mode.

[0135] If the detection command is to perform a fixed-point inspection on a single switchgear, the decision is to use the normal passage mode; if the detection command is to perform a patrol inspection on a group of consecutive switchgear, the decision is to use the wall-mounted detection mode.

[0136] For example, this embodiment of the invention incorporates a task semantic parser to interpret received detection instructions. The parser identifies key information in the instructions, such as the number of target switchgear cabinets and the task type. If the instruction explicitly requires fixed-point detection of a specific switchgear cabinet (e.g., cabinet 101), the parser determines this is an independent task and decides to adopt the normal passage mode. In this mode, the robot's primary goal is to move safely and efficiently along the main aisle to the target location. Conversely, if the instruction requires continuous inspection of a group of spatially adjacent switchgear cabinets (e.g., cabinets 101 to 105), the parser determines this is a sequential task and decides to activate the wall-hugging detection mode. In this mode, the robot's core task is to generate a path that continuously approaches each target cabinet to optimize the detection process.

[0137] In some embodiments, a task semantic parser is a software module used to understand and parse natural language or structured task instructions and extract key task parameters. Continuous inspection is a work mode that requires the robot to perform uninterrupted continuous inspection of multiple devices in sequence to improve work efficiency.

[0138] S402. Based on the traffic pattern, dynamically reconstruct the multi-objective cost function.

[0139] Specifically, when the decision is to use the wall-hugging detection mode, the weight coefficient of the sub-objective that maintains the minimum safe distance from the cabinet in the multi-objective cost function is set to a negative value, and the weight coefficient of the accessibility towards the operating surface is set to the first value, guiding the robot to move close to the cabinet; when the decision is to use the normal passage mode, the weight coefficient of the sub-objective that maintains the minimum safe distance from the cabinet in the multi-objective cost function is set to a positive value, and the weight coefficient of the accessibility towards the operating surface is set to the second value, with the first value being greater than the second value.

[0140] For example, in this embodiment of the invention, the value orientation of path planning—that is, the multi-objective cost function—is dynamically reconstructed based on the determined traffic pattern. In the wall-hugging detection mode, to achieve close-to-the-cabinet patrolling, the cost function is inverted: the weight coefficient of the sub-objective of maintaining the minimum safe distance from the cabinet is set to a negative value. This means that in the planning algorithm, the closer the path is to the cabinet (i.e., violating the traditional minimum safe distance constraint), the lower its total cost, thereby guiding the algorithm to actively generate paths close to the cabinet. Simultaneously, to ensure that the robot faces each cabinet surface correctly, the weight of accessibility towards the operating surface is increased to a very high first value.

[0141] In normal passage mode, a standard strategy is adopted: the weight of maintaining the minimum safe distance from the cabinet is positive to ensure a safe interval; the weight of accessibility toward the operating surface is set to a relatively low second value, because rapid arrival is far more important than end posture at this time.

[0142] In some embodiments, cost function inversion: a special weighting strategy that, by setting the weights of specific constraints to negative values, shifts the optimization objective from avoiding constraint violations to actively seeking constraint violations to achieve specific behavioral goals. Guidance: adjusting the algorithm's internal evaluation criteria to make its output tend to satisfy preset behavioral characteristics.

[0143] S403. Based on the reconstructed multi-objective cost function, a globally optimal path is generated by planning in the enhanced environment map, continuously covering the target switchgear operation surface.

[0144] For example, this embodiment of the invention uses a bidirectional A* algorithm for global path planning in an enhanced environment map. This algorithm starts searching simultaneously from both the starting point and the target point, resulting in high efficiency. Since the cost function has been reconstructed, the algorithm is naturally drawn towards the cabinets during the search process, ensuring that the path continuously passes directly in front of the operating surfaces of all target switch cabinets, ultimately outputting a globally optimal path that meanders like a snake in front of the cabinets. This path eliminates the need for the robot to repeatedly perform approach-away actions in front of each cabinet, thus streamlining the detection process.

[0145] In some embodiments, the bidirectional A* algorithm is a graph search algorithm that improves search efficiency by simultaneously searching from both the start and end points and meeting in the middle. The globally optimal path is the best path evaluated across the entire map according to the current cost function criterion.

[0146] S404. The planned global optimal path is smoothed using curve interpolation, and the robot is controlled to move along the smoothed path.

[0147] For example, the paths directly generated by the search algorithm in this embodiment of the invention are typically polylines composed of a series of discrete points, containing sharp corners, and are unsuitable for direct robot tracking. Therefore, B-spline curve interpolation is used to smooth the original path. This method can generate a smooth curve that passes through or approximates the original path points and has continuous curvature. This smoothed path ensures the stability of the robot's movement, avoiding sharp turns and jitter, which is crucial for ensuring stable acquisition of sensor data during the detection process. Finally, the robot motion controller tracks this smooth path to complete the entire navigation task.

[0148] In some embodiments, B-spline interpolation is a mathematical method for constructing smooth curves, widely used in robot path smoothing. The smoothed path is a smooth trajectory with continuously changing curvature obtained after mathematical optimization, highly suitable for actual robot motion tracking.

[0149] Thus, in the wall-hugging detection mode, the present invention sets the safety distance weight to a negative value, enabling the robot to actively approach the cabinet to generate a continuous detection path, realizing the leap from discrete fixed-point operation to continuous streamlined operation, and effectively solving the problem of low detection efficiency of traditional navigation methods in dense equipment areas.

[0150] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0151] The following are device embodiments of the present invention. For details not described in detail, please refer to the corresponding method embodiments described above.

[0152] Figure 2 A schematic diagram of a switch cabinet inspection robot navigation device based on lidar and vision fusion is shown in an embodiment of the present invention. The navigation device 500 includes a communication module 501 and a processing module 502.

[0153] The communication module 501 is used to acquire laser point cloud data collected by the lidar sensor on the robot and visual image data collected by the vision sensor.

[0154] The processing module 502 is used to perform synchronous localization and map construction based on laser point cloud data to generate a geometric map; perform visual feature recognition and semantic segmentation based on visual image data to determine pixel-level semantic labels; fuse the pixel-level semantic labels and geometric map at the feature layer, and generate an enhanced environment map that combines geometric and semantic information through coordinate transformation and data association; and perform global path planning based on the enhanced environment map and the detection instructions received in real time from the target switch cabinet to obtain the optimal path to the target switch cabinet, thereby realizing robot navigation.

[0155] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 600 includes: a processor 601, a memory 602, and a computer program 603 stored in the memory 602 and executable on the processor 601. When the processor 601 executes the computer program 603, it implements the steps in the above-described method embodiments. Alternatively, when the processor 601 executes the computer program 603, it implements the functions of each module / unit in the above-described device embodiments.

[0156] For example, the computer program 603 may be divided into one or more modules / units, which are stored in the memory 602 and executed by the processor 601 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 603 in the electronic device 600.

[0157] The processor 601 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0158] The memory 602 can be an internal storage unit of the electronic device 600, such as a hard disk or memory of the electronic device 600. The memory 602 can also be an external storage device of the electronic device 600, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD) card, flash card, etc., equipped on the electronic device 600. Furthermore, the memory 602 can include both internal and external storage units of the electronic device 600. The memory 602 is used to store the computer program and other programs and data required by the terminal. The memory 602 can also be used to temporarily store data that has been output or will be output.

[0159] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A navigation method for a switchgear inspection robot based on the fusion of lidar and vision, characterized in that, include: Acquire laser point cloud data collected by the lidar sensor on the robot, and visual image data collected by the vision sensor; Based on laser point cloud data, synchronous positioning and map building are performed to generate a geometric map; Based on visual image data, visual feature recognition and semantic segmentation are performed to determine pixel-level semantic labels; The pixel-level semantic labels and the geometric map are fused at the feature layer, and an enhanced environmental map with both geometric and semantic information is generated by coordinate transformation and data association. Based on the enhanced environment map and the real-time received detection commands from the target switch cabinet, global path planning is performed to obtain the optimal path to the target switch cabinet, thereby realizing robot navigation.

2. The navigation method for switch cabinet inspection robot based on lidar and vision fusion according to claim 1, characterized in that, After the robot navigation is achieved by performing global path planning based on the enhanced environment map and the real-time received detection commands from the target switchgear to obtain the optimal path to the target switchgear, the following steps are also included: When the robot moves to the preset adjacent area of ​​the target switch cabinet according to the optimal path, the visual servo controller is activated, and the visual sensor continuously collects real-time images of the target switch cabinet. Extract at least one visual feature from the real-time image; The visual features are compared with the expected state of the visual features under the preset accurate detection pose to determine the six-degree-of-freedom pose error between the robot's current pose and the accurate detection pose. Based on the six-degree-of-freedom pose error and combined with the image Jacobian matrix, the linear velocity and angular velocity control commands of the robot are determined through the algorithm rules of visual servo control. The control commands are sent to the robot's motion chassis and / or actuators to drive the robot to move. During the movement, real-time images of the target switch cabinet are repeatedly acquired and control commands are generated until the magnitude of the six-degree-of-freedom pose error is less than a preset threshold, so that the robot can achieve the precise detection pose.

3. The navigation method for switch cabinet inspection robot based on lidar and vision fusion according to claim 1, characterized in that, The process of synchronously locating and building a geometric map based on laser point cloud data includes: Cluster analysis is performed on laser point cloud data to identify and filter out dynamic point clouds generated by moving people or equipment in the environment, and to generate static environmental point cloud data. Multi-scale geometric features are extracted from the static environmental point cloud data. The multi-scale geometric features include global point cloud features for loop closure detection and local line and surface features for inter-frame matching. The local line and surface features in the multi-scale geometric features are tightly coupled and optimized with the pre-integrated data of the inertial measurement unit to output the robot pose estimate; Based on the robot pose estimation and global point cloud features, a globally consistent 3D point cloud semantic map skeleton is constructed using a pose graph optimization method to obtain the geometric map.

4. The navigation method for switch cabinet inspection robot based on lidar and vision fusion according to claim 1, characterized in that, The process of performing visual feature recognition and semantic segmentation based on visual image data to determine pixel-level semantic labels includes: Adaptive histogram equalization and homomorphic filtering are performed on the visual image data to suppress the interference of uneven lighting and cabinet door reflection on the recognition effect, resulting in a processed visual image. The processed visual image is input into a deep learning model for pixel-level segmentation to generate a segmentation result that includes both semantic categories and instance identifiers. The semantic categories include cabinet doors, indicator lights, meters, and status signs of the switch cabinet. Based on a predefined spatial topology rule library for switchgear components, the segmentation results are subjected to contextual logic verification and repair to determine the repaired segmentation results. Based on the repaired segmentation results from different perspectives, pixel-level semantic labels are obtained by fusing them through three-dimensional spatial geometric consistency.

5. The navigation method for switch cabinet inspection robot based on lidar and vision fusion according to claim 1, characterized in that, The step of fusing the pixel-level semantic labels and the geometric map at the feature layer, and generating an enhanced environment map that combines geometric and semantic information through coordinate transformation and data association, includes: Based on the pixel-level semantic tags, precise alignment is performed in terms of timestamps and space, and the data is projected onto the 3D point cloud semantic map skeleton of the geometric map to obtain a geometric map with semantic tags. The three-dimensional space of the geometric map is divided into multiple voxel grids; Based on the geometric map with semantic labels, Bayesian probability fusion is performed on the semantic labels falling within the same voxel to calculate the probability distribution of each voxel belonging to different semantic categories. Based on the probability distribution of each voxel belonging to different semantic categories, the visual semantic information and the geometric features of the laser point cloud are cross-validated to determine the validated geometric map. Based on the verified geometric map, isosurfaces of voxel probability distribution are extracted to generate a layered 3D semantic map containing occupancy probability, semantic category probability, and component instances, which serves as an enhanced environment map.

6. The navigation method for switch cabinet inspection robot based on lidar and vision fusion according to claim 1, characterized in that, The process of performing global path planning based on the enhanced environment map and the real-time received detection commands from the target switchgear to obtain the optimal path to the target switchgear and achieve robot navigation includes: The target switchgear number in the detection command is parsed; and the geometric location and operable area of ​​the target switchgear are obtained from the layered three-dimensional semantic map. Based on the geometric position and operable area of ​​the target switch cabinet, the robot's position data, and a multi-objective cost function, path optimization is performed to obtain the optimal movement path. The objectives of the multi-objective cost function include path length, smoothness, safe distance from energized equipment, and accessibility toward the operating surface. Control the robot to travel along the optimal path to the operable area of ​​the target switch cabinet.

7. The navigation method for switch cabinet inspection robot based on lidar and vision fusion according to claim 6, characterized in that, The controlled robot travels along the optimal movement path to the operable area of ​​the target switchgear, including: As the robot travels along the optimal movement path, it uses the robot's lidar and vision sensors to perceive the local environment in front of it in real time and obtain local environment data. Based on the local environmental data, dynamic and static obstacles are identified in real time through clustering algorithms, and temporary dynamic obstacles are identified in combination with the enhanced environmental map. Based on the temporary dynamic obstacles and the preset multi-obstacle semantic classifier, an obstacle avoidance strategy is determined. The obstacle avoidance strategy includes reducing the safety distance weight in the multi-objective cost function if the obstacle is classified as a flexible obstacle that can be temporarily traversed; and triggering local path replanning while maintaining the safety distance weight if the obstacle is classified as a rigid obstacle that must be strictly avoided. If the obstacle avoidance strategy is local path replanning, then based on the updated multi-objective cost function, multiple local trajectories are generated, and the trajectory with the lowest cost value among the multiple local trajectories is selected as the detour path. Control the robot to move along the detour path and continuously evaluate the robot's relative position to the optimal movement path; when it is confirmed that the temporary dynamic obstacle has been removed or the robot has been bypassed and the safety conditions are met, smoothly guide the robot back to the optimal movement path until it reaches the operable area of ​​the target switch cabinet.

8. The navigation method for a switch cabinet inspection robot based on lidar and vision fusion according to any one of claims 1 to 7, characterized in that, The method further includes: During robot navigation, the Euclidean distance from the robot's current position to the target operable area is calculated in real time. Based on the Euclidean distance and the piecewise linear function between the weights of each objective and the Euclidean distance in the multi-objective cost function, the dynamic weights of each objective are obtained. The multi-objective cost function is updated based on the dynamic weights of each objective; Based on the updated multi-objective cost function, during the robot's movement, the path ahead is locally replanned, and the robot is controlled to move along the replanned path.

9. The navigation method for a switch cabinet inspection robot based on lidar and vision fusion according to any one of claims 1 to 7, characterized in that, The method further includes: Based on the detection instructions for the target switchgear, the task objectives are analyzed, and the passage mode is determined. Specifically, if the detection instruction is to perform fixed-point detection on a single switchgear, the normal passage mode is determined; if the detection instruction is to perform inspection on a group of consecutive switchgear, the wall-mounted detection mode is determined. Based on the aforementioned passage mode, the multi-objective cost function is dynamically reconstructed. Specifically, when the decision is a wall-hugging detection mode, the weight coefficient of the sub-objective maintaining the minimum safe distance from the cabinet in the multi-objective cost function is set to a negative value, and the weight coefficient of accessibility towards the operating surface is set to a first value, guiding the robot to move close to the cabinet. When the decision is a normal passage mode, the weight coefficient of the sub-objective maintaining the minimum safe distance from the cabinet in the multi-objective cost function is set to a positive value, and the weight coefficient of accessibility towards the operating surface is set to a second value, where the first value is greater than the second value. Based on the reconstructed multi-objective cost function, a globally optimal path is generated by planning in the enhanced environment map, continuously covering the target switch cabinet operation surface; The globally optimal path is smoothed using curve interpolation, and the robot is controlled to move along the smoothed path.

10. A live-line detection system for switchgear based on a humanoid robot, characterized in that, The system includes an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor is configured to invoke and run the computer program stored in the memory to perform the method as described in any one of claims 1 to 9.