Composite inspection robot global positioning method based on mechanical arm collaborative vision
By constructing multi-layer reference maps and visual information matching with multi-sensor fusion, the global positioning problem in highly similar environments such as data centers is solved, and the stable and efficient positioning of the robot in complex environments is achieved.
Patent Information
- Application Number
- CN202510378104.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-22
AI Technical Summary
In high-similar environments such as data centers, the success rate of global positioning is low, the distinction between lidar characteristics is low, and the matching error is easily generated. Pure visual positioning is easily restricted by light changes and viewing angles, resulting in low stability of the robot.
The global positioning method of composite patrol robot based on robotic arm collaborative vision is adopted. By constructing a multi-layer reference map with multi-sensor fusion, including a two-dimensional occupancy raster map, label topology map and visual point cloud map, and performing joint optimization and data post-processing of multi-map, combining multi-resolution map matching of lidar scanning point clouds and visual information matching with active observation angle adjustment.
It improves the robot's global positioning search ability in high-similar environments, enhances the adaptability of multi-scene environments, and solves the limitations of sensor perspectives and the uncertainty of rotation process in narrow environments.
Smart Images

Figure CN120347730A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot positioning, and specifically to a global positioning method for a composite inspection robot based on the collaboration of a robotic arm and vision. Background Art
[0002] With the rapid development of artificial intelligence, machine learning, and sensor technologies, robot technology has made remarkable progress in various fields such as industry, healthcare, agriculture, and services. Especially in the fields of industrial automation and intelligent manufacturing, robots have been widely used in tasks such as automated production, material handling, inspection, and patrol. As a highly integrated intelligent inspection device of artificial intelligence technology, the inspection robot for data centers is commonly used for intelligent operation and maintenance of computer rooms, and can efficiently and accurately detect abnormal conditions in the computer room environment and its own abnormal states. Global reliable positioning is the foundation for a robot to perform complex inspection tasks, directly affecting the execution of robot autonomous navigation and path planning. In common data center environments, they are mostly composed of rows of cabinets arranged, and their projection planes are mostly highly similar and narrow rectangular frames, with fewer effective unique features in the planar environment, resulting in a low success rate of global positioning and often requiring human intervention. Therefore, breaking through the global positioning and robot kidnapping problems in this environment is the key to improving the high robustness and intelligence of robots.
[0003] In the prior art, global positioning methods mostly rely on lidar or cameras for feature matching and recognition to obtain the position of the robot in the global coordinate system. However, lidar has low feature discrimination in highly similar environments and is prone to matching errors. Common solutions are to rotate and splice features to enhance features. However, the computer room environment is relatively narrow, not only unable to offset the mirror symmetry error, but also there are safety hazards during the rotation process. Pure vision positioning is easily affected by changes in lighting and viewing angles, resulting in low stability of the robot in complex and similar environments. Summary of the Invention
[0004] In view of the requirements and deficiencies in the current technological development, the present invention provides a global positioning method for a composite inspection robot based on the collaboration of a robotic arm and vision. By combining lidar, a robotic arm, and a binocular camera for positioning, it solves the limitations of the sensor viewing angle and the limitations of pure lidar in highly similar environments such as data centers, as well as the uncertainty during the rotation process in narrow environments.
[0005] The technical solution adopted by the global positioning method for a composite inspection robot based on the collaboration of a robotic arm and vision of the present invention to solve the above technical problems is as follows:
[0006] A global positioning method for a composite inspection robot based on the collaboration of a robotic arm and vision includes the following steps:
[0007] Step 1: Construct a multi-layer reference map based on multi-sensor fusion, which includes constructing a two-dimensional occupancy grid map, a label topology map, and a visual point cloud map, followed by multi-map joint optimization and data post-processing;
[0008] Step 2: Based on the output data of step 1, a global positioning search is performed, which specifically includes multi-resolution map matching based on the lidar scanning point cloud and visual information matching based on active observation angle adjustment.
[0009] Optionally, executing step 1, the specific operations of constructing a two-dimensional occupancy grid map include:
[0010] Step 1.1.1. The robot is controlled remotely to collect data within the operating range of the data center computer room, and preparations are made for data collection to ensure that the laser radar, odometer, and IMU sensors on the robot are operating normally;
[0011] Step 1.1.2: Use the laser radar to scan the robot's surrounding environment to obtain the distance information of the environment; use the odometer and IMU to obtain the robot's motion information, and combine the laser scanning matching results to determine the robot's position in the environment;
[0012] Step 1.1.3: Use the ray projection algorithm to map the environmental data within a radius of 5m with the robot as the center to a two-dimensional plane grid to complete the construction of a 5cm grid map, which is used for subsequent precise posture iteration of the robot;
[0013] Step 1.1.4: Use sliding windows of 2×2, 4×4, 8×8, and 16×16 to pool the 5 cm grid map, and generate coarse resolution grid maps of 10 cm, 20 cm, 40 cm, and 80 cm respectively and save them locally; in this process, the robot uses a binocular camera to collect image data in the computer room environment, and uses the ORB algorithm to extract features from the image data to obtain the ORB feature descriptor, and accurately records the robot's posture information at the moment of collecting each frame of image data, associates the ORB feature descriptor with the corresponding posture information and stores it, and constructs a local "ORB feature-pose" dictionary; the "ORB feature-pose" dictionary is loaded when the program is initialized to improve the efficiency and accuracy of lidar matching, and complete the construction of grid maps of different resolutions and the generation of the "ORB feature-pose" dictionary.
[0014] Optionally, step 1 is performed to construct a label topology map, and specific operations include:
[0015] Step 1.2.1: Each cabinet in the data center computer room environment has a UUID label, and the UUID label of the cabinet reflects the specific arrangement number of the cabinet in the computer room. During the movement of the robot, the binocular camera installed on the end effector is used to identify the UUID label of the cabinet, and the position of the channel where it is located can be estimated.
[0016] Step 1.2.2: Raise the robotic arm of the composite inspection robot by 50 cm so that the end camera coordinate system is on the central axis of the robot chassis and 1.75 m away from the chassis base_link coordinate system. At the same time, ensure that the binocular camera's field of view is consistent with the direction of the robot's vehicle head.
[0017] Step 1.2.3: Manually remote-control the robot to make it walk centrally to the first group of cabinets in the cold aisle of the computer room. After reaching the position of the first group of cabinets, create a semantic node by clicking on the screen. This semantic node independently records the serial number and records it in a negative form to distinguish it from the nodes independently created during the construction of the grid map.
[0018] Step 1.2.4: Continue to remote-control the robot to make it walk centrally to the last group of cabinets in the same cold aisle. When reaching the position of the last group of cabinets, perform the recording operation again.
[0019] Step 1.2.5: For all aisles in the computer room, repeat the operations in Step 1.2.3 and Step 1.2.4. After the data collection and semantic node creation operations for all aisles, integrate and process the recorded node information and the associated position and environmental data to construct a complete label topology map.
[0020] Further optionally, the specific operations for performing Step 1 to construct the visual point cloud map include:
[0021] Step 1.3.1: Set the position of the end effector of the composite inspection robot to be the same as when constructing the label topology map, that is, raise the robotic arm by 50 cm, the end camera coordinate system is on the central axis of the robot chassis and 1.75 m away from the chassis base_link coordinate system, and the camera's field of view is consistent with the direction of the robot's vehicle head.
[0022] Step 1.3.2: When independently creating nodes during the construction of the 2D occupancy grid map, the binocular camera collects the dense point cloud data of the current environment, and extracts ORB (Oriented FAST and Rotated BRIEF) feature points from the collected dense point cloud data.
[0023] Step 1.3.3: Bind the extracted ORB feature points to the currently autonomously created nodes. By comparing the ORB feature points extracted from the image with the existing feature points in the local "ORB feature - pose" dictionary, find the most similar feature points according to the descriptor matching degree, and use the pose information associated with these feature points in the local "ORB feature - pose" dictionary as the reference pose information for the new nodes, so that the pose of the current nodes is consistent with the environmental descriptors (i.e., the descriptive information represented by the ORB feature points), thereby establishing the association between the nodes and the environmental features;
[0024] Step 1.3.4: Construct a visual point cloud map by stitching the ORB descriptors corresponding to multiple nodes.
[0025] Further optionally, when performing Step 1, the specific operations for multi - map joint optimization include:
[0026] Step 1.4.1: Construct and optimize a multi - layer reference map based on the graph optimization method to ensure the synchronous optimization of the nodes in the construction of the 2D occupancy grid map and the label topology map;
[0027] Step 1.4.2: Process and fuse the data with different resolutions of the 2D occupancy grid map, the node sequences in the label topology map, and the dictionary of orb feature points and corresponding node poses in the visual point cloud map;
[0028] Step 1.4.3: Finally, output the optimized dictionary of orb feature points and corresponding node poses, the node sequence of the topology map, and the occupancy grid map.
[0029] Further optionally, when performing Step 1, the specific operations in the data post - processing process include:
[0030] Step 1.5.1: Extract the node sequence from the constructed label topology map. These nodes contain label information and the position information of the corresponding first and last points; determine the number of cabinets in the data center computer room and the width of the cold aisle for subsequent calculations;
[0031] Step 1.5.2: For each UUID label, use the position of its corresponding first and last points and the number of cabinets, and calculate the specific position coordinates of the UUID label in the map coordinate system through known geometric relationships or proportional allocation methods; combine the width of the cold aisle to further accurately adjust the position of the UUID label;
[0032] Step 1.5.3: Associate the calculated position coordinates of each UUID tag with the corresponding tag name to form key-value pairs in the form of "tag name - position". Organize all the "tag name - position" key-value pairs into a search dictionary data structure for subsequent quick query and retrieval using this search dictionary in global positioning.
[0033] Further optionally, perform Step 2: Multi-resolution map matching based on lidar scan point clouds. This process specifically includes:
[0034] Step 2.1.1: The robot starts working at the known initial pose and inputs this pose into an 80-cm grid map. Apply Ceres for scan_to_map non-linear optimization. If the calculated score is less than 0.6, enter the visual information matching process based on the adjusted active observation angle. Otherwise, input the iterative pose to a 40-cm grid map for the next optimization;
[0035] Step 2.1.2: Apply Ceres for scan_to_map non-linear optimization. If the calculated score is less than 0.6, enter the visual information matching process based on the adjusted active observation angle. Otherwise, input the iterative pose to a 20-cm grid map for the next optimization;
[0036] Step 2.1.3: Apply Ceres for scan_to_map non-linear optimization. If the calculated score is less than 0.6, enter the visual information matching process based on the adjusted active observation angle. Otherwise, input the iterative pose to a 10-cm grid map for the next optimization;
[0037] Step 2.1.4: Apply Ceres for scan_to_map non-linear optimization. If the calculated score is less than 0.6, enter the visual information matching process based on the adjusted active observation angle. Otherwise, input the iterative pose to a 5-cm grid map for the next optimization;
[0038] Step 2.1.5: Apply Ceres for scan_to_map non-linear optimization to obtain the accurate global pose of the robot, and this global positioning search ends.
[0039] Further optionally, perform Step 2: Visual information matching based on the adjusted active observation angle. This process specifically includes:
[0040] Step 2.2.1: The robot cancels all task operations, adjusts the robotic arm to the standby mode. After releasing the occupancy of the binocular camera, readjust the robotic arm to the scanning mode, and restore the end effector pose to the mapping mode in Step 1. The end effector rotates counterclockwise and extracts data frames at 0°, 90°, 180°, and 270° respectively for subsequent operations;
[0041] Step 2.2.2: Extract ORB descriptors from the multi-angle data frame, search in the local "ORB feature - pose" dictionary output during the construction of the 2D occupancy grid map. If the matching score is higher than 0.75, use the pose after coordinate transformation as the reference pose input for the multi-resolution map matching process based on the lidar scan point cloud; otherwise, proceed to Step 2.2.3.
[0042] Step 2.2.3: Perform OCR recognition on the multi-angle data frame to extract possible cabinet UUID tags.
[0043] i) If there are multiple UUID tags, select the one with the closest Euclidean distance to the robot as the criterion, obtain the rough pose estimate of the robot in the local "ORB feature - pose" dictionary stored in Step 1, and then use it as the reference pose input for the multi-resolution map matching process based on the lidar scan point cloud.
[0044] ii) If no cabinet UUID tag is extracted, it is considered that the global positioning search fails this time.
[0045] Optionally, the involved composite inspection robot includes a mobile chassis, a 6-degree-of-freedom robotic arm, an end effector, and an edge computing unit, where:
[0046] The mobile chassis, as the basic support and moving component of the robot, integrates a high-precision odometer and an inertial measurement unit, which are used to provide preliminary motion estimation and dynamic information to ensure stable motion data acquisition when the robot is moving quickly or rotating. At the same time, it is equipped with a single-line lidar for providing scan information in a planar environment.
[0047] The 6-degree-of-freedom robotic arm is installed on the robot chassis and has six degrees of freedom of motion. It can actively adjust the position and pose of the end effector to meet the requirements of different inspection tasks. At the same time, it provides high-precision joint angle data through the built-in encoder and kinematic model.
[0048] The end effector is installed at the end of the robotic arm and is an integration of a temperature and humidity sensor, a noise sensor, an infrared sensor, and a binocular camera. It is used for environmental information collection and server anomaly detection in the data center inspection task, and relies on the equipped binocular camera to collect server nameplate information and environmental feature descriptors to complete environmental scene recognition.
[0049] The edge computing unit is responsible for real-time processing and analysis of the data collected by the robot.
[0050] The beneficial effects of a global positioning method for a composite inspection robot based on robotic arm and vision cooperation of the present invention compared with the prior art are:
[0051] The present invention can achieve the fusion of multiple sensors such as lidar, binocular cameras, and robotic arms, realizing the transition from passive reception to active perception. By using the robotic arm to actively create observation conditions, it breaks through environmental constraints, improves the global positioning and search capabilities of robots in highly similar environments such as data centers, enhances the adaptability to multi-scenario environments, and solves the limitations of sensor perspectives and the limitations of pure lidar in highly similar environments such as data centers, as well as the uncertainty during the rotation process in narrow environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The Figure 1 block diagram shows the implementation of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] In order to make the technical solutions, technical problems solved, and technical effects of the present invention clearer and more understandable, the following describes the technical solutions of the present invention clearly and completely in conjunction with specific embodiments.
[0054] Embodiment:
[0055] Referring to the Figure 1 drawings, this embodiment proposes a global positioning method for a composite inspection robot based on the cooperation of a robotic arm and vision, including the following steps:
[0056] Step 1: Construct a multi-layer reference map based on multi-sensor fusion, specifically including:
[0057] Step 1.1: Construct a two-dimensional occupancy grid map, and the specific operations include:
[0058] Step 1.1.1: Manually remotely control the robot to collect data within the operation range of the data center computer room, and prepare for data collection to ensure the normal operation of various sensors such as lidar, odometer, and IMU on the robot;
[0059] Step 1.1.2: Use the lidar to scan the surrounding environment of the robot to obtain the distance information of the environment; use the odometer and IMU to obtain the motion information of the robot, and at the same time combine the laser scan matching results to determine the position of the robot in the environment;
[0060] Step 1.1.3: Adopt the ray projection algorithm to map the environmental data within a radius of 5m centered on the robot to a two-dimensional plane grid to complete the construction of a 5cm grid map, which is used for the subsequent precise pose iteration of the robot;
[0061] Step 1.1.4: Use sliding windows of 2×2, 4×4, 8×8, and 16×16 to perform pooling on the 5-cm grid map, generating coarse-resolution grid maps of 10 cm, 20 cm, 40 cm, and 80 cm respectively, and save them locally. During this process, the robot uses a binocular camera to collect image data in the computer room environment, extracts features from the image data through the ORB algorithm to obtain ORB feature descriptors, accurately records the pose information of the robot at the moment of collecting each frame of image data, establishes an association between the ORB feature descriptors and the corresponding pose information and stores them, constructing a local "ORB feature - pose" dictionary. This "ORB feature - pose" dictionary is loaded during program initialization to improve the matching efficiency and accuracy of the lidar, and complete the construction of grid maps with different resolutions and the generation of the "ORB feature - pose" dictionary.
[0062] Step 1.2: Construct a labeled topological map. The specific operations include:
[0063] Step 1.2.1: Each cabinet in the computer room environment of the data center has a UUID label, and the UUID label of the cabinet reflects the specific arrangement number of the cabinet in the computer room. During the movement of the robot, the binocular camera installed on the end effector is used to identify the UUID label of the cabinet, and the position of the channel where it is located can be estimated.
[0064] Step 1.2.2: Raise the robotic arm of the composite inspection robot by 50 cm, so that the end camera coordinate system is located on the central axis of the robot chassis and is 1.75 m away from the base_link coordinate system of the chassis. At the same time, ensure that the field of view of the binocular camera is consistent with the direction of the robot's vehicle head.
[0065] Step 1.2.3: Manually remote-control the robot to make it walk centrally to the first group of cabinets in the cold aisle of the computer room. After reaching the position of the first group of cabinets, create a semantic node by clicking on the screen. This semantic node independently records the serial number and records it in a negative form to distinguish it from the nodes independently created during the construction of the grid map.
[0066] Step 1.2.4: Continue to remote-control the robot to make it walk centrally to the last group of cabinets in the same cold aisle. When reaching the position of the last group of cabinets, perform the recording operation again.
[0067] Step 1.2.5: For all aisles in the computer room, repeat the operations in Step 1.2.3 and Step 1.2.4. After the data collection and semantic node creation operations for all aisles, integrate the recorded node information and the associated position and environmental data to construct a complete labeled topological map.
[0068] Step 1.3: Construct a visual point cloud map. The specific operations include:
[0069] Step 1.3.1: Set the position of the end effector of the composite inspection robot to be the same as when constructing the label topological map, that is, raise the robotic arm by 50 cm, the end camera coordinate system is located on the central axis of the robot chassis and 1.75 m away from the base_link coordinate system of the chassis, and the camera field of view is consistent with the direction of the robot's front end;
[0070] Step 1.3.2: When autonomously creating nodes during the construction of the 2D occupancy grid map, the binocular camera collects the dense point cloud data of the current environment, and extracts ORB (Oriented FAST and Rotated BRIEF) feature points from the collected dense point cloud data;
[0071] Step 1.3.3: Bind the extracted ORB feature points to the currently autonomously created nodes. By comparing the ORB feature points extracted from the image with the existing feature points in the local "ORB feature - pose" dictionary, find the most similar feature points according to the descriptor matching degree, and use the pose information associated with these feature points in the local "ORB feature - pose" dictionary as the reference pose information of the new node, so that the pose of the current node is consistent with the environmental descriptor (i.e., the descriptive information represented by the ORB feature points), thereby establishing the association between the node and the environmental features;
[0072] Step 1.3.4: Construct a visual point cloud map by splicing the ORB descriptors corresponding to multiple nodes.
[0073] Step 1.4: Perform multi - map joint optimization. The specific operations include:
[0074] Step 1.4.1: Construct and optimize a multi - layer reference map based on the graph optimization method to ensure the synchronous optimization of the nodes in the 2D occupancy grid map construction and the label topological map construction;
[0075] Step 1.4.2: Use the data with different resolutions of the 2D occupancy grid map, the node sequence in the label topological map, and the dictionary of orb feature points and corresponding node poses in the visual point cloud map for processing and fusion;
[0076] Step 1.4.3: Finally, output the optimized dictionary of orb feature points and corresponding node poses, the topological map node sequence, and the occupancy grid map.
[0077] Step 1.5: Data post - processing. The specific operations in this process include:
[0078] Step 1.5.1: Extract the node sequence from the constructed label topological map. These nodes contain label information and the position information of the corresponding first and last points; determine the number of cabinets and the width of the cold aisle in the data center computer room for subsequent calculations;
[0079] Step 1.5.2: For each UUID tag, use the positions of its corresponding start and end points and the number of cabinets, and calculate the specific position coordinates of the UUID tag in the map coordinate system through known geometric relationships or proportional allocation methods; combine the width of the cold aisle to further precisely adjust the position of the UUID tag.
[0080] Step 1.5.3: Associate the calculated position coordinates of each UUID tag with the corresponding tag name to form key-value pairs in the form of "tag name - position", and organize all the "tag name - position" key-value pairs into a data structure of a search dictionary for subsequent quick query and retrieval using this search dictionary in global positioning.
[0081] Step 2: Based on the output data of Step 1, perform global positioning search, which specifically includes:
[0082] Step 2.1: Multi-resolution map matching based on lidar scan point cloud. This process specifically includes:
[0083] Step 2.1.1: The robot starts working at the known initial pose and inputs this pose into an 80-cm grid map. Apply Ceres for scan_to_map non-linear optimization. If the calculated score is less than 0.6, enter Step 2.2.1; otherwise, input the iterative pose to a 40-cm grid map for the next optimization.
[0084] Step 2.1.2: Apply Ceres for scan_to_map non-linear optimization. If the calculated score is less than 0.6, enter Step 2.2.1; otherwise, input the iterative pose to a 20-cm grid map for the next optimization.
[0085] Step 2.1.3: Apply Ceres for scan_to_map non-linear optimization. If the calculated score is less than 0.6, enter Step 2.2.1; otherwise, input the iterative pose to a 10-cm grid map for the next optimization.
[0086] Step 2.1.4: Apply Ceres for scan_to_map non-linear optimization. If the calculated score is less than 0.6, enter Step 2.2.1; otherwise, input the iterative pose to a 5-cm grid map for the next optimization.
[0087] Step 2.1.5: Apply Ceres for scan_to_map non-linear optimization to obtain the accurate global pose of the robot, and this global positioning search ends.
[0088] Step 2.2: Visual information matching based on active observation angle adjustment. This process specifically includes:
[0089] Step 2.2.1: The robot cancels all task operations, adjusts the robotic arm to the standby mode. After releasing the occupation of the binocular camera, it adjusts the robotic arm back to the scanning mode again, and the posture of the end effector is restored to the mapping mode in Step 1. The end effector rotates counterclockwise and extracts data frames at 0°, 90°, 180°, and 270° respectively for subsequent operations.
[0090] Step 2.2.2: Extract ORB descriptors from the multi-angle data frames and search in the local "ORB feature - pose" dictionary output in the construction of the 2D occupancy grid map. If the matching score is higher than 0.75, the pose is used as the reference pose input for Step 2.1.1 after coordinate transformation. Otherwise, go to Step 2.2.3.
[0091] Step 2.2.3: Perform OCR recognition on the multi-angle data frames to extract the possible cabinet UUID tags.
[0092] i) If there are multiple UUID tags, select the one with the closest distance based on the Euclidean distance from the robot as the criterion. Obtain the rough pose estimate of the robot in the local "ORB feature - pose" dictionary stored in Step 1, and then use it as the reference pose input for the multi-resolution map matching process based on the lidar scan point cloud.
[0093] ii) If no cabinet UUID tags are extracted, it is considered that the global positioning search fails this time.
[0094] It should be added that the composite inspection robot involved in this embodiment includes a mobile chassis, a 6-degree-of-freedom robotic arm, an end effector, and an edge computing unit, where:
[0095] The mobile chassis, as the basic support and moving component of the robot, integrates a high-precision odometer and an inertial measurement unit, which are used to provide preliminary motion estimation and dynamic information to ensure stable motion data acquisition when the robot is moving or rotating rapidly. At the same time, it is equipped with a single-line lidar for providing scanning information in a planar environment.
[0096] The 6-degree-of-freedom robotic arm is installed on the robot chassis and has six degrees of freedom of motion. It can actively adjust the position and posture of the end effector to meet the requirements of different inspection tasks. At the same time, it provides high-precision joint angle data through the built-in encoder and kinematic model.
[0097] The end effector is installed at the end of the robotic arm and is an integrated body of temperature and humidity sensors, noise sensors, infrared sensors, and binocular cameras. It is used for environmental information collection and server anomaly detection in the data center inspection task, and relies on the equipped binocular camera to collect server nameplate information and environmental feature descriptors to complete environmental scene recognition.
[0098] The edge computing unit is responsible for real-time processing and analysis of the data collected by the robot.
[0099] In summary, by adopting the global positioning method of the composite inspection robot based on robotic arm collaborative vision of the present invention, multi-sensor fusion of lidar, binocular camera, robotic arm, etc. can be achieved, realizing the transformation from passive reception to active perception, using the robotic arm to actively create observation conditions, breaking through environmental constraints, improving the global positioning and search ability of the robot in highly similar environments such as data centers, and enhancing the adaptability to multi-scene environments.
[0100] The above specific application examples have elaborated in detail the principle and implementation manner of the present invention. These examples are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, any improvements and modifications made by those skilled in the art of the present technology without departing from the principle of the present invention shall fall within the patent protection scope of the present invention.
Claims
1. A global positioning method for a composite inspection robot based on the cooperation of a robotic arm and vision, characterized in that, The steps include: Step 1: Construct a multi-layer reference map based on multi-sensor fusion, which includes constructing a two-dimensional occupancy grid map, a label topology map, and a visual point cloud map, followed by multi-map joint optimization and data post-processing; Step 2: Based on the output data of step 1, a global positioning search is performed, which specifically includes multi-resolution map matching based on the lidar scanning point cloud and visual information matching based on active observation angle adjustment.
2. The global positioning method of the composite inspection robot based on the cooperation of the robotic arm and vision according to claim 1, characterized in that, The specific operations for constructing a two-dimensional occupancy grid map include: Step 1.1.
1. The robot is controlled remotely to collect data within the operating range of the data center computer room, and preparations are made for data collection to ensure that the laser radar, odometer, and IMU sensors on the robot are operating normally; Step 1.1.2: Use the laser radar to scan the robot's surrounding environment to obtain the distance information of the environment; use the odometer and IMU to obtain the robot's motion information, and combine the laser scanning matching results to determine the robot's position in the environment; Step 1.1.3: Use the ray projection algorithm to map the environmental data within a radius of 5m with the robot as the center to a two-dimensional plane grid to complete the construction of a 5cm grid map, which is used for subsequent precise posture iteration of the robot; Step 1.1.4: Use sliding windows of 2×2, 4×4, 8×8, and 16×16 to pool the 5 cm grid map, and generate coarse resolution grid maps of 10 cm, 20 cm, 40 cm, and 80 cm respectively and save them locally; in this process, the robot uses a binocular camera to collect image data in the computer room environment, extracts features from the image data through the ORB algorithm, obtains the ORB feature descriptor, and accurately records the posture information of the robot at the moment of collecting each frame of image data, associates the ORB feature descriptor with the corresponding posture information and stores it, and constructs a local "ORB feature-pose" dictionary; the "ORB feature-pose" dictionary is loaded when the program is initialized to improve the efficiency and accuracy of lidar matching, and complete the construction of grid maps with different resolutions and the generation of the "ORB feature-pose" dictionary.
3. The global positioning method of the composite inspection robot based on the cooperation of the robotic arm and vision according to claim 2, characterized in that, The specific operations for building a label topology map include: Step 1.2.
1. Each cabinet in the data center computer room has a UUID label. The UUID label of the cabinet reflects the specific arrangement number of the cabinet in the computer room. During the movement of the robot, the binocular camera installed on the end effector can be used to identify the UUID label of the cabinet to estimate the location of the channel. Step 1.2.2, raise the mechanical arm of the composite inspection robot by 50cm, so that the end camera coordinate system is located on the center axis of the robot chassis and 1.75m away from the chassis base_link coordinate system, and at the same time ensure that the binocular camera field of view is consistent with the direction of the robot head; Step 1.2.3: The robot is controlled remotely to move to the first group of cabinets in the cold channel of the computer room. After reaching the first group of cabinets, a semantic node is created by clicking on the screen. The semantic node records the serial number independently and in the form of a negative number to distinguish it from the nodes created autonomously during the grid map construction process. Step 1.2.4: Continue to remotely control the robot to make it walk in the same cold aisle to the last set of cabinets; when reaching the position of the last set of cabinets, perform the recording operation again; Step 1.2.5: For all aisles in the computer room, repeat the operations in Step 1.2.3 and Step 1.2.
4. After the data collection and semantic node creation operations for all aisles, integrate and process the recorded node information and the associated position and environmental data to construct a complete labeled topological map.
4. The global positioning method of the composite inspection robot based on the cooperation of the robotic arm and vision according to claim 3, characterized in that, The specific operations for constructing the visual point cloud map include: Step 1.3.1: Set the position of the end effector of the composite inspection robot to be the same as when constructing the labeled topological map, that is, raise the robotic arm by 50 cm, the end camera coordinate system is on the central axis of the robot chassis and 1.75 m away from the base_link coordinate system of the chassis, and the camera field of view is consistent with the direction of the robot's vehicle head; Step 1.3.2: When autonomously creating nodes during the construction of the 2D occupancy grid map, the binocular camera collects the dense point cloud data of the current environment, and extracts ORB (Oriented FAST and Rotated BRIEF) feature points from the collected dense point cloud data; Step 1.3.3: Bind the extracted ORB feature points to the currently autonomously created nodes. By comparing the ORB feature points extracted from the image with the existing feature points in the local "ORB feature - pose" dictionary, find the most similar feature points according to the descriptor matching degree, and use the pose information associated with these feature points in the local "ORB feature - pose" dictionary as the reference pose information of the new node, so that the pose of the current node is consistent with the environmental descriptor (i.e., the description information represented by the ORB feature points), thereby establishing the association between the node and the environmental features; Step 1.3.4: Construct the visual point cloud map by stitching the ORB descriptors corresponding to multiple nodes.
5. The global positioning method of the composite inspection robot based on the cooperation of robotic arms and vision according to claim 4, characterized in that, The specific operations for multi - map joint optimization include: Step 1.4.1: Based on the graph optimization method, construct and optimize the multi - layer reference map to ensure the synchronous optimization of the nodes in the 2D occupancy grid map construction and the labeled topological map construction; Step 1.4.2: Use the data with different resolutions of the 2D occupancy grid map, the node sequence in the labeled topological map, and the dictionary of orb feature points and corresponding node poses in the visual point cloud map for processing and fusion; Step 1.4.3: Finally, output the optimized dictionary of orb feature points and corresponding node poses, the topological map node sequence, and the occupancy grid map.
6. The global positioning method of the composite inspection robot based on the cooperation of the robotic arm and vision according to claim 5, characterized in that, The specific operations in the data post - processing process include: Step 1.5.1: Extract the node sequence from the constructed labeled topological map. These nodes contain the label information and the position information of the corresponding start and end points; determine the number of cabinets in the data center computer room and the width of the cold aisle for subsequent calculations; Step 1.5.2: For each UUID tag, use the positions of its corresponding start and end points and the number of cabinets, and calculate the specific position coordinates of the UUID tag in the map coordinate system through known geometric relationships or proportional allocation methods; combine the width of the cold aisle to further precisely adjust the position of the UUID tag. Step 1.5.3: Associate the calculated position coordinates of each UUID tag with the corresponding tag name to form key-value pairs in the form of "tag name - position". Organize all the "tag name - position" key-value pairs into a data structure of a search dictionary for quick query and retrieval using this search dictionary in subsequent global positioning.
7. The global positioning method of the composite inspection robot based on the cooperation of the robotic arm and vision according to claim 6, characterized in that Multi-resolution map matching based on lidar scan point clouds, and this process specifically includes: Step 2.1.1: The robot starts working at the known initial pose and inputs this pose into an 80-cm grid map. Apply Ceres for scan_to_map non-linear optimization. If the calculated score is less than 0.6, enter the visual information matching process based on the adjusted active observation angle; otherwise, input the iterative pose to a 40-cm grid map for the next optimization. Step 2.1.2: Apply Ceres for scan_to_map non-linear optimization. If the calculated score is less than 0.6, enter the visual information matching process based on the adjusted active observation angle; otherwise, input the iterative pose to a 20-cm grid map for the next optimization. Step 2.1.3: Apply Ceres for scan_to_map non-linear optimization. If the calculated score is less than 0.6, enter the visual information matching process based on the adjusted active observation angle; otherwise, input the iterative pose to a 10-cm grid map for the next optimization. Step 2.1.4: Apply Ceres for scan_to_map non-linear optimization. If the calculated score is less than 0.6, enter the visual information matching process based on the adjusted active observation angle; otherwise, input the iterative pose to a 5-cm grid map for the next optimization. Step 2.1.5: Apply Ceres for scan_to_map non-linear optimization to obtain the accurate global pose of the robot, and this global positioning search ends.
8. The global positioning method of the composite inspection robot based on the cooperation of the robotic arm and vision according to claim 7, characterized in that, Visual information matching based on the adjusted active observation angle, and this process specifically includes: Step 2.2.1: The robot cancels all task operations, adjusts the robotic arm to the standby mode. After releasing the occupation of the binocular camera, readjust the robotic arm to the scanning mode, and restore the posture of the end effector to the mapping mode in Step 1. The end effector rotates counterclockwise to extract data frames at 0°, 90°, 180°, and 270° respectively for subsequent operations. Step 2.2.2: Extract ORB descriptors from the multi-angle data frames and search in the local "ORB feature - pose" dictionary output in the constructed two-dimensional occupancy grid map. If the matching score is higher than 0.75, input the pose after coordinate transformation as the reference pose for the multi-resolution map matching process based on lidar scan point clouds; otherwise, enter Step 2.2.
3. Step 2.2.3: Perform OCR recognition on the multi-angle data frame to extract possible cabinet UUID tags. i) If there are multiple UUID tags, select the one with the closest Euclidean distance to the robot as the criterion, obtain the rough pose estimation of the robot from the local "ORB feature - pose" dictionary stored in Step 1, and then use it as the reference pose input for the multi-resolution map matching process based on the lidar scan point cloud. ii) If no cabinet UUID tag is extracted, it is considered that the global positioning search fails this time.
9. The global positioning method of the composite inspection robot based on the cooperation of the robotic arm and vision according to claim 1, characterized in that, The composite inspection robot includes a mobile chassis, a 6-degree-of-freedom robotic arm, an end effector, and an edge computing unit, where: The mobile chassis, as the basic support and moving component of the robot, integrates a high-precision odometer and an inertial measurement unit, which are used to provide preliminary motion estimation and dynamic information, ensure stable motion data acquisition when the robot is moving or rotating rapidly, and at the same time carry a single-line lidar for providing scan information in a planar environment. The 6-degree-of-freedom robotic arm is installed on the robot chassis, has six degrees of freedom of motion, can actively adjust the position and pose of the end effector to adapt to the requirements of different inspection tasks, and at the same time provides high-precision joint angle data through the built-in encoder and kinematic model. The end effector is installed at the end of the robotic arm and is an integration of a temperature and humidity sensor, a noise sensor, an infrared sensor, and a binocular camera, which is used for environmental information collection and server anomaly detection in the data center inspection task, and depends on the carried binocular camera to collect server nameplate information and environmental feature descriptors to complete environmental scene recognition. The edge computing unit is responsible for real-time processing and analysis of the data collected by the robot.