Robot positioning verification and correction method and device based on semantic grid map

Through the method based on semantic raster map, a map containing landmark information is constructed using depth cameras and lidar, and semantic objects are identified for positioning checks and corrections, which solves the problem of reduced positioning accuracy of robots in complex environments, achieving high-precision positioning and rapid map construction.

CN120445202APending Publication Date: 2025-08-08WUHAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510529883.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Existing robot positioning methods are prone to errors in complex environments and lack a deep understanding of environmental semantics, resulting in reduced positioning accuracy and accumulated errors, especially in dynamic and degraded environments, or serious accumulation of errors.

Method used

Using a method based on semantic raster map, a semantic object is recognized through a depth camera, a two-dimensional raster map is constructed with lidar, a semantic raster map containing landmark information is generated, and a three-dimensional coordinates and categories of semantic objects are used for positioning checksum correction, and a loop optimization is used to update the landmark information to improve positioning accuracy.

Benefits of technology

It realizes high-precision robot positioning in complex environments, reduces positioning errors, improves positioning accuracy and stability, and optimizes the speed of map construction and navigation processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120445202A_ABST
    Figure CN120445202A_ABST
Patent Text Reader

Abstract

The invention provides a robot positioning verification and correction method and device based on a semantic grid map, and the method comprises the steps: receiving real-time image data obtained by a depth camera on a robot, and carrying out the recognition based on the real-time image data, so as to obtain at least two continuous semantic objects; determining the positions of the two identified semantic objects in a semantic grid map, and determining the position and orientation of the depth camera in the semantic grid map according to the positions of the at least two semantic objects in the semantic grid map; and determining the position and orientation of the robot in the semantic grid map based on the position and orientation of the depth camera in the semantic grid map so as to complete positioning verification of the robot. According to the method, the landmark information is generated and published in the two-dimensional grid map, and the semantic grid map with the semantic information of the object is obtained and used for robot navigation, so that robot positioning verification is realized based on the semantic object in the robot positioning process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot navigation, and in particular to a robot positioning verification and correction method and device based on a semantic grid map. Background Art

[0002] With the continuous development of robotics, the accuracy of robot positioning and navigation in indoor environments has become a key issue. Traditional positioning methods, such as LiDAR-based SLAM technology, can provide high accuracy in static environments. However, in dynamic and degraded environments, factors such as obstacles, sensor noise, and environmental changes can lead to a significant decrease in positioning accuracy, or even loss of positioning.

[0003] Furthermore, in similar dynamic environments, robots are prone to accumulation of positioning errors and repeated map matching problems, which can lead to the system being unable to correctly recover its positioning. Advances in depth cameras and object detection technology have enabled robots to perceive more environmental information.

[0004] However, existing positioning methods often fail to effectively integrate these multi-source information and lack a deep understanding of environmental semantics, thus failing to solve the problems of positioning loss and error accumulation in complex scenarios. Summary of the Invention

[0005] The present invention provides a robot positioning verification and correction method and device based on a semantic grid map, which is used to solve the defect in the prior art that robot navigation is prone to errors when defined in complex scenes, and to achieve a robot positioning verification and correction method and device based on a semantic grid map with higher positioning accuracy.

[0006] The present invention provides a robot positioning verification and correction method based on a semantic grid map, comprising: receiving real-time image data acquired by a depth camera on the robot, and identifying at least two continuous semantic objects based on the real-time image data; Determining positions of the two identified semantic objects in the semantic grid map, and determining a position and orientation of the depth camera in the semantic grid map based on the positions of the at least two semantic objects in the semantic grid map; Determining the position and orientation of the robot in the semantic grid map based on the position and orientation of the depth camera in the semantic grid map to complete the positioning verification of the robot; Among them, the semantic object is an object that the robot recognizes and publishes landmark information during the mapping process, and the semantic grid map is a two-dimensional grid map constructed by the robot that contains semantic object landmark information. The landmark information includes the three-dimensional coordinates and category of the semantic object.

[0007] According to a robot positioning verification and correction method based on a semantic grid map provided by the present invention, before the step of receiving real-time image data acquired by a depth camera on the robot, the method further includes: Control the robot to move in the target area, and receive environmental information of the target area obtained by the lidar sensor on the robot during the movement to build a two-dimensional grid map of the target area; During the movement process, the target detection model is used to identify objects in the environment, determine the three-dimensional coordinates and categories of the objects, and generate landmark information corresponding to the objects. During the mapping process, the landmark information is published in the two-dimensional grid map, so that after the mapping is completed, a two-dimensional grid map containing landmark information of the target area is obtained as the semantic grid map.

[0008] According to a robot positioning verification and correction method based on a semantic grid map provided by the present invention, the steps of using a target detection model to identify objects in the environment during movement, determining the three-dimensional coordinates and categories of the objects, and generating landmark information corresponding to the objects specifically include: Identify objects in the environment using a target detection model and obtain a two-dimensional detection frame of the object and the category of the object; Selecting the center point of the two-dimensional detection frame of the object as a representative point representing the position of the object, and mapping the representative point to the depth map obtained by the depth camera to obtain the three-dimensional coordinates of the object based on the depth information; The landmark information of the object is generated based on the three-dimensional coordinates of the object and the category of the object.

[0009] According to a robot positioning verification and correction method based on a semantic grid map provided by the present invention, before the step of receiving real-time image data acquired by a depth camera on the robot, the method further includes: During the mapping process, loop closure optimization is used to correct and update the landmark information corresponding to the object.

[0010] According to a robot positioning verification and correction method based on a semantic grid map provided by the present invention, the step of determining the position and orientation of the depth camera in the semantic grid map according to the positions of at least two semantic objects in the semantic grid map specifically includes: Draw two circles with the position of the semantic object in the semantic grid map as the center and the projection length of the line from the semantic object to the depth camera on the XOZ plane of the camera coordinate system as the radius, and determine the coordinates of the two intersection points between the two circles; Determine a first vector pointing from one semantic object to another semantic object, and a second vector pointing from the one semantic object to an intersection point; Calculating the vector cross product of the first vector and the second vector, filtering out an intersection point representing the depth camera from the two intersection points based on the right-hand rule, and using the filtered intersection coordinates as the position of the depth camera in the semantic grid map; The orientation of the depth camera is determined based on the position of the depth camera in the semantic grid map and the position of the semantic object in the camera coordinate system.

[0011] According to a method for robot positioning verification and correction based on a semantic grid map provided by the present invention, after the step of determining the position and orientation of the robot in the semantic grid map based on the position and orientation of the depth camera in the semantic grid map, the method further includes: Comparing the position of the robot in the semantic grid map calculated based on the semantic object with the published robot position to obtain a position difference; Comparing the orientation of the robot in the semantic grid map calculated based on the semantic object with the published orientation of the robot to obtain an orientation difference; In a case where the position difference and / or the orientation difference is greater than corresponding preset thresholds, the published robot position and orientation are corrected using the position and orientation of the robot in the semantic grid map.

[0012] The present invention also provides a robot positioning verification and correction device based on a semantic grid map, comprising: A recognition module, configured to receive real-time image data acquired by a depth camera on the robot, and to recognize at least two continuous semantic objects based on the real-time image data; a determination module, configured to determine positions of at least two semantic objects obtained by identification in a semantic grid map, and determine a position and orientation of the depth camera in the semantic grid map according to the positions of the at least two semantic objects in the semantic grid map; A positioning module, configured to determine the position and orientation of the robot in the semantic grid map based on the position and orientation of the depth camera in the semantic grid map, so as to complete the positioning verification of the robot; a correction module, configured to correct the published position and orientation of the robot using the position and orientation of the robot in the semantic grid map; Among them, the semantic object is an object that the robot recognizes and publishes landmark information during the mapping process, and the semantic grid map is a two-dimensional grid map constructed by the robot that contains semantic object landmark information. The landmark information includes the three-dimensional coordinates and category of the semantic object.

[0013] The present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the robot positioning verification and correction method based on the semantic grid map as described above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described robot positioning verification and correction methods based on semantic grid maps.

[0015] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described robot positioning verification and correction methods based on semantic grid maps.

[0016] The present invention provides a robot positioning verification and correction method and device based on a semantic grid map. During the robot mapping process, the positions and categories of objects are identified, and landmark information is generated and published on a two-dimensional grid map. A semantic grid map with object semantic information is obtained for robot navigation. In this way, robot positioning verification is achieved based on semantic objects during the robot positioning process, and high-precision robot positioning is achieved by utilizing the prior position information of semantic objects, thereby realizing the integrated use of multi-source environmental information. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 1 is a flow chart of a robot positioning verification and correction method based on a semantic grid map provided by the present invention; Figure 2 Schematic diagram of a two-dimensional grid map in the robot positioning verification and correction method based on a semantic grid map provided by the present invention; Figure 3 Schematic diagram of the semantic grid map in the robot positioning verification and correction method based on the semantic grid map provided by the present invention; Figure 4 It is a schematic diagram of two consecutive semantic objects identified by the target detection model in the robot positioning verification and correction method based on the semantic grid map provided by the present invention; Figure 5 Schematic diagram of determining the depth camera pose in the robot positioning verification and correction method based on the semantic grid map provided by the present invention; Figure 6 Schematic diagram of the robot positioning verification and correction method based on the semantic grid map provided by the present invention for screening intersection points representing depth cameras; Figure 7 (a) is a schematic diagram of the positioning result before positioning correction in the robot positioning verification and correction method based on the semantic grid map provided by the present invention; Figure 7 (b) is a schematic diagram of the positioning result after positioning correction in the robot positioning verification and correction method based on the semantic grid map provided by the present invention; Figure 8 Schematic diagram of the structure of the robot positioning verification and correction device based on the semantic grid map provided by the present invention; Figure 9 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0020] The following combination Figures 1 to 7 The present invention introduces a robot positioning verification and correction method based on semantic grid map, such as Figure 1 Shown, including: Step 101: receiving real-time image data acquired by a depth camera on a robot, and identifying at least two continuous semantic objects based on the real-time image data; Among them, the semantic object is an object that the robot recognizes and publishes landmark information during the mapping process, and the semantic grid map is a two-dimensional grid map constructed by the robot that contains semantic object landmark information. The landmark information includes the three-dimensional coordinates and category of the semantic object.

[0021] It is understandable that the robot relies on a pre-built map of the target area during the positioning and navigation process.

[0022] Typically, a robot moves in a target area beforehand, using its mounted LiDAR sensor to obtain environmental information about the target area. This raster map is then constructed and used for subsequent positioning and navigation. However, this raster map lacks a deep understanding of the environmental semantics and fails to integrate multi-source information about the target area, which can easily lead to positioning errors during the robot's positioning and navigation process.

[0023] To this end, in this embodiment, during the robot's mapping process, various objects in the target area's environment are identified, and the three-dimensional coordinates and category of each identified object are obtained as the object's semantic information. This object's semantic information is then packaged to generate the object's landmark information, i.e., the object's landmark information. This object's landmark information is then published on a two-dimensional grid map constructed based on LiDAR data, resulting in a two-dimensional grid map containing semantic object landmark information, known as a semantic grid map. Objects identified and published with landmark information during the mapping process are considered semantic objects.

[0024] In a feasible implementation, the constructed two-dimensional grid map is as follows: Figure 2 As shown, Figure 3 for Figure 2 The corresponding partial enlarged image of the semantic grid map, Figure 3 The text labels such as Exit_4 and Carefully_3 are landmark information published on the two-dimensional raster map, which together with the two-dimensional raster map constitute a semantic raster map.

[0025] On this basis, during the robot navigation process, the robot's main control computer receives real-time image data obtained by a depth camera installed on the robot to identify at least two continuous semantic objects from a frame of real-time image obtained during the navigation process.

[0026] In a specific embodiment, the host computer uses real-time image data as input to the target detection model, and based on the target detection model, realizes detection and recognition of semantic objects in the same way as in the mapping process.

[0027] Specifically, since the robot's mapping process also requires detecting and identifying objects in the environment, thereby generating and publishing landmark information for these objects, the robot's navigation process uses the same method to identify semantic objects in the target area, thereby representing the 3D coordinates of semantic objects at the same location during both mapping and navigation.

[0028] For example, during the mapping process, the target detection model generates a target detection frame for each identified object, uses the center position of the target detection frame to represent the position of the object, and determines the three-dimensional coordinates of the object; during the navigation process, the same target detection model is used, and the center position of the target detection frame is also used to represent the position of the semantic object, and the coordinates of the identified semantic object in the semantic grid map are determined.

[0029] On this basis, the main control machine combines the landmark information of each semantic object pre-stored in the semantic grid map, and performs multiple constraint matching according to the category, spatial relationship and confidence of the semantic objects identified during the navigation process, thereby realizing object detection and recognition based on landmark information in the semantic grid map.

[0030] That is, the landmark information corresponding to the current two continuous semantic objects is identified in the semantic grid map, and the three-dimensional coordinates in the landmark information are used as the three-dimensional coordinates of the determined semantic object in the semantic grid map to realize the utilization of high-precision object positioning prior data.

[0031] Optionally, the target detection model is a YOLO model.

[0032] Optionally, the master machine implements multi-constraint matching based on the consistency of object categories, the matching degree of spatial relationships, and a confidence threshold.

[0033] In a specific embodiment, during navigation, two consecutive semantic objects identified by the target detection model are Figure 4 shown.

[0034] Step 102: determining positions of at least two semantic objects obtained by identification in a semantic grid map, and determining a position and orientation of the depth camera in the semantic grid map according to the positions of the at least two semantic objects in the semantic grid map; Even if the 3D coordinates of a semantic object are known, the robot's depth camera's relative position to the semantic object can be calculated based on the depth information obtained by the depth camera and the semantic object's 3D coordinates in the semantic grid map, resulting in the depth camera's current 3D coordinates in the semantic grid map. However, positioning using a single semantic object can also lead to positioning errors.

[0035] Therefore, in this embodiment, when at least two consecutive semantic objects can be identified in one frame of image, the positioning of the robot is verified based on the relative positions of the at least two semantic objects and the camera.

[0036] Specifically, when observing two semantic objects and the depth camera from a bird's-eye view, the depth camera must be located at one of the intersections of two circles with the semantic object as the center and the projection distance of the line connecting the semantic object and the depth camera on the horizontal plane as the radius.

[0037] Based on this, the projected distance between the depth camera and each semantic object on the horizontal plane is calculated based on the real-time data acquired by the depth camera. This distance is used as the radius of the circle. A circle is then drawn with the (x, y) coordinates of the 3D landmark information corresponding to the semantic object as the center. The intersection of the two circles is the location of the depth camera.

[0038] Furthermore, if there are three semantic objects, a circle may be drawn for the third semantic object in the same manner, and the common intersection of the three circles is determined as the position of the depth camera.

[0039] If there are only two semantic objects, we can also introduce the actual orientation relationship between the two semantic objects. For example, semantic object A is located on the left side of semantic object B. Based on this orientation relationship, we can verify the two intersection points and select the intersection point representing the position of the depth camera. Its three-dimensional coordinates can be used as the position of the depth camera in the semantic grid map.

[0040] The depth camera position calculated in the above manner effectively utilizes the three-dimensional coordinates pre-stored in the landmark information of the semantic object as prior information, and uses it as the basis for calculating the depth camera position, effectively improving the accuracy of depth camera positioning.

[0041] On this basis, the orientation of the depth camera in the semantic grid map can be calculated based on the angle between the camera and any semantic object, and the angle between the semantic object and the X-axis of the semantic grid map coordinate system.

[0042] Step 103: determining the position and orientation of the robot in the semantic grid map based on the position and orientation of the depth camera in the semantic grid map to complete the positioning verification of the robot; Since the relative position relationship between the depth camera and the robot is fixed, once the position and orientation of the depth camera in the semantic grid map are determined, the position and orientation of the robot in the semantic grid map can be determined through coordinate transformation.

[0043] Specifically, the main control computer calculates the precise position and orientation of the robot body based on the tf data (transform data) between the camera and the body, using the known transformation matrix, and obtains the body's pose information, providing an accurate reference for subsequent positioning and navigation tasks: ; Where TF is the transformation matrix, 、 and is the 3D coordinate of the depth camera in the semantic grid map, 、 and is the 3D coordinate of the robot in the semantic grid map.

[0044] Specifically, the above formula can be used to realize the posture conversion between the depth camera and the robot, and provide the information to the subsequent robot positioning and navigation system. By comparing it with the robot position and orientation published by the system, the robot positioning verification can be completed.

[0045] Optionally, in this embodiment, Cartographer is used to complete the task of building the robot's semantic grid map. Correspondingly, the position and orientation of the robot in the semantic grid map calculated based on semantic objects are compared and verified with the robot's position and orientation published by Cartographer to achieve accurate positioning of the robot.

[0046] The present invention identifies the position and category of objects during the robot mapping process, generates landmark information and publishes it on a two-dimensional grid map, thereby obtaining a semantic grid map with object semantic information for robot navigation. In this way, robot positioning verification is achieved based on semantic objects during the robot positioning process, and high-precision robot positioning is achieved by utilizing the prior position information of semantic objects, thereby realizing the integrated use of multi-source environmental information.

[0047] It should be noted that, since the identified object categories and location information are published as landmark information on the two-dimensional grid map in the present invention, a semantic grid map with semantic information can be obtained at the same time as the grid map is constructed, without the need for an additional semantic matching process, which effectively reduces the construction time of the map containing semantic information.

[0048] In the robot positioning verification and correction method based on the semantic grid map of the present invention, before the step of receiving real-time image data acquired by the depth camera on the robot, the method further includes: Control the robot to move in the target area, and receive environmental information of the target area obtained by the lidar sensor on the robot during the movement to build a two-dimensional grid map of the target area; During the robot's mapping process, the main control computer controls the robot to move in the target area. During the movement, it receives environmental information of the target area obtained by the lidar sensor installed on the robot. Using the Cartographer algorithm, the lidar scanning data is rasterized to generate a two-dimensional raster map of the target area.

[0049] During the movement process, the target detection model is used to identify objects in the environment, determine the three-dimensional coordinates and categories of the objects, and generate landmark information corresponding to the objects. During the mapping process, the landmark information is published in the two-dimensional grid map, so that after the mapping is completed, a two-dimensional grid map containing landmark information of the target area is obtained as the semantic grid map.

[0050] During the movement of the robot, the depth camera installed on the robot simultaneously obtains depth map data and color map data.

[0051] The color image data is input into the target detection model. In this embodiment, the YOLO model is used to identify objects in the environment and determine the category of the objects. The depth information of the identified objects is further accurately calculated to determine the three-dimensional coordinates of the objects.

[0052] The three-dimensional coordinates and categories of objects are published as landmark information on a two-dimensional raster map. During the construction of the raster map, the semantic information of the objects is coupled, and ultimately a two-dimensional raster map containing landmark information is directly constructed. In other words, a map with coupled semantic information is directly constructed without the need for an additional matching process, which improves the processing speed of the mapping task.

[0053] In the robot positioning verification and correction method based on semantic grid maps of the present invention, the steps of using the target detection model to identify objects in the environment during movement, determining the three-dimensional coordinates and categories of the objects, and generating landmark information corresponding to the objects specifically include: Identify objects in the environment using a target detection model and obtain a two-dimensional detection frame of the object and the category of the object; Specifically, during the recognition process of the target detection model, the color image data obtained by the depth camera as it moves with the robot is first received, and the two-dimensional detection frame (frame 1) corresponding to the object and the category of the object are identified.

[0054] Selecting the center point of the two-dimensional detection frame of the object as a representative point representing the position of the object, and mapping the representative point to the depth map obtained by the depth camera to obtain the three-dimensional coordinates of the object based on the depth information; Furthermore, in order to define the three-dimensional coordinates of the object based on the two-dimensional detection frame, frame 1 is mapped to the depth image corresponding to the color image. Since each pixel in the color image corresponds to a pixel in the depth image, the two-dimensional detection frame (frame 2) corresponding to frame 1 in the depth image can be obtained.

[0055] The main control can extract the depth data of the identified object according to the position of frame 2 in the depth image.

[0056] Optionally, the center point of frame 2 is used as a representative point.

[0057] In a specific embodiment, in order to avoid directly obtaining the center point as a noise point, the central area of box 2 is selected to represent the representative point. The main control computer calculates the depth value of the central area of box 2 and checks whether the depth value of the area is consistent. If the depth value is consistent, the average value of the depth value of the area can be used as the depth value of the representative point.

[0058] If the depth value in the central area is large or contains invalid values (for example, the depth value is zero or abnormally large), the master computer considers the data in this area to be unreliable. At this time, the master computer uses robust fitting methods such as the RANSAC algorithm to filter out valid values from the depth data, eliminate noise points, and finally calculate a reasonable depth value as the depth value of the representative point.

[0059] On this basis, the main control machine can calculate the depth value of the representative point to obtain the accurate three-dimensional coordinates of the object in the depth camera coordinate system: ; Where Xpixel and Ypixel are the pixel coordinates of the representative point in the depth image, Cx and Cy are the principal points of the depth camera, fx and fy are the focal lengths of the depth camera, and Z is the depth value of the representative point.

[0060] It is understandable that since the positions of the robot and the depth camera installed on the robot are clear during the mapping stage, the three-dimensional coordinates of the object in the two-dimensional grid map can be obtained through coordinate conversion based on the three-dimensional coordinates of the object in the depth camera coordinate system.

[0061] The landmark information of the object is generated based on the three-dimensional coordinates of the object and the category of the object.

[0062] In this implementation, the mapping task is performed based on Cartographer. Therefore, after the three-dimensional coordinates of the object are calculated, the host generates landmark information for each object in Cartographer and publishes it to the LandmarkList topic: Landmark=(ID,Position(x,y,z),Type); Where ID represents the unique identifier of the object, Position (x, y, z) represents the three-dimensional coordinates of the object in the two-dimensional grid map, and Type represents the object category.

[0063] That is, each landmark information contains the category, location and spatial relationship of the object relative to the robot platform.

[0064] In a specific implementation, the document storage format of landmark information is as follows: Carefully_1(12.676 -2.22101 0.198115) Carefully_2(12.4609 -2.21169 0.020299) Carefully_3(-3.94586 -1.3645 0.072590) Exit_1(47.6122 29.2615 -0.0736509) Exit_2(-25.5827 98.9289 0.00529441) Exit_3(-28.1265 58.4405 -0.0280291) Exit_4(-12.136 7.53363 0.0915586) Exit_5(-4.29852 -1.33737 0.0888682) Red_1(3.98059 -1.71261 0.0928893) Slow_1(4.36664 -1.71067 0.0707456) Slow_2(-12.1551 7.04758 0.0893255) In this implementation, the target area is selected as a parking lot. Therefore, Carefully, Exit, Red, and Slow in the landmark are different symbols of the identified parking lot, indicating caution, exit, attention, and slowdown, that is, the categories of the identified objects. The number after each category represents the object's unique identifier, and the content in brackets represents the object's three-dimensional coordinates in the two-dimensional grid map.

[0065] It is understandable that if other algorithms are used to build maps, the three-dimensional coordinates and categories of objects can also be used to generate landmark information corresponding to other algorithms and published.

[0066] Through the above method, in the process of controlling the robot's movement and mapping, the location and category information of the identified objects are packaged as landmark information and published on the two-dimensional grid map, thereby realizing the direct construction of the semantic grid map and optimizing the mapping process of the grid map containing semantic information.

[0067] It is worth mentioning that since this embodiment uses the representative point positions of the two-dimensional detection box of the target detection model to represent the position of the object during the mapping process, the target detection model detects and identifies semantic objects in the same way during the robot navigation process. That is, during the mapping and navigation processes, the corresponding positions are used as the three-dimensional coordinates of the object, thereby optimizing the computational complexity of semantic object matching during the navigation process, eliminating the need for real-time processing of a large amount of point cloud information, and further optimizing the speed of robot positioning during navigation.

[0068] In the robot positioning verification and correction method based on the semantic grid map of the present invention, before the step of receiving real-time image data acquired by the depth camera on the robot, the method further includes: During the mapping process, loop closure optimization is used to correct and update the landmark information corresponding to the object.

[0069] During the robot's mapping process, loop optimization is used, that is, the robot is controlled to return to the position it has passed before, identify loops in the environment, and optimize the global map and trajectory.

[0070] In this embodiment, loop closure optimization is not only used to improve the mapping accuracy of the two-dimensional raster map, but also to perform loop closure optimization on the landmark information of the identified semantic objects, specifically including loop closure optimization of the object category, unique identifier and location information in the landmark information, so as to improve the accuracy of the three-dimensional coordinates of the semantic objects in the landmark information.

[0071] The landmark information after loopback is also saved in the document text for subsequent map updates and positioning verification.

[0072] Optimized landmark information can better reflect the location changes of objects in complex environments, improving the map's real-time update capabilities. Semantic object information is stored in text format (txt), enabling high-precision semantic mapping.

[0073] It should be noted that in this embodiment, since the category and location of the object are packaged as landmark information and published on the two-dimensional grid map, the loop optimization of the landmark information is also achieved during the loop optimization of the two-dimensional grid map. When the loop of the two-dimensional grid map is completed, the loop of the landmark information is completed synchronously without the need for additional processing, which further optimizes the speed of the robot's mapping.

[0074] In the robot positioning verification and correction method based on the semantic grid map of the present invention, the step of determining the position and orientation of the depth camera in the semantic grid map according to the positions of at least two semantic objects in the semantic grid map specifically includes: Draw two circles with the position of the semantic object in the semantic grid map as the center and the projection length of the line from the semantic object to the depth camera on the XOZ plane of the camera coordinate system as the radius, and determine the coordinates of the two intersection points between the two circles; In this embodiment, in order to accurately obtain the coordinates of the depth camera during navigation, the recognized semantic object pair is first determined. The semantic object pair includes two semantic objects recognized continuously. If multiple semantic objects are recognized, any two of them can be selected and calculated in the same way.

[0075] It is understandable that after two continuous semantic objects are identified during navigation by the target recognition model, the accurate three-dimensional coordinates of the two semantic objects in the semantic grid map can be obtained based on the pre-stored landmark information of the two semantic objects.

[0076] In this embodiment, if Figure 5 As shown, Figure 5 The two consecutive blue dots in the image represent two consecutive semantic objects that have been identified. The calculation is performed from a bird's-eye view on a plane parallel to the ground (usually corresponding to the XOZ plane of the camera coordinate system). Therefore, only the X and Z coordinates of the semantic objects are used in the calculation.

[0077] Take the (X, Z) of the semantic object as the center of the circle, and use the projection length of the line connecting the semantic object and the depth camera on the XOZ plane of the camera coordinate system, that is, the distance between the semantic object and the depth camera in a plane parallel to the ground as the radius, and draw two circles, such as Figure 5 As shown in the red circle.

[0078] The projection length of the line connecting the semantic object and the depth camera on the XOZ plane of the camera coordinate system is calculated as follows: During the navigation process, the target detection model identifies the two-dimensional detection frame of the semantic object, determines the position of the representative point of the two-dimensional detection frame, and calculates the three-dimensional coordinates of the representative point in the camera coordinate system in combination with the depth image of the depth camera as the position of the identified semantic object in the camera coordinate system (Xcamera, Ycamera, Zcamera). Then, the projection length d of the line connecting the object and the depth camera on the camera's XOZ plane is calculated using the Euclidean formula. XZ : ; Where Xcamera and Zcamera are the X and Z coordinates of the semantic object in the camera coordinate system.

[0079] Furthermore, the intersection of the two circles drawn satisfies the following formula: ; Where, (X A, Z A ) and (X B , Z B ) is the coordinate of the two consecutive semantic objects A and B in the semantic grid map, r A and r B is the radius of the circle drawn by semantic objects A and B.

[0080] The two circles drawn intersect at two points C1 and C2, one of which is the real map coordinate of the camera.

[0081] Determine a first vector pointing from one semantic object to another semantic object, and a second vector pointing from the one semantic object to an intersection point; Calculating the vector cross product of the first vector and the second vector, filtering out an intersection point representing the depth camera from the two intersection points based on the right-hand rule, and using the filtered intersection coordinates as the position of the depth camera in the semantic grid map; In order to accurately find the intersection point corresponding to the depth camera among the two intersection points, the right-hand rule is used for screening in this embodiment.

[0082] Specifically, for two consecutive semantic objects, the semantic object on the left is identified as object A, the semantic object on the right is identified as object B, and the intersection is identified as point C.

[0083] Then the vector pointing from object A to object B is is the first vector ( ), the vector of object A pointing to the intersection point C is the second vector ( ).

[0084] Then calculate the cross product of the first vector and the second vector: ; In this case, the two intersection points and Substituting the coordinates of and into the right-hand rule, there must be a cross product result greater than zero at one intersection and a cross product result less than zero at another intersection. The intersection corresponding to the cross product result greater than zero is determined as the intersection corresponding to the depth camera, and the coordinates of the intersection are used as the position of the depth camera in the semantic grid map.

[0085] like Figure 6 As shown, the green and blue dots represent two consecutive semantic objects A and B, respectively. Indicates the projection length of A from the depth camera on the camera XOZ plane, that is, the radius of circle A; Indicates the continuous projection length of B from the depth camera on the camera XOZ plane, which is the radius of circle B; circle A and circle B intersect at two points and , calculate the vector and The cross product of is the actual depth camera position, are the intersection points that are screened out.

[0086] In another feasible implementation, the vector pointing from object B to object A can be used as the first vector, and the vector pointing from object B to intersection C can be used as the second vector. In this case, the intersection where the cross product result is less than zero is used as the position of the depth camera.

[0087] In this embodiment, Figure 5 The black dot at the intersection of the red circle is the intersection point representing the depth camera position selected by the above method.

[0088] The orientation of the depth camera is determined based on the position of the depth camera in the semantic grid map and the position of the semantic object in the camera coordinate system.

[0089] After determining the position of the depth camera, the orientation of the depth camera can be further determined based on the relative position relationship between the depth camera and any semantic object.

[0090] In this embodiment, the calculation is still performed by taking the semantic object A on the relatively left side as an example.

[0091] First, determine the angle between the semantic object A and the Z axis of the camera coordinate system : ; Where, 、 is the coordinate of semantic object A in the XOZ plane of the camera coordinate system.

[0092] Secondly, calculate the angle between the line connecting the semantic object A and the depth camera and the X-axis of the semantic grid map : ; Where, and is the coordinate of semantic object A in the semantic grid map, and is the coordinate of the depth camera in the semantic grid map.

[0093] The angle representing the camera position relative to the semantic grid map coordinate system. Therefore, the host adds these two angles to obtain the camera's heading angle. : ; Through the above method, the position and orientation of the depth camera in the semantic grid map can be determined in sequence, realizing robot positioning detection integrating semantic information.

[0094] In the robot positioning verification and correction method based on the semantic grid map of the present invention, after the step of determining the position and orientation of the robot in the semantic grid map based on the position and orientation of the depth camera in the semantic grid map, the method further includes: Comparing the position of the robot in the semantic grid map calculated based on the semantic object with the published robot position to obtain a position difference; Comparing the orientation of the robot in the semantic grid map calculated based on the semantic object with the published orientation of the robot to obtain an orientation difference; In this embodiment, the mapping task is completed in advance by Cartographer, so the published robot position is the current posture of the robot published by Cartographer.

[0095] The robot's posture includes the robot's position and orientation. Figure 5 As shown, Figure 5 The green dots and green arrows in the figure represent the published robot positions and orientations, and the black dots and red arrows represent the calculated robot positions and orientations.

[0096] It is understandable that the robot position and orientation calculated based on the landmark information of the semantic object is more accurate robot pose information. Therefore, the calculated robot position is compared with the published robot position to obtain the position difference (specifically, the distance between the two coordinates can be calculated).

[0097] Compare the calculated robot orientation with the published robot orientation to obtain the orientation difference (specifically, the angle difference).

[0098] In a case where the position difference and / or the orientation difference is greater than corresponding preset thresholds, the published robot position and orientation are corrected using the position and orientation of the robot in the semantic grid map.

[0099] Based on the actual calculation method of the position difference and the orientation difference, corresponding preset thresholds, such as a distance threshold and an angle threshold, are determined in advance according to experience.

[0100] If the position difference and / or orientation difference is greater than the corresponding preset threshold, the robot is considered to have lost its positioning, the positioning verification mechanism is activated, the positioning loss detection and emergency braking functions are activated, and the published robot position and orientation are updated to the calculated position and orientation to ensure that the robot can restore accurate positioning. After restoring positioning, it can continue with subsequent navigation tasks to achieve robot positioning correction.

[0101] Specifically, the system uses the calculated position and orientation of the semantic object as pose information, executes Cartographer's relocalization algorithm, and performs positioning correction.

[0102] In a specific embodiment, Figure 7 As shown, Figure 7 (a) indicates that the robot has lost its positioning and some semantic objects are outside the target area; Figure 7 (b) is the image after the robot posture is corrected using semantic objects. It is not difficult to find that some semantic objects outside the target area have basically returned to the correct position.

[0103] The following describes the robot positioning verification and correction device based on the semantic grid map provided by the present invention. The robot positioning verification and correction device based on the semantic grid map described below and the robot positioning verification and correction method based on the semantic grid map described above can be referenced to each other.

[0104] like Figure 8 As shown, the robot positioning verification and correction device based on the semantic grid map includes an identification module 801, a determination module 802, a positioning module 803 and a correction module 804; The recognition module 801 is configured to receive real-time image data acquired by a depth camera on the robot, and to recognize at least two continuous semantic objects based on the real-time image data; It is understandable that the robot relies on a pre-built map of the target area during the positioning and navigation process.

[0105] Typically, a robot moves in a target area beforehand, using its mounted LiDAR sensor to obtain environmental information about the target area. This raster map is then constructed and used for subsequent positioning and navigation. However, this raster map lacks a deep understanding of the environmental semantics and fails to integrate multi-source information about the target area, which can easily lead to positioning errors during the robot's positioning and navigation process.

[0106] To this end, in this embodiment, during the robot's mapping process, various objects in the target area's environment are identified, and the three-dimensional coordinates and category of each identified object are obtained as the object's semantic information. This object's semantic information is then packaged to generate the object's landmark information, i.e., the object's landmark information. This object's landmark information is then published on a two-dimensional grid map constructed based on LiDAR data, resulting in a two-dimensional grid map containing semantic object landmark information, known as a semantic grid map. Objects identified and published with landmark information during the mapping process are considered semantic objects.

[0107] In a feasible implementation, the constructed two-dimensional grid map is as follows: Figure 2 As shown, Figure 3 for Figure 2 The corresponding partial enlarged image of the semantic grid map, Figure 3 The text labels such as Exit_4 and Carefully_3 are landmark information published on the two-dimensional raster map, which together with the two-dimensional raster map constitute a semantic raster map.

[0108] On this basis, during the robot navigation process, the robot's main control computer receives real-time image data obtained by a depth camera installed on the robot to identify at least two continuous semantic objects from a frame of real-time image obtained during the navigation process.

[0109] In a specific embodiment, the host computer uses real-time image data as input to the target detection model, and based on the target detection model, realizes detection and recognition of semantic objects in the same way as in the mapping process.

[0110] Specifically, since the robot's mapping process also requires detecting and identifying objects in the environment, thereby generating and publishing landmark information for these objects, the robot's navigation process uses the same method to identify semantic objects in the target area, thereby representing the 3D coordinates of semantic objects at the same location during both mapping and navigation.

[0111] For example, during the mapping process, the target detection model generates a target detection frame for each identified object, uses the center position of the target detection frame to represent the position of the object, and determines the three-dimensional coordinates of the object; during the navigation process, the same target detection model is used, and the center position of the target detection frame is also used to represent the position of the semantic object, and the coordinates of the identified semantic object in the semantic grid map are determined.

[0112] On this basis, the main control machine combines the landmark information of each semantic object pre-stored in the semantic grid map, and performs multiple constraint matching according to the category, spatial relationship and confidence of the semantic objects identified during the navigation process, thereby realizing object detection and recognition based on landmark information in the semantic grid map.

[0113] That is, the landmark information corresponding to the current two continuous semantic objects is identified in the semantic grid map, and the three-dimensional coordinates in the landmark information are used as the three-dimensional coordinates of the determined semantic object in the semantic grid map to realize the utilization of high-precision object positioning prior data.

[0114] Optionally, the target detection model is a YOLO model.

[0115] Optionally, the master machine implements multi-constraint matching based on the consistency of object categories, the matching degree of spatial relationships, and a confidence threshold.

[0116] In a specific embodiment, during navigation, two consecutive semantic objects identified by the target detection model are Figure 4 shown.

[0117] A determination module 802 is configured to determine positions of at least two semantic objects obtained by identification in a semantic grid map, and determine a position and orientation of the depth camera in the semantic grid map based on the positions of the at least two semantic objects in the semantic grid map; In this embodiment, when at least two consecutive semantic objects can be identified in one frame of image, the positioning of the robot is verified based on the relative positions of the at least two semantic objects and the camera.

[0118] Specifically, when observing two semantic objects and the depth camera from a bird's-eye view, the depth camera must be located at one of the intersections of two circles with the semantic object as the center and the projection distance of the line connecting the semantic object and the depth camera on the horizontal plane as the radius.

[0119] Based on this, the projected distance between the depth camera and each semantic object on the horizontal plane is calculated based on the real-time data acquired by the depth camera. This distance is used as the radius of the circle. A circle is then drawn with the (x, y) coordinates of the 3D landmark information corresponding to the semantic object as the center. The intersection of the two circles is the location of the depth camera.

[0120] Furthermore, if there are three semantic objects, a circle may be drawn for the third semantic object in the same manner, and the common intersection of the three circles is determined as the position of the depth camera.

[0121] If there are only two semantic objects, we can also introduce the actual orientation relationship between the two semantic objects. For example, semantic object A is located on the left side of semantic object B. Based on this orientation relationship, we can verify the two intersection points and select the intersection point representing the position of the depth camera. Its three-dimensional coordinates can be used as the position of the depth camera in the semantic grid map.

[0122] The depth camera position calculated in the above manner effectively utilizes the three-dimensional coordinates pre-stored in the landmark information of the semantic object as prior information, and uses it as the basis for calculating the depth camera position, effectively improving the accuracy of depth camera positioning.

[0123] On this basis, the orientation of the depth camera in the semantic grid map can be calculated based on the angle between the camera and any semantic object, and the angle between the semantic object and the X-axis of the semantic grid map coordinate system.

[0124] A positioning module 803 is configured to determine the position and orientation of the robot in the semantic grid map based on the position and orientation of the depth camera in the semantic grid map, so as to complete the positioning verification of the robot; Since the relative position relationship between the depth camera and the robot is fixed, once the position and orientation of the depth camera in the semantic grid map are determined, the position and orientation of the robot in the semantic grid map can be determined through coordinate transformation.

[0125] a correction module 804 for correcting the published robot position and orientation using the robot's position and orientation in the semantic grid map; Among them, the semantic object is an object that the robot recognizes and publishes landmark information during the mapping process, and the semantic grid map is a two-dimensional grid map constructed by the robot that contains semantic object landmark information. The landmark information includes the three-dimensional coordinates and category of the semantic object.

[0126] Based on the actual calculation method of the position difference and the orientation difference, corresponding preset thresholds, such as a distance threshold and an angle threshold, are determined in advance according to experience.

[0127] If the position difference and / or orientation difference is greater than the corresponding preset threshold, the robot is considered to have lost its positioning, the positioning verification mechanism is activated, the positioning loss detection and emergency braking functions are activated, and the published robot position and orientation are updated to the calculated position and orientation to ensure that the robot can restore accurate positioning. After restoring positioning, it can continue with subsequent navigation tasks to achieve robot positioning correction.

[0128] The present invention identifies the position and category of objects during the robot mapping process, generates landmark information and publishes it on a two-dimensional grid map, thereby obtaining a semantic grid map with object semantic information for robot navigation. In this way, robot positioning verification is achieved based on semantic objects during the robot positioning process, and high-precision robot positioning is achieved by utilizing the prior position information of semantic objects, thereby realizing the integrated use of multi-source environmental information.

[0129] Figure 9 An example of a physical structure diagram of an electronic device is shown below. Figure 9As shown, the electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940, wherein the processor 910, the communication interface 920, and the memory 930 communicate with each other via the communication bus 940. The processor 910 may call the logic instructions in the memory 930 to execute a robot positioning verification and correction method based on a semantic grid map, the method comprising: receiving real-time image data acquired by a depth camera on the robot, and identifying at least two continuous semantic objects based on the real-time image data; determining the positions of the two identified semantic objects in the semantic grid map, and determining the position and orientation of the depth camera in the semantic grid map based on the positions of the at least two semantic objects in the semantic grid map; and determining the position and orientation of the robot in the semantic grid map based on the position and orientation of the depth camera in the semantic grid map to complete the positioning verification of the robot.

[0130] Furthermore, the logic instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0131] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the robot positioning verification and correction method based on the semantic grid map provided by the above methods, the method including: receiving real-time image data acquired by a depth camera on the robot to identify at least two continuous semantic objects based on the real-time image data; determining the positions of the two identified semantic objects in the semantic grid map, and determining the position and orientation of the depth camera in the semantic grid map based on the positions of at least two semantic objects in the semantic grid map; determining the position and orientation of the robot in the semantic grid map based on the position and orientation of the depth camera in the semantic grid map to complete the positioning verification of the robot.

[0132] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the robot positioning verification and correction method based on the semantic grid map provided by the above-mentioned methods, the method comprising: receiving real-time image data acquired by a depth camera on the robot to identify at least two continuous semantic objects based on the real-time image data; determining the positions of the two identified semantic objects in the semantic grid map, and determining the position and orientation of the depth camera in the semantic grid map based on the positions of the at least two semantic objects in the semantic grid map; determining the position and orientation of the robot in the semantic grid map based on the position and orientation of the depth camera in the semantic grid map to complete the positioning verification of the robot.

[0133] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0134] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A robot positioning verification and correction method based on semantic grid map, characterized in that: include: receiving real-time image data acquired by a depth camera on the robot, and identifying at least two continuous semantic objects based on the real-time image data; Determining positions of the two identified semantic objects in the semantic grid map, and determining a position and orientation of the depth camera in the semantic grid map based on the positions of at least two semantic objects in the semantic grid map; Determining the position and orientation of the robot in the semantic grid map based on the position and orientation of the depth camera in the semantic grid map to complete the positioning verification of the robot; Among them, the semantic object is an object that the robot recognizes and publishes landmark information during the mapping process, and the semantic grid map is a two-dimensional grid map constructed by the robot that contains semantic object landmark information. The landmark information includes the three-dimensional coordinates and category of the semantic object.

2. The robot positioning verification and correction method based on semantic grid map according to claim 1 is characterized in that: Before the step of receiving real-time image data acquired by the depth camera on the robot, the method further includes: Control the robot to move in the target area, and receive environmental information of the target area obtained by the lidar sensor on the robot during the movement to build a two-dimensional grid map of the target area; During the movement process, the target detection model is used to identify objects in the environment, determine the three-dimensional coordinates and categories of the objects, and generate landmark information corresponding to the objects. During the mapping process, the landmark information is published in the two-dimensional grid map, so that after the mapping is completed, a two-dimensional grid map containing landmark information of the target area is obtained as the semantic grid map.

3. The robot positioning verification and correction method based on semantic grid map according to claim 2 is characterized in that: The step of using the target detection model to identify objects in the environment during movement, determining the three-dimensional coordinates and categories of the objects, and generating landmark information corresponding to the objects specifically includes: Identify objects in the environment using a target detection model and obtain a two-dimensional detection frame of the object and the category of the object; Selecting the center point of the two-dimensional detection frame of the object as a representative point representing the position of the object, and mapping the representative point to the depth map obtained by the depth camera to obtain the three-dimensional coordinates of the object based on the depth information; The landmark information of the object is generated based on the three-dimensional coordinates of the object and the category of the object.

4. The robot positioning verification and correction method based on semantic grid map according to claim 2 is characterized in that: Before the step of receiving real-time image data acquired by the depth camera on the robot, the method further includes: During the mapping process, loop closure optimization is used to correct and update the landmark information corresponding to the object.

5. The robot positioning verification and correction method based on semantic grid map according to any one of claims 1 to 4, characterized in that: The step of determining the position and orientation of the depth camera in the semantic grid map according to the positions of at least two semantic objects in the semantic grid map specifically includes: Draw two circles with the position of the semantic object in the semantic grid map as the center and the projection length of the line from the semantic object to the depth camera on the XOZ plane of the camera coordinate system as the radius, and determine the coordinates of the two intersection points between the two circles; Determine a first vector pointing from one semantic object to another semantic object, and a second vector pointing from the one semantic object to an intersection point; Calculating the vector cross product of the first vector and the second vector, filtering out an intersection point representing the depth camera from the two intersection points based on the right-hand rule, and using the filtered intersection coordinates as the position of the depth camera in the semantic grid map; The orientation of the depth camera is determined based on the position of the depth camera in the semantic grid map and the position of the semantic object in the camera coordinate system.

6. The robot positioning verification and correction method based on semantic grid map according to any one of claims 1 to 4, characterized in that: After the step of determining the position and orientation of the robot in the semantic grid map based on the position and orientation of the depth camera in the semantic grid map, the method further includes: Comparing the position of the robot in the semantic grid map calculated based on the semantic object with the published robot position to obtain a position difference; Comparing the orientation of the robot in the semantic grid map calculated based on the semantic object with the published orientation of the robot to obtain an orientation difference; In a case where the position difference and / or the orientation difference is greater than corresponding preset thresholds, the published robot position and orientation are corrected using the position and orientation of the robot in the semantic grid map.

7. A robot positioning verification and correction device based on semantic grid map, characterized in that: include: A recognition module, configured to receive real-time image data acquired by a depth camera on the robot, and to recognize at least two continuous semantic objects based on the real-time image data; a determination module, configured to determine positions of at least two semantic objects obtained by identification in a semantic grid map, and determine a position and orientation of the depth camera in the semantic grid map according to the positions of the at least two semantic objects in the semantic grid map; A positioning module, configured to determine the position and orientation of the robot in the semantic grid map based on the position and orientation of the depth camera in the semantic grid map, so as to complete the positioning verification of the robot; a correction module, configured to correct the published position and orientation of the robot using the position and orientation of the robot in the semantic grid map; Among them, the semantic object is an object that the robot recognizes and publishes landmark information during the mapping process, and the semantic grid map is a two-dimensional grid map constructed by the robot that contains semantic object landmark information. The landmark information includes the three-dimensional coordinates and category of the semantic object.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the robot positioning verification and correction method based on the semantic grid map as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the robot positioning verification and correction method based on the semantic grid map as described in any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the robot positioning verification and correction method based on the semantic grid map as described in any one of claims 1 to 6 is implemented.