Multi-sensor fusion positioning method and system based on semantic instance

By constructing semantic dimension chains and multi-sensor fusion positioning methods, the problem of insufficient positioning accuracy and robustness of traditional SLAM in dynamic environments is solved, and high precision and high robustness positioning in complex environments is achieved.

CN120403595APending Publication Date: 2025-08-01WUHAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510462296.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Traditional SLAM methods are difficult to accurately extract stable semantic features in dynamic environments, resulting in reduced positioning accuracy and insufficient robustness, and lack of deep understanding of semantic information in the environment, making it difficult to establish stable pose references in scenarios with highly similar or structural changes.

Method used

The multi-sensor fusion positioning method based on semantic instances is adopted, and the semantic dimension chain is obtained, the corner category is determined, the dynamic obstacle laser points are eliminated, the local search window is constructed, and the robot's precise positioning is obtained, and the multi-sensor cumulative error compensation mechanism and abduction detection mechanism are used to achieve high robust global positioning.

Benefits of technology

Achieve accurate global positioning in milliseconds in dynamic and degraded environments, eliminate dynamic obstacles, eliminate cumulative errors of odometers and IMUs, quickly restore robot posture, and improve positioning accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120403595A_ABST
    Figure CN120403595A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic instance-based multi-sensor fusion positioning method and system, a storage medium and electronic equipment. The method comprises the following steps: acquiring a semantic instance map; extracting a boundary contour and a semantic instance of the semantic instance map, and constructing a semantic size chain based on the boundary contour and the semantic instance; determining wall corner categories; and performing semantic attitude estimation on the candidate region in the semantic size chain to obtain a prior pose of the robot, removing dynamic obstacle laser points, and constructing a local search window around the prior pose to obtain an accurate pose of the robot. According to the invention, the positioning performance of the robot in a complex environment can be improved, semantic information can be fused, and the robustness and real-time performance of positioning are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of indoor mobile robot mapping, positioning and navigation, and in particular to a multi-sensor fusion positioning method, system, storage medium and electronic device based on semantic instances. Background Art

[0002] In the field of indoor robot navigation and localization, the challenges brought by dynamic environments and sensor degradation have always been a research hotspot.

[0003] Traditional SLAM methods often suffer from reduced accuracy and insufficient robustness when dealing with dynamic obstacles, sensor noise, and accumulated errors. While multi-sensor fusion methods based on lidar, IMU, and RGB-D cameras can improve environmental perception to a certain extent, accurately extracting stable semantic features in dynamic scenes and combining them with historical information for highly robust global positioning remains a challenge.

[0004] In addition, existing positioning methods generally rely on geometric features and lack a deep understanding of semantic information in the environment, making it difficult for robots to establish stable pose references in scenes that are highly similar or have large structural changes. Summary of the Invention

[0005] The embodiments of the present application provide a multi-sensor fusion positioning method, system, storage medium and electronic device based on semantic instances, which can improve the positioning performance of robots in complex environments, and can fuse semantic information to improve positioning accuracy.

[0006] The present invention provides a multi-sensor fusion positioning method based on semantic instances, including: Get semantic instance map; Extracting boundary contours and semantic instances of the semantic instance map, and constructing a semantic dimension chain based on the boundary contours and the semantic instances; Determine the corner category; Semantic pose estimation is performed on the candidate areas in the semantic dimension chain to obtain the robot's prior pose, dynamic obstacle laser points are eliminated, and a local search window is constructed around the prior pose to obtain the robot's precise pose.

[0007] Furthermore, the multi-sensor fusion positioning method based on semantic instances, wherein the step of extracting the boundary contours and semantic instances of the semantic instance map and constructing a semantic dimension chain based on the boundary contours and the semantic instances, includes: Denoising the semantic instance map and filling semantic gaps through a morphological expansion operation; Extract the boundary contours and semantic instances of the semantic instance map after filling the semantic gap, explore a clockwise or counterclockwise path from each of the boundary contours, and calculate the boundary points of each semantic instance that are closest to the boundary contour; Sort the semantic instances according to the order of the boundary points, and calculate the category and distance relationship between the upstream and downstream semantic instances to construct a semantic dimension chain.

[0008] Further, in the above multi-sensor fusion localization method based on semantic instances, wherein, the determining the corner category includes: When a corner is detected within a preset heading angle range, determine the corner category based on the heading angle range; When the heading angle exceeds the preset heading angle, calculate the direction angle of the corner to determine the corner category; Verify the corner category based on the structural constraint theory.

[0009] Further, in the above multi-sensor fusion localization method based on semantic instances, wherein, the calculating the direction angle to determine the corner category when the heading angle exceeds the preset heading angle includes: Calculate the direction angles of the left and right walls of the corner and :

[0010] Wherein, represents the heading angle, and respectively represent the angles between the heading of the robot and the left and right walls of the concave angle of is 1 for the concave angle and is -1 for the convex angle, and are parameters for angle adjustment:

[0011]

[0012] Wherein, is the direction angle of the corner, represents the maximum vertical error.

[0013] Further, in the above multi-sensor fusion localization method based on semantic instances, wherein, the performing semantic pose estimation on the candidate regions in the semantic dimension chain to obtain the prior pose of the robot includes: Determine the candidate regions in the semantic dimension chain according to the category of the detection frame; Calculate the position matching degree scores between all instance sets in each of the candidate regions; Perform semantic pose estimation on the candidate region with the highest position matching score to obtain the prior pose of the robot.

[0014] Furthermore, in the above multi-sensor fusion localization method based on semantic instances, where the semantic pose estimation is performed on the candidate region with the highest position matching score to obtain the prior pose of the robot, which is calculated by the following formula:

[0015]

[0016] Where, is the heading angle, is the x coordinate of the center of the semantic feature in the detection box, is the width of the camera image, represents the horizontal field of view of the camera, ([[]] xc , yc ), ( xr , yr ) and ( x 3, y 3) represent the third element on the semantic dimension chain of the camera coordinates, the robot, and the map, d A is the depth value of semantic feature A, a constant and represent the distance and angle between the camera and the robot coordinate system.

[0017] Furthermore, in the above multi-sensor fusion localization method based on semantic instances, where the laser points of dynamic obstacles are removed, including: Segment the laser point cloud data based on the clustering algorithm of Euclidean distance to obtain multiple clusters; Calculate the centroid of each cluster; Calculate the position change value of the centroid on two consecutive frames, calculate the velocity vector of the corresponding cluster based on the position change value, if the velocity vector exceeds the preset velocity threshold, then the cluster is the laser point of the dynamic obstacle, and the cluster is removed.

[0018] Furthermore, in the above multi-sensor fusion localization method based on semantic instances, where a local search window is constructed around the prior pose to obtain the accurate pose of the robot, including: Form a search window around the prior pose W :

[0019]

[0020]

[0021] Among them, , , respectively represent the search steps in the x, y, directions, , , and are preset hyperparameters, represents the total set of candidate poses, ([[]] ) represents a certain candidate pose in the set ; Calculate the scores of each candidate pose in the search window:

[0022] Among them, the matrix represents converting the laser point into the mapped coordinate system, represents the probability value of retrieving the corresponding grid point.

[0023] Furthermore, for the above-mentioned multi-sensor fusion localization method based on semantic instances, after the step of obtaining the accurate pose of the robot, it includes: Obtain odometer change data and IMU data; Predict the next pose based on the accurate pose, the odometer change data, and the IMU data; Update the pose based on the laser information and correct the errors of the odometer change data and the IMU data.

[0024] The embodiment of the present application also provides a multi-sensor fusion localization system based on semantic instances, including: An acquisition module for acquiring a semantic instance map; A semantic dimension chain construction module for extracting the boundary contour and semantic instances of the semantic instance map and constructing a semantic dimension chain based on the boundary contour and the semantic instances; A corner classification module for determining the corner category; A pose matching module for performing semantic pose estimation on the candidate regions in the semantic dimension chain to obtain the prior pose of the robot, removing the laser points of dynamic obstacles, and constructing a local search window around the prior pose to obtain the accurate pose of the robot.

[0025] The embodiment of the present application also provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded by a processor to execute any one of the above-mentioned multi-sensor fusion localization methods based on semantic instances.

[0026] An embodiment of the present application also provides an electronic device, including a processor and a memory. The processor is electrically connected to the memory. The memory is used to store instructions and data, and the processor is used for the steps in the multi-sensor fusion positioning method based on semantic instances described in any one of the above.

[0027] The multi-sensor fusion positioning method, system, storage medium and electronic device provided by the present application model the prior semantic instance map using a path exploration model and construct a semantic size chain. In terms of global positioning, a robust semantic size chain pre-matching algorithm and a reasonable matching score mechanism can quickly obtain the prior pose, and dynamic obstacle laser points are removed through a continuous frame clustering algorithm. When this algorithm is integrated with the multi-resolution laser Scan-to-map matching technology, even in dynamic and degraded environments, millisecond-level accurate global positioning can be achieved. In addition, during the long-term positioning process, the present application uses a multi-sensor cumulative error compensation mechanism and a kidnapping detection mechanism to effectively eliminate the cumulative errors of the odometer and IMU, and can also accurately detect abnormal robot positioning and quickly recover the pose. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The technical solutions and other beneficial effects of the present application will become obvious by describing the specific embodiments of the present application in detail with reference to the accompanying drawings.

[0029] Figure 1 It is a flowchart of the multi-sensor fusion positioning method based on semantic instances provided by an embodiment of the present application.

[0030] Figure 2 It is another flowchart of the multi-sensor fusion positioning method based on semantic instances provided by an embodiment of the present application.

[0031] Figure 3 It is a flowchart of constructing a semantic size chain provided by an embodiment of the present application. <W

[0032] Figure 4 It is a semantic instance map and a semantic size chain constructed in scenario A provided by an embodiment of the present application.

[0033] Figure 5 It is a semantic instance map and a semantic size chain constructed in scenario B provided by an embodiment of the present application.

[0034] Figure 6 It is a schematic diagram for determining the corner direction provided by an embodiment of the present application.

[0035] Figure 7 It is a schematic diagram of a multi-corner classification strategy based on structural constraints provided by an embodiment of the present application.

[0036] Figure 8Schematic diagram of semantic dimension chain pre-matching provided by an embodiment of the present application.

[0037] Figure 9 Flowchart of multi-resolution laser Scan-to-map matching provided by an embodiment of the present application.

[0038] Figure 10 Flowchart of global positioning provided by an embodiment of the present application.

[0039] Figure 11 Schematic diagram of the structure of a multi-sensor fusion positioning system based on semantic instances provided by an embodiment of the present application.

[0040] Figure 12 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0041] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0042] To solve the above problems, an embodiment of the present application provides a multi-sensor fusion positioning method, system, storage medium, and electronic device based on semantic instances. A multi-sensor fusion positioning system based on semantic instances provided by an embodiment of the present application can be integrated into an electronic device, and the electronic device can be a device such as a terminal, a server, etc. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a micro-processing box, or other devices, etc.

[0043] Please refer to Figure 1 And Figure 2 , Figure 1 Flowchart of the multi-sensor fusion positioning method based on semantic instances provided by an embodiment of the present application. Figure 2 Another flowchart of the multi-sensor fusion positioning method based on semantic instances provided by an embodiment of the present application, which is applied to an electronic device. The multi-sensor fusion positioning method based on semantic instances includes the following steps: S1, obtain a semantic instance map.

[0044] S2, extract the boundary contour and semantic instances of the semantic instance map, and construct a semantic dimension chain based on the boundary contour and semantic instances.

[0045] In one embodiment, step S2 includes the following steps: S21, Denoise the semantic instance map and fill the semantic gaps through morphological expansion operations; Specifically, first extract the pixel sets of all semantic instances , and calculate the average distance :

[0046] where is the coordinate of the i-th pixel, and n is the number of pixels.

[0047] Then, delete from the map those semantic instance features whose average distance exceeds the threshold :

[0048] where represents the geometric diameter of the semantic instance on the horizontal plane, is a scaling factor.

[0049] Next, discard the pixels with a distance greater than from the remaining semantic instances , as well as the pixels whose 8-neighborhoods are all empty. Finally, semantic feature expansion and gap filling are achieved through morphological expansion operations.

[0050] S22, Extract the boundary contours and semantic instances of the semantic instance map after filling the semantic gaps, explore a clockwise or counterclockwise path from each boundary contour, and calculate the boundary point of each semantic instance closest to the boundary contour; S23, Sort the semantic instances according to the boundary point order, and calculate the category and distance relationships of upstream and downstream semantic instances to construct a semantic dimension chain.

[0051] Figure 3 This is the flowchart for constructing a semantic dimension chain provided by the embodiments of the present application, as shown in Figure 3As shown, steps (a - e) detail the process of finding a closed - loop contour. The red color represents unvisited boundary points, the green color represents visited boundary points, the blue color represents the current position, the purple color represents invalid path points deleted by the branching mechanism, S and E represent the starting and ending points of the path, F1 is the branch point closest to the current position, the yellow arrow points to the next path point, (c) demonstrates a scenario where the next point cannot be found, thus prompting a return to the nearest branch point, and (f) shows the complete boundary closed - loop path and the semantic dimension chain. First, sort the map boundary contours by area to determine their order. For each contour, the point closest to the map origin is designated as the starting point. Then, by selecting adjacent unmarked points, when marking the current point, the closed - loop path is traced clockwise. If multiple unmarked points are found and not all of them are horizontally or vertically adjacent, the current point is also marked as a branch point. If the next point is not found, the algorithm backtracks to the nearest branch point and deletes the invalid path segments. Repeat this process until the end point is reached. Finally, according to the nearest boundary points, semantic instance objects are organized in a clockwise order along the map boundary. Figure 4 The semantic instance map and semantic dimension chain constructed in scenario A provided by the embodiment of the present application Figure 5 The semantic instance map and semantic dimension chain constructed in scenario B provided by the embodiment of the present application.

[0052] S3. Determine the corner category.

[0053] Specifically, the indoor corner category is determined by the corner concavity - convexity (A, V) and corner directionality (1, 2, 3, 4), that is, the corner category includes A1, A2, A3, A4, V1, V2, V3, V4, a total of 8 categories.

[0054] The corner concavity - convexity can be achieved through object detection, and the corner directionality can be achieved through the following strategies (one step corresponds to one strategy): S31. When a corner is detected within a preset heading - angle range, determine the corner category based on the heading - angle range.

[0055] Specifically, strategy 1 is: Considering that the maximum field of view of the camera is 70, when the robot detects a corner within a specific heading - angle range, it can directly determine the category of the corner. Figure 6 The schematic diagram of corner directionality determination (strategy 1 and strategy 2) provided by the embodiment of the present application Figure 6 As shown, assuming the north - facing heading angle is 0°, when the heading angle of the robot is between 3° and 55°, the detected corner must belong to A1 or V1, and when the heading angle of the robot is between 125° and 145°, the detected corner must belong to A4 or V4.

[0056] S32. When the heading angle exceeds the preset heading angle, calculate the direction angle of the corner to determine the corner type.

[0057] Specifically, Strategy 2 is as follows: If the heading angle of the robot exceeds the defined range, determine the type of the corner by calculating its direction angle. First, calculate the direction angles of the left and right walls of the corner. and :

[0058] where represents the heading angle, and respectively represent the angles between the heading of the robot and the left and right walls of the concave corner. is 1 for the concave corner, is -1 for the convex corner, and are parameters for angle adjustment:

[0059]

[0060] where is the direction angle of the corner, represents the maximum vertical error. Considering the sensor error, ~10 degrees are allowed when determining the corner type.

[0061] S33. Verify the corner type based on the structural constraint theory.

[0062] Although Strategy 1 or Strategy 2 can accurately determine the corner type in most cases, the limitations of the vision and laser sensor accuracy may lead to misclassification in extreme cases. Use a multi-corner classification strategy based on structural constraints to further verify the corner type judged by Strategy 1 or Strategy 2. Once the type of one corner and the concave-convex attribute of another corner are determined, and these two corners share the same laser segment, the directional property of the latter can be directly inferred using the structural constraint characteristics of adjacent corners. Figure 7 is a schematic diagram of the multi-corner classification strategy based on structural constraints (Strategy 3) provided by the embodiment of the present application. All 16 possible structural constraint cases are shown in Figure 7 .

[0063] Among the three strategies mentioned above, the robot's heading angle is determined based on the camera's viewing angle range and the conditions of the indoor corners. The corner direction is determined by calculating the direction angles of the two walls that form the corner and taking the average as the direction angle of the corner. Structural constraint judgment is achieved by utilizing the structural constraint characteristics of adjacent corners. If the category of one corner is determined, the concavity of the other is determined, and the two corners share the same wall, the directionality of the other corner can be directly inferred.

[0064] S4, performs semantic pose estimation on the candidate areas in the semantic dimension chain to obtain the robot's prior pose, removes dynamic obstacle laser points, and constructs a local search window around the prior pose to obtain the robot's precise pose.

[0065] In one embodiment, Figure 8 The schematic diagram of semantic dimension chain pre-matching provided in the embodiment of the present application is as follows: Step S4 of performing semantic pose estimation on the candidate region in the semantic dimension chain to obtain the robot's prior pose includes the following steps: S41, determining a candidate region in the semantic size chain according to the category of the detection box.

[0066] S42, calculating the position matching scores between all instance sets in each candidate region.

[0067] Calculated by the following formula:

[0068] in, and They represent the distance between instances calculated in the current frame and the distance between instances of the corresponding category calculated in the semantic dimension chain, respectively.

[0069] S43, performing semantic pose estimation on the candidate region with the highest position matching score to obtain the prior pose of the robot.

[0070] Specifically, the prior pose is calculated by the following formula:

[0071]

[0072] in, is the heading angle, is the x-coordinate of the center of the semantic feature in the detection box, is the width of the camera image, represents the horizontal field of view of the camera, ( xc , yc ), ( xr , yr )and( x 3, y3) Represents the coordinates of the camera, the robot, and the third element on the semantic dimension chain on the map. d A is the depth value of semantic feature A, a constant and Represents the distance and angle between the camera and the robot coordinate system.

[0073] In one embodiment, removing the dynamic obstacle laser points in step S4 includes the following steps: S44, Segment the laser point cloud data using a clustering algorithm based on Euclidean distance to obtain multiple clusters.

[0074] Before laser matching, use a clustering algorithm based on Euclidean distance to segment the laser point cloud to obtain multiple clusters. For each cluster .

[0075] S45, Calculate the centroid of each cluster.

[0076] The calculation method of the centroid is:

[0077] where represents the number of points in the cluster .

[0078] S46, Calculate the position change value of the centroid on two consecutive frames, calculate the velocity vector of the corresponding cluster based on the position change value, and if the velocity vector exceeds a preset velocity threshold, the cluster is a dynamic obstacle laser point and the cluster is removed.

[0079] Specifically, track the position change of the centroid of each cluster on two consecutive frames. Let and respectively represent the centroid positions at times and . Then estimate the velocity vector of the cluster:

[0080] If exceeds a predefined threshold , the cluster will be classified as a dynamic obstacle and subsequently removed from the laser point cloud. Figure 9 This is the multi-resolution laser Scan-to-map matching flowchart provided by the embodiment of the present application. This process is shown in Figure 9 .

[0081] In one embodiment, constructing a local search window around the prior pose in step S4 to obtain the accurate pose of the robot includes the following steps: Obtain the prior pose from semantic matchingx p After that, a search window is formed around it. W :

[0082]

[0083]

[0084] Among them, , , respectively represent the search steps in the x, y, direction, , , and are preset hyperparameters, represents the total set of candidate poses, ([[]] ) represents a certain candidate pose in the set .

[0085] Calculate the scores of each candidate pose in the search window:

[0086] Among them, the matrix represents converting the laser point into the mapping coordinate system, represents the probability value of retrieving the corresponding grid point, and the candidate pose with the highest score is the true pose. Figure 10 This is the global positioning flowchart provided by the embodiment of the present application. As Figure 10 shown, (a) shows the target detection result, (b) represents the depth value after clustering, (c) illustrates the pre-positioning based on the semantic dimension chain, where the magenta rectangle represents the local position search window, and the two red dotted lines represent the pose search window, and (d) represents the best robot pose after the search.

[0087] Furthermore, during the local positioning process of the robot, there will be cumulative errors in the long-term operation of the odometer and IMU sensors. After the step of obtaining the accurate pose of the robot, the following steps are further included: S51, Obtain the odometer change data and IMU data.

[0088] S52, Predict the next pose based on the accurate pose, odometer change data, and IMU data.

[0089] Predict the next pose through the odometer change data and IMU data:

[0090]

[0091] Among them, and are the predicted position and heading at time t+1, represents the initial attitude of global positioning (i.e., the accurate pose obtained in step S4). The transformation matrix projects the odometer change onto the map coordinate system, and represent the compensation values of the position and heading at time t.

[0092] S53. Update the attitude based on the laser information, and correct the errors of the odometer change data and IMU data.

[0093] Subsequently, update the attitude using the laser information and correct the errors of the odometer and IMU:

[0094]

[0095] Among them, , the function generates a local search window W centered on and defines the search range and , the function filters out the area in W that is almost perpendicular to the motion direction, while the function identifies the attitude with the highest score within the search window.

[0096] Furthermore, use the local estimated pose, odometer change, IMU change, and semantic matching pose to perform kidnapping detection. If any of the following conditions are met for m consecutive frames, it is considered that the robot is in an abnormal positioning state and global relocalization is required: (1) The local attitude estimation score is lower than the threshold ; (2) The odometer change amount is extremely small, but the IMU change amount is normal; (3) The semantic pose differs significantly from the pose at the previous moment ; (4) There is odometer movement, but the local estimated pose hardly changes or varies greatly.

[0097]

[0098] Among them , , are hyperparameters, and the superscript c represents the current frame. When is zero and is not zero, the functioncom Returns true, the function is a distance function.

[0099] The present invention proposes a multi-sensor fusion localization method based on semantic instances, which uses image morphological operations to optimize the semantic instance map, remove noise, and constructs a semantic size chain using a path exploration model to orderly associate scattered semantic instance features. During global localization, a robust localization algorithm combining pre-matching based on the semantic size chain and multi-resolution Scan-to-map precise matching is used. Through an efficient and accurate multi-strategy corner classification model and a reasonable matching score mechanism, possible prior poses can be quickly calculated, and dynamic obstacle laser points can be removed through a continuous frame clustering algorithm. Even in dynamic and degraded environments, millisecond-level accurate global localization can be achieved. At the same time, during long-term localization, a multi-sensor cumulative error compensation mechanism and a kidnapping detection mechanism are fully utilized, which not only effectively eliminates the long-term cumulative errors of the odometer and IMU, but also can accurately detect abnormal robot localization in real time and quickly recover the pose.

[0100] According to the method described in the above embodiments, this embodiment will further describe from the perspective of a multi-sensor fusion localization system based on semantic instances. The multi-sensor fusion localization system based on semantic instances can be specifically implemented as an independent entity or integrated in an electronic device. The electronic device can be a terminal, a server, or other devices. Among them, the terminal can include a tablet computer, a laptop computer, a personal computer (PC), a microprocessing box, or other devices, etc.

[0101] Please refer to Figure 11 , Figure 11 which specifically describes the multi-sensor fusion localization system based on semantic instances provided in the embodiments of the present application, applied in an electronic device. The multi-sensor fusion localization system based on semantic instances can include: An acquisition module, configured to acquire a semantic instance map; A semantic size chain construction module, configured to extract the boundary contour and semantic instances of the semantic instance map, and construct a semantic size chain based on the boundary contour and the semantic instances; A corner classification module, configured to determine the corner category; A pose matching module, configured to perform semantic pose estimation on candidate regions in the semantic size chain to obtain the prior pose of the robot, remove dynamic obstacle laser points, and construct a local search window around the prior pose to obtain the accurate pose of the robot.

[0102] The hardware platform relied on by this system includes: a main control computer, a robot platform chassis, a depth camera, an IMU, and a lidar. The main control computer is sequentially connected to the robot platform chassis, the depth camera, the IMU, and the lidar.

[0103] In one embodiment, the main control computer selected is the Intel NUC6I7KYK mini computer; the robot platform chassis selected is the drive board chassis of the Lunqu brand; the depth camera selected is the Microsoft Kinect V2; the IMU selected is the HWT605 of Weite Intelligence; the lidar sensor selected is the SICK lms111 lidar with stable performance.

[0104] In specific implementation, each of the above modules and / or units can be implemented as an independent entity, or can be combined arbitrarily and implemented as the same or several entities. For the specific implementation of each of the above modules and / or units, reference can be made to the method embodiments above. For the beneficial effects that can be specifically achieved, please also refer to the beneficial effects in the method embodiments above, which will not be elaborated here.

[0105] In addition, the embodiment of the present application also provides an electronic device, which can be a device such as a computer or a tablet computer. This electronic device can implement the steps in any of the embodiments of the multi-sensor fusion positioning method based on semantic instances provided by the embodiment of the present application. Therefore, it can achieve the beneficial effects that can be achieved by any of the multi-sensor fusion positioning methods based on semantic instances provided by the embodiments of the present invention. For details, please refer to the previous embodiments, which will not be elaborated here.

[0106] Figure 12 The specific structural block diagram of the electronic device provided by the embodiment of the present invention is shown. This electronic device can be used to implement the multi-sensor fusion positioning method based on semantic instances provided in the above embodiments. The electronic device 500 can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a micro-processing box, or other devices, etc.

[0107] The RF circuit 510 is used to receive and transmit electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices. The RF circuit 510 may include various existing circuit components for performing these functions. For example, an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, a memory, and so on. The RF circuit 510 can communicate with various networks such as the Internet, an enterprise intranet, a wireless network or communicate with other devices through a wireless network. The above-mentioned wireless network may include a cellular phone network, a wireless local area network or a metropolitan area network. The above-mentioned wireless network can use various communication standards, protocols and technologies, including but not limited to the Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as the Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE 802.11g and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging and short messages, and any other suitable communication protocols, and may even include those protocols that have not yet been developed currently.

[0108] The memory 520 can be used to store software programs and modules, such as the corresponding program instructions / modules in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, that is, to implement functions such as taking pictures with the front camera, processing the captured images, and switching the display colors of the display content on the display screen. The memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 520 may further include a memory remotely disposed relative to the processor 580, and these remote memories can be connected to the electronic device 500 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0109] The input unit 530 can be used to receive input digital or character information, and generate a keyboard and a mouse related to user settings and function controls. The display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, and these graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 540 may include a display panel 541. Optionally, the display panel 541 can be configured in the form of an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode).

[0110] The audio circuit 560, the speaker 561, and the microphone 562 can provide an audio interface between the user and the electronic device 500. The audio circuit 560 can transmit the electrical signal converted from the received audio data to the speaker 561, and the speaker 561 converts it into a sound signal for output; on the other hand, the microphone 562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 560 and then converted into audio data. After the audio data is output to the processor 580 for processing, it is sent to another terminal, for example, through the RF circuit 510, or the audio data is output to the memory 520 for further processing. The audio circuit 560 may also include an earphone jack to provide communication between the peripheral earphone and the electronic device 500.

[0111] The electronic device 500 can help the user receive requests, send information, etc. through the transmission module 570 (such as a Wi-Fi module), and it provides the user with wireless broadband Internet access. Although the transmission module 570 is shown in the figure, it can be understood that it does not belong to the essential components of the electronic device 500 and can be omitted completely within the scope of not changing the essence of the invention according to needs.

[0112] Processor 580 is the control center of electronic device 500. It connects all components of the phone using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 520 and accessing data stored in memory 520, it executes various functions of electronic device 500 and processes data, thereby providing overall monitoring of the electronic device. Optionally, processor 580 may include one or more processing cores. In some embodiments, processor 580 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 580.

[0113] Electronic device 500 also includes a power supply 590 (e.g., a battery) for powering various components. In some embodiments, the power supply can be logically connected to processor 580 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. Power supply 590 can also include any components, such as one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0114] Although not shown, the electronic device 500 also includes a camera (such as a front camera and a rear camera), a Bluetooth module, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations: Get semantic instance map; Extracting boundary contours and semantic instances of the semantic instance map, and constructing a semantic dimension chain based on the boundary contours and the semantic instances; Determine the corner category; Semantic pose estimation is performed on the candidate areas in the semantic dimension chain to obtain the robot's prior pose, dynamic obstacle laser points are eliminated, and a local search window is constructed around the prior pose to obtain the robot's precise pose.

[0115] During specific implementation, the above modules can be implemented as independent entities, or can be arbitrarily combined and implemented as the same or several entities. The specific implementation of the above modules can be found in the previous method embodiments and will not be repeated here.

[0116] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. To this end, an embodiment of the present invention provides a storage medium in which multiple instructions are stored. The instructions can be loaded by a processor to execute the steps of any one of the embodiments of the multi-sensor fusion positioning method based on semantic instances provided by the embodiments of the present invention.

[0117] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0118] Since the instructions stored in the storage medium can execute the steps in any one of the embodiments of the multi-sensor fusion positioning method based on semantic instances provided by the embodiments of the present invention, the beneficial effects achievable by any of the multi-sensor fusion positioning methods based on semantic instances provided by the embodiments of the present invention can be realized. See the previous embodiments for details and will not be elaborated here.

[0119] The above has introduced in detail a multi-sensor fusion positioning method, system, storage medium, and electronic device based on semantic instances provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A multi-sensor fusion localization method based on semantic instances, characterized in that The method includes: Obtaining a semantic instance map; Extracting the boundary contour and semantic instances of the semantic instance map, and constructing a semantic dimension chain based on the boundary contour and the semantic instances; Determining the corner category; Performing semantic pose estimation on candidate regions in the semantic dimension chain to obtain the prior pose of the robot, removing laser points of dynamic obstacles, and constructing a local search window around the prior pose to obtain the accurate pose of the robot.

2. The multi-sensor fusion localization method based on semantic instances according to claim 1, wherein The extracting the boundary contour and semantic instances of the semantic instance map, and constructing a semantic dimension chain based on the boundary contour and the semantic instances includes: Denosing the semantic instance map and filling semantic gaps through morphological expansion operations; Extracting the boundary contour and semantic instances of the semantic instance map after filling the semantic gaps, exploring a clockwise or counterclockwise path from each boundary contour, and calculating the boundary point of each semantic instance closest to the boundary contour; Sorting the semantic instances according to the order of boundary points, and calculating the category and distance relationship between upstream and downstream semantic instances to construct a semantic dimension chain.

3. The multi-sensor fusion positioning method based on semantic instances according to claim 1, characterized in that The determining the corner category includes: When detecting an angle within a preset heading angle range, determining the corner category based on the heading angle range; When the heading angle exceeds the preset heading angle, calculating the direction angle of the corner to determine the corner category; Verifying the corner category based on the structural constraint theory.

4. The multi-sensor fusion localization method based on semantic instances according to claim 3, characterized in that The when the heading angle exceeds the preset heading angle, calculating the direction angle to determine the corner category includes: Calculate the direction angles of the left and right walls at the corner and : Among them, represents the heading angle, and respectively represent the angles between the heading of the robot and the left and right walls, is 1 for concave angles, is -1 for convex angles, and are parameters for angle adjustment: where, for the direction angle of the corner, represents the maximum vertical error.

5. The multi-sensor fusion positioning method based on semantic instances according to claim 1, characterized in that The performing semantic pose estimation on candidate regions in the semantic dimension chain to obtain the prior pose of the robot includes: Determining candidate regions in the semantic dimension chain according to the category of the detection frame; Calculating the position matching degree score between all instance sets in each candidate region; Performing semantic pose estimation on the candidate region with the highest position matching degree score to obtain the prior pose of the robot.

6. The multi-sensor fusion positioning method based on semantic instances according to claim 5, wherein, The performing semantic pose estimation on the candidate region with the highest position matching degree score to obtain the prior pose of the robot is calculated by the following formula: Among them, is the heading angle, is the x - coordinate of the semantic feature center in the detection frame, is the width of the camera image, represents the horizontal field of view of the camera, ([[]] xc , yc ), ( xr , yr ) and ( x 3, y ) represent the coordinates of the camera, the robot, and the third element on the semantic dimension chain on the map, d A is the depth value of semantic feature A, a constant and represent the distance and angle between the camera and the robot coordinate systems.

7. The multi-sensor fusion positioning method based on semantic instances according to claim 1, wherein The removing laser points of dynamic obstacles includes: Segmenting the laser point cloud data based on the clustering algorithm based on the Euclidean distance to obtain multiple clusters; Calculating the centroid of each cluster; Calculating the position change value of the centroid on two consecutive frames, calculating the velocity vector of the corresponding cluster based on the position change value, if the velocity vector exceeds the preset velocity threshold, then the cluster is the laser point of the dynamic obstacle, and removing the cluster.

8. The multi-sensor fusion positioning method based on semantic instances according to claim 1, wherein Constructing a local search window around the prior pose to obtain the accurate pose of the robot includes: Form a search window around the prior pose W : Among them, , , respectively represent the search steps in the x, y, directions, , , and are preset hyperparameters, represents the total set of candidate poses, ( ) represents a certain candidate pose in the set ; Calculating the score of each candidate pose in the search window: Among them, the matrix represents converting the laser point into a mapped coordinate system, represents retrieving the probability value of the corresponding grid point.

9. The multi-sensor fusion positioning method based on semantic instances according to claim 1, characterized in that After the step of obtaining the accurate pose of the robot, it includes: Obtaining odometer change data and IMU data; Predicting the next pose based on the accurate pose, the odometer change data and the IMU data; Updating the pose based on the laser information and correcting the errors of the odometer change data and the IMU data.

10. A multi-sensor fusion positioning system based on semantic instances, characterized in that, Includes: An acquisition module for acquiring a semantic instance map; A semantic dimension chain construction module, which is used to extract the boundary contours and semantic instances of the semantic instance map, and construct a semantic dimension chain based on the boundary contours and the semantic instances; A corner classification module, which is used to determine the corner category; A pose matching module, which is used to perform semantic pose estimation on the candidate regions in the semantic dimension chain to obtain the prior pose of the robot, eliminate the laser points of dynamic obstacles, and construct a local search window around the prior pose to obtain the accurate pose of the robot.