Accurate sensing and positioning method, device and equipment for sweeping robot and medium
By constructing semantic maps and visual repositioning of the sweeping robot, the positioning drift problem of the sweeping robot in similar scenarios is solved, and a higher precision and stable positioning effect is achieved.
Patent Information
- Application Number
- CN202510448833.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Sweeping robots are prone to positioning drift in similar scenarios, resulting in inaccurate positioning.
By semantic classification of visual feature points in scene image frames, a semantic map is constructed, and visual repositioning is performed based on the spatial association relationship and visual constraints of semantic feature points, and positioning correction is performed in combination with visual perception parameters and constraints to avoid positioning drift.
It effectively avoids positioning drift caused by similar scenes, improves the positioning accuracy and stability of the sweeping robot in similar environments, enhances the recognition ability of similar indoor objects, and reduces the risk of mismatch.
Smart Images

Figure CN120284162A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of perception and positioning, and more specifically, to a precise perception and positioning method, device, equipment, and medium for a sweeping robot. Background Art
[0002] Perception and positioning determines the position of the robot in space by obtaining environmental information, uses the sensors carried by the robot to perceive the surrounding environment in real time, and analyzes and processes the sensor data to achieve precise positioning of the robot. Perception and positioning not only depends on the sensor data, but also improves the positioning accuracy through continuous updating and correction. Especially in complex or dynamic environments, it can effectively identify and avoid obstacles, and plan the best path. With the development of technology, perception and positioning has been able to handle challenges in different scenarios, such as problems like light changes, environmental changes, and repetitive scenarios. By continuously optimizing algorithms and enhancing perception capabilities, perception and positioning technology has become an important basis for the autonomous navigation of robots, providing reliable spatial cognition capabilities for robots to perform tasks.
[0003] In a home environment, many rooms may contain a large number of similar or repetitive scene elements, such as symmetrical furniture layouts, similar wall decorations, and long corridors. These features may cause misjudgments when the sweeping robot performs positioning. When the robot passes through these areas, the visual sensor may incorrectly identify similar scenes at different positions as the same place, resulting in feature matching errors. Without timely correction, this error will cause the positioning error to gradually accumulate, leading to positioning drift. As the task progresses, the robot may think it is still at a certain known position, but the actual position has shifted, resulting in inaccurate positioning of the sweeping robot. Therefore, how to avoid positioning drift under the influence of similar scenes has become an urgent problem to be solved in the industry. Summary of the Invention
[0004] This application provides a precise perception and positioning method, device, equipment, and medium for a sweeping robot, which can avoid positioning drift under the influence of similar scenes.
[0005] In a first aspect, this application provides a precise perception and positioning method for a sweeping robot, including the following steps:
[0006] The sweeping robot rotates to obtain each scene image frame around the current position;
[0007] Semantically classify each visual feature point in each scene image frame to obtain multiple semantic feature points for each scene image frame, and then construct a semantic map for the sweeping robot to perform positioning at the current position based on all the semantic feature points;
[0008] Determine the spatial association relationship between each semantic feature point in each scene image frame according to the feature descriptors of each semantic feature point, and determine the visual constraint conditions when the sweeping robot perceives each similar indoor object around it at the current position according to the semantic map and all the spatial association relationships;
[0009] Perform similarity matching on adjacent scene image frames to obtain the scene overlap degree between adjacent scene image frames, and determine the visual perception parameters of the sweeping robot for similar indoor objects at the current position according to all the scene overlap degrees;
[0010] Perform visual relocalization on the sweeping robot according to the visual perception parameters and the visual constraint conditions when the sweeping robot perceives each indoor object around it, and obtain the localization confidence of the sweeping robot at the current position.
[0011] In some embodiments, performing semantic classification on each visual feature point in each scene image frame to obtain multiple semantic feature points in each scene image frame specifically includes:
[0012] Determine each visual feature point in each scene image frame;
[0013] Extract the semantic labels of each visual feature point in each scene image frame from each visual feature point in each scene image frame based on a pre-trained semantic segmentation model;
[0014] Perform semantic screening on the semantic labels of each visual feature point in each scene image frame to obtain multiple semantic feature points in each scene image frame.
[0015] In some embodiments, constructing a semantic map for the sweeping robot to perform localization at the current position according to all the semantic feature points specifically includes:
[0016] Map each semantic feature point of each scene image frame to the reference coordinate system of the sweeping robot at the current position to obtain the three-dimensional spatial coordinates of each semantic feature point in each scene image frame in the reference coordinate system;
[0017] Aggregate the three-dimensional spatial coordinates of each semantic feature point in all scene image frames to obtain different semantic entities in the reference coordinate system;
[0018] Determine the semantic map for the sweeping robot to perform localization at the current position according to all the semantic entities in the reference coordinate system.
[0019] In some embodiments, determining the spatial association relationship between each semantic feature point in each scene image frame according to the feature descriptors of each semantic feature point specifically includes:
[0020] For each scene image frame;
[0021] Determine the feature descriptors of each semantic feature point in the scene image frame;
[0022] Determine the feature correlation degree between each semantic feature point in the scene image frame according to the feature descriptors of each semantic feature point in the scene image frame;
[0023] Obtain the pixel coordinates of each semantic feature point within the scene image frame;
[0024] Determine the spatial correlation degree between each semantic feature point in the scene image frame according to the pixel coordinates of each semantic feature point within the scene image frame;
[0025] Determine the spatial correlation relationship between each semantic feature point in the scene image frame through the feature correlation degree and the spatial correlation degree.
[0026] In some embodiments, determining the visual constraint conditions when the sweeping robot senses each similar indoor object around at the current position according to the semantic map and all the spatial correlation relationships specifically includes:
[0027] Extract the interference influence coefficient when the sweeping robot senses each similar indoor object around at the current position from the semantic map;
[0028] Determine the visual error when the sweeping robot senses each similar indoor object around at the current position according to all the spatial correlation relationships;
[0029] Determine the visual constraint conditions when the sweeping robot senses each similar indoor object around at the current position through the interference influence coefficient and the visual error when the sweeping robot senses each similar indoor object around at the current position.
[0030] In some embodiments, performing similarity matching on adjacent scene image frames to obtain the scene overlap degree between adjacent scene image frames specifically includes:
[0031] For every two adjacent scene image frames;
[0032] Determine each pair of matching feature points between two adjacent scene image frames;
[0033] Perform similarity evaluation on each pair of matching feature points between two adjacent scene image frames to obtain the scene overlap degree between adjacent scene image frames.
[0034] In some embodiments, performing visual relocalization on the sweeping robot according to the visual perception parameters and the visual constraint conditions when the sweeping robot senses each indoor object around to obtain the localization confidence of the sweeping robot at the current position specifically includes:
[0035] Initialize the position estimation of the sweeping robot;
[0036] Relocate the position estimation of the sweeping robot according to the visual perception parameters and the visual constraint conditions when the sweeping robot perceives each indoor object around it, and obtain the positioning vector of the sweeping robot at the current position;
[0037] Determine the positioning confidence of the sweeping robot at the current position through the positioning vector.
[0038] In a second aspect, the present application provides a precise perception and positioning device for a sweeping robot, including:
[0039] An acquisition module, configured to rotate the sweeping robot to acquire various scene image frames around the current position;
[0040] A processing module, configured to perform semantic classification on each visual feature point in each scene image frame to obtain multiple semantic feature points of each scene image frame, and further construct a semantic map for the sweeping robot to perform positioning at the current position based on all the semantic feature points;
[0041] The processing module is further configured to determine the spatial association relationship between each semantic feature point in each scene image frame according to the feature descriptors of each semantic feature point in each scene image frame, and determine the visual constraint conditions when the sweeping robot perceives each similar indoor object around it at the current position according to the semantic map and all the spatial association relationships;
[0042] The processing module is further configured to perform similarity matching on adjacent scene image frames to obtain the scene coincidence degree between adjacent scene image frames, and determine the visual perception parameters of the sweeping robot for similar indoor objects at the current position according to all the scene coincidence degrees;
[0043] An execution module, configured to perform visual relocalization on the sweeping robot according to the visual perception parameters and the visual constraint conditions when the sweeping robot perceives each indoor object around it, and obtain the positioning confidence of the sweeping robot at the current position.
[0044] In a third aspect, the present application provides a computer device, where the computer device includes a memory and a processor, the memory stores code, and the processor is configured to obtain the code and execute the above-mentioned precise perception and positioning method for the sweeping robot.
[0045] In a fourth aspect, the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned precise perception and positioning method for the sweeping robot is implemented.
[0046] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects:
[0047] In a method, device, equipment and medium for accurate perception and positioning of a floor cleaning robot provided by this application, the floor cleaning robot rotates to obtain various scene image frames around the current position; semantic classification is performed on each visual feature point in each scene image frame to obtain multiple semantic feature points of each scene image frame, and then a semantic map for positioning the floor cleaning robot at the current position is constructed based on all the semantic feature points; the spatial association relationship between each semantic feature point in each scene image frame is determined according to the feature descriptor of each semantic feature point in each scene image frame, and the visual constraint condition for the floor cleaning robot to perceive each similar indoor object around at the current position is determined according to the semantic map and all the spatial association relationships; similarity matching is performed on adjacent scene image frames to obtain the scene coincidence degree between adjacent scene image frames, and the visual perception parameter of the floor cleaning robot for similar indoor objects at the current position is determined according to all the scene coincidence degrees; the floor cleaning robot is visually repositioned according to the visual perception parameter and the visual constraint condition for the floor cleaning robot to perceive each indoor object around, and the positioning confidence of the floor cleaning robot at the current position is obtained.
[0048] It can be seen that in the present application, the visual relocalization of the sweeping robot can be performed according to the visual perception parameters and the visual constraint conditions when the sweeping robot perceives each indoor object around it, and the localization confidence of the sweeping robot at the current position can be obtained. Among them, first, semantic classification is performed on each visual feature point in each scene image frame, and the traditional low-level features can be upgraded to semantic feature points with semantic discrimination ability, thereby enhancing the sweeping robot's ability to identify key objects in similar scenes. Secondly, constructing a semantic map when the sweeping robot locates at the current position helps to introduce higher-level semantic information in environmental perception, enabling the sweeping robot to identify and distinguish different categories of indoor objects and structural features. Compared with traditional maps based only on geometry or texture, the semantic map has stronger discriminability and can effectively reduce the risk of false matching in indoor scenes with repetitiveness or similarity, thereby enhancing the sweeping robot's environmental understanding ability and effectively avoiding the localization drift problem caused by scene similarity. Furthermore, determining the spatial association relationship between each semantic feature point in each scene image frame can establish structural constraints at the feature level, enhance the spatial consistency between semantic information, and help the sweeping robot distinguish position differences in a similar visual information environment, avoiding false matching caused by local visual similarity, thereby reducing the risk of localization drift. Further, by combining the spatial association relationship between the semantic map and the semantic feature points, the position and semantic features of each similar indoor object in the overall spatial structure can be accurately identified. This visual constraint condition can effectively limit the matching ambiguity caused by scene similarity during the perception process of the sweeping robot, enhance the robot's ability to identify subtle differences between different positions, and enable it to maintain stable and accurate localization judgment when facing repetitive or symmetric layouts. Then, through comprehensive analysis of all scene overlaps, the features of the robot's perception stability in similar scenes can be effectively extracted, and the visual perception parameters can be determined accordingly, thereby enhancing the sweeping robot's ability to distinguish repetitive visual information. Finally, by combining the visual perception parameters with the visual constraint conditions when the sweeping robot perceives the indoor objects around it, the initial position estimate can be corrected under the interference of similar scenes to achieve more robust visual relocalization. This process can effectively identify the perception errors caused by repeated structures or similar objects in the environment and dynamically compensate for the errors through the relocalization algorithm, thereby improving the accuracy and stability of localization. By evaluating the confidence of the localization result, the sweeping robot can determine whether the current localization is reliable. If the confidence is insufficient, a further relocalization operation can be triggered to avoid continuous localization drift in similar scenes. In summary, the solution of the present application can achieve the avoidance of localization drift under the influence of similar scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is an exemplary flowchart of a method for precise perception and localization of a sweeping robot according to some embodiments of the present application;
[0050] Figure 2 is a schematic flowchart of determining semantic feature points as shown in some embodiments of the present application;
[0051] Figure 3 is a schematic flowchart of determining the spatial association relationship between each semantic feature point as shown in some embodiments of the present application;
[0052] Figure 4 is a schematic structural diagram of a precise perception and positioning device for a floor cleaning robot as shown in some embodiments of the present application;
[0053] Figure 5 is a schematic structural diagram of a computer device for implementing the precise perception and positioning method of a floor cleaning robot as shown in some embodiments of the present application. Detailed implementation manners
[0054] To better understand the technical solution of the present application, the technical solution of the present application will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0055] Refer to Figure 1 , this figure is an exemplary flowchart of the precise perception and positioning method of a floor cleaning robot as shown in some embodiments of the present application. The method of this embodiment can be used for the floor cleaning robot to perform repositioning after power-on and startup in a home environment. The precise perception and positioning method 100 of the floor cleaning robot mainly includes the following steps:
[0056] In step 101, the floor cleaning robot rotates to obtain each scene image frame around the current position.
[0057] It should be noted that in the present application, when the floor cleaning robot is powered on and started, if it detects that the current position cannot be matched with the existing map, the "repositioning" mode is triggered.
[0058] Specifically, when implemented, the floor cleaning robot controls the drive wheels at the current position to rotate around its vertical axis at a slower speed, and the rotation angle can be set to collect the scene image frames of the surrounding environment every fixed angle (such as 10° or 15°), so as to obtain each scene image frame around the current position.
[0059] In step 102, semantic classification is performed on each visual feature point in each scene image frame to obtain multiple semantic feature points of each scene image frame, and then a semantic map of the floor cleaning robot at the current position is constructed based on all the semantic feature points.
[0060] In some embodiments, refer to Figure 2As shown in the figure, this is a schematic flowchart of determining semantic feature points in some embodiments of the present application. In this embodiment, semantic classification is performed on each visual feature point in each scene image frame, and multiple semantic feature points of each scene image frame can be obtained by the following steps:
[0061] First, in step 1021, each visual feature point in each scene image frame is determined;
[0062] Secondly, in step 1022, based on the pre-trained semantic segmentation model, semantic labels of each visual feature point in each scene image frame are extracted from each visual feature point in each scene image frame;
[0063] Then, in step 1023, semantic screening is performed on the semantic labels of each visual feature point in each scene image frame to obtain multiple semantic feature points of each scene image frame.
[0064] When specifically implemented, each visual feature point in each scene image frame can be determined in the following manner, that is: for each scene image frame, the scene image frame is grayscale processed to improve image contrast and edge sharpness. Then, a feature point detection algorithm (such as the Scale-Invariant Feature Transform algorithm) is used to detect pixel points of texture, corner points or edge structures from the grayscale scene image frame, and the obtained pixel points are all used as visual feature points of the scene image frame, so as to obtain each visual feature point in each scene image frame.
[0065] It should be noted that in the present application, a visual feature point refers to a pixel point with significant visual features (such as corner points, edges, textures, etc.). The visual feature points have rotation and scale invariance and can maintain a stable recognition effect under different perspectives and lighting conditions.
[0066] It should be noted that the semantic segmentation model in the present application is a type of deep learning model (such as DeepLabv3+) used to assign each pixel in an image to a specific semantic category (such as "ground", "wall", "furniture", etc.). This deep learning model is based on a convolutional neural network, extracts multi-scale semantic features of the image through an encoder, and then restores the feature map to a pixel-level classification result with the same size as the input image through a decoder. This deep learning model has been trained on a large-scale indoor scene dataset (such as the ADE20K dataset), and can accurately capture the boundaries and semantic structures of different indoor objects in the indoor scene, and achieve high-precision semantic segmentation of different regions in the complex indoor scene.
[0067] In specific implementation, extracting the semantic labels of each visual feature point in each scene image frame from a pre-trained semantic segmentation model can be achieved in the following way, that is: for each scene image frame, use the pre-trained semantic segmentation model to perform pixel-level semantic segmentation processing on each visual feature point in the scene image frame, generate the semantic categories of each visual feature point in the scene image frame, and use the obtained semantic categories as semantic labels, so as to obtain the semantic labels of each visual feature point in each scene image frame.
[0068] It should be noted that in this application, the semantic label represents identification information (such as "ground", "wall", "table", "door", etc.) describing the semantic category name of the visual feature point, and the semantic label reflects the semantic attributes of each visual feature point in the scene image frame.
[0069] In specific implementation, semantic screening of the semantic labels of each visual feature point in each scene image frame to obtain multiple semantic feature points of each scene image frame can be achieved in the following way, that is: for each scene image frame, according to a pre-set semantic label set (such as "wall", "furniture", "door frame", etc.), traverse the semantic labels of each visual feature point extracted in the scene image frame, compare the semantic label of each visual feature point with the pre-set semantic label set, and judge whether the visual feature point belongs to a valid semantic category. For example, if the set valid categories include "wall" and "furniture", then when the semantic label of a visual feature point is "wall", this visual feature point will be retained; and if the semantic label of this visual feature point is "carpet" and this category is not included in the semantic label set, then this visual feature point will be filtered out, so that the visual feature points obtained after screening each visual feature point in the scene image frame are used as semantic feature points, and then multiple semantic feature points of each scene image frame are obtained.
[0070] It should be noted that in this application, the semantic feature point represents a pixel point with semantic identification and significant visual features.
[0071] In some embodiments, constructing a semantic map for the floor cleaning robot to perform positioning at the current position based on all the semantic feature points can be achieved through the following steps:
[0072] Map each semantic feature point of each scene image frame to the reference coordinate system of the floor cleaning robot at the current position to obtain the three-dimensional spatial coordinates of each semantic feature point in each scene image frame in the reference coordinate system;
[0073] Aggregate the three-dimensional spatial coordinates of each semantic feature point in all scene image frames to obtain different semantic entities in the reference coordinate system;
[0074] Determine the semantic map of the sweeping robot when positioning at the current position according to all semantic entities in the reference coordinate system.
[0075] When specifically implemented, map each semantic feature point of each scene image frame to the reference coordinate system of the sweeping robot at the current position. To obtain the three-dimensional spatial coordinates of each semantic feature point in each scene image frame in the reference coordinate system, the following method can be used, that is: the sweeping robot uses Visual Simultaneous Localization and Mapping (VSLAM) technology to estimate the three-dimensional position of the current sweeping robot relative to the indoor environment. Then, using the internal and external parameters of the stereo camera carried by the sweeping robot (obtained through camera calibration) and the three-dimensional position, the coordinate transformation is used to convert each semantic feature point of each scene image frame into the three-dimensional spatial coordinates in the reference coordinate system of the sweeping robot, so as to obtain the three-dimensional spatial coordinates of each semantic feature point in each scene image frame in the reference coordinate system.
[0076] When specifically implemented, aggregate the three-dimensional spatial coordinates of each semantic feature point in all scene image frames to obtain different semantic entities in the reference coordinate system. The following method can be used, that is: use a clustering algorithm (such as the Euclidean clustering algorithm) to cluster the three-dimensional spatial coordinates of each semantic feature point in all scene image frames, so as to obtain multiple data clusters. Then, use a point cloud registration algorithm (such as the ICP algorithm) to spatially align the semantic feature points from different scene image frames in each data cluster, so that the data clusters obtained after spatial alignment are all used as semantic entities in the reference coordinate system, and then different semantic entities in the reference coordinate system are obtained.
[0077] It should be noted that the semantic entity described in this application represents an indoor object with a semantic label and multiple three-dimensional spatial coordinates in the reference coordinate of the sweeping robot.
[0078] It should be noted that in this application, by uniformly organizing and structuring all semantic entities into a map model that can express spatial geometric relationships and high-level semantic relationships, the robot is enabled to have semantic cognitive ability and spatial positioning ability for the environment.
[0079] In specific implementation, the semantic map for the sweeping robot to perform positioning at the current position can be implemented in the following manner based on all semantic entities in the reference coordinate system, that is: First, use a spatial index structure (such as KD-Tree) to retrieve all semantic entities, and use the set of retrieved semantic entities as the set of semantic entities in the reference coordinate system. Then, group the set of semantic entities according to semantic labels, and further classify each semantic entity in combination with the geometric attributes of each semantic entity (such as spatial extension direction, size range) to eliminate redundant representations or overlapping entities caused by multi-frame observations. On this basis, use a sparse point map structure to register each classified semantic entity in the map according to its spatial distribution, record the position, category, and adjacency relationship of each semantic entity in the map, so as to obtain an initial map. Finally, use a graph optimization algorithm (such as the G2O algorithm) to globally adjust the position relationship between each semantic entity in the initial map, and use the initial map obtained after global adjustment as the semantic map for the sweeping robot to perform positioning at the current position.
[0080] It should be noted that the semantic map described in this application represents a map that fuses the spatial information and semantic information of indoor objects.
[0081] In step 103, according to the feature descriptors of each semantic feature point in each scene image frame, determine the spatial association relationship between each semantic feature point in each scene image frame, and determine the visual constraint conditions for the sweeping robot to perceive each similar indoor object around it at the current position according to the semantic map and all spatial association relationships.
[0082] In some embodiments, refer to Figure 3 As shown, this figure is a schematic flow chart for determining the spatial association relationship between each semantic feature point in some embodiments of this application. In this embodiment, the spatial association relationship between each semantic feature point in each scene image frame can be implemented by the following steps:
[0083] For each scene image frame;
[0084] Determine the feature descriptors of each semantic feature point in the scene image frame;
[0085] Determine the feature association degree between each semantic feature point in the scene image frame according to the feature descriptors of each semantic feature point in the scene image frame;
[0086] Obtain the pixel coordinates of each semantic feature point within the scene image frame;
[0087] Determine the spatial association degree between each semantic feature point in the scene image frame according to the pixel coordinates of each semantic feature point within the scene image frame;
[0088] Determine the spatial association relationship between each semantic feature point in the scene image frame based on the feature association degree and the spatial association degree.
[0089] When specifically implemented, the feature descriptors of each semantic feature point in the scene image frame can be determined in the following manner: for each semantic feature point in the scene image frame, use the scale-invariant feature transform descriptor as the feature descriptor of the semantic feature point, so as to obtain the feature descriptors of each semantic feature point in the scene image frame.
[0090] It should be noted that in this application, the feature descriptor represents a vector describing the features of the local area of the semantic feature point.
[0091] When specifically implemented, the feature association degree between each semantic feature point in the scene image frame can be determined according to the feature descriptors of each semantic feature point in the following manner: for the feature descriptors of every two semantic feature points in the scene image frame, calculate the cosine similarity of the feature descriptors of every two semantic feature points, and use the obtained cosine similarity as the correlation parameter for measuring the relationship between the two semantic feature points. Then, select the smallest correlation parameter from all the correlation parameters as the feature association degree between each semantic feature point in the scene image frame. Among them, the smallest correlation parameter usually means that the two semantic feature points are closest in the feature space, indicating that the visual similarity between them is the highest.
[0092] It should be noted that in this application, the feature association degree represents the similarity degree of each semantic feature point in visual information expression.
[0093] When specifically implemented, the spatial association degree between each semantic feature point in the scene image frame can be determined according to the pixel coordinates of each semantic feature point in the scene image frame in the following manner: for the pixel coordinates of every two semantic feature points in the scene image frame, calculate the Euclidean distance of the pixel coordinates of every two semantic feature points. Then, use the Gaussian function as the weight function to assign weights to the Euclidean distances between the pixel coordinates of every two semantic feature points. The semantic feature points with closer distances have higher weights, and the semantic feature points with farther distances have lower weights. For example, the weight function can be where d represents the Euclidean distance, and σ is a pre-set parameter that can be adjusted according to experiments. Further, sum up all the weighted Euclidean distances, and use the obtained sum value as the spatial association degree between each semantic feature point in the scene image frame.
[0094] It should be noted that in this application, the spatial association degree represents the degree of closeness of the distribution of each semantic feature point in the scene image frame.
[0095] In addition, it should be noted that the spatial association relationship represents the characteristic parameters in which each semantic feature point forms a visually expressed information interdependence in terms of spatial position. As a preferred embodiment, the spatial association relationship between each semantic feature point in the scene image frame can be determined by the feature association degree and the spatial association degree in the following manner, that is: taking the quotient of the spatial association degree and the feature association degree as the spatial association relationship between each semantic feature point in the scene image frame. Among them, taking the quotient of the feature association degree and the spatial association degree as the spatial association relationship. In this way, the balance between visual similarity and spatial relationship can be emphasized, and the excessive influence of extreme values in a certain dimension on the final result can be reduced.
[0096] In some embodiments, the visual constraint conditions when the sweeping robot perceives each similar indoor object around it at the current position can be determined according to the semantic map and all the spatial association relationships by the following steps:
[0097] Extract the interference influence coefficient when the sweeping robot perceives each similar indoor object around it at the current position from the semantic map;
[0098] Determine the visual error when the sweeping robot perceives each similar indoor object around it at the current position according to all the spatial association relationships;
[0099] Determine the visual constraint conditions when the sweeping robot perceives each similar indoor object around it at the current position through the interference influence coefficient and the visual error when the sweeping robot perceives each similar indoor object around it at the current position.
[0100] Specifically, when implemented, the interference influence coefficient when the sweeping robot perceives each similar indoor object around it at the current position can be extracted from the semantic map in the following manner, that is:
[0101] It should be noted that in the real indoor environment, objects such as chairs, tables, and lockers may be highly similar in geometric appearance. By quantifying the risk of confusion caused by similar indoor objects in the visual perception task, the positioning robustness and accuracy of the sweeping robot in similar scenarios can be assisted in improving.
[0102] When specifically implemented, the interference influence coefficient when the floor sweeping robot perceives each surrounding similar indoor object at the current position can be extracted from the semantic map in the following way, that is: First, use word embedding technology (such as Word2Vec) to convert the semantic labels of all semantic entities in the semantic map into vectors, and regard the obtained vectors as the label vectors of the semantic entities. For every two semantic entities, calculate the cosine similarity between the label vectors of the two semantic entities. Obtain the similar indoor object threshold from the database of the floor sweeping robot, and compare all the cosine similarities with the similar indoor object threshold. If the cosine similarity is greater than or equal to the similar indoor object threshold, then regard the two semantic entities corresponding to the cosine similarity as similar indoor objects, so as to obtain multiple similar indoor objects. Then, for each similar indoor object, obtain all the three-dimensional space coordinates of the semantic entity corresponding to the similar indoor object from the semantic map, then form a vector with all the three-dimensional space coordinates, and regard the obtained vector as the three-dimensional space coordinate vector of the similar indoor object. Then obtain the orientation vector of the floor sweeping robot at the current position, and then use the cosine theorem to calculate the cosine value of the angle between the three-dimensional space coordinate vector and the orientation vector, and then divide the cosine value of the angle by the distance between the floor sweeping robot and the similar indoor object, and regard the obtained value as the interference influence coefficient when the floor sweeping robot perceives the similar indoor object at the current position, and further obtain the interference influence coefficient when the floor sweeping robot perceives each surrounding similar indoor object at the current position.
[0103] It should be noted that the interference influence coefficient described in this application represents the interference degree of the similar indoor object on the visual perception of the floor sweeping robot.
[0104] In addition, it should also be noted that the similar indoor object threshold in this application can be obtained by vectorizing the names of all indoor objects, calculating the cosine similarity of all the obtained vectors, and then using a clustering algorithm (such as DBSCAN) to cluster all the cosine similarities to obtain multiple clustering clusters. Select the smallest cosine similarity from each cluster, and then regard the mean value of all the smallest cosine similarities as the similar indoor object threshold.
[0105] In specific implementation, to determine the visual error of the sweeping robot when perceiving each similar indoor object around it at the current position according to all the spatial association relationships, the following method can be adopted, that is: obtain all the scene image frames. For each similar indoor object, use an existing image recognition model (such as the MobileNet model) to screen out the scene image frames in which the similar indoor object appears from each scene image frame. Then, perform differential processing on the spatial association relationships between each semantic feature point in all the screened scene image frames, and then sum all the values obtained after the differential processing, and use the sum value as the visual error of the sweeping robot when perceiving the similar indoor objects around it at the current position. Furthermore, the visual error of the sweeping robot when perceiving each similar indoor object around it at the current position is obtained.
[0106] It should be noted that the visual error described in this application represents the recognition error of the sweeping robot when visually perceiving similar indoor objects.
[0107] In addition, it should also be noted that the visual constraint condition described in this application represents the parameter restricted by the similarity visual information during the perception process of the sweeping robot for similar indoor objects.
[0108] In specific implementation, to determine the visual constraint condition of the sweeping robot when perceiving each similar indoor object around it at the current position through the interference influence coefficient and the visual error of the sweeping robot when perceiving each similar indoor object around it at the current position, the following method can be adopted, that is: for each similar indoor object, the quotient of the visual error of the sweeping robot when perceiving the similar indoor objects around it at the current position and the interference influence coefficient can be used as the visual constraint condition of the sweeping robot when perceiving the similar indoor objects around it at the current position. Furthermore, the visual constraint condition of the sweeping robot when perceiving each similar indoor object around it at the current position is obtained.
[0109] It should be noted that when the visual error is relatively large and the interference influence coefficient is relatively small, the quotient value will be relatively large, which indicates that although the interference of external objects is not large, due to the relatively large error of the robot's own vision system, there is a large uncertainty when perceiving similar indoor objects. Therefore, more stringent visual constraint conditions are required to guide the robot's behavior. For example, more caution should be exercised when making positioning decisions. On the contrary, when the visual error is relatively small and the interference influence coefficient is relatively large, the quotient value is relatively small, indicating that although the interference of external objects is relatively strong, the robot's vision system is relatively accurate and can offset the influence brought by part of the interference to a certain extent. At this time, the visual constraint conditions are relatively loose. And when both are relatively large or relatively small, the quotient value will correspondingly reflect the comprehensive influence degree.
[0110] In step 104, adjacent scene image frames are subjected to similarity matching to obtain the scene overlap degree between adjacent scene image frames, and visual perception parameters of similar indoor objects of the sweeping robot at the current position are determined based on all the scene overlap degrees.
[0111] In some embodiments, performing similarity matching on adjacent scene image frames to obtain the scene overlap degree between adjacent scene image frames can be implemented by the following steps:
[0112] For every two adjacent scene image frames;
[0113] Determine each pair of matching feature points between two adjacent scene image frames;
[0114] Perform similarity evaluation on each pair of matching feature points between two adjacent scene image frames to obtain the scene overlap degree between adjacent scene image frames.
[0115] Specifically, determining each pair of matching feature points between two adjacent scene image frames can be implemented in the following way, that is: First, obtain each visual feature point of two adjacent scene image frames, then use the fast nearest neighbor search algorithm to match each visual feature point of two adjacent scene image frames, and take all the similar points obtained by matching each two scene image frames as pairs of matching feature points, so as to obtain each pair of matching feature points between two adjacent scene image frames.
[0116] It should be noted that in this application, the pair of matching feature points refers to a pair of pixel points with similar local features between two adjacent scene image frames, where the pair of matching feature points consists of two pixel points, and the two pixel points are within two scene image frames.
[0117] Specifically, performing similarity evaluation on each pair of matching feature points between two adjacent scene image frames to obtain the scene overlap degree between adjacent scene image frames can be implemented in the following way, that is: For each pair of matching feature points, select one pixel point from the pair of matching feature points as the selected pixel point, subtract the sum of all pixel values adjacent to the position of the selected pixel point in the scene image frame where the selected pixel point is located from the sum of all pixel values adjacent to the position of the other pixel point in the scene image frame of the pair of matching feature points, and take the value obtained by the subtraction as the pixel similarity of the pair of matching feature points, and further obtain the pixel similarity of each pair of matching feature points. Further, take the sum of the pixel similarities of all pairs of matching feature points as the scene overlap degree between adjacent scene image frames.
[0118] It should be noted that in this application, the scene overlap degree represents the similarity degree of the indoor scene in adjacent scene image frames.
[0119] In some embodiments, the visual perception parameters of the sweeping robot for similar indoor objects at the current position can be determined according to the overall scene overlap degree by the following steps:
[0120] Perform linear fitting on all the scene overlap degrees to obtain a fitting curve of the scene overlap degree;
[0121] Determine the visual perception parameters of the sweeping robot for similar indoor objects at the current position according to the fitting curve of the scene overlap degree.
[0122] Specifically, performing linear fitting on all the scene overlap degrees to obtain a fitting curve of the scene overlap degree can be achieved in the following way, that is: use an existing linear fitting algorithm (such as the least squares support vector machine algorithm) to perform linear fitting on all the scene overlap degrees, so as to obtain a fitting curve of the scene overlap degree.
[0123] Specifically, determining the visual perception parameters of the sweeping robot for similar indoor objects at the current position according to the fitting curve of the scene overlap degree can be achieved in the following way, that is: First, calculate the slope at each position on the fitting curve, and take the maximum slope as the visual perception parameter of the sweeping robot for similar indoor objects at the current position. Among them, the larger the slope, the faster the change of the scene overlap degree, indicating that the influence of similar indoor objects on visual perception is more unstable.
[0124] It should be noted that the visual perception parameters described in this application refer to the parameters relied on by the visual system of the sweeping robot to perceive the indoor environment position under the influence of similar indoor objects.
[0125] In step 105, perform visual relocalization on the sweeping robot according to the visual perception parameters and the visual constraint conditions when the sweeping robot perceives each surrounding indoor object, to obtain the localization confidence of the sweeping robot at the current position.
[0126] In some embodiments, performing visual relocalization on the sweeping robot according to the visual perception parameters and the visual constraint conditions when the sweeping robot perceives each surrounding indoor object, to obtain the localization confidence of the sweeping robot at the current position can be achieved by the following steps:
[0127] Initialize the position estimation of the sweeping robot;
[0128] Perform relocalization on the position estimation of the sweeping robot according to the visual perception parameters and the visual constraint conditions when the sweeping robot perceives each surrounding indoor object, to obtain the localization vector of the sweeping robot at the current position;
[0129] Determine the localization confidence of the sweeping robot at the current position through the localization vector.
[0130] In specific implementation, the initialization position estimation of the sweeping robot can be achieved in the following manner, namely: the sweeping robot uses visual simultaneous localization and mapping (VSLAM) technology to estimate the three-dimensional position of the sweeping robot at the current position, and uses the obtained three-dimensional position as the position estimation of the sweeping robot.
[0131] It should be noted that the position estimation described in this application represents the three-dimensional spatial coordinates estimated by the sweeping robot in the current indoor scene.
[0132] In a specific implementation, the position estimation of the sweeping robot is relocated according to the visual perception parameters and the visual constraints when the sweeping robot perceives each indoor object around it, and the positioning vector of the sweeping robot at the current position can be obtained in the following manner, namely: first, the sweeping robot obtains environmental perception data of the current position, wherein the environmental perception data includes the relative position, field of view angle and relative distance of each indoor object perceived by the robot around it, and then, the position is relocated based on the particle filtering algorithm. In the particle filtering algorithm, each particle contains the position estimation, orientation and other state information of the sweeping robot, and then the visual perception parameters are set as the initial weight of each particle, and then the visual constraints are set as additional constraints of the motion model in the particle filtering algorithm to constrain the movement range and position update of the particles, and through multiple resampling steps of the particle filtering algorithm, the spatial position coordinates of the sweeping robot are recorded in turn at each resampling, and finally the vector composed of all the spatial position coordinates obtained after the resampling is completed is used as the positioning vector of the sweeping robot at the current position.
[0133] It should be noted that the positioning vector described in the present application represents a vector composed of different spatial position coordinates of the sweeping robot at the current position. The positioning vector can reflect the different spatial distributions of the sweeping robot at the current position.
[0134] In specific implementation, the positioning confidence of the sweeping robot at the current position can be determined by the positioning vector in the following manner, namely: obtaining the position estimate of the sweeping robot, calculating the distance between each spatial position coordinate in the positioning vector and the position estimate of the sweeping robot, and then normalizing all distances, and further taking the average of all values obtained after normalization as the positioning confidence of the sweeping robot at the current position.
[0135] It should be noted that the positioning confidence is a measure of the reliability of the sweeping robot's estimation result of the current spatial position.
[0136] In addition, it should be noted that the last space in the positioning vector of the floor cleaning robot can be used as the positioning coordinates of the floor cleaning robot at the current position, and the positioning confidence is used to measure the credibility of the positioning coordinates. If the positioning confidence is lower than the preset positioning confidence threshold (the positioning confidence threshold can be obtained from the positioning system of the floor cleaning robot), the floor cleaning robot continues to execute the repositioning process until the positioning confidence is greater than or equal to the preset positioning confidence threshold to ensure the accuracy of the current spatial position.
[0137] In addition, on the other hand of the present application, in some embodiments, the present application provides a precise perception and positioning device for a floor cleaning robot. Referring to Figure 4 , which is a schematic structural diagram of the precise perception and positioning device for a floor cleaning robot according to some embodiments of the present application. The precise perception and positioning device 400 for a floor cleaning robot includes: an acquisition module 401, a processing module 402, and an execution module 403, which are described as follows:
[0138] The acquisition module 401 is mainly used in the present application for the floor cleaning robot to rotate to obtain various scene image frames around the current position;
[0139] The processing module 402 is used in the present application to perform semantic classification on each visual feature point in each scene image frame to obtain multiple semantic feature points of each scene image frame, and then construct a semantic map for the floor cleaning robot to perform positioning at the current position based on all the semantic feature points;
[0140] It should be noted that the processing module 402 in the present application is further used to determine the spatial association relationship between each semantic feature point in each scene image frame according to the feature descriptors of each semantic feature point in each scene image frame, and determine the visual constraint conditions when the floor cleaning robot perceives each similar indoor object around at the current position according to the semantic map and all the spatial association relationships;
[0141] In addition, the processing module 402 in the present application is further used to perform similarity matching on adjacent scene image frames to obtain the scene coincidence degree between adjacent scene image frames, and determine the visual perception parameters of the floor cleaning robot for similar indoor objects at the current position according to all the scene coincidence degrees;
[0142] The execution module 403 is mainly used in the present application to perform visual repositioning on the floor cleaning robot according to the visual perception parameters and the visual constraint conditions when the floor cleaning robot perceives each indoor object around to obtain the positioning confidence of the floor cleaning robot at the current position.
[0143] In addition, the present application also provides a computer device, which includes a memory and a processor. The memory stores code, and the processor is configured to obtain the code and execute the above-mentioned accurate perception and positioning method for the floor cleaning robot.
[0144] In some embodiments, referring to Figure 5 ..., this figure is a schematic structural diagram of a computer device for implementing the accurate perception and positioning method of the floor cleaning robot according to some embodiments of the present application. The above-mentioned accurate perception and positioning method for the floor cleaning robot in the embodiments can be implemented by Figure 5 the computer device shown below. The computer device 500 includes at least one processor 501, a communication bus 502, a memory 503, and at least one communication interface 504.
[0145] The processor 501 can be a general-purpose central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more for controlling the execution of the accurate perception and positioning method of the floor cleaning robot in the present application.
[0146] The communication bus 502 can be used to transfer information between the above components.
[0147] The memory 503 can be a read-only memory (ROM), or other types of static storage devices that can store static information and instructions, a random access memory (RAM), or other types of dynamic storage devices that can store information and instructions. It can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disks, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 503 can exist independently and be connected to the processor 501 through the communication bus 502. The memory 503 can also be integrated with the processor 501.
[0148] Among them, the memory 503 is used to store the program code for executing the solution of this application, and is controlled by the processor 501 for execution. The processor 501 is used to execute the program code stored in the memory 503. The program code may include one or more software modules. The methods described in the above method embodiments can be implemented by one or more software modules in the processor 501 and the program code in the memory 503.
[0149] The communication interface 504 uses any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0150] In a specific implementation, as an embodiment, the computer device may include multiple processors, and each of these processors may be a single-CPU processor or a multi-CPU processor. Here, the processor may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0151] The above computer device may be a general-purpose computer device or a special-purpose computer device. In a specific implementation, the computer device may be a desktop computer, a laptop computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of this application do not limit the type of the computer device.
[0152] In addition, this application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned precise perception and positioning method for a sweeping robot is implemented.
[0153] Although the preferred embodiments of this application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments and all changes and modifications falling within the scope of this application.
[0154] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. A precise perception and positioning method for a floor-sweeping robot, characterized in that, It includes the following steps: The sweeping robot rotates to obtain various scene image frames around the current position; Semantically classify each visual feature point in each scene image frame to obtain multiple semantic feature points for each scene image frame, and then construct a semantic map for the sweeping robot to perform positioning at the current position based on all the semantic feature points; Determine the spatial association relationship between each semantic feature point in each scene image frame according to the feature descriptor of each semantic feature point in each scene image frame, and determine the visual constraint conditions when the sweeping robot perceives each similar indoor object around at the current position according to the semantic map and all the spatial association relationships; Perform similarity matching on adjacent scene image frames to obtain the scene overlap degree between adjacent scene image frames, and determine the visual perception parameters of the sweeping robot for similar indoor objects at the current position according to all the scene overlap degrees; Perform visual relocalization on the sweeping robot according to the visual perception parameters and the visual constraint conditions when the sweeping robot perceives each indoor object around to obtain the positioning confidence of the sweeping robot at the current position.
2. The method according to claim 1, characterized in that, Semantically classifying each visual feature point in each scene image frame to obtain multiple semantic feature points for each scene image frame specifically includes: Determine each visual feature point in each scene image frame; Extract the semantic labels of each visual feature point in each scene image frame from each visual feature point in each scene image frame based on a pre-trained semantic segmentation model; Perform semantic screening on the semantic labels of each visual feature point in each scene image frame to obtain multiple semantic feature points for each scene image frame.
3. The method according to claim 1, characterized in that, Constructing a semantic map for the sweeping robot to perform positioning at the current position based on all the semantic feature points specifically includes: Map each semantic feature point of each scene image frame to the reference coordinate system of the sweeping robot at the current position to obtain the three-dimensional spatial coordinates of each semantic feature point in each scene image frame in the reference coordinate system; Aggregate the three-dimensional spatial coordinates of each semantic feature point in all scene image frames to obtain different semantic entities in the reference coordinate system; Determine the semantic map for the sweeping robot to perform positioning at the current position according to all the semantic entities in the reference coordinate system.
4. The method according to claim 1, characterized in that, Determining the spatial association relationship between each semantic feature point in each scene image frame according to the feature descriptor of each semantic feature point in each scene image frame specifically includes: For each scene image frame; Determine the feature descriptor of each semantic feature point in the scene image frame; Determine the feature association degree between each semantic feature point in the scene image frame according to the feature descriptor of each semantic feature point in the scene image frame; Obtain the pixel coordinates of each semantic feature point in the scene image frame; Determine the spatial association degree between each semantic feature point in the scene image frame according to the pixel coordinates of each semantic feature point in the scene image frame; Determine the spatial association relationship between each semantic feature point in the scene image frame through the feature association degree and the spatial association degree.
5. The method according to claim 1, characterized in that, Determining the visual constraint conditions when the sweeping robot perceives each similar indoor object around it at the current position according to the semantic map and all spatial association relationships specifically includes: Extracting the interference influence coefficient when the sweeping robot perceives each similar indoor object around it at the current position from the semantic map; Determining the visual error when the sweeping robot perceives each similar indoor object around it at the current position according to all spatial association relationships; Determining the visual constraint conditions when the sweeping robot perceives each similar indoor object around it at the current position through the interference influence coefficient and visual error when the sweeping robot perceives each similar indoor object around it at the current position.
6. The method according to claim 1, wherein Performing similarity matching on adjacent scene image frames to obtain the scene overlap degree between adjacent scene image frames specifically includes: For every two adjacent scene image frames; Determining each pair of matching feature points between two adjacent scene image frames; Performing similarity evaluation on each pair of matching feature points between two adjacent scene image frames to obtain the scene overlap degree between adjacent scene image frames.
7. The method according to claim 1, wherein Performing visual relocalization on the sweeping robot according to the visual perception parameters and the visual constraint conditions when the sweeping robot perceives each indoor object around it to obtain the localization confidence of the sweeping robot at the current position specifically includes: Initializing the position estimation of the sweeping robot; Performing relocalization on the position estimation of the sweeping robot according to the visual perception parameters and the visual constraint conditions when the sweeping robot perceives each indoor object around it to obtain the localization vector of the sweeping robot at the current position; Determining the localization confidence of the sweeping robot at the current position through the localization vector.
8. A precise perception and positioning device for a floor cleaning robot, characterized in that, Including: An acquisition module, configured to rotate the sweeping robot to acquire each scene image frame around the current position; A processing module, configured to perform semantic classification on each visual feature point in each scene image frame to obtain multiple semantic feature points of each scene image frame, and further construct a semantic map for the sweeping robot to perform localization at the current position according to all semantic feature points; The processing module is further configured to determine the spatial association relationship between each semantic feature point in each scene image frame according to the feature descriptor of each semantic feature point in each scene image frame, and determine the visual constraint conditions when the sweeping robot perceives each similar indoor object around it at the current position according to the semantic map and all spatial association relationships; The processing module is further configured to perform similarity matching on adjacent scene image frames to obtain the scene overlap degree between adjacent scene image frames, and determine the visual perception parameters of the sweeping robot for similar indoor objects at the current position according to all scene overlap degrees; An execution module, configured to perform visual relocalization on the sweeping robot according to the visual perception parameters and the visual constraint conditions when the sweeping robot perceives each indoor object around it to obtain the localization confidence of the sweeping robot at the current position.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the accurate perception and localization method of the sweeping robot according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the accurate perception and localization method of the sweeping robot according to any one of claims 1 to 7.
Citation Information
Patent Citations
Visual positioning method and system and computer readable storage medium
CN111046125A
Scene re-identification method and device, electronic equipment and storage medium
CN114627365A
Positioning method and device based on visual map
CN115265544A
Robot positioning method, device and equipment and storage medium
CN115597585A
Visual repositioning method and device based on scene semantic graph and computer equipment
CN118196448A