Map optimization method and device and electronic equipment

Through the feature point processing method of multi-camera, dynamic feature points are identified and eliminated, which solves the problems of errors and ghosting in robot mapping and improves the accuracy of mapping and positioning.

CN120800342AActive Publication Date: 2025-10-17HANGZHOU EZVIZ SOFTWARE CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510823132.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-17
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

During the robot mapping process, positioning based on the feature points of moving objects in the environment leads to errors and ghosting, which reduces the accuracy of the map and robot positioning.

Method used

A multi-camera system equipped with at least three cameras is used to identify and eliminate dynamic feature points through feature point extraction, residual calculation and accumulation, and update map data and robot posture.

Benefits of technology

It improves the accuracy of the map and positioning constructed by the robot, reduces the probability of errors and ghosting, and reduces the misjudgment of feature points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120800342A_ABST
    Figure CN120800342A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an optimization method and device for a map and electronic equipment, and relates to the technical field of robots. The optimization method for the map comprises the steps of performing feature point extraction in response to each frame of image acquired by a multi-view camera at the same moment, constructing a current feature point set based on the extracted feature points, and for each feature point in the current feature point set, respectively calculating a residual error of the feature point relative to each camera group, and accumulating the obtained residual errors to obtain an accumulated value corresponding to the feature point, identifying whether the feature point belongs to a dynamic feature point based on the accumulated value corresponding to the feature point, determining a target feature point set in response to the fact that each feature point in the current feature point set is identified, and determining a target feature point based on the feature points in the target feature point set. And updating the map data and the pose of the robot. According to the scheme, the accuracy of the map constructed by the robot and the accuracy of robot positioning can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robots, and in particular to a map optimization method and device and electronic equipment. BACKGROUND

[0002] In recent years, various types of robots have developed rapidly. A robot is a machine device that automatically performs work and is a machine that relies on its own power and control ability to achieve various functions. Map construction (referred to as mapping) can be a function implemented by a robot. The robot can collect data required for mapping, and complete the construction of a map based on the data required for mapping. Mapping can usually be performed with the aid of SLAM (Simultaneous Localization And Mapping, simultaneous localization and mapping) technology, and the SLAM technology needs to rely on feature points in the environment to achieve positioning of the robot during mapping.

[0003] However, during mapping, if mapping is performed based on feature points of moving objects in the environment, errors can occur in positioning of the robot, and ghosting can occur in the constructed map. As a result, the accuracy of the map constructed by the robot and the positioning of the robot is low. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide a map optimization method, device and electronic equipment to improve the accuracy of the map constructed by the robot and the positioning of the robot. The specific technical solutions are as follows:

[0005] In a first aspect, the embodiments of the present application provide a map optimization method applied to a robot. The robot is loaded with a multi-view camera. The number of cameras of the multi-view camera is at least three, and the at least three cameras have a common view area. The method comprises:

[0006] After initializing map data of a target environment scene and a pose of the robot based on images collected by the multi-view camera, in response to each frame of image collected by the multi-view camera at the same time, feature points are extracted based on each frame of image currently collected.

[0007] Based on the extracted feature points, a current feature point set is constructed. The current feature point set is a set of feature points located in the common view area.

[0008] For each feature point in the current feature point set, respectively, calculate the residual of the feature point with respect to each camera group; accumulate the obtained residuals to obtain an accumulated value corresponding to the feature point; and based on the accumulated value corresponding to the feature point, identify whether the feature point is a dynamic feature point; wherein each camera group includes any two cameras of the at least three cameras; the dynamic feature point is a feature point of a moving object in the target environment scene;

[0009] In response to all feature points in the current feature point set being identified, a target feature point set is determined; wherein the target feature point set is a set obtained by excluding target feature points that are dynamic feature points from the current feature point set;

[0010] Based on the feature points in the target feature point set, map data of the target environment scene and the position and posture of the robot are updated.

[0011] In a second aspect, an embodiment of the present application provides a map optimization device, which is applied to a robot, wherein the robot is equipped with a multi-camera, the number of cameras of the multi-camera is at least three, and the at least three cameras have a common viewing area; the device includes:

[0012] A feature point extraction module is configured to, after initializing the map data of the target environment scene and the robot's posture based on the images captured by the multi-camera, extract feature points based on each frame of the image currently captured in response to the multi-camera capturing each frame at the same time;

[0013] A construction module, configured to construct a current feature point set based on the extracted feature points; wherein the current feature point set is a set of feature points located in the common view area;

[0014] a calculation module configured to calculate, for each feature point in the current feature point set, a residual of the feature point with respect to each camera group; accumulate the obtained residuals to obtain an accumulated value corresponding to the feature point; and identify, based on the accumulated value corresponding to the feature point, whether the feature point is a dynamic feature point; wherein each camera group includes any two cameras from the at least three cameras; and the dynamic feature point is a feature point of a moving object in the target environment scene;

[0015] a determination module configured to determine a target feature point set in response to all feature points in the current feature point set being identified; wherein the target feature point set is a set obtained by excluding target feature points that are dynamic feature points from the current feature point set;

[0016] An updating module is configured to update map data of a target environment scene and a pose of the robot based on a feature point in the target feature point set.

[0017] In a third aspect, an electronic device is provided, and the electronic device comprises:

[0018] A memory is configured to store a computer program.

[0019] A processor is configured to execute the computer program stored in the memory, and implement any of the above optimization methods for a map.

[0020] In a fourth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. When the computer program is executed by a processor, any of the above optimization methods for a map is implemented.

[0021] In a fifth aspect, a computer program product is provided, and the computer program product comprises a computer program. When the computer program is executed by a processor, any of the above optimization methods for a map is implemented.

[0022] The embodiments of the present application have the following beneficial effects:

[0023] The method for optimizing a map provided in the embodiments of the present application, in response to the multi-camera capturing each frame of image at the same time, performs feature point extraction based on each frame of image currently captured, and constructs a current feature point set based on the extracted feature points, for each feature point in the current feature point set, respectively calculates the residual of the feature point with respect to each camera group, and accumulates the obtained residual to obtain the accumulated value corresponding to the feature point, and based on the accumulated value corresponding to the feature point, identifies whether the feature point belongs to a dynamic feature point, in response to each feature point in the current feature point set being identified, determines a target feature point set, the target feature point set determined has eliminated the target feature point belonging to the dynamic feature point, which can ensure that the target feature point set based on which the map data of the target environment scene and the pose of the robot are updated does not contain the target feature point belonging to the dynamic feature point, compared with the related art, in the present application, each time the multi-camera captures each frame of image at the same time, it can be determined whether each feature point in the common view area belongs to a dynamic feature point, and the target feature point belonging to the dynamic feature point is eliminated, thereby realizing continuous updating of the current feature point set, so the embodiments of the present application will not map based on the feature points of the moving objects in the environment, which reduces the probability of error in positioning the robot, also reduces the probability of ghost appearing in the constructed map, thereby improving the accuracy of the map constructed by the robot and the positioning of the robot. And, by the accumulated value corresponding to the feature point, it can also be determined whether the feature point belongs to a dynamic feature point, which can also avoid the problem of small residual caused by the moving direction of the moving object being parallel to the direction of the epipolar line in a single camera group, thereby reducing the probability of misjudgment when judging the feature point.

[0024] Of course, implementing any product or method of the present application does not necessarily require all the advantages described above. BRIEF DESCRIPTION OF DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other embodiments can also be obtained by those skilled in the art based on these drawings.

[0026] Figure 1 A schematic diagram of the principle of epipolar constraint provided in the embodiments of the present application;

[0027] Figure 2 A flowchart of a method for optimizing a map provided in the embodiments of the present application;

[0028] Figure 3A flowchart of another map optimization method provided in an embodiment of the present application;

[0029] Figure 4 A flowchart of another map optimization method provided in an embodiment of the present application;

[0030] Figure 5 A schematic diagram of the structure of a map optimization device provided in an embodiment of the present application;

[0031] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field based on this application are within the scope of protection of this application.

[0033] First, some professional terms in the embodiments of this application are briefly introduced:

[0034] Epipolar constraint: When the same point is projected onto two images with different perspectives, the constraint formed by the image point and the optical center of the camera under the projection model. Figure 1 As shown, the line O1O2 connecting the optical centers of the two cameras is called the baseline. The intersection points e1 and e2 of the baseline with pixel plane 1 and pixel plane 2 (which can also be considered as two images) are called cardinal points. Plane O1O2P can be called the epipolar plane, and the intersection lines e1p1 and e2p2 of the epipolar plane with pixel plane 1 and pixel plane 2, respectively, can both be called epipolar lines. The image point (i.e., pixel point) of point P on pixel plane 1 is p1, and the image point (i.e., pixel point) of point P on pixel plane 2 is p2. Image point p2 must lie on the intersection line e2p2 between planes O1O2P and pixel plane 2. This can also be called an epipolar constraint; the image points can also be called feature points.

[0035] Dynamic map: A map that can continuously update the location of objects. The location of objects can be updated in real time during the map update process.

[0036] EKF filter: Extended Kalman Filte, a highly efficient recursive filter (autoregressive filter), also known as an extended Kalman filter.

[0037] ORB: Oriented FAST and Rotated BRIEF, a computer vision algorithm designed to achieve real-time feature extraction and matching.

[0038] Optical flow: also known as optical flow or optic flow, is a concept in the detection of object motion in a field of view; used to describe the motion of an observed target, surface or edge caused by motion relative to the observer.

[0039] Of course, the above-mentioned epipolar constraint can also be embodied in the form of a formula, and the formula form of the epipolar constraint will be introduced below by means of the image points P1 and P2 introduced above: Figure 1

[0040]

[0041] Wherein, p1 is the image point p1, whose coordinates are represented as (u1, v1), p2 is the image point p2, whose coordinates are represented as (u2, v2), F 12 is the transformation matrix between camera 1 and camera 2, camera 1 can be the camera of pixel plane 1, camera 2 can be the camera of pixel plane 2, T is the transpose, for representing the vector of image point p1 projected by camera 2 is 0, that is, the image point p1 is located on the epipolar line; and, K2 is the intrinsic matrix of camera 2, t 12 is the displacement (i.e. displacement in space) of camera 1 and camera 2 when the current image is collected, relative to when the last frame of image is collected, R 12 is the angle of rotation (i.e. the angle of rotation in space) of camera 1 and camera 2 when the current image is collected, relative to when the last frame of image is collected, K1 is the intrinsic matrix of camera 1, -T represents the inverse transpose. Of course, the above is only an exemplary introduction, and the embodiments of the present application do not make specific limitations thereto.

[0042] In order to better understand the present scheme, the related technologies will be introduced as follows:

[0043] In related technology 1, multiple cameras can be set in the environment scene, so as to increase the camera field of view, so that the range of the constructed map is wider; however, the related technology does not consider the influence of dynamic obstacles in the environment scene on pose estimation, and the estimated pose may not be accurate;

[0044] ​In the related art 2, whether any feature point belongs to a dynamic feature point can be determined by recognizing facial features with the help of an AI (Artificial Intelligence) target detection technology; however, the detection object of this related technology can only be a living being, that is, only the facial features of a living being can be recognized, and when the detection object is an object, the object cannot be recognized, and the object cannot be constructed on a map.

[0045] Based on the above-mentioned problems, the embodiment of the present application provides a map optimization method, device and electronic equipment.

[0046] Secondly, the map optimization method provided by the embodiment of the present application is introduced.

[0047] Among them, the map optimization method provided by the embodiment of the present application can be applied to a robot, which is a robot with motion ability, and the robot is loaded with a multi-camera; for example, the robot can be a sweeping robot, a UAV robot, an AGV (Automated Guided Vehicle) with a multi-camera, etc., which is not limited by the embodiment of the present application. For example, in an implementation manner, the execution subject of the embodiment of the present application can also be the main control module / control module of the robot, which is not limited by the embodiment of the present application. In addition, the number of cameras of the multi-camera loaded by the robot in the embodiment of the present application is at least three, and the at least three cameras have a common viewing area; of course, any multi-camera with at least three cameras can be used as the multi-camera in the embodiment of the present application, and the multi-camera in the present application is a visual system (which can also be called a visual component or a visual sensor component) containing at least three cameras, which is not limited by the embodiment of the present application. In addition, the map optimization method provided by the embodiment of the present application has no requirements for the detection object in the map, and the detection object can be an object or a living being, which is not limited by the embodiment of the present application.

[0048] In addition, the application scenario of the map optimization method provided by the embodiment of the present application can be various map construction scenarios, and any scenario of using a robot to construct a map is applicable to the embodiment of the present application, which is not limited by the embodiment of the present application. For example, the map optimization method can be applied to the scenario of constructing a map by a sweeping robot, and can also be applied to the scenario of constructing a map by a UAV robot. In addition, the constructed map can be a dynamic map, which is not limited by the embodiment of the present application.

[0049] One of the optimization methods for the map is applied to a robot loaded with a multi-view camera, the number of cameras of the multi-view camera is at least three, and the at least three cameras have a common view area; the method comprises:

[0050] After the map data of the target environment scene and the pose of the robot are initialized based on the images collected by the multi-view camera, in response to the multi-view camera collecting each frame of image at the same time, based on the current collected each frame of image, feature point extraction is performed;

[0051] Based on the extracted feature points, a current feature point set is constructed; wherein the current feature point set is a set of feature points located in the common view area;

[0052] For each feature point in the current feature point set, the residual error of the feature point with respect to each camera group is calculated respectively; the obtained residual error is accumulated to obtain the accumulated value corresponding to the feature point, and based on the accumulated value corresponding to the feature point, it is identified whether the feature point belongs to a dynamic feature point; wherein each camera group includes any two cameras in the at least three cameras; the dynamic feature point is the feature point of the moving object in the target environment scene;

[0053] In response to each feature point in the current feature point set being identified, a target feature point set is determined; wherein the target feature point set is obtained by removing the target feature points belonging to the dynamic feature points from the current feature point set;

[0054] Based on the feature points in the target feature point set, the map data of the target environment scene and the pose of the robot are updated.

[0055] The optimization method for a map provided in the embodiments of the present application can be used to determine a target feature point set, and the target feature points belonging to dynamic feature points have been removed from the target feature point set. When the map data of a target environment scene and the pose of a robot are updated, the target feature point set used as a basis does not contain target feature points belonging to dynamic feature points. Compared with related technologies, the embodiments of the present application can determine whether each feature point in the common view area belongs to a dynamic feature point and remove the target feature points belonging to dynamic feature points whenever the multi-view camera collects each frame of image at the same time. Therefore, the current feature point set can be continuously updated. The embodiments of the present application do not build a map based on the feature points of a moving object in the environment, thereby reducing the probability of errors in the positioning of the robot and the probability of ghosting in the constructed map, and improving the accuracy of the map constructed by the robot and the positioning of the robot. In addition, the determination of whether the feature point belongs to a dynamic feature point based on the accumulated value of the feature point can also avoid the problem that the residual error is small due to the parallelism between the moving direction of the moving object and the direction of the epipolar line in a single camera group, thereby reducing the probability of misjudgment when the feature point is determined.

[0056] A method for optimizing a map is provided in the embodiments of the present application.

[0057] As shown in Figure 2 A method for optimizing a map is provided in the embodiments of the present application. The method comprises the following steps.

[0058] S201, after the map data of a target environment scene and the pose of a robot are initialized based on the images collected by a multi-view camera, in response to the multi-view camera collecting each frame of image at the same time, feature points are extracted based on the currently collected each frame of image.

[0059] It can be understood that the initialization can be considered as a process of constructing a map data of a target environment scene based on images collected by the multi-view camera, and preliminarily determining a pose of the robot; for example, in an implementation manner, the initialization can include: acquiring each frame of image collected by the multi-view camera and having the same timestamp, so that the consistency of time can be ensured; performing image correction on each frame of image acquired by using a calibration parameter (an intrinsic matrix, an extrinsic matrix, a distortion coefficient, etc.) of the multi-view camera, wherein the image correction can include image de-distortion, polar correction, etc., and the embodiments of the present application do not make specific limitation thereto; for each camera in the multi-view camera, extracting feature points in each frame of image collected by the camera, matching the feature points in the current frame with the feature points in the last frame, so as to estimate a fundamental matrix between adjacent frames, and obtain a rotation vector and a translation vector of the camera; transforming the pose of the camera to a body coordinate system (a coordinate system with the robot as the origin) according to an extrinsic matrix of the camera in the body coordinate system, and completing the initialization for the camera; after the initialization for each camera is completed, the poses of the cameras and the map points can be jointly optimized to complete the overall pose initialization of the multi-view camera and the map construction, and the initial pose of the robot can also be determined based on the determined pose of the multi-view camera, wherein the feature points can also be referred to as ORB feature points, and the feature point matching can be used to match the feature points belonging to the same object, and the extraction of the feature points and the feature point matching can be realized by using the ORB technology, and the embodiments of the present application do not make specific limitation thereto. Of course, the above is only an example, and any initialization manner in the related art can be used to realize the initialization of the embodiments of the present application, and the embodiments of the present application do not make specific limitation thereto.

[0060] After the initialization, the initial map data of the target environment scene and the initial pose of the robot can be determined, and then, in response to that each frame of image collected by the multi-view camera at the same time (each frame of image collected by the multi-view camera and having the same timestamp), feature points in the current collected each frame of image can be extracted; wherein the feature point extraction can be considered as a process of extracting feature points of an object in each frame of image; in addition, the extracted feature points can also be referred to as feature points of the current frame collected by the multi-view camera. Of course, any feature point extraction manner in the related art is applicable to the embodiments of the present application, and the embodiments of the present application do not make specific limitation thereto.

[0061] S202, constructing a current feature point set based on the extracted feature points;

[0062] The current feature point set is a set of feature points located in the common view area.

[0063] It can be understood that after feature points are extracted based on the current collected frames of images, considering that there can be feature points of moving objects in the extracted feature points, the application does not directly use the extracted feature points to update the map data of the target environment scene and the pose of the robot; but after the feature points are extracted, the feature points located in the common view area are determined from the extracted feature points, and the current feature point set is constructed; then, the feature points in the current feature point set are all located in the common view area of the multi-camera, and the current feature point set can include feature points of moving objects located in the common view area of the multi-camera, or can include feature points of non-moving objects located in the common view area of the multi-camera, which is not limited in the embodiments of the application. For example, based on the current collected frames of images, 100 feature points can be extracted, and 50 feature points located in the common view area can be extracted, and the current feature point set Q can be constructed based on the 50 feature points located in the common view area.

[0064] In order to keep the layout clear, the process of constructing the current feature point set based on the extracted feature points will be introduced in other embodiments, and will not be described in detail here.

[0065] S203, for each feature point in the current feature point set, the residual of the feature point with respect to each camera group is calculated respectively; the obtained residual is accumulated to obtain the accumulated value corresponding to the feature point, and whether the feature point belongs to a dynamic feature point is identified based on the accumulated value corresponding to the feature point.

[0066] Each camera group includes any two cameras in the at least three cameras, and the cameras included in different camera groups are different; the dynamic feature point is a feature point of a moving object in the target environment scene.

[0067] It can be understood that the camera group can be pre-set, the multi-camera can be traversed to determine a plurality of camera groups, each camera group includes two cameras, and the cameras included in different camera groups are different, which is not limited in the embodiments of the application; for example, if the multi-camera includes three cameras, camera 1, camera 2 and camera 3, camera 1 and camera 2 can form a camera group, camera 1 and camera 3 can form another camera group, and camera 2 and camera 3 can form another camera group, thereby obtaining three camera groups. In addition, when the residual of the feature point with respect to any camera group is small, the feature point can not belong to a dynamic feature point (a feature point of a moving object), and when the residual of the feature point with respect to any camera group is large, the feature point belongs to a dynamic feature point.

[0068] It is emphasized that when the moving direction of any object is parallel to the direction of the epipolar line of the cameras in a camera group, the calculated residual of the feature point of the object with respect to the camera group can also be small, so there can be a misjudgment when only the residual of a single camera group is used to determine whether the feature point belongs to a dynamic feature point; then, the embodiment of the application can divide the multi-view camera into multiple camera groups, and calculate the residual of the feature point with respect to each camera group, and then the obtained residuals can be accumulated to obtain the accumulated value corresponding to the feature point, and based on the accumulated value corresponding to the feature point, it is identified whether the feature point belongs to a dynamic feature point, thereby reducing the probability of misjudgment when the feature point is determined. For example, the residual of feature point a with respect to camera group 1 is 0, the residual of feature point a with respect to camera group 2 is 1, and the residual of feature point a with respect to camera group 3 is 0.6, so the accumulated value corresponding to feature point a can be obtained as 1.6.

[0069] In addition, the residual of the feature point with respect to any camera group is used to represent the difference between the position of the feature point predicted based on the image of the camera group and the actual position of the feature point captured by the camera group. For example, in an implementation, if the calculated residual is small, it can also be considered that the difference between the position of the feature point predicted by the image of the camera group and the actual position of the feature point captured by the camera group is small, and if the calculated residual is large, it can be considered that the difference between the position of the feature point predicted by the image of the camera group and the actual position of the feature point captured by the camera group is large, and the embodiment of the application does not make a specific limitation.

[0070] In order to clearly layout, the process of calculating the residual of the feature point with respect to each camera group, the process of accumulating the obtained residuals to obtain the accumulated value corresponding to the feature point, and the process of identifying whether the feature point belongs to a dynamic feature point based on the accumulated value corresponding to the feature point are introduced in other embodiments, and will not be described in detail here.

[0071] S204, in response to each feature point in the current feature point set being identified, determining a target feature point set;

[0072] The target feature point set is a set obtained by removing the target feature points belonging to dynamic feature points from the current feature point set;

[0073] It can be understood that after each feature point in the current feature point set is identified, it can be considered that the target feature points belonging to dynamic feature points in the current feature point set are all eliminated, and the feature point set after eliminating the target feature points is taken as a target feature point set, and the target feature point set only contains feature points belonging to static feature points (feature points of non-moving objects in the target environment scene), which is not limited by the embodiments of the present application. For example, there are 50 feature points in the current feature point set, and after identification, 20 target feature points belonging to dynamic feature points are eliminated, and the remaining 30 feature points can be determined as the target feature point set.

[0074] In S205, the map data of the target environment scene and the pose of the robot are updated based on the feature points in the target feature point set.

[0075] It can be understood that the process of updating the map data of the target environment scene and the pose of the robot can include: through the optical flow method, the feature points in the target feature point set are associated and matched with the last frame of image collected by the camera group, and the feature points in the target feature point set are projected and matched with the visible map points (3D points) in the map data of the target environment scene to establish a corresponding relationship; based on the corresponding relationship between the visible map points and the feature points in the target feature point set, the pose of the current frame collected by any camera in the camera group is calculated, and the pose of the current frame collected by another camera in the camera group is determined; the mis-matched feature point pairs are eliminated, and the pose of the current frame is optimized; based on the time interval, the motion amplitude, the number of new feature points, etc., it is determined whether the current frame can be used as a key frame, if so, it is determined as a current key frame, and the current key frame and the co-view key frame (the key frame having co-view relationship with the current key frame) are feature-matched, and based on the unmatched feature points, new feature points are determined; from the determined feature points, the feature points located behind the cameras in the camera group are eliminated, based on the eliminated feature points, the calculated pose of the current frame is updated, and based on the feature points not in the co-view area among the extracted feature points, the map data of the target environment scene is updated.

[0076] Of course, any method for updating the map data and the pose of the robot based on feature points in the related art is applicable to the embodiments of the present application, which is not limited by the embodiments of the present application.

[0077] In addition, after responding to the instruction of stopping updating issued by the user for the robot, the robot can control the multi-view camera to stop image acquisition, so as to stop executing the optimization method for the map provided by the embodiments of the present application, which is not limited by the embodiments of the present application.

[0078] The method for optimizing a map provided in the embodiments of the present application can respond to the fact that the multi-view camera collects each frame of image at the same time, perform feature point extraction based on each frame of image collected at the present time, construct a current feature point set based on the extracted feature points, calculate the residual of each feature point in the current feature point set with respect to each camera group respectively, accumulate the obtained residual to obtain an accumulated value corresponding to the feature point, and identify whether the feature point belongs to a dynamic feature point based on the accumulated value corresponding to the feature point. After each feature point in the current feature point set is identified, a target feature point set is determined. The target feature point set determined has eliminated the target feature points belonging to dynamic feature points. This can ensure that the target feature point set based on which the map data of a target environment scene and the pose of a robot are updated does not contain target feature points belonging to dynamic feature points. Compared with related technologies, in the present application, each time the multi-view camera collects each frame of image at the same time, it can be determined whether each feature point in the common view area belongs to a dynamic feature point, and the target feature points belonging to dynamic feature points are eliminated, so that the current feature point set is continuously updated. Therefore, the embodiments of the present application do not perform mapping based on the feature points of moving objects in the environment, which reduces the probability of errors in positioning the robot and reduces the probability of ghosting in the constructed map, thereby improving the accuracy of the map constructed by the robot and the positioning of the robot. In addition, by using the accumulated value corresponding to the feature point, it can be determined whether the feature point belongs to a dynamic feature point, which can also avoid the problem that the residual is small due to the fact that the moving direction of the moving object is parallel to the direction of the epipolar line in a single camera group, thereby reducing the probability of misjudgment when judging the feature point.

[0079] In addition, the principle of the method for optimizing a map provided in the embodiments of the present application is relatively simple, and a small amount of logic code is added to the robot, which is easy to implement, has low cost, and is quickly landed.

[0080] Optionally, in another embodiment, as shown in Figure 3 calculating the residual of the feature point with respect to each camera group includes:

[0081] S301, mapping the feature point to the pixel plane of each camera to obtain the position point of the feature point in the image collected by each camera at the present time;

[0082] It can be understood that mapping the feature point to the pixel plane of each camera can realize the conversion of the feature point from a three-dimensional point to a two-dimensional point, that is, the conversion of the feature point in a three-dimensional coordinate system to a pixel coordinate system to obtain the position point of the feature point in the current image collected by each camera; and after obtaining the position point of the feature point in the current image collected by each camera, the coordinate information of the feature point in the current image collected by each camera can be determined; of course, any conversion manner in the related art is applicable to the embodiments of the present application, for example, the three-dimensional coordinates of the feature point can be input into a neural network model by means of the neural network model, so that the neural network model outputs the position point of the feature point in the current image collected by each camera, and the embodiments of the present application do not make specific limitation thereon. For example, if the camera group includes camera 1 and camera 2, any feature point P is mapped to the pixel plane of each camera, the image point p1 of the feature point P mapped to the camera 1 can be obtained, the coordinate information of p1 is (u1, v1), and the image point p2 of the feature point P mapped to the camera 2 can also be obtained, the coordinate information of p2 is (u2, v2).

[0083] In S302, based on the position point of the feature point in the current image collected by each camera, the residual of the feature point with respect to each camera group is determined.

[0084] It can be understood that based on the position point of the feature point in the current image collected by each camera, the residual of the feature point with respect to each camera group can be calculated by means of a preset residual calculation formula; wherein the residual of the feature point with respect to each camera group can represent the difference between the position of the feature point predicted based on the image of the camera group and the actual position of the feature point collected by the camera group.

[0085] It can be seen that the embodiments of the present application can map the feature point to the pixel plane of each camera to obtain the position point of the feature point in the current image collected by each camera, based on the position point of the feature point in the current image collected by each camera, the residual of the feature point with respect to each camera group can be determined, so that the feature point can be judged to belong to a dynamic feature point based on the calculated residual, which provides an implementation basis for subsequent updating of the map data of the target environment scene and the pose of the robot, thereby improving the accuracy of the map constructed by the robot and the robot positioning.

[0086] For example, in one implementation, the calculation manner of the residual of the feature point with respect to any camera group includes manner A1:

[0087] In manner A1, the residual of the feature point with respect to the camera group is calculated based on a preset residual calculation formula; wherein the preset residual calculation formula is:

[0088]

[0089] Among them, r ij is the residual of the feature point with respect to the camera group containing camera i and camera j, p j ′ is the predicted position of the feature point in the image currently captured by camera j, and p j ′ =p j +Δp j , p j is the position of the feature point in the previous frame image captured by camera j, Δp j is the pixel distance determined by the displacement of the current image captured by camera j relative to the previous image captured, p i is the position of the feature point in the previous frame image captured by camera i, F ij is the transformation matrix between camera i and camera j, F ij p i is the feature point p i The projection on camera j, (F ij p i ) x is the feature point p i The displacement of the projection on camera j relative to the optical center of camera j in the x direction, (F ij p i ) y is the feature point p i The displacement of the projection on camera j relative to the optical center of camera j in the y direction, Characterizing feature points p i The absolute value of the vector projected on camera j; T represents the transpose.

[0090] It can be understood that camera i and camera j can be a group of camera groups. The numerator of the formula can represent the position of the feature point obtained by image prediction of the camera group, and the denominator of the formula can represent the actual position of the feature point collected by the camera group. Then, the residual of the feature point with respect to each camera group can represent: the difference between the position of the feature point obtained based on the image prediction of the camera group and the actual position of the feature point collected by the camera group.

[0091] And, for the numerator of the formula, p j ′ =p j +Δp jIt can be considered that camera j is displaced after the position point in the previous frame image captured by camera j. The displacement distance is the displacement of the current image captured by camera j relative to the previous frame image captured. Based on this displacement, the corresponding pixel distance can be determined, that is, Δp j ; The position of the feature point in the previous frame of image captured by camera j plus the determined corresponding pixel distance can obtain the predicted position of the feature point in the image currently captured by camera j; then, it can be considered that when the feature point does not belong to a dynamic feature point, the displacement of the feature point between the position of the current captured image and the position of the previous frame of image captured is consistent with the pixel distance determined based on the displacement between the position of the current image captured by camera j and the position of the previous frame of image captured, that is, the object corresponding to the feature point has not moved, it is only because the camera has shifted that the feature point has also shifted in the adjacent image frame. The embodiment of the present application does not make specific restrictions on this. And, if the feature point does not belong to a dynamic feature point, it can be considered that the predicted position of the feature point in the image currently captured by camera j is located on the polar line; and, Can represent the feature point p i The absolute value of the vector projected on camera j. If the feature point is not a dynamic feature point, then Therefore, the calculated residual can also be 0; of course, there may be errors in the actual calculation. is not equal to 0, but if the extreme line constraint is satisfied, The result should also be extremely small; in addition, based on the calculation result of the numerator of the formula, the size of the residual of the feature point with respect to the camera group can be determined. If the calculation result of the numerator of the formula is large, it can be considered that the feature point does not satisfy the epipolar constraint. Accordingly, the calculated residual of the feature point with respect to the camera group is also large. Conversely, if the calculation result of the numerator of the formula is small, it can be considered that the feature point satisfies the epipolar constraint. Accordingly, the calculated residual of the feature point with respect to the camera group is also small. The embodiments of the present application do not make specific limitations on this.

[0092] Correspondingly, for the denominator of the formula, F ij p i is the feature point p i The projection on camera j, (F ij p i ) x is the feature point p i The displacement of the projection on camera j relative to the optical center of camera j in the x direction, (F ii p i ) y is the feature point p iThe displacement of the projection on the camera j relative to the optical center of the camera j in the y direction is denoted as yj, and the displacement of the projection on the camera i relative to the optical center of the camera i in the y direction is denoted as yi. It can be considered that F ij p i is normalized to calculate p i The actual distance of the projection on the camera j is not specifically limited in the embodiments of the application.

[0093] In addition, F ij It can be considered that the transformation matrix between camera i and camera j is specifically represented as K j is the intrinsic matrix of camera j, t ij is the displacement of camera i and camera j relative to the displacement when the previous frame of image is collected (i.e., the displacement in space) when the current image is collected, R ij is the angle of rotation of camera i and camera j relative to the angle of rotation when the previous frame of image is collected (i.e., the angle of rotation in space) when the current image is collected, K i is the intrinsic matrix of camera i, which is not specifically limited in the embodiments of the application. Of course, the above-mentioned calculation formula is only exemplary, and the residual of the feature point relative to any camera group can also be calculated by means of a neural network model, which is not specifically limited in the embodiments of the application.

[0094] It can be seen that based on the above-mentioned preset residual calculation formula, the residual of the feature point relative to the camera group can be calculated, the residual of the feature point relative to the camera group is larger when the epipolar constraint is not satisfied, and the residual of the feature point relative to the camera group is smaller when the epipolar constraint is satisfied. Based on the calculated residual, it can be judged whether the feature point belongs to a dynamic feature point, which provides an implementation basis for subsequent updating of the map data of the target environment scene and the pose of the robot, thereby improving the accuracy of the map constructed by the robot and the robot positioning.

[0095] Optionally, in another embodiment, as shown in Figure 4 the obtained residual is accumulated to obtain an accumulated value corresponding to the feature point, comprising:

[0096] S401, determining a weight corresponding to each camera group; wherein the weight corresponding to any camera group is positively correlated with the baseline between the two cameras in the camera group;

[0097] It can be understood that the weight corresponding to each camera group can be determined in advance, and the weight corresponding to any camera group is positively correlated with the baseline between the two cameras in the camera group. The longer the baseline between the two cameras in the camera group, the greater the weight corresponding to the camera group. For example, if there are camera group 1, camera group 2 and camera group 3, the baseline between camera 1 and camera 2 in camera group 1 is 15 cm, the weight corresponding to camera group 1 can be set to 0.5, the baseline between camera 1 and camera 3 in camera group 2 is 10 cm, the weight corresponding to camera group 2 can be set to 0.33, and the baseline between camera 2 and camera 3 in camera group 3 is 5 cm, the weight corresponding to camera group 3 can be set to 0.17. It should be emphasized that the greater the baseline between the two cameras, the greater the range (field of view) of the two cameras, so the weight of the camera group containing the two cameras can also be set to be larger. Of course, the above is only an example, and the weight corresponding to each camera group is not limited in the embodiment of the application.

[0098] S402, according to the weight corresponding to each camera group, the residual of the feature point with respect to each camera group is weighted and accumulated to obtain the accumulated value corresponding to the feature point.

[0099] It can be understood that according to the weight corresponding to each camera group, the residual of the feature point with respect to each camera group can be weighted and calculated, and the calculation result can be accumulated to obtain the accumulated value corresponding to the feature point. For example, the weight corresponding to camera group 1 is set to 0.5, the residual of feature point a with respect to camera group 1 is 0.2, the weight corresponding to camera group 2 is set to 0.33, the residual of feature point a with respect to camera group 2 is 0.3, the weight corresponding to camera group 3 is set to 0.17, and the residual of feature point a with respect to camera group 1 is 0.1. The accumulated value corresponding to feature point a can be obtained as 0.216.

[0100] In order to better understand the process of weighted accumulation of the residual of each camera group, the following formula is introduced:

[0101]

[0102] wherein r total is the accumulated value corresponding to the feature point, r ij is the residual of the feature point with respect to the camera group containing camera i and camera j, w ij is the weight corresponding to the camera group containing camera i and camera j, and N is the total number of cameras of the multi-camera.

[0103] It can be seen that the embodiments of the present application can set a weight for each camera group, and add up the residuals of the feature point with respect to each camera group according to the corresponding weight of each camera group to obtain the accumulated value corresponding to the feature point. Subsequently, whether the feature point belongs to a dynamic feature point is determined by using the accumulated value corresponding to the feature point, which can avoid the problem of small residuals caused by the moving direction of the moving object being parallel to the direction of the epipolar line in a single camera group, thereby reducing the probability of misjudgment when judging the feature point.

[0104] Optionally, in another embodiment, the identifying whether the feature point belongs to a dynamic feature point based on the accumulated value corresponding to the feature point comprises steps B1-B3:

[0105] Step B1, detecting the size relationship between the accumulated value corresponding to the feature point and a predetermined residual threshold; wherein the predetermined residual threshold is a residual threshold determined based on the sparsity of the objects in the target environment scene;

[0106] It can be understood that the predetermined residual value is a residual threshold determined based on the sparsity of the objects in the target environment scene. The sparser the objects in the target environment scene, the larger the value of the predetermined residual threshold, and vice versa. Specifically, the predetermined residual threshold can also be represented as Y dynamic The empirical value of the predetermined residual threshold can be 0.1, which is not limited in the embodiments of the present application. After the accumulated value corresponding to the feature point is calculated, the size relationship between the accumulated value corresponding to the feature point and the predetermined residual threshold can be detected. If the accumulated value corresponding to the feature point is greater than the predetermined residual threshold, step B2 is performed, otherwise, step B3 is performed.

[0107] Step B2, if the accumulated value corresponding to the feature point is greater than the predetermined residual threshold, it is determined that the feature point belongs to a dynamic feature point;

[0108] It can be understood that if the accumulated value corresponding to the feature point is greater than the predetermined residual threshold, it can be determined that the feature point belongs to a dynamic feature point, i.e., a feature point of a moving object in the target environment scene. Subsequently, the feature point can be excluded from the current feature point set. For example, the accumulated value corresponding to the feature point a is 0.216, and the predetermined residual threshold is 0.1. At this time, the accumulated value corresponding to the feature point a is greater than the predetermined residual threshold, and it can be determined that the feature point a belongs to a dynamic feature point. Subsequently, the feature point a can be excluded from the current feature point set. In addition, the accumulated value corresponding to the feature point being greater than the predetermined residual threshold can also be represented as r total > Y dynamic which is not limited in the embodiments of the present application.

[0109] Step B3, if the accumulated value corresponding to the feature point is not greater than the residual threshold, it is determined that the feature point does not belong to the dynamic feature point.

[0110] It can be understood that if the accumulated value corresponding to the feature point is not greater than the predetermined residual threshold, it can be determined that the feature point does not belong to the dynamic feature point, that is, the feature point does not belong to the moving object in the target environment scene, and the feature point does not need to be removed from the current feature point set. For example, the accumulated value corresponding to the feature point b is 0.1, and the predetermined residual threshold is 0.1. At this time, the accumulated value corresponding to the feature point b is not greater than the predetermined residual threshold, and it can be determined that the feature point b does not belong to the dynamic feature point. In addition, the accumulated value corresponding to the feature point is not greater than the predetermined residual threshold can also be represented as r total ≤Υ dynamic The embodiments of the present application are not limited in this regard.

[0111] It can be seen that the embodiments of the present application can determine whether the feature point belongs to the dynamic feature point based on the size relationship between the accumulated value corresponding to the feature point and the predetermined residual threshold. Compared with the related art, the embodiments of the present application can determine whether the feature point belongs to the dynamic feature point based on the accumulated value corresponding to the feature point, which can reduce the probability of misjudgment when determining the feature point, and improve the accuracy of identifying whether the feature point belongs to the dynamic feature point.

[0112] Optionally, in another embodiment, based on the extracted feature points, a current feature point set is constructed, including steps C1-C2:

[0113] Step C1, determining the feature points in the extracted feature points located in the common view area to obtain a plurality of initial feature points;

[0114] Step C2, removing a specified feature point from the obtained plurality of initial feature points, and constructing the remaining initial feature points as the current feature point set; wherein the position of the specified feature point matches the predicted position obtained based on position prediction of any auxiliary feature point; the auxiliary feature point is a feature point belonging to the dynamic feature point in the auxiliary feature point set; the auxiliary feature point set is a feature point set constructed based on the extracted feature points after feature point extraction on the images collected before the current images;

[0115] Correspondingly, the method further includes step D1:

[0116] Step D1, for each feature point belonging to the dynamic feature point in the current feature point set, performing position prediction of the next frame based on the position of the feature point to obtain the predicted position corresponding to the feature point.

[0117] It can be understood that, for step C1, the feature points located in the common view area can be determined from the extracted feature points to obtain a plurality of initial feature points. For example, based on the current acquisition of each frame image, 100 feature points are extracted, and from the 100 feature points, 50 feature points located in the common view area are determined to obtain 50 initial feature points.

[0118] It can be understood that, for step C2, in response to the multi-view camera acquiring each frame image at the same time, after the feature point extraction of the image acquired before the current frame image, the feature point set constructed based on the extracted feature points is determined as the auxiliary feature point set, which can include feature points belonging to dynamic feature points and feature points not belonging to dynamic feature points, wherein the feature points belonging to dynamic feature points can also be called auxiliary feature points. Then, from the obtained plurality of initial feature points, the specified feature points are removed, and the remaining initial feature points are constructed into the current feature point set; wherein the position of the specified feature point matches the predicted position obtained after the position prediction based on any auxiliary feature point, that is, the predicted position obtained after the position prediction based on any auxiliary feature point can be the same as the position of the specified feature point; when the position of any feature point in the plurality of initial feature points matches the predicted position obtained after the position prediction of the auxiliary feature point (same), it can be determined that the feature point belongs to the dynamic feature point, and the feature point is taken as the specified feature point, and the specified feature point is removed from the plurality of initial feature points. For example, the initial feature points include 4 feature points, which are feature points a-feature points d, wherein feature point a is determined as the specified feature point, so feature point a is removed from the initial feature points, and the remaining feature points b-feature points d are constructed into the current feature point set.

[0119] It can be understood that, for each feature point belonging to the dynamic feature point in the current feature point set, a position prediction of the next frame can be performed based on the position of the feature point to obtain a predicted position corresponding to the feature point. When the multi-view camera captures new frames of images, the current feature point set in step D1 can be determined as an auxiliary feature point set, and each feature point belonging to the dynamic feature point can be determined as an auxiliary feature point. Based on the predicted position of the auxiliary feature point, it can be determined whether there is a specified feature point in the extracted feature points from the captured new frames of images. If there is, the specified feature point is removed. For example, the dynamic feature point in the current feature point set includes feature point a. The position prediction of the next frame can be performed based on the position of the feature point a to obtain the predicted position corresponding to the feature point a. In addition, if any feature point belonging to the dynamic feature point in the current feature point set is located at the edge of the image, the position prediction of the next frame for the feature point can not be performed.

[0120] For example, in an implementation, a prediction model can be established for each feature point belonging to the dynamic feature point in the current feature point set, where the prediction model can be a prediction model in the EKF filter. Any feature point belonging to the dynamic feature point is input into the prediction model, so that the prediction model outputs the position of the next frame of the feature point to obtain the predicted position corresponding to the feature point. Of course, any method for predicting the position in the related art is applicable to the embodiments of the present application, which are not limited in this regard.

[0121] It can be seen that the embodiments of the present application can determine the feature points located in the common view area in the extracted feature points to obtain a plurality of initial feature points. The specified feature points are removed from the obtained plurality of initial feature points. The remaining initial feature points are constructed into the current feature point set. Subsequently, for each feature point belonging to the dynamic feature point in the current feature point set, the position prediction of the next frame can be performed based on the position of the feature point to obtain the predicted position corresponding to the feature point. Then, for the feature point matching the predicted position obtained after the position prediction of any auxiliary feature point, the feature point can be directly determined as the feature point belonging to the dynamic feature point and removed from the plurality of initial feature points, without the need for further judgment of the feature point. Therefore, the current feature point set can be adjusted in real time, the feature points belonging to the dynamic feature point are not identified multiple times, the efficiency of identifying the feature points is improved, and the efficiency of optimizing the map is also improved.

[0122] Based on the above method embodiments, the embodiments of the present application also provide an optimization device for a map, which is applied to a robot. The robot is loaded with a multi-view camera. The number of cameras of the multi-view camera is at least three, and the at least three cameras have a common view area.

[0123] Figure 5 A structural schematic diagram of an optimization device for a map provided by an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, the device comprises: Figure 5

[0124] The feature point extraction module 510 is configured to, after initializing the map data of the target environment scene and the pose of the robot based on the images collected by the multi-view camera, respond to the multi-view camera collecting each frame of image at the same time, and perform feature point extraction based on each frame of image collected at the current time.

[0125] The construction module 520 is configured to construct a current feature point set based on the extracted feature points, wherein the current feature point set is a set of feature points located in the common view area.

[0126] The calculation module 530 is configured to calculate, for each feature point in the current feature point set, a residual error of the feature point with respect to each camera group, respectively; accumulate the obtained residual errors to obtain an accumulated value corresponding to the feature point, and identify whether the feature point belongs to a dynamic feature point based on the accumulated value corresponding to the feature point; wherein each camera group comprises any two cameras in the at least three cameras; and the dynamic feature point is a feature point of a moving object in the target environment scene.

[0127] The determination module 540 is configured to, in response to each feature point in the current feature point set being identified, determine a target feature point set; wherein the target feature point set is obtained by removing the target feature points belonging to the dynamic feature points from the current feature point set.

[0128] The update module 550 is configured to update the map data of the target environment scene and the pose of the robot based on the feature points in the target feature point set.

[0129] Optionally, the calculation module is specifically configured to:

[0130] map the feature point to the pixel plane of each camera to obtain a position point of the feature point in the current collected image of each camera;

[0131] determine the residual error of the feature point with respect to each camera group based on the position point of the feature point in the current collected image of each camera.

[0132] Optionally, the calculation module is further configured to:

[0133] determine a weight corresponding to each camera group; wherein the weight corresponding to any camera group is positively correlated with the baseline between the two cameras in the camera group. ​

[0134] According to the weight corresponding to each camera group, the residual of the feature point with respect to each camera group is weighted and accumulated to obtain an accumulated value corresponding to the feature point.

[0135] Optionally, the computing module is further configured to:

[0136] detect a size relationship between the accumulated value corresponding to the feature point and a predetermined residual threshold value; wherein the predetermined residual threshold value is a residual threshold value determined based on a sparsity of the object in the target environment scene;

[0137] if the accumulated value corresponding to the feature point is greater than the predetermined residual threshold value, it is determined that the feature point belongs to a dynamic feature point;

[0138] if the accumulated value corresponding to the feature point is not greater than the residual threshold value, it is determined that the feature point does not belong to a dynamic feature point.

[0139] Optionally, the constructing module is specifically configured to:

[0140] determine a feature point in the extracted feature points located in the common view area to obtain a plurality of initial feature points;

[0141] from the obtained plurality of initial feature points, a designated feature point is removed, and the remaining initial feature points are constructed as a current feature point set; wherein the position of the designated feature point matches a predicted position obtained based on position prediction of any auxiliary feature point; the auxiliary feature point is a feature point belonging to a dynamic feature point in an auxiliary feature point set; the auxiliary feature point set is a feature point set constructed based on the extracted feature points after feature point extraction on images collected before the current images;

[0142] The device further comprises:

[0143] a position prediction module configured to, for each feature point belonging to a dynamic feature point in the current feature point set, perform position prediction of the next frame based on the position of the feature point to obtain a predicted position corresponding to the feature point.

[0144] Optionally, the calculation method of the residual of the feature point with respect to any camera group comprises:

[0145] calculating the residual of the feature point with respect to the camera group based on a preset residual calculation formula; wherein the preset residual calculation formula is:

[0146]

[0147] wherein r ij is the residual of the feature point with respect to the camera group comprising camera i and camera j, pj ′ To predict the position point of the feature point in the image currently captured by camera j, and p j ′ = p j + Δp j , p j is the position point of the feature point in the last frame of image captured by camera j, Δp j is the pixel distance determined based on the displacement between the position of the image currently captured by camera j and the position of the last frame of image captured, p i is the position point of the feature point in the last frame of image captured by camera i, F ij is the transformation matrix between camera i and camera j, F ij p i is the projection of feature point p i on camera j, (F ij p i ) x is the displacement of the projection of feature point p i on camera j relative to the optical center of camera j in the x direction, (F ij p i ) y is the displacement of the projection of feature point p i on camera j relative to the optical center of camera j in the y direction, characterizes the absolute value of the vector of the projection of feature point p i on camera j; T characterizes transposition.

[0148] In the technical solution of the present application, the operations of obtaining, storing, using, processing, transmitting, providing and disclosing of user personal information are all performed under the authorization of the user.

[0149] The present application also provides an electronic device, as shown in the figure, comprising: Figure 6

[0150] a memory 601 for storing computer programs;

[0151] a processor 602 for executing the programs stored in the memory 601 to realize any of the above-mentioned optimization methods for maps.

[0152] And the above-mentioned electronic device can also include a communication bus and / or a communication interface, and the processor 602, the communication interface and the memory 601 complete mutual communication through the communication bus.

[0153] Optionally, the electronic device further comprises a multi-view camera, the number of cameras of the multi-view camera is at least three, and the at least three cameras have a common view area.​

[0154] Optionally, the electronic device is an autonomous mobile device. For example, the autonomous mobile device can be a robot as described above.

[0155] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The communication bus can be divided into an address bus, a data bus, a control bus, and the like. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.

[0156] The communication interface is used for communication between the electronic device and other devices.

[0157] The memory can include a Random Access Memory (RAM) and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located away from the processor.

[0158] The processor described above can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), and the like; can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.

[0159] In another embodiment provided in the present application, a computer readable storage medium is also provided, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement any of the above methods for optimizing a map.

[0160] In another embodiment provided in the present application, a computer program product containing instructions is also provided, and when the computer program product is run on a computer, the computer is caused to execute any of the above methods for optimizing a map in the above embodiments.

[0161] In the embodiments described above, all or some of the steps can be implemented by using software, hardware, firmware or any combination thereof. When implemented by using software, all or some of the steps can be implemented in the form of one or more computer programs. The computer program is stored in a computer readable medium, and can be executed by a computer to perform all or some of the steps described above. The computer readable medium can be a computer program product in the form of a memory, such as a read-only memory (ROM), a flash memory, a random access memory (RAM), a programmable read-only memory (PROM) or an electrically programmable read-only memory (EPROM). The computer readable medium can also be a removable storage medium, such as a floppy disk, a flexible disk, an optical disk, a magnetic disk, a memory card or a memory stick. The computer readable medium can also be a computer database, a computer network or a computer server. The computer program can be executed by a computer to perform all or some of the steps described above.

[0162] It should be noted that, in the present document, the terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Also, the terms "comprising", "containing", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus that comprises a list of elements does not include those elements only, but can also include other elements not expressly listed, or other elements inherent in such process, method, article or apparatus. Without more limitations, an element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.

[0163] Each of the embodiments in the present document is described in a related manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the description of the method embodiments.

[0164] The above only describes the preferred embodiments of the present application, and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A map optimization method, characterized in that: Applied to a robot, the robot is equipped with a multi-camera, the number of cameras of the multi-camera is at least three, and the at least three cameras have a common viewing area; the method includes: After initializing the map data of the target environment scene and the posture of the robot based on the images captured by the multi-camera, in response to the multi-camera capturing each frame of images at the same time, extracting feature points based on each frame of images currently captured; Based on the extracted feature points, a current feature point set is constructed; wherein the current feature point set is a set of feature points located in the common view area; For each feature point in the current feature point set, respectively, calculate the residual of the feature point with respect to each camera group; accumulate the obtained residuals to obtain an accumulated value corresponding to the feature point; and based on the accumulated value corresponding to the feature point, identify whether the feature point is a dynamic feature point; wherein each camera group includes any two cameras of the at least three cameras; the dynamic feature point is a feature point of a moving object in the target environment scene; In response to all feature points in the current feature point set being identified, a target feature point set is determined; wherein the target feature point set is a set obtained by excluding target feature points that are dynamic feature points from the current feature point set; Based on the feature points in the target feature point set, map data of the target environment scene and the position and posture of the robot are updated.

2. The method according to claim 1, characterized in that Calculate the residual of the feature point for each camera group separately, including: Mapping the feature point to the pixel plane of each camera to obtain the position of the feature point in the image currently captured by each camera; Based on the position of the feature point in the image currently captured by each camera, the residual of the feature point with respect to each camera group is determined.

3. The method according to claim 1, characterized in that The step of accumulating the obtained residuals to obtain the accumulated value corresponding to the feature point includes: Determine a weight corresponding to each camera group; wherein the weight corresponding to any camera group is positively correlated with the baseline between the two cameras in the camera group; According to the weight corresponding to each camera group, the residual of the feature point with respect to each camera group is weightedly accumulated to obtain the accumulated value corresponding to the feature point.

4. The method according to claim 1, wherein The identifying whether the feature point is a dynamic feature point based on the accumulated value corresponding to the feature point includes: Detecting the relationship between the accumulated value corresponding to the feature point and a predetermined residual threshold; wherein the predetermined residual threshold is a residual threshold determined based on the sparsity of objects in the target environment scene; If the accumulated value corresponding to the feature point is greater than the predetermined residual threshold, it is determined that the feature point is a dynamic feature point; If the accumulated value corresponding to the feature point is not greater than the residual threshold, it is determined that the feature point is not a dynamic feature point.

5. The method according to claim 1, wherein Based on the extracted feature points, the current feature point set is constructed, including: Determining feature points located in the common view area among the extracted feature points to obtain a plurality of initial feature points; Eliminating designated feature points from the obtained multiple initial feature points and constructing the remaining initial feature points into a current feature point set; wherein the position of the designated feature point matches the predicted position obtained after position prediction based on any auxiliary feature point; the auxiliary feature point is a feature point that is a dynamic feature point in the auxiliary feature point set; and the auxiliary feature point set is a feature point set constructed based on feature points extracted from images acquired before each frame of the current image. The method further comprises: For each dynamic feature point in the current feature point set, the position of the next frame is predicted based on the position of the feature point to obtain the predicted position corresponding to the feature point.

6. The method according to claim 1, characterized in that The calculation method of the residual of the feature point with respect to any camera group includes: Based on a preset residual calculation formula, the residual of the feature point with respect to the camera group is calculated; wherein the preset residual calculation formula is: Among them, r ij is the residual of the feature point with respect to the camera group containing camera i and camera j, p j ′ is the predicted position of the feature point in the image currently captured by camera j, and p j ′ =p j +Δp j , p j is the position of the feature point in the previous frame image captured by camera j, Δp j is the pixel distance determined based on the displacement between the position of the current image captured by camera j and the position of the previous frame image captured, p i is the position of the feature point in the previous frame image captured by camera i, F ij is the transformation matrix between camera i and camera j, F ij p i is the feature point p i The projection on camera j, (F ij p i ) x is the feature point p i The displacement of the projection on camera j relative to the optical center of camera j in the x direction, (F ij p i ) y is the feature point p i The displacement of the projection on camera j relative to the optical center of camera j in the y direction, Characterizing feature points p i The absolute value of the vector projected on camera j; T represents the transpose.

7. A map optimization device, characterized in that: Applied to a robot, the robot is equipped with a multi-camera, the number of the multi-camera is at least three, and the at least three cameras have a common viewing area; the device comprises: A feature point extraction module is configured to, after initializing the map data of the target environment scene and the robot's posture based on the images captured by the multi-camera, extract feature points based on each frame of the image currently captured in response to the multi-camera capturing each frame at the same time; A construction module, configured to construct a current feature point set based on the extracted feature points; wherein the current feature point set is a set of feature points located in the common view area; a calculation module configured to calculate, for each feature point in the current feature point set, a residual of the feature point with respect to each camera group; accumulate the obtained residuals to obtain an accumulated value corresponding to the feature point; and identify, based on the accumulated value corresponding to the feature point, whether the feature point is a dynamic feature point; wherein each camera group includes any two cameras from the at least three cameras; and the dynamic feature point is a feature point of a moving object in the target environment scene; a determination module configured to determine a target feature point set in response to all feature points in the current feature point set being identified; wherein the target feature point set is a set obtained by excluding target feature points that are dynamic feature points from the current feature point set; An updating module is used to update the map data of the target environment scene and the position and posture of the robot based on the feature points in the target feature point set.

8. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 6 when executing a program stored in a memory.

9. The electronic device according to claim 8, wherein: Also includes: A multi-camera, wherein the number of cameras of the multi-camera is at least three, and the at least three cameras have a common viewing area.

10. The electronic device according to claim 9, characterized in that The electronic device is an autonomous mobile device.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

12. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Robot RGB-D SLAM method based on geometric and motion constraints in dynamic environment

    CN112378409A

  • Semantic map construction method and device oriented to dynamic environment

    CN113570713A

  • Minimum dynamic SLAM method based on semantics and multi-object errors and robot

    CN115326073A

  • Instant positioning and mapping method and device in dynamic environment and electronic equipment

    CN116412809A

  • Robot positioning method, device and equipment, robot and storage medium

    CN117392229A