Visual slam method, electronic device, storage medium, and product

By unifying the expression of point features, line features, surface features, and Manhattan world features in visual SLAM methods, the problems of inconsistent feature expression and insufficient robustness in visual SLAM methods are solved, and a higher-precision visual SLAM system is realized.

CN114972491BActive Publication Date: 2026-01-02MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210530318.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-16
Publication Date
2026-01-02
Estimated Expiration
2042-05-16

AI Technical Summary

Technical Problem

Existing visual SLAM methods typically use only one feature, resulting in insufficient robustness in low-texture scenes. Furthermore, the inconsistent representation of different features makes optimization complex and leads to low accuracy.

Method used

A unified expression method is adopted to represent point features, line features, surface features and Manhattan world features. By combining the origin, direction vector, normal vector and orthogonal vector of the sub-coordinate system, a unified expression for each feature is formed, and visual SLAM is performed based on this.

Benefits of technology

It improves the robustness and accuracy of the visual SLAM system, enables consistent representation across different types of features, and enhances the accuracy and stability of visual SLAM.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972491B_ABST
    Figure CN114972491B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of robots, and provides a visual SLAM method, an electronic device, a storage medium and a product, the method comprising: acquiring multiple types of SLAM features, the types of SLAM features including at least two types of point features, line features, surface features and Manhattan world features; uniformly expressing each type of SLAM feature based on a unified expression mode; and performing visual SLAM based on the uniformly expressed different types of SLAM features to obtain visual SLAM result data. The visual SLAM method provided in the present application embodiment uniformly expresses different types of visual SLAM features, which is conducive to improving the expression uniformity of different types of visual SLAM features, and helps to improve the robustness and accuracy of the visual SLAM system when performing visual SLAM on multiple visual SLAM features.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robots, and in particular to a visual SLAM method, an electronic device, a storage medium and a product. BACKGROUND

[0002] Simultaneous Localization and Mapping (SLAM) is a key technology for robots to realize autonomous positioning, navigation planning and task execution. Visual SLAM is mainly divided into two categories: direct method based on grayscale and indirect method based on features. The features used by the feature-based visual SLAM method include point features, line features, surface features and Manhattan world features, etc. However, the current visual SLAM method usually only uses one of the features, which can easily lead to insufficient system robustness. For example, in a low-texture scene, the visual SLAM based on point features cannot work normally, and in a scene facing only one plane, the visual SLAM based on surface features often loses positioning due to too few constraints. Some visual SLAM methods using multiple features have inconsistent expression methods, complex optimization, and cannot utilize the constraint relationship between features, resulting in low visual SLAM accuracy. SUMMARY

[0003] The present application aims to at least solve one of the technical problems existing in the prior art.

[0004] To this end, the present application provides a visual SLAM method, obtaining multiple types of SLAM features corresponding to a SLAM system, uniformly expressing each type of SLAM feature based on a unified expression method, so that each visual SLAM feature has its corresponding expression, even if the visual SLAM features of different types, which is conducive to improving the expression uniformity of visual SLAM features of different types, and can more accurately express the pose of each visual SLAM feature. When performing visual SLAM, multiple visual SLAM features after uniform expression can be used for visual SLAM based on multiple visual SLAM features, which helps to improve the robustness and accuracy of the visual SLAM system.

[0005] The present application also provides a visual SLAM device.

[0006] The present application also provides an electronic device.

[0007] The present application also provides a non-transitory computer-readable storage medium.

[0008] The present application also provides a computer program product.

[0009] According to the visual SLAM method of the first aspect of the present application, the method comprises:

[0010] acquire a plurality of types of SLAM features, the types of the SLAM features including at least two types of point features, line features, surface features and Manhattan world features;

[0011] unified expression of each type of the SLAM features based on a unified expression manner;

[0012] visual SLAM based on the different types of the SLAM features after the unified expression, to obtain visual SLAM result data.

[0013] According to the visual SLAM method of the embodiment of the application, the plurality of types of SLAM features such as point features, line features, surface features and Manhattan world features are unified expressed, so that each visual SLAM feature has its corresponding unified expression, even the visual SLAM features of different types can be expressed through the unified expression, which is beneficial to improve the expression uniformity of the visual SLAM features of different types, and can more accurately express the pose of each visual SLAM feature. When the visual SLAM is performed, the visual SLAM based on the plurality of visual SLAM features after the unified expression can be performed, which is helpful to improve the robustness and accuracy of the visual SLAM system.

[0014] According to an embodiment of the application, the unified expression of each type of the SLAM features based on the unified expression manner comprises:

[0015] unified expression of each type of the SLAM features based on a sub-coordinate unified expression manner to obtain a unified expression, the unified expression including the sub-coordinate system origin, the sub-coordinate system direction vector, the sub-coordinate system normal vector and the sub-coordinate system orthogonal vector corresponding to each type of the SLAM features.

[0016] According to the visual SLAM method of the embodiment of the application, the sub-coordinate system origin, the sub-coordinate system direction vector, the sub-coordinate system normal vector and the sub-coordinate system orthogonal vector of each visual SLAM feature are obtained, so that each visual SLAM feature has its corresponding expression, even the visual SLAM features of different types are expressed through the expression with the four parameters of the sub-coordinate system origin, the sub-coordinate system direction vector, the sub-coordinate system normal vector and the sub-coordinate system orthogonal vector, which is beneficial to improve the expression uniformity of the visual SLAM features of different types, and can more accurately express the pose of each visual SLAM feature.

[0017] According to an embodiment of the application, the unified expression of each type of the SLAM features based on the unified expression manner comprises:

[0018] When the visual SLAM feature is a point feature, the unified expression manner is used to respectively perform unified expression on each type of the SLAM feature, including:

[0019] a three-dimensional coordinate of the point feature is taken as an origin of a sub-coordinate system of the point feature;

[0020] a direction vector of the sub-coordinate system of the point feature is taken as a direction vector of the sub-coordinate system of the point feature;

[0021] a normal vector of a plane formed by the direction vector of the sub-coordinate system of the point feature and a Z axis of the reference coordinate system is taken as a normal vector of the sub-coordinate system of the point feature;

[0022] a product of the direction vector of the sub-coordinate system of the point feature and the normal vector of the sub-coordinate system of the point feature is taken as an orthogonal vector of the sub-coordinate system of the point feature;

[0023] the point feature is unified expressed in combination with the origin of the sub-coordinate system of the point feature, the direction vector of the sub-coordinate system of the point feature, the normal vector of the sub-coordinate system of the point feature and the orthogonal vector of the sub-coordinate system of the point feature.

[0024] According to the visual SLAM method, for the point feature, a three-dimensional coordinate of the point feature is taken as an origin of a sub-coordinate system of the point feature, a direction vector of the sub-coordinate system of the point feature is taken as a direction vector of the sub-coordinate system of the point feature, a normal vector of a plane formed by the direction vector of the sub-coordinate system of the point feature and a Z axis of the reference coordinate system is taken as a normal vector of the sub-coordinate system of the point feature, a product of the direction vector of the sub-coordinate system of the point feature and the normal vector of the sub-coordinate system of the point feature is taken as an orthogonal vector of the sub-coordinate system of the point feature, and the point feature is unified expressed in combination with the origin of the sub-coordinate system of the point feature, the direction vector of the sub-coordinate system of the point feature, the normal vector of the sub-coordinate system of the point feature and the orthogonal vector of the sub-coordinate system of the point feature, so that the pose of the point feature in the reference coordinate system can be more accurately expressed, and the accuracy of subsequent visual SLAM is improved.

[0025] When the visual SLAM feature is a line feature, the unified expression manner is used to respectively perform unified expression on each type of the SLAM feature, including:

[0026] a foot point of the line feature from an origin of a reference coordinate system of a SLAM system is taken as an origin of a sub-coordinate system of the line feature;

[0027] a direction vector of the line feature is taken as a direction vector of the sub-coordinate system of the line feature;

[0028] a normal vector of a plane formed by the direction vector of the sub-coordinate system of the line feature and a Z axis of the reference coordinate system is taken as a normal vector of the sub-coordinate system of the line feature;

[0029] The product of the sub-coordinate system direction vector of the line feature and the sub-coordinate system normal vector of the line feature is taken as the sub-coordinate system orthogonal vector of the line feature.

[0030] The line feature is uniformly expressed by combining the sub-coordinate system origin, sub-coordinate system direction vector, sub-coordinate system normal vector, and sub-coordinate system orthogonal vector.

[0031] According to an embodiment of the present invention, a visual SLAM method for line features takes the foot of the perpendicular from the origin of the reference coordinate system of the SLAM system to the line feature as the origin of the sub-coordinate system of the line feature, takes the direction vector of the line feature as the direction vector of the sub-coordinate system of the line feature, takes the normal vector of the plane formed by the direction vector of the sub-coordinate system of the line feature and the Z-axis of the reference coordinate system as the normal vector of the sub-coordinate system of the line feature, and takes the product of the direction vector of the sub-coordinate system of the line feature and the normal vector of the sub-coordinate system of the line feature as the orthogonal vector of the sub-coordinate system of the line feature. Combining the origin, direction vector, normal vector and orthogonal vector of the sub-coordinate system of the line feature, an expression for the line feature is formed, which can more accurately express the pose of the line feature in the reference coordinate system and helps to improve the accuracy of subsequent visual SLAM.

[0032] When the visual SLAM feature is a surface feature, the unified expression of each type of SLAM feature based on a unified expression method includes:

[0033] The foot of the perpendicular from the origin of the SLAM system's reference coordinate system to the surface feature is taken as the origin of the surface feature's sub-coordinate system.

[0034] The projection direction of the Z-axis of the reference coordinate system onto the surface feature is taken as the sub-coordinate system direction vector of the surface feature.

[0035] The normal vector of the surface feature is used as the normal vector of the sub-coordinate system of the surface feature;

[0036] The product of the sub-coordinate system direction vector of the surface feature and the sub-coordinate system normal vector of the surface feature is taken as the sub-coordinate system orthogonal vector of the surface feature.

[0037] The surface feature is uniformly expressed by combining the origin, direction vector, normal vector, and orthogonal vector of the sub-coordinate system of the surface feature.

[0038] According to the visual SLAM method, for a plane feature, the origin of a reference coordinate system of a SLAM system to a foot point of the plane feature is taken as an origin of a sub-coordinate system of the plane feature, a projection direction of a Z-axis of the reference coordinate system on the plane feature is taken as a direction vector of the sub-coordinate system of the plane feature, a normal vector of the plane feature is taken as a normal vector of the sub-coordinate system of the plane feature, and a product of the direction vector of the sub-coordinate system of the plane feature and the normal vector of the sub-coordinate system of the plane feature is taken as an orthogonal vector of the sub-coordinate system of the plane feature, and then the expression of the plane feature is formed by combining the origin of the sub-coordinate system of the plane feature, the direction vector of the sub-coordinate system of the plane feature, the normal vector of the sub-coordinate system of the plane feature and the orthogonal vector of the sub-coordinate system of the plane feature, so that the pose of the plane feature in the reference coordinate system can be more accurately expressed, and the accuracy of subsequent visual SLAM is improved.

[0039] When the visual SLAM feature is a Manhattan world feature, the unified expression of each type of the SLAM feature based on the unified expression manner comprises:

[0040] The intersection of the three orthogonal planes of the Manhattan world feature is taken as the origin of the sub-coordinate system of the Manhattan world feature.

[0041] The normal vectors corresponding to the three orthogonal planes of the Manhattan world feature are taken as the direction vector of the sub-coordinate system of the Manhattan world feature, the normal vector of the sub-coordinate system of the Manhattan world feature and the orthogonal vector of the sub-coordinate system of the Manhattan world feature, respectively.

[0042] The Manhattan world feature is expressed by combining the origin of the sub-coordinate system of the Manhattan world feature, the direction vector of the sub-coordinate system of the Manhattan world feature, the normal vector of the sub-coordinate system of the Manhattan world feature and the orthogonal vector of the sub-coordinate system of the Manhattan world feature.

[0043] According to the visual SLAM method, for a Manhattan world feature, the intersection of the three orthogonal planes of the Manhattan world feature is taken as the origin of the sub-coordinate system of the Manhattan world feature, the normal vectors corresponding to the three orthogonal planes of the Manhattan world feature are taken as the direction vector of the sub-coordinate system of the Manhattan world feature, the normal vector of the sub-coordinate system of the Manhattan world feature and the orthogonal vector of the sub-coordinate system of the Manhattan world feature, respectively, and then the expression of the Manhattan world feature is formed by combining the origin of the sub-coordinate system of the Manhattan world feature, the direction vector of the sub-coordinate system of the Manhattan world feature, the normal vector of the sub-coordinate system of the Manhattan world feature and the orthogonal vector of the sub-coordinate system of the Manhattan world feature, so that the pose of the Manhattan world feature in the reference coordinate system can be more accurately expressed, and the accuracy of subsequent visual SLAM is improved.

[0044] According to an embodiment of the present application, the visual SLAM is performed according to the expression of the visual SLAM feature to obtain visual SLAM result data, which comprises:

[0045] The actual visual SLAM feature is matched with the pre-stored visual SLAM feature to obtain a visual SLAM feature pair.

[0046] According to the expression of the actual visual SLAM feature and the expression of the pre-stored visual SLAM feature, the distance between the two visual SLAM features of the visual SLAM feature pair is obtained.

[0047] According to the distance between the two visual SLAM features of the visual SLAM feature pair, visual SLAM is performed to obtain visual SLAM result data.

[0048] According to the visual SLAM method of the embodiment of the application, the actual visual SLAM feature obtained by the robot is matched with the pre-stored visual SLAM feature in the map one by one to obtain a visual SLAM feature pair, the distance between the two visual SLAM features of the visual SLAM feature pair is accurately calculated according to the expression of the actual visual SLAM feature and the expression of the pre-stored visual SLAM feature, and then the visual SLAM is sufficiently robust according to the distance between the two visual SLAM features of the visual SLAM feature pair, which helps to improve the accuracy of the obtained visual SLAM result data.

[0049] According to an embodiment of the application, the distance between the two visual SLAM features of the visual SLAM feature pair is obtained according to the expression of the actual visual SLAM feature and the expression of the pre-stored visual SLAM feature, and specifically:

[0050] According to the expression of the actual visual SLAM feature and the expression of the pre-stored visual SLAM feature, the distance between the origin points of the sub-coordinate systems corresponding to the actual visual SLAM feature and the pre-stored visual SLAM feature, the distance between the direction vectors of the sub-coordinate systems, and the distance between the normal vectors of the sub-coordinate systems are obtained.

[0051] According to the distance between the origin points of the sub-coordinate systems corresponding to the actual visual SLAM feature and the pre-stored visual SLAM feature, the distance between the direction vectors of the sub-coordinate systems, and the distance between the normal vectors of the sub-coordinate systems, the distance between the two visual SLAM features of the visual SLAM feature pair is obtained.

[0052] According to the visual SLAM method of the embodiment of the application, according to the feature distance calculation formula, the distance between the two visual SLAM features of each visual SLAM feature pair can be accurately calculated based on the origin points of the sub-coordinate systems, the direction vectors of the sub-coordinate systems, the normal vectors of the sub-coordinate systems of the actual visual SLAM feature, and the origin points of the sub-coordinate systems, the direction vectors of the sub-coordinate systems, and the normal vectors of the sub-coordinate systems of the pre-stored visual SLAM feature, which helps to improve the accuracy of visual SLAM and the accuracy of the obtained visual SLAM result data.

[0053] According to one embodiment of the present application, the visual SLAM according to the distance between the two visual SLAM features of the visual SLAM feature pair further comprises:

[0054] The distance between the two visual SLAM features of the visual SLAM feature pair is adjusted by an optimization function, to obtain an optimized distance between the two visual SLAM features of the visual SLAM feature pair, wherein the optimization function is determined by the visual SLAM feature pair, the feature parameters of the actual visual SLAM feature, the feature parameters of the pre-stored visual SLAM feature, and the robot pose.

[0055] According to the visual SLAM method of the embodiment of the present application, the distance between the two visual SLAM features of the visual SLAM feature pair is adjusted by the visual SLAM feature pair, the feature parameters of the actual visual SLAM feature, the feature parameters of the pre-stored visual SLAM feature, and the robot pose, which can further improve the accuracy of the distance between the two visual SLAM features of the visual SLAM feature pair.

[0056] According to one embodiment of the present application, the adjusting the distance between the two visual SLAM features of the visual SLAM feature pair further comprises:

[0057] The loop closure detection is performed on the visual SLAM features obtained by the robot in the closed loop by the bag-of-words model, to reduce the cumulative error of the robot pose and the pre-stored visual SLAM feature.

[0058] According to the visual SLAM method of the embodiment of the present application, the loop closure detection is performed on the visual SLAM features obtained by the robot in the closed loop by the bag-of-words model, which can significantly reduce the cumulative error of the robot pose and the pre-stored visual SLAM feature, and help the robot to more accurately and quickly perform obstacle avoidance navigation, thereby improving the accuracy of the obtained visual SLAM result data.

[0059] According to the visual SLAM device of the second aspect of the embodiment of the present application, comprising:

[0060] The feature acquisition module is configured to acquire a plurality of types of SLAM features, and the types of SLAM features include at least two types of point features, line features, surface features, and Manhattan world features.

[0061] The unified expression module is configured to perform unified expression on each type of SLAM feature based on a unified expression manner.

[0062] The visual SLAM module is configured to perform visual SLAM based on the different types of SLAM features after unified expression, to obtain visual SLAM result data.

[0063] According to the visual SLAM device provided in the embodiment of the present application, the expression uniformity of different types of visual SLAM features is improved, the pose of each visual SLAM feature can be more accurately expressed, and when visual SLAM is performed by the visual SLAM module, visual SLAM based on multiple visual SLAM features can be performed according to the multiple visual SLAM features expressed uniformly, which helps to improve the robustness and accuracy of the visual SLAM system.

[0064] According to the electronic device provided in the third aspect of the present application, the electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the visual SLAM method according to any one of the above aspects when executing the program.

[0065] According to the non-transitory computer readable storage medium provided in the fourth aspect of the present application, the computer program is stored on the non-transitory computer readable storage medium, and the processor implements the visual SLAM method according to any one of the above aspects when executing the program.

[0066] According to the computer program product provided in the fifth aspect of the present application, the computer program product comprises a computer program, and the computer program implements the visual SLAM method according to any one of the above aspects when executed by the processor.

[0067] The one or more technical solutions described above in the embodiments of the present application have at least one of the following technical effects:

[0068] By uniformly expressing multiple types of SLAM features such as point features, line features, surface features and Manhattan world features, each visual SLAM feature has its corresponding uniform expression, even if the visual SLAM features are of different types, they can be expressed by the uniform expression, which helps to improve the expression uniformity of different types of visual SLAM features, can more accurately express the pose of each visual SLAM feature, and when visual SLAM is performed, visual SLAM based on multiple visual SLAM features can be performed according to the multiple visual SLAM features expressed uniformly, which helps to improve the robustness and accuracy of the visual SLAM system.

[0069] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0070] In order to make the technical solutions of the embodiments of the present application or the prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description only need to be some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort.

[0071] Figure 1 is a flowchart of a visual SLAM method provided by an embodiment of the present application;

[0072] Figure 2 is a SLAM feature type diagram provided by an embodiment of the present application;

[0073] Figure 3 is a visual SLAM framework diagram provided by an embodiment of the present application;

[0074] Figure 4 is a structural diagram of a visual SLAM device provided by an embodiment of the present application;

[0075] Figure 5 is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0076] The embodiments of the present application will be further described in detail below with reference to the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but cannot be used to limit the scope of the present application.

[0077] Simultaneous localization and mapping (SLAM) is a core key technology for robots to work autonomously in an unknown environment, and is a research focus in the field of robot automation. In an unknown environment, based on the environment perception data obtained by the external sensor of the robot, SLAM constructs the surrounding environment map for the robot, provides the position of the robot in the environment map, and performs incremental construction of the environment map and continuous positioning of the robot as the robot moves. It is the basis for realizing the environment perception and automated operation of the robot. In SLAM, a distance sensor is generally used as the data source for environment perception. Compared with radar, sonar and other range measuring instruments, a visual sensor has the characteristics of small size, low power consumption and rich information acquisition, and can provide rich external environment texture information for various robots. Therefore, visual SLAM has become a research hotspot. Because the visual information obtained by the camera is easily disturbed by the environment and has large noise, the processing of visual SLAM is difficult and complex. At present, with the continuous development of computer vision technology, the technical level of visual SLAM has also been improved, and has been preliminarily applied in the fields of indoor autonomous navigation, VR / AR, etc.

[0078] The visual SLAM method, device, electronic equipment, storage medium and product provided by the embodiments of the present application are described in detail below.

[0079] Figure 1 A flowchart of a visual SLAM method provided by an embodiment of the present application.

[0080] With reference to Figure 1 The visual SLAM method provided by the embodiments of the present application can include:

[0081] In step 110, a plurality of types of SLAM features are acquired, and the types of the SLAM features include at least two types of point features, line features, surface features and Manhattan world features.

[0082] In step 120, the SLAM features of each type are uniformly expressed based on a unified expression mode.

[0083] In step 130, visual SLAM is performed based on the uniformly expressed SLAM features of different types, and visual SLAM result data is obtained.

[0084] It should be noted that the visual SLAM method provided by the present application can be applied to various products, such as robots, unmanned aerial vehicles or mobile terminals, etc., and the corresponding products can be equipped with 3D scanning devices and processors. The robot can be a household service robot, a cleaning robot, a wheeled mobile robot, a biped or multi-legged mobile robot, etc. The mobile terminal can be a mobile phone with a laser radar sensor, etc.

[0085] In some embodiments of the present application, a robot is taken as an example for description. The robot is provided with a sensor combination capable of detecting environmental data of a space, and the sensor combination includes at least one or more 3D scanning devices such as laser radar sensors. The robot is also provided with a processor capable of transmitting and receiving instructions and processing data information. The type of the laser radar sensor can be a single-line laser radar, a multi-line laser radar or a solid-state laser radar, etc.

[0086] The robot can acquire point cloud data of a navigation area through the 3D scanning device. The point cloud data refers to a collection of a plurality of points in a three-dimensional coordinate system. In addition to geometric positions, some point cloud data also has color information. The color information is usually obtained from a color image of a depth camera, and then the color information of the pixels corresponding to the positions is assigned to the corresponding points in the point cloud. In a specific implementation, the robot performs data processing based on the acquired point cloud data or depth image, and then acquires a plurality of types of SLAM features.

[0087] In step 110, the robot acquires multiple types of SLAM features, including point features, line features, surface features, and Manhattan world features, etc. For different types of visual SLAM features, they can be expressed by sub-coordinate systems according to their specific poses. Then, the robot can perform visual SLAM according to the expressions of different types of visual SLAM features, avoiding the robustness deficiency of visual SLAM system caused by single feature.

[0088] In step 120, the robot uniformly expresses each type of SLAM feature based on the unified expression method to obtain uniformly expressed visual SLAM features. The unified expression means that different SLAM features are uniformly described by the same expression form.

[0089] In step 130, the robot performs visual SLAM based on the uniformly expressed different types of SLAM features to obtain visual SLAM result data.

[0090] It should be noted that the expression methods of the multiple visual SLAM features after unified expression are consistent, which can support the robot to perform any visual SLAM method based on the expressed visual SLAM features, and the present application does not limit this.

[0091] The visual SLAM method of the embodiment of the present application uniformly expresses multiple types of SLAM features such as point features, line features, surface features, and Manhattan world features, so that each visual SLAM feature has its corresponding unified expression. Even different types of visual SLAM features can be expressed by unified expressions, which is beneficial to improve the expression uniformity of different types of visual SLAM features, can more accurately express the pose of each visual SLAM feature, and can perform visual SLAM based on multiple visual SLAM features according to the multiple visual SLAM features after unified expression, which helps to improve the robustness and accuracy of the visual SLAM system.

[0092] In some embodiments of the present application, the unified expression of each type of SLAM feature based on the unified expression method includes: uniformly expressing each type of SLAM feature based on the unified expression method of the sub-coordinate system to obtain a unified expression, which includes the origin of the sub-coordinate system, the direction vector of the sub-coordinate system, the normal vector of the sub-coordinate system, and the orthogonal vector of the sub-coordinate system corresponding to each type of SLAM feature.

[0093] It should be noted that, since the visual SLAM is based on the reference coordinate system of the visual SLAM system, the expression can be defined as the expression of the sub-coordinate system, for example, L=(O,Vd ,V n ,V o ), L represents a target expression, O represents a sub-coordinate system origin, V d represents a sub-coordinate system direction vector, V n represents a sub-coordinate system normal vector, and V o represents a sub-coordinate system orthogonal vector. When forming an expression of a visual SLAM feature, the sub-coordinate system origin, the sub-coordinate system direction vector, the sub-coordinate system normal vector, and the sub-coordinate system orthogonal vector of the corresponding sub-coordinate system can be obtained according to different types of visual SLAM features, so that each visual SLAM feature is expressed through the corresponding sub-coordinate system, the pose relationship between the visual SLAM feature and the reference coordinate system can be better expressed, and the accuracy of the visual SLAM can be improved.

[0094] Figure 2 is a SLAM feature type schematic diagram provided by an embodiment of the present application, referring to Figure 2 It can be seen that the point feature alpha, the line feature beta, the surface feature delta, and the Manhattan world feature epsilon are related in the sub-coordinate system with O as the origin.

[0095] In some embodiments of the present application, each type of SLAM feature is uniformly expressed based on a unified expression method, including:

[0096] When the visual SLAM feature is a point feature, the robot can first extract a two-dimensional planar point feature through an ORB (Oriented Fast and Rotated Brief, feature extraction algorithm) method or the like, then obtain a three-dimensional coordinate of the point feature by using depth information, and then take the three-dimensional coordinate of the point feature as a sub-coordinate system origin O1 of the point feature, take a direction of a line connecting the reference coordinate system of the visual SLAM system and the three-dimensional coordinate of the point feature as a sub-coordinate system direction vector V d1 of the point feature, take a normal vector of a plane formed by the sub-coordinate system direction vector of the point feature and the Z-axis of the reference coordinate system as a sub-coordinate system normal vector V n1 of the point feature, take a product V d1 of the sub-coordinate system direction vector of the point feature and the sub-coordinate system normal vector V n1 of the point feature as a sub-coordinate system orthogonal vector V d1 of the point feature, and then form an expression of the point feature in combination with the sub-coordinate system origin, the sub-coordinate system direction vector, the sub-coordinate system normal vector, and the sub-coordinate system orthogonal vector of the point feature, so that the expression of the point feature is: L1=(O1, V n1 ,V o1 . d1 ,V n1 ,V o1The sub-coordinate system of the point feature can more accurately express the pose of the point feature in the reference coordinate system of the visual SLAM system, and helps to improve the accuracy of subsequent visual SLAM.

[0097] It should be noted that ORB is the abbreviation of Oriented Fast and Rotated Brief, which can be used to quickly create feature vectors for key points in an image. These feature vectors can be used to identify objects in the image. Fast and Brief are feature detection algorithms and vector creation algorithms, respectively. ORB first searches for special areas in the image, called key points. Key points are small areas in the image that stand out, such as corners, for example, they have a feature of a sharp change in pixel value from light to dark. Then ORB calculates the corresponding feature vector for each key point. The feature vector created by the ORB algorithm only contains 1 and 0, which is called a binary feature vector. The order of 1 and 0 changes according to the specific key point and its surrounding pixel area. The vector represents the intensity pattern around the key point, so multiple feature vectors can be used to identify larger areas, or even specific objects in the image. Point features can be quickly extracted by ORB, and to some extent, are not affected by noise and image transformations, such as rotation and scaling transformations.

[0098] When the visual SLAM feature is a line feature, the robot can first extract a two-dimensional straight line in the plane by using a method such as CannyLines (line segment detection), and then obtain the three-dimensional coordinates of the line end points by using line fitting. Then, the foot of the perpendicular from the origin of the reference coordinate system of the visual SLAM system to the line feature is taken as the origin O2 of the sub-coordinate system of the line feature, and the direction vector of the line feature is taken as the direction vector V d2 of the sub-coordinate system of the line feature. d2 The normal vector of the plane formed by the Z-axis of the reference coordinate system and the direction vector V n2 of the sub-coordinate system of the line feature is taken as the normal vector V d2 of the sub-coordinate system of the line feature. n2 The product V d2 ×V n2 is taken as the orthogonal vector V o2 of the sub-coordinate system of the line feature. d2 The expression of the line feature is formed by combining the origin of the sub-coordinate system of the line feature, the direction vector of the sub-coordinate system of the line feature, the normal vector of the sub-coordinate system of the line feature, and the orthogonal vector of the sub-coordinate system of the line feature, so that the expression of the line feature is: L2=(O2,V n2 ,V o2 ). Through the sub-coordinate system of the line feature, the pose of the line feature in the reference coordinate system of the visual SLAM system can be more accurately expressed, which helps to improve the accuracy of subsequent visual SLAM.

[0099] It should be noted that the robot can detect and extract edge images using the CannyLines line segment detection method based on gradient magnitude in existing technology, collect collinear point groups from the edge images, and fit the collinear point groups into straight lines using the least squares fitting method to obtain two-dimensional plane straight lines.

[0100] When visual SLAM features are surface features, the robot can first extract the 3D plane using methods such as AHC (hierarchical clustering) to obtain surface features. Then, the foot of the perpendicular from the origin of the visual SLAM system's reference coordinate system to the surface feature is taken as the origin O3 of the surface feature's sub-coordinate system, and the projection direction of the Z-axis of the reference coordinate system onto the surface feature is taken as the sub-coordinate system direction vector V of the surface feature. d3 The normal vector of the surface feature is used as the normal vector V of the sub-coordinate system of the surface feature. n3 The sub-coordinate system direction vector V of the surface feature d3 The sub-coordinate normal vector V of the surface feature n3 The product V d3 ×V n3 V, as the orthogonal vector of the sub-coordinate system of the surface feature o3 Then, combining the origin, direction vector, normal vector, and orthogonal vector of the sub-coordinate system of the surface feature, an expression for the surface feature is formed, such that the expression for the surface feature is: L3=(O3,V d3 V n3 V o3 By using a sub-coordinate system for surface features, the pose of surface features in the reference coordinate system of the visual SLAM system can be expressed more accurately, which helps to improve the accuracy of subsequent visual SLAM.

[0101] In some embodiments, the implementation process of the AHC hierarchical clustering method can be as follows:

[0102] 1) First, establish a min-heap data structure to more effectively find the data with the minimum mean square error for fusion;

[0103] 2) Calculate the mean square error of the fused plane fitting again, and find the two corresponding nodes with the minimum mean square error;

[0104] 3) If the mean square error exceeds a pre-set threshold (not fixed), a plane segmentation node is found, and it is extracted from the graph to obtain the surface feature. Otherwise, the two nodes are merged and added back to the constructed graph, the min-heap is updated, and steps 2) and 3) are repeated.

[0105] When the visual SLAM feature is a Manhattan world feature, the robot can directly take the intersection of the three orthogonal planes of the Manhattan world feature as the origin O4 of the sub-coordinate system of the Manhattan world feature, and take the normal vectors corresponding to the three orthogonal planes of the Manhattan world feature as the direction vectors V d4 of the sub-coordinate system of the Manhattan world feature n4 of the sub-coordinate system of the Manhattan world feature o4 Then, the origin of the sub-coordinate system, the direction vectors, the normal vectors and the orthogonal vectors of the sub-coordinate system of the Manhattan world feature are combined to form the expression of the Manhattan world feature, so that the expression of the Manhattan world feature is: L4=(O4,V d4 ,V n4 ,V o4 The sub-coordinate system of the Manhattan world feature can more accurately express the pose of the Manhattan world feature in the reference coordinate system of the visual SLAM system, which helps to improve the accuracy of subsequent visual SLAM.

[0106] Specifically, step 130 can include:

[0107] The robot matches the actual visual SLAM feature with the pre-stored visual SLAM feature to obtain a visual SLAM feature pair;

[0108] The robot obtains the distance between the two visual SLAM features of the visual SLAM feature pair according to the expression of the actual visual SLAM feature and the expression of the pre-stored visual SLAM feature;

[0109] The robot performs visual SLAM according to the distance between the two visual SLAM features of the visual SLAM feature pair to obtain visual SLAM result data.

[0110] It should be noted that in visual SLAM, image observation information needs to be associated with the environment, i.e., to determine the correspondence between the sequence image content and the real environment. In current visual SLAM, corner features are often used for correlation between sequence images. By extracting and tracking feature points between images, the correspondence between spatial points and homonymous image points is formed between multiple images. Due to the different positions and angles of the camera when acquiring sequence images, combined with the changes in environmental lighting, the appearance of homonymous points on sequence images must change, which requires that the feature point expression should not be affected by image geometric changes such as rotation, scaling, inclination and lighting changes.

[0111] It should be noted that the actual visual SLAM feature is actually obtained by the robot during walking, and after the actual visual SLAM feature is obtained, the robot can first form the expression of the actual visual SLAM feature by taking the sub-coordinate system origin point, the sub-coordinate system direction vector, the sub-coordinate system normal vector and the sub-coordinate system orthogonal vector of the actual visual SLAM feature. The pre-stored visual SLAM feature can be pre-stored in the map by the robot in advance, and the pre-stored visual SLAM feature in the map is expressed by its corresponding expression. When performing visual SLAM, the robot can first match the actual visual SLAM feature with the pre-stored visual SLAM feature one by one, and associate the data to obtain a visual SLAM feature pair, and then the robot can calculate the distance between the two visual SLAM features based on the expression of the actual visual SLAM feature in the visual SLAM feature pair and the expression of the pre-stored visual SLAM feature, and then perform visual SLAM according to the distance between the two visual SLAM features of the visual SLAM feature pair to obtain accurate visual SLAM result data.

[0112] According to the visual SLAM method of the embodiment of the application, each visual SLAM feature has its corresponding expression by obtaining the sub-coordinate system origin point, the sub-coordinate system direction vector, the sub-coordinate system normal vector and the sub-coordinate system orthogonal vector of each visual SLAM feature, that is, even different types of visual SLAM features are expressed by the expression with the four parameters of the sub-coordinate system origin point, the sub-coordinate system direction vector, the sub-coordinate system normal vector and the sub-coordinate system orthogonal vector, which is beneficial to improve the expression uniformity of different types of visual SLAM features and can more accurately express the pose of each visual SLAM feature. When performing visual SLAM, visual SLAM based on multiple visual SLAM features can be performed according to the uniformly expressed multiple visual SLAM features, which is helpful to improve the robustness and accuracy of the visual SLAM system.

[0113] Further, according to an embodiment of the application, the distance between the two visual SLAM features of the visual SLAM feature pair is obtained according to the expression of the actual visual SLAM feature and the expression of the pre-stored visual SLAM feature, and specifically can be:

[0114] According to the expression of the actual visual SLAM feature and the expression of the pre-stored visual SLAM feature, the distance between the two visual SLAM features of the visual SLAM feature pair is obtained by a feature distance calculation formula, and the feature distance calculation formula is:

[0115]

[0116] wherein e(L', L") represents the distance between the actual visual SLAM feature and the pre-stored visual SLAM feature, e O represents the distance between the sub-coordinate system origin of the actual visual SLAM feature and the sub-coordinate system origin of the pre-stored visual SLAM feature, e d represents the distance between the sub-coordinate system direction vector of the actual visual SLAM feature and the sub-coordinate system direction vector of the pre-stored visual SLAM feature, e n represents the distance between the sub-coordinate system normal vector of the actual visual SLAM feature and the sub-coordinate system normal vector of the pre-stored visual SLAM feature, L O ' represents the sub-coordinate system origin of the actual visual SLAM feature, L O " represents the sub-coordinate system origin of the pre-stored visual SLAM feature, represents the sub-coordinate system direction vector of the actual visual SLAM feature, represents the sub-coordinate system direction vector of the pre-stored visual SLAM feature, represents the sub-coordinate system normal vector of the actual visual SLAM feature, represents the sub-coordinate system normal vector of the pre-stored visual SLAM feature.

[0117] According to the visual SLAM method of the embodiment of the present application, the distance between the two visual SLAM features of each pair of visual SLAM feature pairs can be accurately calculated based on the sub-coordinate system origin, the sub-coordinate system direction vector, the sub-coordinate system normal vector of the actual visual SLAM feature, and the sub-coordinate system origin, the sub-coordinate system direction vector, the sub-coordinate system normal vector of the pre-stored visual SLAM feature according to the feature distance calculation formula, which helps to improve the accuracy of visual SLAM and the accuracy of the obtained visual SLAM result data.

[0118] Further, according to the embodiment of the present application, the visual SLAM is performed according to the distance between the two visual SLAM features of the visual SLAM feature pair, and before that, the method further comprises:

[0119] optimizing and adjusting the distance between the two visual SLAM features of the visual SLAM feature pair.

[0120] According to the visual SLAM method of the embodiment of the present application, the distance between the two visual SLAM features of the visual SLAM feature pair is optimized and adjusted, which can further improve the accuracy of the distance between the two visual SLAM features of each pair of visual SLAM feature pairs, thereby ensuring the accuracy of visual SLAM.

[0121] According to an embodiment of the present application, the optimization adjustment refers to adjusting the distance between the two visual SLAM features of the visual SLAM feature pair by an optimization function, to obtain the distance between the two visual SLAM features of the visual SLAM feature pair after optimization.

[0122] The optimization function is:

[0123]

[0124] Wherein, T represents the robot pose, L W represents the feature parameters of the pre-stored visual SLAM feature, represents the visual SLAM feature pair, L C represents the feature parameters of the actual visual SLAM feature.

[0125] According to the visual SLAM method of the embodiment of the present application, the distance between the two visual SLAM features of the visual SLAM feature pair is adjusted by the visual SLAM feature pair, the feature parameters of the actual visual SLAM feature, the feature parameters of the pre-stored visual SLAM feature, and the robot pose, which can further improve the accuracy of the distance between the two visual SLAM features of the visual SLAM feature pair.

[0126] Further, according to an embodiment of the present application, the adjusting the distance between the two visual SLAM features of the visual SLAM feature pair further comprises:

[0127] The loop closure detection is performed on the visual SLAM features obtained by the robot in the closed loop according to the bag-of-words model, so as to reduce the cumulative error of the robot pose and the pre-stored visual SLAM feature.

[0128] It should be noted that the bag-of-words model can perform loop closure detection on the visual SLAM features obtained by the robot in the closed loop and uniformly expressed according to the target expression. The loop closure detection by the bag-of-words model can be implemented by any method in the prior art, which is not limited in this article.

[0129] According to the visual SLAM method of the embodiment of the present application, the loop closure detection is performed on the visual SLAM features obtained by the robot in the closed loop according to the bag-of-words model, which can significantly reduce the cumulative error of the robot pose and the pre-stored visual SLAM feature, and help the robot to more accurately and quickly perform obstacle avoidance navigation work, thereby improving the accuracy of the obtained visual SLAM result data.

[0130] On the other hand, a typical visual SLAM system includes sensor data, visual odometry, backend optimization, loop closure detection, and mapping. Regarding sensor data, in visual SLAM, this mainly involves reading and preprocessing camera image information. In robots, it may also involve reading and synchronizing information from encoders, inertial sensors, etc. Regarding visual odometry, its main task is to estimate camera motion between adjacent images and the appearance of a local map; the simplest example is the motion relationship between two images. How does a computer determine camera motion from images? On an image, we can only see individual pixels, knowing they are the result of projections of certain spatial points onto the camera's imaging plane. Therefore, we must first understand the geometric relationship between the camera and spatial points. The odometry (also known as the front end) can estimate camera motion from images between adjacent frames and reconstruct the spatial structure of the scene; this is called odometry. It's called odometry because it only calculates motion at adjacent moments and has no relation to past information. The motion at adjacent moments is chained together to form the robot's trajectory, thus solving the localization problem. Based on the camera position at each moment, the position of the spatial point corresponding to each pixel is calculated, resulting in a map. Regarding backend optimization, it mainly deals with noise issues during the SLAM process. All sensors have noise, so in addition to dealing with "how to estimate camera motion from an image," we also need to consider how much noise this estimate contains. The front end provides the back end with the data to be optimized, along with the initial values ​​of that data. The back end is responsible for the overall optimization process; it often only deals with the data and doesn't need to care where that data comes from. In visual SLAM, the front end is more relevant to computational and visual research fields, such as image feature extraction and matching, while the back end mainly involves filtering and nonlinear optimization algorithms. Regarding loop closure detection, also known as loop shut-off detection, it refers to the robot's ability to recognize scenes it has visited. Successful detection can significantly reduce accumulated errors. Loop closure detection is essentially an algorithm for detecting the similarity of observed data. For visual SLAM, most systems use the relatively mature Bag-of-Words (BoW) model. The Bag-of-Words model clusters visual features (SIFT, SURF, etc.) in an image, then builds a dictionary to find which "words" are contained in each image. Alternatively, traditional pattern recognition methods can be used to construct loop closure detection as a classification problem, training a classifier to perform the classification. Regarding mapping, mapping primarily involves creating a map that corresponds to the task requirements based on the estimated trajectory. In robotics, map representation mainly includes four types: grid maps, direct representation methods, topological maps, and feature point maps. Feature point maps represent the environment using relevant geometric features (such as points, lines, and surfaces), and are commonly used in visual SLAM technology. Figure 1Generally produced by vSLAM algorithm in sparse mode such as GPS, UWB and camera, the advantage is relatively small data storage and operation amount, and is mostly seen in the earliest SLAM algorithm.

[0131] To solve the above SLAM problem, the embodiment of the application provides a visual SLAM system, referring to Figure 3 , the input of the SLAM system is point cloud data information or an RGB-D image (RGB image and depth map), then point features, line features, surface features or Manhattan world features are extracted according to the RGB-D image, respectively, and each type of SLAM feature is uniformly expressed based on a unified expression mode, after obtaining the unified expression of multiple types of features, data association is performed on the extracted features to obtain feature pairs, and a new key frame and landmark are created after rough optimization and fine optimization are sequentially performed according to the initial pose of the robot and the feature pairs, so as to construct a landmark map and a key frame map. While constructing the landmark map and the key frame map, local optimization is performed through loop detection, and finally the map and the robot pose are overall optimized to output the map and the robot pose.

[0132] It should be noted that the working mode of most visual SLAM systems is to track the key points through continuous camera frames, locate the 3D position of the key points by using a triangular algorithm, and use the information to approximate the pose of the camera. In simple terms, the goal of these systems is to draw a map of the environment related to the position of the robot. This map can be used for the robot system to navigate in the environment. Unlike other forms of SLAM technology, this can be done with only a 3D vision camera. By tracking a sufficient number of key points in the video frames of the camera, the direction of the sensor and the structure of the surrounding physical environment can be quickly understood. All visual SLAM systems are constantly working to minimize the reprojection error or the difference between the projected points and the actual points, usually by a solution called Bundle Adjustment (BA). Visual SLAM systems need to operate in real time, which involves a large amount of calculation, so the position data and the mapping data are often separately subjected to Bundle Adjustment, but at the same time, it is convenient to speed up the processing speed before the final merging.

[0133] Further, in the above embodiments, the visual SLAM system includes a MonoSLAM system, a PTAM system, an ORB-SLAM system, and an ORB-SLAM2 system. The MonoSLAM system is the first real-time monocular visual SLAM system. The MonoSLAM uses an EKF (extended Kalman filter) as a backend, tracks sparse feature points in a front end, and updates a mean and a covariance of a current state of a camera and all landmark points. In the EKF, a position of each feature point is subject to a Gaussian distribution, and an ellipsoid can be used to represent the mean and uncertainty thereof, and the longer they are in a certain direction, the more unstable they are in the direction. Disadvantages of the method include a narrow scene, a limited number of landmarks, and easy loss of sparse feature points.

[0134] The PTAM system proposes and implements parallelization of tracking and mapping, and first distinguishes a front end and a back end (tracking needs to respond to image data in real time, and map optimization is performed in the back end), and many subsequent visual SLAM system designs also adopt a similar method. The PTAM is the first scheme using a nonlinear optimization as a back end, rather than a filter back end scheme. A key frame mechanism is proposed, that is, instead of processing each image in detail, several key images are strung together to optimize the trajectory and map thereof. Disadvantages of the method include a small scene and easy loss of tracking.

[0135] The ORB-SLAM system surrounds ORB feature calculation, and includes an ORB dictionary for visual odometry and loop detection. The ORB feature calculation has higher efficiency than SIFT or SURF, and has good rotation and scaling invariance. The ORB-SLAM innovatively uses three threads to complete SLAM, and the three threads are a Tracking thread for tracking feature points in real time, an optimization thread of local Bundle Adjustment, and a loop detection and optimization thread of a global Pose Graph. Disadvantages of the method include that calculation of ORB features for each image is very time-consuming, and the three-thread structure brings a heavy burden to the CPU. A sparse feature point map can only meet positioning needs, and cannot provide navigation, obstacle avoidance, and the like.

[0136] The ORB-SLAM2 system makes the following contributions based on the monocular ORB-SLAM: the first open-source SLAM system for monocular, stereo, and RGB-D, including loop closure, relocalization, and map reuse; RGB-D results show that, by using bundle adjustment, higher accuracy is obtained than the most advanced method based on iterative closest point (ICP) or photometric and depth error minimization; by using close and far stereo points and monocular observation results, the stereo effect is more accurate than the most advanced direct stereo SLAM; a lightweight localization mode can effectively reuse a map when mapping is unavailable.

[0137] The visual SLAM device provided by the embodiment of the present application is described below. The visual SLAM device described below can be referred to the visual SLAM method described above.

[0138] Figure 4 FIG. 1 is a structural schematic diagram of a visual SLAM device provided by an embodiment of the present application.

[0139] Referring to Figure 4 The visual SLAM device provided by the embodiment of the present application can include:

[0140] The visual SLAM device provided by the embodiment of the present application can include:

[0141] The feature acquisition module 410 is configured to acquire a plurality of types of SLAM features, and the types of SLAM features include at least two types of point features, line features, surface features, and Manhattan world features.

[0142] The unified expression module 420 is configured to perform unified expression on each type of SLAM feature based on a unified expression manner.

[0143] The visual SLAM module 430 is configured to perform visual SLAM based on the different types of SLAM features after unified expression to obtain visual SLAM result data.

[0144] The visual SLAM device provided by the embodiment of the present application can obtain a plurality of types of SLAM features through the feature acquisition module, and then perform unified expression on each type of SLAM feature through the unified expression module. The visual SLAM device performs visual SLAM on the unified expression of the plurality of types of SLAM features through the visual SLAM module. This is beneficial to improve the unified expression of different types of visual SLAM features, and can more accurately express the pose of each visual SLAM feature. When the visual SLAM module performs visual SLAM, the visual SLAM can be performed based on the plurality of visual SLAM features after unified expression, which is helpful to improve the robustness and accuracy of the visual SLAM system.

[0145] Further, according to an embodiment of the present application, the unified expression module 420 performs unified expression on each type of SLAM feature based on a unified expression manner, including:

[0146] The unified expression module 420 performs unified expression on each type of SLAM feature based on a unified expression manner, including:

[0147] Further, according to an embodiment of the present application, the unified expression module 420 can include:

[0148] a point feature expression sub-module: for, in the case that the visual SLAM feature is a point feature, taking the three-dimensional coordinates of the point feature as the origin of the sub-coordinate system of the point feature; taking the direction of the line connecting the reference coordinate system of the SLAM system and the three-dimensional coordinates of the point feature as the direction vector of the sub-coordinate system of the point feature; taking the normal vector of the plane formed by the direction vector of the coordinate system of the point feature and the Z-axis of the reference coordinate system as the normal vector of the sub-coordinate system of the point feature; taking the product of the direction vector of the sub-coordinate system of the point feature and the normal vector of the sub-coordinate system of the point feature as the orthogonal vector of the sub-coordinate system of the point feature; and combining the origin of the sub-coordinate system of the point feature, the direction vector of the sub-coordinate system of the point feature, the normal vector of the sub-coordinate system of the point feature, and the orthogonal vector of the sub-coordinate system of the point feature to form the expression of the point feature.

[0149] Further, according to an embodiment of the present application, the unified expression module 420 can include:

[0150] a line feature expression sub-module: for, in the case that the visual SLAM feature is a line feature, taking the foot point of the line feature from the origin of the reference coordinate system of the SLAM system as the origin of the sub-coordinate system of the line feature; taking the direction vector of the line feature as the direction vector of the sub-coordinate system of the line feature; taking the normal vector of the plane formed by the direction vector of the sub-coordinate system of the line feature and the Z-axis of the reference coordinate system as the normal vector of the sub-coordinate system of the line feature; taking the product of the direction vector of the sub-coordinate system of the line feature and the normal vector of the sub-coordinate system of the line feature as the orthogonal vector of the sub-coordinate system of the line feature; and combining the origin of the sub-coordinate system of the line feature, the direction vector of the sub-coordinate system of the line feature, the normal vector of the sub-coordinate system of the line feature, and the orthogonal vector of the sub-coordinate system of the line feature to form the expression of the line feature.

[0151] Further, according to an embodiment of the present application, the unified expression module 420 can include:

[0152] a surface feature expression sub-module: for, in the case that the visual SLAM feature is a surface feature, taking the foot point of the surface feature from the origin of the reference coordinate system of the SLAM system as the origin of the sub-coordinate system of the surface feature; taking the projection direction of the Z-axis of the reference coordinate system on the surface feature as the direction vector of the sub-coordinate system of the surface feature; taking the normal vector of the surface feature as the normal vector of the sub-coordinate system of the surface feature; taking the product of the direction vector of the sub-coordinate system of the surface feature and the normal vector of the sub-coordinate system of the surface feature as the orthogonal vector of the sub-coordinate system of the surface feature; and combining the origin of the sub-coordinate system of the surface feature, the direction vector of the sub-coordinate system of the surface feature, the normal vector of the sub-coordinate system of the surface feature, and the orthogonal vector of the sub-coordinate system of the surface feature to form the expression of the surface feature.

[0153] Further, according to an embodiment of the present application, the unified expression module 420 can include:

[0154] a Manhattan world feature expression sub-module, configured to: in a case where the visual SLAM feature is a Manhattan world feature, take an intersection of three orthogonal planes of the Manhattan world feature as an origin of a sub-coordinate system of the Manhattan world feature; take normal vectors corresponding to the three orthogonal planes of the Manhattan world feature as a direction vector, a normal vector and an orthogonal vector of the sub-coordinate system of the Manhattan world feature, respectively; and form an expression of the Manhattan world feature in combination with the origin of the sub-coordinate system of the Manhattan world feature, the direction vector, the normal vector and the orthogonal vector of the sub-coordinate system of the Manhattan world feature.

[0155] Further, according to an embodiment of the present application, the unified expression module 420 can include:

[0156] a visual SLAM feature matching sub-module, configured to: match the actual visual SLAM feature with the pre-stored visual SLAM feature to obtain a visual SLAM feature pair.

[0157] a visual SLAM feature distance calculation sub-module, configured to: obtain a distance between two visual SLAM features of the visual SLAM feature pair according to the expression of the actual visual SLAM feature and the expression of the pre-stored visual SLAM feature.

[0158] Further, according to an embodiment of the present application, the visual SLAM module 430 can include:

[0159] a visual SLAM sub-module, configured to: perform visual SLAM according to the distance between the two visual SLAM features of the visual SLAM feature pair to obtain visual SLAM result data.

[0160] Further, according to an embodiment of the present application, the visual SLAM feature distance calculation sub-module is specifically configured to:

[0161] obtain distances between origins of sub-coordinate systems, distances between direction vectors of sub-coordinate systems and distances between normal vectors of sub-coordinate systems corresponding to the actual visual SLAM feature and the pre-stored visual SLAM feature according to the expression of the actual visual SLAM feature and the expression of the pre-stored visual SLAM feature;

[0162] obtain the distance between the two visual SLAM features of the visual SLAM feature pair according to the distances between the origins of the sub-coordinate systems, the distances between the direction vectors of the sub-coordinate systems and the distances between the normal vectors of the sub-coordinate systems corresponding to the actual visual SLAM feature and the pre-stored visual SLAM feature.

[0163] Further, according to an embodiment of the present application, the visual SLAM module 430 can further comprise:

[0164] an adjusting sub-module, configured to adjust the distance between the two visual SLAM features of the visual SLAM feature pair before visual SLAM according to the distance between the two visual SLAM features of the visual SLAM feature pair, to obtain the distance between the two visual SLAM features of the visual SLAM feature pair after optimization, wherein the optimization function is determined by the visual SLAM feature pair, the feature parameters of the actual visual SLAM feature, the feature parameters of the pre-stored visual SLAM feature and the robot pose.

[0165] Further, according to an embodiment of the present application, the visual SLAM module 430 can further comprise:

[0166] a second adjusting sub-module, configured to perform loop closure detection according to the visual SLAM features obtained by the robot in the closed loop by the bag-of-words model, to reduce the cumulative error of the robot pose and the pre-stored visual SLAM feature.

[0167] The visual SLAM device provided by the embodiment of the present application can perform visual SLAM based on multiple visual SLAM features, and different types of visual SLAM features are tightly coupled, which can effectively solve the problems of sharp decline in precision of a visual SLAM system based on visual SLAM features in a low-texture environment, system failure and the like; can well realize positioning of a mobile platform and construction of surrounding environment features with structural information according to an indoor structural environment, can well obtain high-precision results on a plurality of public experimental data sets, can utilize matched point, line, surface and Manhattan world features to construct a map of a mobile platform and a surrounding environment in real time and efficiently, and perform loop closure detection processing, and fully utilize loop closure detection to reduce cumulative error.

[0168] Figure 5 An example of a schematic diagram of a physical structure of an electronic device is shown in Figure 3 The electronic device can include a processor 810, a communications interface 820, a memory 830 and a communications bus 840, wherein the processor 810, the communications interface 820 and the memory 830 complete mutual communication through the communications bus 840. The processor 810 can invoke a logical instruction in the memory 830 to execute the following method:

[0169] Obtaining a plurality of types of SLAM features, the types of the SLAM features including at least two types of point features, line features, surface features and Manhattan world features.

[0170] The SLAM features of each type are uniformly expressed based on a uniform expression manner.

[0171] The visual SLAM is performed based on the uniformly expressed SLAM features of different types, to obtain visual SLAM result data.

[0172] In addition, the logic instructions in the memory 830 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0173] Further, the embodiment of the present application discloses a computer program product, the computer program product includes a computer program stored on a non-transitory computer readable storage medium, the computer program includes program instructions, when the program instructions are executed by a computer, the computer can execute the method provided by the above-mentioned method embodiments, for example, including:

[0174] Obtain a plurality of types of SLAM features, the types of the SLAM features including at least two types of point features, line features, surface features, and Manhattan world features.

[0175] The SLAM features of each type are uniformly expressed based on a uniform expression manner.

[0176] The visual SLAM is performed based on the uniformly expressed SLAM features of different types, to obtain visual SLAM result data.

[0177] On the other hand, the embodiment of the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the transmission method provided by the above-mentioned embodiments, for example, including:

[0178] Obtain a plurality of types of SLAM features, the types of the SLAM features including at least two types of point features, line features, surface features, and Manhattan world features.

[0179] The SLAM features of each type are uniformly expressed based on a uniform expression mode.

[0180] The different types of SLAM features after uniform expression are used for visual SLAM to obtain visual SLAM result data.

[0181] The apparatus embodiments described above are merely illustrative, wherein the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0182] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0183] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

[0184] The above embodiments are only used to illustrate the present application, and not to limit the present application. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that various combinations, modifications or equivalent replacements of the technical solutions of the present application do not deviate from the spirit and scope of the technical solutions of the present application, and should be covered in the scope of the technical solutions of the present application.

Claims

1. A visual SLAM method, characterized in that, include: Acquire multiple types of SLAM features, including at least two types of point features, line features, area features, and Manhattan world features; Each type of SLAM feature is uniformly expressed based on a unified expression method; Visual SLAM is performed based on the different types of SLAM features after unified representation to obtain visual SLAM result data; The step of uniformly expressing the SLAM features of each type based on a unified expression method includes: Based on the unified expression method of sub-coordinates, a unified expression is obtained by uniformly expressing the SLAM features of each type. The unified expression includes the origin of the sub-coordinate system, the direction vector of the sub-coordinate system, the normal vector of the sub-coordinate system, and the orthogonal vector of the sub-coordinate system corresponding to the SLAM features of each type. Each visual SLAM feature is based on the same reference coordinate system of the visual SLAM system and is used to express the pose relationship between the visual SLAM feature and the reference coordinate system.

2. The visual SLAM method according to claim 1, characterized in that, The unified expression of each type of SLAM feature based on a unified expression method includes: When the visual SLAM feature is a point feature, the unified expression of each type of SLAM feature based on a unified expression method includes: The three-dimensional coordinates of the point feature are used as the origin of the sub-coordinate system of the point feature; The direction of the line connecting the reference coordinate system of the SLAM system to the three-dimensional coordinates of the point feature is used as the sub-coordinate system direction vector of the point feature. The normal vector of the plane formed by the coordinate system direction vector of the point feature and the Z-axis of the reference coordinate system is taken as the sub-coordinate system normal vector of the point feature. The product of the sub-coordinate system direction vector of the point feature and the sub-coordinate system normal vector of the point feature is taken as the sub-coordinate system orthogonal vector of the point feature. By combining the origin of the sub-coordinate system, the direction vector of the sub-coordinate system, the normal vector of the sub-coordinate system, and the orthogonal vector of the sub-coordinate system of the point features, the point features are uniformly expressed; When the visual SLAM feature is a line feature, the unified expression of each type of SLAM feature based on a unified expression method includes: The foot of the perpendicular from the origin of the SLAM system's reference coordinate system to the line feature is taken as the origin of the line feature's sub-coordinate system. The direction vector of the line feature is used as the sub-coordinate system direction vector of the line feature; The normal vector of the plane formed by the sub-coordinate system direction vector of the line feature and the Z-axis of the reference coordinate system is taken as the sub-coordinate system normal vector of the line feature. The product of the sub-coordinate system direction vector of the line feature and the sub-coordinate system normal vector of the line feature is taken as the sub-coordinate system orthogonal vector of the line feature. By combining the sub-coordinate system origin, sub-coordinate system direction vector, sub-coordinate system normal vector, and sub-coordinate system orthogonal vector of the line feature, the line feature is uniformly expressed; When the visual SLAM feature is a surface feature, the unified expression of each type of SLAM feature based on a unified expression method includes: The foot of the perpendicular from the origin of the SLAM system's reference coordinate system to the surface feature is taken as the origin of the surface feature's sub-coordinate system. The projection direction of the Z-axis of the reference coordinate system onto the surface feature is taken as the sub-coordinate system direction vector of the surface feature. The normal vector of the surface feature is used as the normal vector of the sub-coordinate system of the surface feature; The product of the sub-coordinate system direction vector of the surface feature and the sub-coordinate system normal vector of the surface feature is taken as the sub-coordinate system orthogonal vector of the surface feature. By combining the sub-coordinate system origin, sub-coordinate system direction vector, sub-coordinate system normal vector, and sub-coordinate system orthogonal vector of the surface feature, the surface feature is uniformly expressed; When the visual SLAM features are Manhattan world features, the unified expression of each type of SLAM feature based on a unified expression method includes: The intersection point of the three orthogonal planes of the Manhattan world feature is taken as the origin of the sub-coordinate system of the Manhattan world feature; The normal vectors corresponding to the three orthogonal planes of the Manhattan world feature are respectively used as the sub-coordinate system direction vector, sub-coordinate system normal vector, and sub-coordinate system orthogonal vector of the Manhattan world feature; By combining the origin, direction vector, normal vector, and orthogonal vector of the sub-coordinate system of the Manhattan world features, the Manhattan world features are expressed in a unified manner.

3. The visual SLAM method according to any one of claims 1-2, characterized in that, The visual SLAM result data obtained by performing visual SLAM based on the different types of SLAM features after unified representation includes: The observed visual SLAM features are matched with the preset, stored visual SLAM features to obtain visual SLAM feature pairs. Based on the expression of the actual visual SLAM feature and the expression of the pre-stored visual SLAM feature, the distance between the two visual SLAM features of the visual SLAM feature pair is obtained; Visual SLAM is performed based on the distance between two visual SLAM features of the visual SLAM feature pair to obtain visual SLAM result data.

4. The visual SLAM method according to claim 3, characterized in that, The step of obtaining the distance between two visual SLAM features of the visual SLAM feature pair based on the expression of the actual visual SLAM feature and the expression of the pre-stored visual SLAM feature is specifically as follows: Based on the expressions of the actual visual SLAM features and the expressions of the pre-stored visual SLAM features, the distances between the origins of the sub-coordinate systems, the distances between the direction vectors of the sub-coordinate systems, and the distances between the normal vectors of the sub-coordinate systems corresponding to the actual visual SLAM features and the pre-stored visual SLAM features are obtained respectively. The distance between two visual SLAM features of the visual SLAM feature pair is obtained based on the distances between the origins of the sub-coordinate systems, the distances between the direction vectors of the sub-coordinate systems, and the distances between the normal vectors of the sub-coordinate systems corresponding to the actual visual SLAM features and the pre-stored visual SLAM features, respectively.

5. The visual SLAM method according to claim 3, characterized in that, The step of performing visual SLAM based on the distance between two visual SLAM features of the visual SLAM feature pair, prior to which includes: The distance between the two visual SLAM features of the visual SLAM feature pair is adjusted by an optimization function to obtain the optimized distance between the two visual SLAM features of the visual SLAM feature pair. The optimization function is determined by the visual SLAM feature pair, the feature parameters of the actual visual SLAM features, the feature parameters of the pre-stored visual SLAM features, and the robot pose.

6. The visual SLAM method according to claim 5, characterized in that, The adjustment of the distance between the two visual SLAM features of the visual SLAM feature pair further includes: By using a bag-of-words model to perform loop closure detection based on the visual SLAM features obtained by the robot in a closed loop, the cumulative error of the robot pose and the pre-stored visual SLAM features can be reduced.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the visual SLAM method as described in any one of claims 1 to 6.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the visual SLAM method as described in any one of claims 1 to 6.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the visual SLAM method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Spatial multivariate feature registration optimization method and device based on unified residual model

    CN111815684A