Visual map construction method and device, equipment and storage medium

By dividing the image frames captured by the visual sensor into image groups and constructing map submaps, the problem of low loop detection accuracy during the construction of global maps of visual SLAM in the prior art is solved, and higher accuracy of map construction and pose optimization is achieved.

CN120014185APending Publication Date: 2025-05-16ZHUHAI MOJIE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311535281.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, when building a visual SLAM global map, fewer spatial coordinate points can be compared with the current keyframe during loopback detection, resulting in lower accuracy.

Method used

By dividing multiple image frames of different poses captured by the visual sensor into different image groups, and constructing corresponding map submaps based on the poses and feature points of the image frames in each image group, multiple map submaps together form a complete visual map.

Benefits of technology

It improves the accuracy of map construction and subsequent pose optimization, increases the map points of map submaps that can be compared during loop detection, and improves the accuracy of pose optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014185A_ABST
    Figure CN120014185A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of SLAM mapping, and provides a visual map construction method and device, equipment and a storage medium, and the method comprises the steps: obtaining image frames shot by a visual sensor at all poses; according to the poses corresponding to the image frames, the image frames are divided into at least one image group, and the range of the pose corresponding to each image frame in each image group meets a preset condition; according to the poses and the feature points corresponding to the image frames in the image group, map sub-graphs corresponding to the image group are constructed, the map sub-graphs at least comprise map points corresponding to the feature points, and the visual map comprises at least one map sub-graph corresponding to the image group. Compared with the construction of a global map, the method has the advantages that the plurality of local map sub-maps are firstly constructed, and then the plurality of local map sub-maps form the overall visual map, so that the accuracy of map construction is improved, and the map points in the map sub-maps are more, that is, the map points of the map sub-maps which can be compared during loopback detection are more, and the accuracy of map construction is improved. And the accuracy of pose optimization is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of SLAM mapping technology, and in particular to a method, device, equipment and storage medium for constructing a visual map. Background Art

[0002] SLAM (Simultaneous Localization and Mapping) refers to the technology of positioning based on self-measurement data and external observation data while performing incremental mapping in the process of moving in an unknown environment. It is widely used in fields such as autonomous driving and intelligent robots. Visual maps have the advantages of low cost, high accuracy, and strong updatability, and are the most commonly used map form in SLAM technology.

[0003] In the related technology, visual SLAM Figure 1 Generally, the overall spatial coordinate points are stored, and the overall spatial coordinate points are associated with the feature points on the image frame to build a global map. Then, during loop detection, the spatial coordinate points corresponding to the key frame that is most similar to the current key frame are searched in the global map to correct the pose. However, the related technology constructs a visual SLAM global map. Since there is only one global map, there are relatively few spatial coordinate points that can be compared with the current key frame during loop detection, resulting in low accuracy in subsequent loop detection. Summary of the invention

[0004] The main purpose of this application is to provide a method, device, equipment and storage medium for constructing a visual map, which is conducive to improving the accuracy of map construction and subsequent posture optimization.

[0005] In a first aspect, the present application provides a method for constructing a visual map, comprising:

[0006] Obtain image frames captured by the visual sensor at each posture;

[0007] According to the postures corresponding to the plurality of image frames, the plurality of image frames are divided into at least one image group, and a range of postures corresponding to the image frames in each of the image groups satisfies a preset condition;

[0008] According to the posture and feature points corresponding to each of the image frames in the image group, a map sub-graph corresponding to the image group is constructed, wherein the map sub-graph at least includes map points corresponding to the feature points, and the visual map includes at least one map sub-graph corresponding to the image group.

[0009] In a second aspect, the present application also provides a device for constructing a visual map, comprising:

[0010] An acquisition module is used to acquire image frames captured by the visual sensor at various postures;

[0011] An image division module, used for dividing the plurality of image frames into at least one image group according to the postures corresponding to the plurality of image frames, wherein a range of postures corresponding to the image frames in each of the image groups satisfies a preset condition;

[0012] A map construction module is used to construct a map sub-graph corresponding to the image group according to the posture and feature points corresponding to each image frame in the image group, wherein the map sub-graph at least includes map points corresponding to the feature points, and the visual map includes at least one map sub-graph corresponding to the image group.

[0013] In a third aspect, the present application further provides a computer device, the computer device comprising a memory and a processor;

[0014] Memory for storing computer programs;

[0015] A processor is used to execute a computer program and implement the method for constructing a visual map as described above when executing the computer program.

[0016] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for constructing a visual map as described above are implemented.

[0017] The present application provides a method, apparatus, device and storage medium for constructing a visual map, wherein the method comprises: obtaining image frames taken by a visual sensor at various postures; dividing the multiple image frames into at least one image group according to the postures corresponding to each of the multiple image frames, wherein the range of the postures corresponding to each image frame in each of the image groups satisfies a preset condition; constructing a map subgraph corresponding to the image group according to the postures and feature points corresponding to each of the image frames in the image group, wherein the map subgraph at least includes the map points corresponding to the feature points, and the visual map includes a map subgraph corresponding to at least one image group. The present application divides multiple image frames of different postures taken by a visual sensor into different image groups, and constructs corresponding map sub-graphs according to the postures and feature points of the image frames in each image group. The multiple map sub-graphs together constitute a complete visual map. Compared with the construction of a global map in the related art, the present application first constructs multiple local map sub-graphs, and then combines the multiple local map sub-graphs into an overall visual map, which is beneficial to improving the accuracy of map construction. The map points corresponding to different feature points can correspond to different map sub-graphs. There are more map points in the map sub-graph, which is different from the relatively small number of spatial coordinate points for loop detection in the related art. Due to the increase in map points, the present application has more map points in the map sub-graph that can be compared during loop detection, which improves the accuracy of loop detection, thereby improving the accuracy of posture optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 A flowchart of a method for constructing a visual map provided in an embodiment of the present application;

[0020] Figure 2 A schematic diagram of a map sub-map construction process provided in an embodiment of the present application;

[0021] Figure 3 A schematic block diagram of a visual map construction device provided in an embodiment of the present application;

[0022] Figure 4 A schematic block diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0024] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.

[0025] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0026] See also Figure 1 , Figure 1It is a flow chart of a method for constructing a visual map provided in an embodiment of the present application. It should be noted that the method for constructing a visual map provided in an embodiment of the present application can be used for a movable electronic device, which can be a mobile phone, a tablet computer, a laptop computer, smart glasses, a robot and other devices. Of course, the method for constructing a visual map can also be used for a server, where a movable electronic device collects data and sends the collected data to a server, which analyzes and processes the data to construct a visual map. In addition, the server can send the visual map obtained by the method for constructing a visual map to a movable electronic device. The server can be a separate server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0027] It should be noted that SLAM (Simultaneous Localization and Mapping) refers to the process in which a moving object calculates its own position based on the information from the sensor while building an environmental map, solving the positioning and mapping problems of movable electronic devices when moving in an unknown environment. SLAM is mainly used in the fields of robots, drones, unmanned driving, AR glasses, VR glasses, etc. SLAM technology is divided into two categories: if the sensor is a lidar, it is called laser SLAM; if the sensor is a camera, it is called visual SLAM.

[0028] Visual SLAM can be divided into four parts: front-end, back-end, mapping, and loop detection. The main function of the front-end is to calculate the relative position relationship between the image frames and the corresponding camera poses between the image frames. The front-end includes feature extraction, feature matching, and pose calculation using matched features. The main function of the back-end is to optimize the output results of the front-end to obtain the optimal pose estimation. Loop detection can also be called closed loop detection, which refers to the ability of the robot to recognize the scene it has reached. Loop detection provides the association between the current data and all historical data. After the tracking algorithm is lost, loop detection can also be used for repositioning.

[0029] like Figure 1 As shown, the method for constructing the visual map includes steps S101 to S103.

[0030] Step S101: Obtain image frames captured by the visual sensor at various postures.

[0031] It should be noted that a visual sensor is a sensor that can simulate the human visual system. It can convert optical signals into digital signals to realize functions such as object recognition, tracking and measurement. Visual sensors mainly include optical devices, image sensors, digital signal processors and other parts. In addition, visual sensors are important devices for visual SLAM map construction. Posture refers to the position and posture of an object, robot or person in a specified coordinate system. Obviously, different postures correspond to different positions and / or postures. In practical applications, a mobile electronic device equipped with a visual sensor can control the visual sensor to capture the external environment from different angles and positions to obtain multiple image frames with different postures during movement.

[0032] Step S102: divide the multiple image frames into at least one image group according to the postures corresponding to the multiple image frames, and the range of the postures corresponding to the image frames in each image group meets a preset condition.

[0033] It should be noted that the image frames include key frames, which refer to key image frames in the video and can be used to represent important events, actions and transitions in the video. Image frames that are key frames generally have corresponding postures. For example, among the multiple image frames obtained by indoor shooting, the first image frame corresponds to the tea room on the right, the second image frame corresponds to the corridor in front, and the third image frame corresponds to the office area on the left. Based on this, multiple image frames can be divided, that is, the image frames corresponding to the postures in the range that meet the preset conditions are classified into the same image group. It can be understood that the posture corresponding to the image frame, that is, the posture and the position can be regarded as an area, and the range of the posture area of ​​each image frame in the same image group should meet the preset conditions. Specifically, the room includes multiple rooms in different positions, such as rooms in the four positions of southeast, northwest, and northeast. Therefore, the postures corresponding to the image frames obtained by shooting different rooms are generally different. At this time, the spatial range of the room can be used as a preset condition. If the range of the posture corresponding to the image frame meets the preset condition, the image frames that meet the preset condition are classified into the same image group.

[0034] Step S103: construct a map sub-graph corresponding to the image group according to the posture and feature points corresponding to each image frame in the image group, the map sub-graph at least including map points corresponding to the feature points, and the visual map includes at least one map sub-graph corresponding to the image group.

[0035] There are feature points on each image frame in the image group, and feature points refer to the key points of the image frame. For example, there are many pixels on an image with a resolution of 640*400. If these pixels are directly extracted, the amount of data processing will be very large, so pixels with relatively obvious features can be extracted as feature points to describe this image frame through these feature points. The pose corresponding to each image frame, that is, the position and posture corresponding to each image frame can be reflected by spatial coordinates. Each map sub-graph in the present application has a corresponding coordinate system. In the process of constructing a map sub-graph, the three-dimensional space corresponding to the image frame can be determined according to the pose corresponding to each image frame, that is, the position and posture, and then the spatial coordinates of the feature points on the image frame corresponding to the three-dimensional space, that is, the map points, can be determined according to the brief descriptor corresponding to the feature points on each image frame, so that the map sub-graph can be constructed according to the three-dimensional space corresponding to the pose of the image frame and the map points corresponding to the feature points of the image frame. Based on this, the feature points on the image frames in the same image group are bound to the feature points corresponding to the map points in the corresponding map sub-graph, that is, the feature points on the image frame are bound to the map points in the corresponding map sub-graph, so that the corresponding map sub-graph in the same image group can be constructed. Finally, the map sub-images corresponding to multiple image groups can form an overall visual map.

[0036] For example, after constructing different map sub-graphs, the present application can also extract more feature points from the image frame to determine the map points according to the brief descriptors corresponding to these feature points, and bind the map points to the corresponding map sub-graphs to increase the map points of each map sub-graph. As the map points increase, more map points of the map sub-graphs can be compared during loop detection, so that the constructed map sub-graphs are more accurate and the accuracy of pose optimization is improved. Figure 2 As shown, Figure 2 There are four coordinate systems in the figure, which correspond to four map sub-graphs. Different black dots represent different map points, and multiple purple dots represent the poses of image frames as key frames. The black lines connect the coordinate systems and map points corresponding to the map sub-graphs, which means that different map points are bound to the coordinate systems corresponding to different map sub-graphs.

[0037] The method for constructing a visual map provided in the above embodiment includes: obtaining image frames taken by a visual sensor at various postures; dividing the multiple image frames into at least one image group according to the postures corresponding to each of the multiple image frames, and the range of the postures corresponding to each image frame in each image group meets a preset condition; according to the postures and feature points corresponding to each image frame in the image group, constructing a map subgraph corresponding to the image group, the map subgraph at least including map points corresponding to the feature points, and the visual map includes a map subgraph corresponding to at least one image group. The present application divides multiple image frames of different postures taken by a visual sensor into different image groups, and constructs corresponding map sub-graphs according to the postures and feature points of the image frames in each image group. The multiple map sub-graphs together constitute a complete visual map. Compared with the construction of a global map in the related art, the present application first constructs multiple local map sub-graphs, and then combines the multiple local map sub-graphs into an overall visual map, which is beneficial to improving the accuracy of map construction. The map points corresponding to different feature points can correspond to different map sub-graphs. There are more map points in the map sub-graph, which is different from the relatively small number of spatial coordinate points for loop detection in the related art. In the embodiments of the present application, due to the increase in map points, there are more map points in the map sub-graph that can be compared during loop detection, which improves the accuracy of loop detection, thereby improving the accuracy of posture optimization.

[0038] In an exemplary embodiment, step S102 specifically involves dividing the image frames whose postures are within a preset range into the same image group to obtain at least one image group when the number of image frames whose postures are within a preset range is greater than or equal to a quantity threshold, and the lower limit value of the preset range is greater than the preset lower limit value, and the upper limit value of the preset range is less than the preset upper limit value.

[0039] In general, in order to facilitate the construction of different map sub-maps, the embodiment of the present application can control the data volume of the map sub-map. Specifically, the number of image frames whose postures are within the preset range can be controlled to be greater than or equal to the number threshold. For example, the number threshold is 1000, and the number of image frames within the preset range is 1100. At the same time, the lower limit value of the preset range where the posture is located can be limited to be greater than the preset lower limit value, and the upper limit value of the preset range can be less than the preset upper limit value. For example, the range composed of the preset lower limit value and the preset upper limit value is 0-10 square meters, and the preset range where the posture of the image frame is located is 3-9.5 square meters. Then, the 1100 image frames with postures in 3-9.5 square meters can be divided into the same image group. It can be understood that if the number of image frames is too small and the preset range where the posture is located is larger, it will result in too little data, and the map points for constructing the map sub-map will also be relatively small. Based on this, in the embodiment of the present application, the image frames whose postures are within the preset range and whose number is greater than or equal to the number threshold are divided into the same image group, which can improve the accuracy of constructing the map sub-map and facilitate the construction of multiple different map sub-maps.

[0040] In an exemplary embodiment, the map sub-images corresponding to the two image groups respectively include at least one common view map point, and the common view map point is a map point corresponding to the same feature point in the map sub-images corresponding to the two image groups respectively. After step S103, the following is further included:

[0041] The map sub-images corresponding to the two image groups are associated through at least one common view map point, wherein the distances between the map points in the map sub-images corresponding to the two image groups and the common view map point are both less than or equal to a distance threshold.

[0042] In the embodiment of the present application, the map sub-images corresponding to the two image groups include at least one common view map point, such as Figure 2 As shown, the map points connected by two red lines are the common view map points of the two map sub-images. It should be noted that the common view map points are the map points corresponding to the same feature points in the two map sub-images. Specifically, the image frame for constructing the first map sub-image includes the first image frame, and the image frame for constructing the second map sub-image includes the second image frame, and there is an overlap between the first image frame and the second image frame. Therefore, the first map sub-image and the second map sub-image constructed according to the feature points of the first image frame and the feature points of the second image frame will have the same feature points. It can be understood that the map sub-images corresponding to each of the two image groups can be associated through at least one common view map point, and the distance between any map point in the two map sub-images and the common view map point is less than or equal to the distance threshold. In practical applications, it means that the brief descriptor of the feature point corresponding to any map point in the two map sub-images and the brief descriptor of the feature point corresponding to the common view map point are less than or equal to the distance threshold, so that the coherence between the two map sub-images can be improved. For example, if the two map sub-maps of Room A and Room B are not associated with each other through the common view map points, during the loop detection process, it is possible that the map point originally belonging to Room A will become the map point marking Room C.

[0043] In an exemplary embodiment, the method further includes: if the matching degree between the map sub-images corresponding to at least two image groups is greater than or equal to a matching degree threshold, deleting the map sub-image corresponding to at least one of the at least two image groups from the visual map.

[0044] In the process of constructing a map sub-map, there may be different image groups including the same image frame, resulting in the map sub-maps corresponding to the image groups being relatively similar. However, the device resources are limited. If a large amount of redundant data is stored, the device resources will be occupied. In order to avoid wasting the running memory of the device and ensure smooth operation of the device, similar map sub-maps can be deleted. For example, if there are multiple map sub-maps with a high degree of matching, it means that the similarity of these map sub-maps is high and there are a large number of identical map points. It can be understood that if the feature points corresponding to the map points in two map sub-maps are relatively similar, it means that the similarity of the two map sub-maps is relatively high, that is, the matching degree of the two map sub-maps is relatively high. Based on this, the embodiment of the present application can set a matching degree threshold. When the matching degree between multiple map sub-maps reaches the matching degree threshold, it means that these map sub-maps are almost the same. At this time, one of the map sub-maps can be retained in the visual map, and the other map sub-maps whose matching degree reaches the matching degree threshold can be deleted to solve the data redundancy caused by the map points of similar map sub-maps, avoid wasting the running memory of the device, and ensure smooth operation of the device.

[0045] In an exemplary embodiment, if the matching degree between the map sub-graphs corresponding to at least two image groups is greater than or equal to a matching degree threshold, at least one map sub-graph constructed earlier in the at least two image groups is deleted from the visual map.

[0046] It should be noted that the mobile electronic device can continuously construct a map sub-map according to the image frames obtained by the visual sensor during the movement, that is, the construction of the map sub-map is a time-series increment process. Therefore, each map sub-map has a corresponding construction timing. During the operation of the mobile electronic device, it continuously obtains image frames and then continuously constructs the map sub-map. Therefore, the later the map sub-map is constructed, the later the corresponding construction timing is. In addition, in the process of building a visual map, it is necessary to continuously optimize the posture. Therefore, the map sub-map with a later construction timing can more accurately reflect the current real environment, and the accuracy is generally relatively high. Based on this, when the matching degree between multiple map sub-maps reaches the matching degree threshold, the map sub-maps with earlier construction timings can be deleted from the visual map, and only one map sub-map with the latest construction timing is retained, so that the accuracy of building a visual map can be improved while ensuring the smooth operation of the device.

[0047] In an exemplary implementation, after step S103, the method further includes:

[0048] Get the current key frame;

[0049] Perform similarity detection between the feature points on the current key frame and the map points corresponding to the feature points in each map sub-map;

[0050] The pose of the current key frame is optimized according to the map points corresponding to the feature points on the current key frame and the feature points in the map sub-map whose similarity is greater than or equal to the similarity threshold.

[0051] Constructing a map submap is a time-series incremental process. The mobile electronic device continuously constructs a map submap during movement, and loop detection is also required. Specifically, loop detection is performed based on the new image frame acquired during the movement, that is, the current key frame.

[0052] It can be understood that the brief descriptor of the feature point on the image frame is the same as the brief descriptor of the feature point corresponding to the map point on the map sub-map, that is, the feature point on the image frame is the same as the feature point corresponding to the map point on the map sub-map. Therefore, in the process of constructing the map sub-map, the feature point on the image frame can be bound to the feature point corresponding to the map point on the map sub-map, that is, the feature point on the image frame can be bound to the map point on the map sub-map. Based on this, binding the feature point on the image frame to the map point in the map sub-map can increase the map points in the map sub-map.

[0053] After obtaining the current keyframe, the feature points on the current keyframe can be tested for similarity with the feature points corresponding to the map points in all map sub-maps. In practical applications, it means testing the similarity between the brief descriptors of the feature points on the current keyframe and the brief descriptors of the feature points corresponding to the map points on the map sub-map. If the brief descriptors of the feature points on the current keyframe are similar to the brief descriptors of the feature points corresponding to the map points on the map sub-map, it means that the feature points on the current keyframe are similar to the feature points corresponding to the map points on the map sub-map, that is, the similarity is relatively high. If the similarity is greater than or equal to the similarity threshold, it can be determined that the current keyframe belongs to the map sub-map. Based on this, the map subgraphs with similarity greater than or equal to the similarity threshold can be screened out to determine which one or several map subgraphs the current key frame belongs to, and then the pose of the current key frame is optimized based on the feature points on the current key frame and the feature points corresponding to the map points in the map subgraph with similarity greater than or equal to the similarity threshold. Specifically, the maximum difference between the feature points corresponding to the map points of the map subgraph with similarity greater than or equal to the similarity threshold and the feature points on the current key frame is found, and then the difference is corrected by posegraph (graph optimization) to optimize the pose of the current key frame. This process belongs to loop detection, which refers to determining whether the movable electronic device has returned to its original position and performing loop closure to correct the estimated error. The loop detection is also called closed loop detection, which is a prior art and will not be described here.

[0054] In the related art, the map points corresponding to the feature points of the current key frame are usually compared with the map points corresponding to the feature points of the similar key frame, and a key frame generally only extracts about 1,000 feature points, so the amount of comparison data is small and not accurate enough. In the embodiment of the present application, a map sub-graph can usually include multiple key frames, so the number of feature points is relatively large. Correspondingly, the map points corresponding to the feature points are bound to the corresponding map sub-graph, which is not a global visual map, so that the map points of the map sub-graph are increased, and the locality is better, so that the amount of comparison data between the feature points of the current key frame and the feature points of the map sub-graph is also relatively large, which improves the accuracy of optimization.

[0055] In an exemplary embodiment, the method further comprises:

[0056] storing a map sub-image corresponding to each of at least one image group in the visual map in the cloud;

[0057] According to the target pose of the current target image frame, at least one map sub-image corresponding to the target pose of the target image frame is obtained from the cloud.

[0058] In order to reduce the bandwidth occupation caused by excessive data volume, each map sub-map can be stored in the cloud after the map sub-map corresponding to each image frame is constructed. It is understandable that the pose of the image frame in the image group corresponding to the map sub-map can be stored correspondingly while storing the map sub-map. When it is needed in the future, the target pose of the target image frame is sent to the cloud. The cloud can determine which image frame's pose is closer to the target pose based on the target pose of the current target image frame and the pose of the stored image frame, thereby determining which one or several map sub-maps to return. Therefore, the mobile electronic device can obtain at least one map sub-map corresponding to the target pose of the target image frame from the cloud, so that the mobile electronic device and the cloud can achieve real-time data interaction.

[0059] See also Figure 3 , Figure 3 1 is a schematic block diagram of a visual map construction device provided in an embodiment of the present application. The visual map construction device can be configured in a server or an electronic device to execute the aforementioned visual map construction method.

[0060] like Figure 3 As shown, the visual map construction device includes: a first acquisition module 110, an image division module 120 and a map construction module 130.

[0061] The first acquisition module 110 is used to acquire image frames captured by the visual sensor at various postures.

[0062] The image division module 120 is used to divide the multiple image frames into at least one image group according to the postures corresponding to the multiple image frames, and the range of the postures corresponding to the image frames in each image group meets the preset conditions.

[0063] The map construction module 130 is used to construct a map sub-image corresponding to the image group according to the posture and feature points corresponding to each image frame in the image group. The map sub-image at least includes map points corresponding to the feature points, and the visual map includes at least one map sub-image corresponding to the image group.

[0064] In an exemplary embodiment, the image division module 120 is specifically used to divide the image frames with postures within a preset range into the same image group to obtain at least one image group when the number of image frames with postures within a preset range is greater than or equal to a quantity threshold, and the lower limit value of the preset range is greater than the preset lower limit value, and the upper limit value of the preset range is less than the preset upper limit value.

[0065] In an exemplary embodiment, the map sub-images corresponding to the two image groups respectively include at least one common view map point, which is a map point corresponding to the same feature point in the map sub-images corresponding to the two image groups respectively, and the device also includes an association module.

[0066] The association module is used to associate the map sub-images corresponding to the two image groups through at least one common view map point, wherein the distances between the map points in the map sub-images corresponding to the two image groups and the common view map point are both less than or equal to a distance threshold.

[0067] In an exemplary embodiment, the device further includes a deletion module.

[0068] The deletion module is used to delete the map sub-image corresponding to at least one of the at least two image groups from the visual map if the matching degree between the map sub-images corresponding to the at least two image groups is greater than or equal to the matching degree threshold.

[0069] In an exemplary embodiment, the deletion module is specifically used to delete at least one map sub-image constructed earlier in the at least two image groups from the visual map if the matching degree between the map sub-images corresponding to the at least two image groups is greater than or equal to the matching degree threshold.

[0070] In an exemplary embodiment, the device further includes a second acquisition module, a detection module, and an optimization module.

[0071] The second acquisition module is used to acquire the current key frame.

[0072] The detection module is used to perform similarity detection between the feature points on the current key frame and the map points corresponding to the feature points in each map sub-map.

[0073] The optimization module is used to optimize the position and posture of the current key frame according to the feature points on the current key frame and the map points corresponding to the feature points in the map sub-map whose similarity is greater than or equal to the similarity threshold.

[0074] In an exemplary embodiment, the device further includes a storage module and a third acquisition module.

[0075] The storage module is used to store the map sub-images corresponding to at least one image group in the visual map in the cloud.

[0076] The third acquisition module is used to acquire at least one map sub-image corresponding to the target posture of the target image frame from the cloud according to the target posture of the current target image frame.

[0077] It should be noted that technicians in the relevant field can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described devices and modules and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0078] The method of the present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0079] Exemplarily, the above method and apparatus may be implemented in the form of a computer program, which may be run on a computer device.

[0080] See also Figure 4 , Figure 4 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device may be a server or an electronic device.

[0081] like Figure 4 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a storage medium and an internal memory.

[0082] The storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute the steps of any one of the visual map construction methods.

[0083] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.

[0084] The internal memory provides an environment for the operation of the computer program in the storage medium. When the computer program is executed by the processor, the processor can execute the steps of any method for constructing a visual map.

[0085] The network interface is used for network communication, such as sending assigned tasks.

[0086] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0087] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0088] In one embodiment, the processor is used to execute a computer program and implement the following steps when executing the computer program:

[0089] Obtain image frames captured by the visual sensor at each posture;

[0090] According to the postures corresponding to the multiple image frames, the multiple image frames are divided into at least one image group, and the range of the postures corresponding to the image frames in each image group meets the preset conditions;

[0091] According to the posture and feature points corresponding to each image frame in the image group, a map sub-graph corresponding to the image group is constructed, the map sub-graph at least includes map points corresponding to the feature points, and the visual map includes at least one map sub-graph corresponding to the image group.

[0092] In one embodiment, the plurality of image frames are divided into at least one image group according to the postures corresponding to the plurality of image frames, including:

[0093] When the number of image frames whose postures are within a preset range is greater than or equal to a quantity threshold, and the lower limit value of the preset range is greater than the preset lower limit value, and the upper limit value of the preset range is less than the preset upper limit value, the image frames whose postures are within the preset range are divided into the same image group to obtain at least one image group.

[0094] In one embodiment, the map sub-images corresponding to the two image groups respectively include at least one common view map point, and the common view map point is a map point corresponding to the same feature point in the map sub-images corresponding to the two image groups respectively;

[0095] After constructing the map sub-image corresponding to the image group according to the pose and feature points corresponding to each image frame in the image group, the following is also included:

[0096] The map sub-images corresponding to the two image groups are associated through at least one common view map point, wherein the distances between the map points in the map sub-images corresponding to the two image groups and the common view map point are both less than or equal to a distance threshold.

[0097] In one embodiment, the method further comprises:

[0098] If the matching degree between the map sub-images corresponding to at least two image groups is greater than or equal to the matching degree threshold, the map sub-image corresponding to at least one of the at least two image groups is deleted from the visual map.

[0099] In one embodiment, if the matching degree between the map sub-images corresponding to at least two image groups is greater than or equal to a matching degree threshold, deleting the map sub-image corresponding to at least one of the at least two image groups from the visual map includes:

[0100] If the matching degree between the map sub-images corresponding to at least two image groups is greater than or equal to the matching degree threshold, at least one map sub-image constructed earlier in the at least two image groups is deleted from the visual map.

[0101] In one embodiment, after constructing a map sub-image corresponding to the image group according to the posture and feature points corresponding to each image frame in the image group, the method further includes:

[0102] Get the current key frame;

[0103] Perform similarity detection between the feature points on the current key frame and the map points corresponding to the feature points in each map sub-map;

[0104] The pose of the current key frame is optimized according to the map points corresponding to the feature points on the current key frame and the feature points in the map sub-map whose similarity is greater than or equal to the similarity threshold.

[0105] In one embodiment, the method further comprises:

[0106] storing a map sub-image corresponding to each of at least one image group in the visual map in the cloud;

[0107] According to the target pose of the current target image frame, at least one map sub-image corresponding to the target pose of the target image frame is obtained from the cloud.

[0108] It should be noted that technical personnel in the relevant field can clearly understand that for the convenience and conciseness of description, the specific working process of constructing the visual map described above can refer to the corresponding process in the embodiment of the aforementioned visual map construction method, and will not be repeated here.

[0109] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. The method implemented when the computer program is executed by a processor can refer to the various embodiments of the method for constructing a visual map of the present application.

[0110] The computer-readable storage medium may be an internal storage unit of the computer device of the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device.

[0111] It should be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0112] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system including the element.

[0113] The serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments. The above are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A method for constructing a visual map, characterized in that: include: Obtain image frames captured by the visual sensor at each posture; According to the postures corresponding to the plurality of image frames, the plurality of image frames are divided into at least one image group, and a range of postures corresponding to the image frames in each of the image groups satisfies a preset condition; According to the posture and feature points corresponding to each of the image frames in the image group, a map sub-graph corresponding to the image group is constructed, wherein the map sub-graph at least includes map points corresponding to the feature points, and the visual map includes at least one map sub-graph corresponding to the image group.

2. The method for constructing a visual map according to claim 1, characterized in that: The step of dividing the plurality of image frames into at least one image group according to the postures corresponding to the plurality of image frames respectively comprises: When the number of image frames whose postures are within a preset range is greater than or equal to a quantity threshold, and the lower limit value of the preset range is greater than the preset lower limit value, and the upper limit value of the preset range is less than the preset upper limit value, the image frames whose postures are within the preset range are divided into the same image group to obtain at least one image group.

3. The method for constructing a visual map according to claim 1, characterized in that: The map sub-images corresponding to the two image groups respectively include at least one common view map point, and the common view map point is a map point corresponding to the same feature point in the map sub-images corresponding to the two image groups respectively; After constructing the map sub-image corresponding to the image group according to the posture and feature points corresponding to each of the image frames in the image group, the method further includes: The map sub-images corresponding to the two image groups are associated through at least one of the common view map points, wherein the distances between the map points in the map sub-images corresponding to the two image groups and the common view map point are both less than or equal to a distance threshold.

4. The method for constructing a visual map according to claim 1, characterized in that: The method further comprises: If the matching degree between the map sub-images corresponding to at least two image groups is greater than or equal to the matching degree threshold, the map sub-image corresponding to at least one of the at least two image groups is deleted from the visual map.

5. The method for constructing a visual map according to claim 4, characterized in that: If the matching degree between the map sub-images corresponding to the at least two image groups is greater than or equal to the matching degree threshold, deleting the map sub-image corresponding to at least one of the at least two image groups from the visual map comprises: If the matching degree between the map sub-images corresponding to at least two image groups is greater than or equal to the matching degree threshold, at least one map sub-image constructed earlier in the at least two image groups is deleted from the visual map.

6. The method for constructing a visual map according to any one of claims 1 to 5, characterized in that: After constructing the map sub-image corresponding to the image group according to the posture and feature points corresponding to each of the image frames in the image group, the method further includes: Get the current key frame; Performing similarity detection between the feature points on the current key frame and the map points corresponding to the feature points in each of the map sub-maps; The position and posture of the current key frame are optimized according to the feature points on the current key frame and the map points corresponding to the feature points in the map sub-map whose similarity is greater than or equal to a similarity threshold.

7. The method for constructing a visual map according to claim 1, characterized in that: The method further comprises: storing a map sub-image corresponding to each of at least one image group in the visual map in the cloud; According to the target pose of the current target image frame, at least one map sub-image corresponding to the target pose of the target image frame is obtained from the cloud.

8. A device for constructing a visual map, characterized in that: include: An acquisition module is used to acquire image frames captured by the visual sensor at various postures; An image division module, used for dividing the plurality of image frames into at least one image group according to the postures corresponding to the plurality of image frames, wherein a range of postures corresponding to the image frames in each of the image groups satisfies a preset condition; A map construction module is used to construct a map sub-graph corresponding to the image group according to the posture and feature points corresponding to each image frame in the image group, wherein the map sub-graph at least includes map points corresponding to the feature points, and the visual map includes at least one map sub-graph corresponding to the image group.

9. A computer device, characterized in that: Computer equipment includes memory and processor; Memory for storing computer programs; A processor, configured to execute a computer program and implement the method for constructing a visual map as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for constructing a visual map as described in any one of claims 1 to 7 are implemented.