Map generation method and device, electronic equipment, storage medium, and vehicle

By generating global descriptors and calculating the matching degree between images in the autonomous driving system, and selecting the target image as the starting image, the accuracy and speed problems in high-precision map generation are solved, and faster and more accurate autonomous driving map construction is achieved.

CN116295466BActive Publication Date: 2026-04-24BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2022-03-31
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently generate high-precision autonomous driving maps, especially when the detection range is wide, making it difficult to balance map generation accuracy and speed.

Method used

By determining the set of local feature point descriptors of multiple images, performing aggregation processing to generate a global descriptor, calculating the global matching degree between images, and selecting the target image as the starting image of the map, the target map is constructed using the global descriptor.

Benefits of technology

It improves the accuracy and speed of map generation, supports mapping of unordered images, shortens the association search time of associated frames, and generates target maps that are more consistent with actual environmental information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116295466B_ABST
    Figure CN116295466B_ABST
Patent Text Reader

Abstract

The present disclosure provides a map generation method and device, electronic equipment, storage medium, program product and autonomous vehicle, relates to the technical field of computer vision, and particularly relates to the technical field of autonomous driving. The specific implementation scheme is as follows: determining a local feature point descriptor set of each of a plurality of images; performing aggregation processing on the local feature point descriptor set of each image in the plurality of images to determine a global descriptor of each of the plurality of images; calculating a global matching degree between any two images in the plurality of images based on the global descriptor of each of the plurality of images to obtain a plurality of global matching degrees; and determining a plurality of target images from the plurality of images based on the plurality of global matching degrees, the plurality of target images being starting images for generating a target map.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of application filed on March 31, 2022, with application number 202210352745.9, entitled "Map Generation Method, Apparatus, Electronic Device, Storage Medium, and Vehicle". Technical Field

[0002] This disclosure relates to the field of computer vision technology, and more particularly to the field of autonomous driving technology, specifically to map generation methods, devices, electronic devices, storage media, software products, and autonomous vehicles. Background Technology

[0003] High-precision map-based autonomous driving technology plays a crucial role in enabling safe and reliable autonomous driving of vehicles. High-precision maps are essential for high-precision positioning, environmental perception, path planning, and simulation experiments. The technology for generating high-precision maps with high accuracy requirements and wide detection ranges is receiving increasing attention. Summary of the Invention

[0004] This disclosure provides a map generation method, apparatus, electronic device, storage medium, program product, and autonomous vehicle.

[0005] According to one aspect of this disclosure, a map generation method is provided, comprising: determining a set of local feature point descriptors for each of a plurality of images; for each of the plurality of images, performing aggregation processing on the set of local feature point descriptors of the image to determine a global descriptor for each of the plurality of images; calculating a global matching degree between any two of the plurality of images based on the global descriptors of the plurality of images to obtain a plurality of global matching degrees; and determining a plurality of target images from the plurality of images based on the plurality of global matching degrees, wherein the plurality of target images are starting images for generating a target map.

[0006] According to another aspect of this disclosure, a map generation apparatus is provided, comprising: a first determining module, configured to determine a set of local feature point descriptors for each of a plurality of images; an aggregation module, configured to aggregate the set of local feature point descriptors for each of the plurality of images to determine a global descriptor for each of the plurality of images; a second determining module, configured to calculate a global matching degree between any two of the plurality of images based on the global descriptors for each of the plurality of images to obtain a plurality of global matching degrees; and a third determining module, configured to determine a plurality of target images from the plurality of images based on the plurality of global matching degrees, wherein the plurality of target images are starting images for generating a target map.

[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to said at least one processor; wherein the memory stores instructions executable by said at least one processor, said instructions being executed by said at least one processor to enable said at least one processor to perform a method as disclosed herein.

[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the methods as disclosed herein.

[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method as disclosed herein.

[0010] According to another aspect of this disclosure, an autonomous vehicle is provided, including electronic devices as disclosed herein.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0013] Figure 1 This illustration schematically shows an exemplary system architecture to which map generation methods and apparatus can be applied according to embodiments of the present disclosure;

[0014] Figure 2 A flowchart illustrating a map generation method according to an embodiment of the present disclosure is shown schematically;

[0015] Figure 3 This illustration schematically shows a flowchart of predicting content of interest to a user based on feature information of target content according to an embodiment of the present disclosure;

[0016] Figure 4 This illustration schematically depicts an application scenario where a user takes notes while reading an e-book, according to an embodiment of the present disclosure.

[0017] Figure 5 A block diagram of a map generation apparatus according to an embodiment of the present disclosure is schematically shown; and

[0018] Figure 6 A block diagram of an electronic device suitable for implementing a map generation method according to an embodiment of the present disclosure is shown schematically. Detailed Implementation

[0019] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0020] This disclosure provides a map generation method, apparatus, electronic device, storage medium, program product, and autonomous vehicle.

[0021] According to embodiments of this disclosure, a map generation method is provided, which may include: determining a set of local feature point descriptors for each of a plurality of images; for each of the plurality of images, aggregating the set of local feature point descriptors of the image to determine a global descriptor for each of the plurality of images; calculating a global matching degree between any two of the plurality of images based on the global descriptors of each of the plurality of images to obtain a plurality of global matching degrees; and determining a plurality of target images from the plurality of images based on the plurality of global matching degrees, wherein the plurality of target images are starting images for generating a target map.

[0022] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of users' personal information comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.

[0023] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.

[0024] Figure 1 The illustration schematically shows an exemplary system architecture for applying map generation methods and apparatus according to embodiments of the present disclosure.

[0025] It is important to note that Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0026] like Figure 1 As shown, the system architecture 100 according to this embodiment may include sensors 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the sensors 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0027] Sensors 101, 102, and 103 can interact with server 105 via network 104 to receive or send messages, etc.

[0028] Sensors 101, 102, and 103 can be functional components integrated into the autonomous vehicle 106, such as infrared sensors, ultrasonic sensors, millimeter-wave radar, information acquisition devices, etc. Sensors 101, 102, and 103 can be used to collect environmental information and surrounding road information about the autonomous vehicle 106.

[0029] Server 105 can also be integrated into autonomous vehicle 106, but it is not limited to this. It can also be set up at a remote end that can establish communication with the vehicle terminal. Specifically, it can be implemented as a distributed server cluster composed of multiple servers, or it can be implemented as a single server.

[0030] Server 105 can be a server providing various services. Applications such as map navigation and map generation can be installed on server 105. Taking server 105 running a map generation application as an example: images transmitted from sensors 101, 102, and 103 are received via network 104. Local feature point descriptor sets are determined for each of the multiple images; for each image, the local feature point descriptor sets are aggregated to determine global descriptors for each image; based on the global descriptors, the global matching degree between any two images is calculated to obtain multiple global matching degrees; and based on the multiple global matching degrees, multiple target images are determined from the multiple images so that the multiple target images can be used as starting images to generate a target map.

[0031] It should be noted that the map generation method provided in this embodiment can generally be executed by the server 105. Accordingly, the map generation device provided in this embodiment can also be located in the server 105.

[0032] It should be understood that Figure 1 The number of sensors, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of sensors, networks, and servers can be included.

[0033] It should be noted that the sequence numbers of the operations in the following methods are for descriptive purposes only and should not be considered as indicating the execution order of the operations. Unless explicitly stated otherwise, the method does not need to be executed in the exact order shown.

[0034] Figure 2 A flowchart illustrating a map generation method according to an embodiment of the present disclosure is shown schematically.

[0035] like Figure 2As shown, the method includes operations S210 to S240.

[0036] In operation S210, a set of local feature point descriptors for each of the multiple images is determined.

[0037] In operation S220, for each of the multiple images, the set of local feature point descriptors of the image is aggregated to determine the global descriptor of each of the multiple images.

[0038] In operation S230, based on the global descriptors of each of the multiple images, the global matching degree between any two images is calculated, resulting in multiple global matching degrees.

[0039] In operation S240, multiple target images are determined from multiple images based on multiple global matching degrees, wherein the multiple target images are the starting images used to generate the target map.

[0040] According to embodiments of this disclosure, the set of local feature point descriptors can refer to a set of multiple local feature point descriptors that correspond one-to-one with multiple local feature points. Each local feature point descriptor in the set can refer to a feature point descriptor with scale or rotation invariance. Each local feature point descriptor in the set can be extracted from the image based on rules, such as through rules or algorithms like SIFT (Scale-invariant feature transform) or ORB (Oriented Fast and Rotated Brief). However, it is not limited to this. Deep learning models can also be used to extract the descriptors from the image, such as through Super Point or ASLFeat (Learning Local Features of Accurate Shape and Localization) deep learning models.

[0041] According to optional embodiments of this disclosure, a deep learning model can be used to extract local features from an image to obtain local feature point descriptors. This approach can improve the scene adaptability of local feature point descriptors, thereby expanding the applicability of the map generation method provided in the embodiments of this disclosure.

[0042] According to embodiments of this disclosure, for each of a plurality of images, the aggregation processing of the set of local feature point descriptors of the image to determine the global descriptor of each of the plurality of images may include: for each of the plurality of images, the aggregation processing of multiple local feature point descriptors in the set of local feature point descriptors of the image may be performed to obtain multiple global matching degrees corresponding one-to-one with the plurality of images.

[0043] According to other embodiments of this disclosure, global descriptors can be extracted from images using global feature extraction algorithms. For example, global descriptors in an image can be extracted using DELF (Deep Local Features) model.

[0044] According to embodiments of this disclosure, compared to using global feature extraction algorithms to extract global descriptors from images, using local feature point descriptors of an image to determine global descriptors allows global descriptors to be obtained based on a set of local feature point descriptors. This improves both the computational speed and the extraction accuracy of global descriptors.

[0045] According to embodiments of this disclosure, a target image can be determined from multiple images based on the global matching degree between any two images. The global matching degree can refer to vector similarity, and it can be calculated based on the global descriptors of each of the two images. The global matching degree may include cosine similarity, Euclidean distance, or Manhattan distance, etc., which will not be elaborated further here.

[0046] According to embodiments of this disclosure, the target image may refer to the starting image used to generate the target map. The target image can serve as the initial positioning image for target map construction. The number of target images is not limited; for example, it can be two, but it is not limited to this. It can also be three, four, or five, and can be determined according to the actual situation.

[0047] According to embodiments of this disclosure, the image information shared by multiple target images is clearer and more explicit than the image information of other images in the multiple images besides the multiple target images. By using the image information in the multiple target images to construct a map, the generated target map is more consistent with the actual environmental information expressed, thereby accelerating the construction of the common view plane of the associated frames of the target map.

[0048] According to other embodiments of this disclosure, a target image can be determined from multiple images based on local feature point descriptors in a set of local feature point descriptors, and then the target image can be used as the starting image of a target map.

[0049] According to embodiments of this disclosure, compared to directly determining the target image using local feature point descriptors, determining the target image using global descriptors is possible by comparing global matching degrees. This method is simple to calculate, highly accurate, and supports unordered mapping, thereby improving the speed of target map generation.

[0050] In summary, by using global descriptors to determine the global matching degree between any two images, and then using this global matching degree to identify the target image from multiple images, it is possible to construct a target map supporting multiple unordered images, thus accelerating the association search speed of associated frames. Furthermore, since the global descriptor is determined based on a set of local feature point descriptors, this improves both the computational speed and accuracy of the global descriptor.

[0051] Figure 3 A schematic diagram illustrating the determination of a plurality of target images according to an embodiment of the present disclosure is shown.

[0052] like Figure 3 As shown, the unordered multiple images may include image P_A, image P_B, image P_C, and image P_D. The set of local feature point descriptors for any one of the multiple images may include the local feature point confidence scores of each of the M local feature points, the local feature point location information of each of the M local feature points, and the local feature point descriptors of each of the M local feature points. M is an integer greater than or equal to 2. The number of local feature points in any two of the multiple images may be the same or different, which will not be elaborated further here.

[0053] like Figure 3 As shown, taking image P_A as an example, image P_A includes 6 local feature points, such as local feature points p1, p2, p3, p4, p5, and p6. Each local feature point corresponds one-to-one with a confidence level of that local feature point. A predetermined confidence threshold can be set, selecting local feature points with a confidence level greater than the threshold as target local feature points, and deleting local feature points with a confidence level less than or equal to the threshold. Figure 3 As shown, the confidence levels of local feature points p1 and p2 are less than the confidence threshold, while the confidence levels of local feature points p3, p4, p5, and p6 are greater than the confidence threshold. Therefore, local feature points p1 and p2 can be deleted, and local feature points p3, p4, p5, and p6 can be retained as target local feature points. This yields image P_A', which retains the target local feature points p3, p4, p5, and p6. Similarly, the confidence threshold can be used to filter images P_B, P_C, and P_D, that is, to filter out local feature points with confidence levels less than the confidence threshold, resulting in images P_B', P_C', and P_D'.

[0054] According to embodiments of this disclosure, M local feature points in an image can be filtered based on a confidence threshold, retaining Q target local feature points in the image. The local feature point descriptors of each of the Q target local feature points and their respective positional information, such as two-dimensional coordinates (x, y), can be input into a local aggregation vector model to obtain a global descriptor for the image.

[0055] According to embodiments of this disclosure, the local aggregated vector model may include VLAD (Vector of locally aggregated descriptors), but is not limited thereto. It may also include BOF (bag-of-feature) model, attention mechanism model, or improved VLAD-based model such as NetVLAD.

[0056] According to embodiments of this disclosure, by using a local aggregation vector model to process the positional information of multiple target local feature points and local feature point descriptors, the attention mechanism in the local aggregation vector model can be used to aggregate multiple target local feature points, thereby improving the accuracy of the global descriptor.

[0057] like Figure 3 As shown, the global matching degree between two images can be determined based on the global descriptors of any two images from multiple images. This yields the global matching degrees between images P_A' and P_B', P_A' and P_C', P_A' and P_D', P_B' and P_C', P_B' and P_D', and P_C' and P_D'. These global matching degrees are then sorted, for example, from highest to lowest, to obtain a ranking result. Based on this ranking result, multiple target images are determined from the multiple images. For example, if the global matching degree between images P_B' and P_C' is the highest, then images P_B' and P_C' with the highest global matching degree are selected as target images.

[0058] According to embodiments of this disclosure, a confidence threshold can be used to filter multiple local feature points in an image, removing low-confidence local feature points and retaining high-confidence target local feature points. This improves the accuracy of global descriptor determination while maintaining the computational speed. Consequently, the index position of the target map can be quickly located, allowing multiple target images to be used as the starting image for the target map.

[0059] Figure 4 A flowchart illustrating the determination of target local feature point pairs according to an embodiment of the present disclosure is shown schematically.

[0060] like Figure 4 As shown, determining the target local feature point pair may include operations S410 to S450.

[0061] In operation S410, for any two images among multiple images, based on the multiple local feature point descriptors of each of the two images, N pairs of local feature points and N local matching degrees are determined from the two images.

[0062] According to embodiments of this disclosure, N is an integer greater than or equal to 2. There is a one-to-one correspondence between the N local feature point pairs and the N local matching degrees. The local feature point descriptor set includes local feature point descriptors for each of the M local feature points. M is an integer greater than or equal to 2. The number of local feature points in any two images among multiple images can be the same or different, which will not be elaborated further here.

[0063] According to embodiments of this disclosure, determining N pairs of local feature points and N local matching degrees from two images based on multiple local feature point descriptors of each image may include: calculating the vector similarity between two local feature point descriptors, and determining the local matching degree between the two local feature point descriptors. This vector similarity may include cosine similarity, Euclidean distance, or Manhattan distance, etc., which will not be elaborated further here.

[0064] According to embodiments of this disclosure, multiple initial local feature point pairs can be determined from two images based on vector similarity. For each initial local feature point pair, the fundamental matrix between the two images is calculated using epipolar geometry principles. Through Random Sample Consensus (RANSAC), N local feature point pairs are determined from the multiple initial local feature point pairs. The N initial local matching degrees, corresponding one-to-one with the N local feature point pairs, can be normalized to obtain N local matching degrees.

[0065] According to other embodiments of this disclosure, an image sequence can be generated based on the sorting results of multiple global matching degrees arranged from high to low, and local feature point pairs between two adjacent images in the image sequence can be calculated to reduce the amount of data processing.

[0066] In operation S420, N initial target local matching degrees are obtained based on the global matching degree between two images and the local matching degree of each of the N local feature points of the two images.

[0067] According to embodiments of this disclosure, obtaining N initial target local matching degrees based on the global matching degree between two images and the local matching degrees of N local feature points of the two images may include: for each of the N local feature points, using the global matching degree as a weight, weighting the local matching degree to obtain the initial target local matching degree. However, this is not limited to this. Alternatively, based on the global matching degree, a transformation weight between the two images may be determined according to a predetermined transformation rule. Based on the transformation weight, the local matching degree may be weighted to obtain the initial target local matching degree. For example, based on the global matching degree, the transformation weight may be determined to be a value between 0 and 1 according to a predetermined transformation rule.

[0068] According to embodiments of this disclosure, weighting the global matching degree of the global descriptor onto the local matching degree allows for a more accurate and effective determination of local feature point pairs by combining global and local feature information.

[0069] In operation S430, it is determined whether the initial target local matching degree is greater than a predetermined local matching degree threshold. If it is determined that the initial target local matching degree is greater than the predetermined local matching degree threshold, operation S440 is executed. If it is determined that the initial target local matching degree is less than or equal to the predetermined local matching degree threshold, operation S450 is executed.

[0070] In operation S440, the initial local matching degree that is greater than the predetermined local matching degree threshold is taken as the target local matching degree, and the local feature point pair corresponding to the target local matching degree is taken as the target local feature point pair.

[0071] According to embodiments of this disclosure, operation S440 can be repeatedly performed to obtain multiple target local feature point pairs of multiple images.

[0072] In operation S450, stop further operations on this local feature point pair.

[0073] According to embodiments of this disclosure, a target local feature point pair is determined from multiple local feature points by using a predetermined local matching degree threshold, thereby improving the speed of determining key locations in the target image.

[0074] According to embodiments of this disclosure, the map generation method may further include the operations of: triangulating a pair of target local feature points between any two target images from a plurality of target images to generate spatial three-dimensional points; and generating a map model based on the spatial three-dimensional points.

[0075] According to embodiments of this disclosure, triangulation can be understood as: given the two-dimensional projection points of a spatial three-dimensional point observed at different locations, using triangulation relationships to recover the depth information of that spatial three-dimensional point. Triangulation can be used to recover the spatial three-dimensional point of any pair of local feature points from any two target images in multiple target images by utilizing their respective positional information.

[0076] According to embodiments of this disclosure, the spatial three-dimensional points obtained by triangulation can be used as initial information. The reprojection error can be minimized using bundle adjustment (BA) to obtain target spatial three-dimensional points, so as to generate a map model based on the target spatial three-dimensional points.

[0077] According to embodiments of this disclosure, the map generation method may further include the operation of registering other images into a map model based on other target local feature point pairs to generate a target map.

[0078] According to embodiments of this disclosure, other target local feature point pairs include target local feature point pairs other than any two target images from a plurality of target local feature point pairs, and other images include images other than the target images.

[0079] According to embodiments of this disclosure, generating a map model based on spatial three-dimensional points is the initial process for three-dimensional reconstruction based on target images. The map model can be referred to as an initial point cloud model. The map model can refer to a portion of a three-dimensional map generated using multiple target images as a common-view correlation graph. The map model can be used as the starting point for the target map, and the target map can be generated based on local feature points of other targets in other images.

[0080] According to embodiments of this disclosure, registration can be understood as: fusing local feature point pairs of other targets in other images into the map model based on the ranking results obtained from the global matching degree, which is called registration.

[0081] According to embodiments of this disclosure, based on other target local feature point pairs, other images are registered into the map model one by one to complete the fusion of other target local feature point pairs with the map model until all images in multiple images are registered to obtain the target map.

[0082] According to embodiments of this disclosure, a target map is generated by combining local feature point descriptors and global descriptors, where the global descriptor is obtained through aggregation of local feature point descriptors. This results in fast map construction and high accuracy of the generated target map. Simultaneously, the target map is generated through an associative search method using global and local matching degrees, as well as an incremental method utilizing registration, resulting in robustness and ease of operation in the generated target map.

[0083] Figure 5 A block diagram of a map generation apparatus according to an embodiment of the present disclosure is shown schematically.

[0084] like Figure 5 As shown, the map generation device 500 may include a first determining module 510, an aggregation module 520, a second determining module 530, and a third determining module 540.

[0085] The first determining module 510 is used to determine the local feature point descriptor sets of multiple images respectively.

[0086] The aggregation module 520 is used to aggregate the set of local feature point descriptors of each of the multiple images to determine the global descriptors of each of the multiple images.

[0087] The second determining module 530 is used to calculate the global matching degree between any two images in the multiple images based on the global descriptors of each image, and obtain multiple global matching degrees.

[0088] The third determining module 540 is used to determine multiple target images from multiple images based on multiple global matching degrees, wherein the multiple target images are the starting images used to generate the target map.

[0089] According to embodiments of this disclosure, the local feature point descriptor set includes local feature point descriptors for each of the multiple local feature points.

[0090] According to embodiments of this disclosure, the map generation apparatus may further include a fourth determining module and a fifth determining module.

[0091] The fourth determination module is used to determine N pairs of local feature points and N local matching degrees from any two images among multiple images, based on the multiple local feature point descriptors of each of the two images, where N is an integer greater than or equal to 2.

[0092] The fifth determination module is used to determine the target local feature point pairs from the two images based on the global matching degree and N local matching degrees between the two images, so as to obtain multiple target local feature point pairs from multiple images.

[0093] According to embodiments of this disclosure, the fifth determining module may include a local matching unit, a first filtering unit, and a point-to-point determining unit.

[0094] The local matching unit is used to obtain N initial target local matching degrees based on the global matching degree between two images and the local matching degree of N local feature points of the two images.

[0095] The first filtering unit is used to determine the target local matching degree from N initial target local matching degrees, wherein the target local matching degree is greater than a predetermined local matching degree threshold.

[0096] The point pair determination unit is used to identify the local feature point pairs that correspond to the local matching degree of the target as the target local feature point pairs.

[0097] According to embodiments of this disclosure, the local feature point descriptor subset includes M local feature points, the local feature point confidence scores of each of the M local feature points, and the local feature point location information of each of the M local feature points. M is an integer greater than or equal to 2.

[0098] According to embodiments of this disclosure, the aggregation module may include a second filtering unit and an aggregation unit.

[0099] The second filtering unit is used to determine Q target local feature points from M local feature points in any one of multiple images, based on the confidence scores of M local feature points. The confidence score of any one of the Q target local feature points is greater than a confidence threshold.

[0100] The aggregation unit is used to input the local feature point descriptors of each of the Q target local feature points and the position information of each of the Q target local feature points into the local aggregation vector model to obtain the global descriptor of the image.

[0101] According to embodiments of this disclosure, the map generation apparatus may further include a triangulation module and a generation module.

[0102] The triangulation module is used to triangulate the local feature point pairs of any two target images in multiple target images to generate three-dimensional points in space.

[0103] The generation module is used to generate map models based on three-dimensional points in space.

[0104] According to embodiments of this disclosure, the map generation apparatus may further include a registration module.

[0105] The registration module is used to register other images into the map model based on other target local feature point pairs to generate the target map. The other target local feature point pairs include target local feature point pairs other than any two target images from multiple target local feature point pairs. Other images include images other than the target image.

[0106] According to embodiments of this disclosure, the third determining module may include a sorting unit and a result determining unit.

[0107] The sorting unit is used to sort multiple global matching degrees to obtain the sorting result.

[0108] The result determination unit is used to determine multiple target images from multiple images based on the sorting results.

[0109] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, a computer program product, and an autonomous vehicle.

[0110] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method as described in the embodiments of the present disclosure.

[0111] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform methods as described in embodiments of the present disclosure.

[0112] According to embodiments of the present disclosure, a computer program product includes a computer program that, when executed by a processor, implements the methods as described in embodiments of the present disclosure.

[0113] According to an embodiment of this disclosure, an autonomous vehicle is equipped with the aforementioned electronic equipment, which, when executed by its processor, can implement the map generation method described in the above embodiments.

[0114] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0115] like Figure 6As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded from storage unit 608 into random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0116] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0117] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as map generation methods. For example, in some embodiments, the map generation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the map generation method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform map generation methods by any other suitable means (e.g., by means of firmware).

[0118] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0119] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0120] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0121] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0122] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0123] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.

[0124] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0125] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A map generation method, comprising: Determine the local feature point descriptor subsets for each of the multiple images; For each of the plurality of images, the set of local feature point descriptors of the image is aggregated to determine the global descriptor of each of the plurality of images; Based on the global descriptors of each of the multiple images, the global matching degree between any two images is calculated to obtain multiple global matching degrees; Based on the multiple global matching degrees, multiple target images are determined from the multiple images, wherein the multiple target images are the starting images used to generate the target map; For any two target images among the plurality of target images, determine N pairs of local feature points and N local matching degrees from the two target images, where N is an integer greater than or equal to 2; Using the global matching degree of the two target images as weights, each local matching degree is weighted to obtain the initial target local matching degree; and An initial target local matching degree greater than a predetermined local matching degree threshold is taken as the target local matching degree, and the local feature point pair corresponding to the target local matching degree is taken as the target local feature point pair.

2. The method according to claim 1, wherein, The set of local feature point descriptors includes local feature point descriptors for each of the multiple local feature points; Determining N pairs of local feature points and N local matching degrees from the two target images includes: Based on the multiple local feature point descriptors of each of the two target images, the N pairs of local feature points and the N local matching degrees are determined from the two target images.

3. The method according to claim 1 or 2, wherein, The local feature point descriptor subset includes M local feature points, the local feature point confidence of each of the M local feature points, and the local feature point location information of each of the M local feature points, wherein M is an integer greater than or equal to 2; The aggregation process of the local feature point descriptor set to determine the global descriptor for each of the multiple images includes: For any one of the plurality of images, based on the confidence scores of M local feature points, Q target local feature points are determined from the M local feature points in the image, wherein the confidence score of any one of the Q target local feature points is greater than a confidence threshold; and The local feature point descriptors of each of the Q target local feature points and the position information of each of the Q target local feature points are input into the local aggregation vector model to obtain the global descriptor of the image.

4. The method according to any one of claims 1 to 3, further comprising: Triangulation is performed on any two target images from the plurality of target images to generate three-dimensional points in space. as well as A map model is generated based on the three-dimensional points in the space.

5. The method according to claim 4, further comprising: Based on other target local feature point pairs, other images are registered into the map model to generate the target map. The other target local feature point pairs include target local feature point pairs other than the target local feature points of any two target images from the plurality of target local feature point pairs. The other images include images other than the target image.

6. The method according to any one of claims 1 to 5, wherein, The step of determining multiple target images from the multiple images based on the multiple global matching degrees includes: The multiple global matching degrees are sorted to obtain the sorting results; and Based on the sorting results, the plurality of target images are determined from the plurality of images.

7. A map generation apparatus, comprising: The first determining module is used to determine the local feature point descriptor sets for each of the multiple images; An aggregation module is used to aggregate the set of local feature point descriptors of each of the plurality of images to determine the global descriptor of each of the plurality of images. The second determining module is used to calculate the global matching degree between any two images in the plurality of images based on the global descriptors of each of the plurality of images, and obtain a plurality of global matching degrees. The third determining module is used to determine multiple target images from the multiple images based on the multiple global matching degrees, wherein the multiple target images are the starting images for generating the target map; The device is also used for: For any two target images among the plurality of target images, N pairs of local feature points and N local matching degrees are determined from the two target images, where N is an integer greater than or equal to 2; Using the global matching degree of the two target images as weights, each local matching degree is weighted to obtain the initial target local matching degree; and An initial target local matching degree greater than a predetermined local matching degree threshold is taken as the target local matching degree, and the local feature point pair corresponding to the target local matching degree is taken as the target local feature point pair.

8. The apparatus according to claim 7, wherein, The set of local feature point descriptors includes local feature point descriptors for each of the multiple local feature points; The determination of N local feature point pairs and N local matching degrees from the two target images is used for: Based on the multiple local feature point descriptors of each of the two images, the N pairs of local feature points and the N local matching degrees are determined from the two images.

9. The apparatus according to claim 7 or 8, wherein, The local feature point descriptor subset includes M local feature points, the local feature point confidence of each of the M local feature points, and the local feature point location information of each of the M local feature points, wherein M is an integer greater than or equal to 2; The aggregation module includes: The second filtering unit is configured to, for any one of the plurality of images, determine Q target local feature points from the M local feature points in the image based on the confidence scores of the M local feature points, wherein the confidence score of any one of the Q target local feature points is greater than a confidence threshold; and An aggregation unit is used to input the local feature point descriptors of the Q target local feature points and the position information of the Q target local feature points into a local aggregation vector model to obtain the global descriptor of the image.

10. The apparatus according to any one of claims 7 to 9, further comprising: The triangulation module is used to triangulate the target local feature point pairs of any two target images in the plurality of target images to generate spatial three-dimensional points; as well as The generation module is used to generate a map model based on the three-dimensional points in the space.

11. The apparatus of claim 10, further comprising: The registration module is used to register other images into the map model based on other target local feature point pairs to generate the target map. The other target local feature point pairs include target local feature point pairs other than the target local feature points of any two target images from the plurality of target local feature point pairs. The other images include images other than the target images.

12. The apparatus according to any one of claims 7 to 11, wherein, The third determining module includes: A sorting unit is used to sort the plurality of global matching degrees to obtain a sorting result; and The result determination unit is used to determine the plurality of target images from the plurality of images based on the sorting result.

13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 6.

15. An autonomous vehicle, comprising: The electronic device as claimed in claim 13.

Citation Information

Patent Citations

  • Hierarchical landmark identification method integrating global visual characteristics and local visual characteristics

    CN102542058A

  • An image matching method based on image global features and local features

    CN109447173A