Map building network and method

By designing a map to build a network, using semantic segmentation features to weight initial query, and vector decoding is performed based on the attention mechanism, the problems of inaccurate map construction and information loss in the existing technology are solved, and high-precision and stable vectorized map construction are achieved.

CN120198532APending Publication Date: 2025-06-24TSINGHUA UNIVERSITY +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510227423.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

It is difficult for the prior art to build an accurate map without losing instantiated information. Semantic segmentation maps lack the summary nature of map element instance information, while vectorized maps lose a large amount of original data, resulting in unstable overall shape of the map.

Method used

A map construction network is designed, including feature extraction network, semantic decoder, query weighting module and vector decoder. The initial query is weighted through semantic segmentation features, and vector decoding of optimized query and image features based on attention mechanism to build a vectorized map.

Benefits of technology

In the process of building a vectorized map, the semantic segmentation results are fused into the vector results, thereby improving the accuracy and stability of the map and effectively utilizing visual information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198532A_ABST
    Figure CN120198532A_ABST
Patent Text Reader

Abstract

The invention discloses a map construction network and method, and the network comprises a feature extraction network which is used for carrying out the feature extraction of road image data of a to-be-constructed region, and obtaining image features; the semantic decoder is used for carrying out semantic segmentation on the image features to obtain semantic segmentation features; the semantic segmentation features are used for representing semantic information of each pixel point in the road image data; the query weighting module is used for weighting the initial query based on the semantic segmentation features to obtain a weighted query; and the vector decoder is used for optimizing the weighted query to obtain an optimized query, performing vector decoding on the optimized query and the image features based on an attention mechanism to obtain vector features, and constructing a vectorized map of the to-be-mapped area based on the vector features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of environment construction, and particularly to a map construction network and method. Background Art

[0002] With the development of autonomous driving technology, the extraction technology of map elements and the map construction technology have become research hotspots. Currently, the map construction methods are mainly divided into two types. One is to construct a semantic segmentation map, and the other is to construct a structured vector map. The semantic segmentation map can retain the shape information of map elements in the original image data in units of pixels. However, it lacks the summary nature of map element instance information and is difficult to be used by subsequent modules. The vector map can obtain map element instance information, but a large amount of original data is lost. The position error of a single point coordinate will bring great disturbance to the overall shape of the map and result in an unstable map. In the map construction methods in the related art, it is difficult to construct an accurate map without losing instantiation information. Summary of the Invention

[0003] In a first aspect, an embodiment of this application provides a map construction network, including:

[0004] A feature extraction network, configured to extract features from the road image data of the area to be mapped, and obtain image features;

[0005] A semantic decoder, configured to perform semantic segmentation on the image features to obtain semantic segmentation features; the semantic segmentation features are used to represent the semantic information of each pixel point in the road image data;

[0006] A query weighting module, configured to weight an initial query based on the semantic segmentation features to obtain a weighted query;

[0007] A vector decoder, configured to optimize the weighted query to obtain an optimized query, perform vector decoding on the optimized query and the image features based on an attention mechanism to obtain vector features, and construct a vector map of the area to be mapped based on the vector features.

[0008] In a second aspect, an embodiment of this application provides a map construction method based on the map construction network described in the first aspect. The method includes:

[0009] Extract features from the road image data of the area to be mapped through a feature extraction network to obtain image features;

[0010] Perform semantic segmentation on the image features through a semantic decoder to obtain semantic segmentation features; the semantic segmentation features are used to represent the semantic information of each pixel point in the road image data;

[0011] The initial query is weighted based on the semantic segmentation features by a query weighting module to obtain a weighted query.

[0012] The weighted query is optimized by a vector decoder to obtain an optimized query. Based on the attention mechanism, the optimized query and the image features are vector decoded to obtain vector features, and a vectorized map of the area to be built is constructed based on the vector features.

[0013] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in any embodiment of the present application is implemented.

[0014] In the embodiment of the present application, semantic segmentation features are obtained through a semantic decoder, and then the initial query is weighted using the semantic segmentation features to obtain a weighted query. After the weighted query is optimized by a vector decoder to obtain an optimized query, based on the attention mechanism, the optimized query and the image features are vector decoded to obtain vector features, and a vectorized map is constructed based on the vector features. On the one hand, the present application constructs a vectorized map that can represent the information of map element instances; on the other hand, since the optimized query incorporates semantic segmentation features, the obtained vector features also incorporate semantic segmentation features. In this way, when constructing the vectorized map, the semantic segmentation results can be integrated into the vector results, thereby effectively using visual information to improve the accuracy of the vectorized map.

[0015] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. Description of the Drawings

[0016] The drawings here are incorporated into the specification and form a part of the present application. These drawings illustrate embodiments consistent with the present application and are used together with the specification to explain the technical solutions of the present application.

[0017] Figure 1 It is a schematic diagram of the levels of the map construction network according to an embodiment of the present application.

[0018] Figure 2 It is a schematic diagram of the structure of the map construction network according to an embodiment of the present application.

[0019] Figure 3 It is a schematic diagram of the structure of the temporal information enhancement module according to an embodiment of the present application.

[0020] Figure 4 It is a flowchart of the map construction method according to an embodiment of the present application.

[0021] Figure 5 It is a schematic diagram of a computer device according to an embodiment of the present application. Detailed implementation manners

[0022] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0023] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any or all possible combinations of one or more of the associated listed items. In addition, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality.

[0024] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0025] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the embodiments of the present application and make the above-mentioned objects, features, and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0026] With the rapid development of deep learning technology and the significant improvement of hardware computing power, autonomous driving vehicle hardware can now perform real-time calculations with limited computing resources, efficiently process the real vehicle data captured by vehicle-mounted sensors, and output results that meet the requirements of autonomous driving technology. This series of improvements has significantly promoted the continuous optimization of the performance of autonomous driving technology. Among many sensors, the camera plays a crucial role. The camera can take photos and videos of the driving scene, and these photos and videos can be used to obtain the shape, position, and semantic information of map elements, thereby realizing local map construction.

[0027] However, the current map construction based on in-vehicle cameras faces several challenges. One of the challenges is that there is no unified expression form for autonomous driving maps. Currently, there are mainly two forms of map expression. One is the semantic segmentation map, and the other is the structured vectorized map. The semantic segmentation map can retain the shape information of map elements in the original image data in pixel units. However, it lacks the summary information of map element instances and is difficult to be used by subsequent modules. The vectorized map is usually a set of structured vectorized map element data. However, this expression form loses a large amount of original data, and the position error of a single point coordinate will cause great disturbance to the overall shape of the map, resulting in unstable results for the map, which are often unacceptable. In summary, the related technologies are difficult to construct a map that is accurate and does not lose instantiation information.

[0028] To overcome these challenges, it is urgent to propose an autonomous driving map model that can unify different map expression forms, so as to effectively represent the shape and position information of map elements and make the subsequent online mapping more accurate and stable. The following will describe the specific implementation manners of the present application in detail with reference to the accompanying drawings and embodiments.

[0029] The present application proposes a map hierarchical expression model and a vectorized key point expression manner of map elements according to the spatio-temporal static characteristics of map elements, so as to hierarchically summarize maps with different information dimensions, thereby unifying the semantic segmentation map and the vector map. And due to the spatial invariance of the map, the map information of each layer can interact with each other to act synergistically, improving the information accuracy, richness and reliability provided by the autonomous driving map to subsequent modules. Each level is as Figure 1 shown.

[0030] The first layer of the hierarchical expression model is the pixel-level grid map layer, which is the bottommost map layer. The pixel-level grid map layer includes a semantic decoder, which obtains semantic segmentation feature F mask by performing semantic segmentation on the image features. The semantic segmentation feature F mask is an intermediate feature generated during the process of the semantic decoder outputting the grid-level semantic map M mask . M in the figure mask,1 , ……, M mask,T respectively represent the semantic maps of the 1st image frame M1 to the Tth image frame M T in the road image data. L Mask,1 , ……, L Mask,N respectively represent the local semantic maps of the 1st section to the local semantic maps of the Nth section, and G Mask represents the global semantic map of the road including N sections. The semantic segmentation feature F maskIt is closest to the information of the original road image data, reflects the pixel-level distribution of the semantic map at the current moment, can retain most of the map element category information and the image plane shape information. By performing semantic segmentation calculations at this layer, it can provide weights for subsequent semantic segmentation raster shape constraint operations, classify the map for each pixel, and form a pixel-level raster map mask. Since the data structure of this layer is flat, this layer can be used as a container, and on this basis, map element instances and autonomous driving semantics can be superimposed. Moreover, compared with the raster-level semantic map M mask , the semantic segmentation feature F mask contains more detailed local information in the image, and this information can improve the perception ability of details, enhance diversity and expression ability.

[0031] The second layer of the map hierarchical expression model is the vector instance map element layer. This layer is composed of structured map element key point expression templates, and these templates are vector data M obtained by vector decoding the image features through a vector decoder vector . The M in the figure vector,1 , ……, M vector,T respectively represent the vector data of the 1st image frame M1 to the Tth image frame M in the road image data T , L Vector,1 , ……, L Vector,N respectively represent the local vector data of the 1st section to the local vector data of the Nth section, and G Vector represents the global vector data of the road including N sections. The information of the vector instance map element layer is highly structured and contains the instance individuals, categories and contour information of the map elements. The function of this layer is to express specific map elements, such as lane lines, curbs and crosswalks, in vector form, providing an accurate geometric description for the map. These vector information can not only be converted into the underlying raster map to provide detailed road structure information for autonomous driving vehicles, but also support the construction of upper-layer understanding semantics by analyzing the mutual relationships between elements.

[0032] Based on the above hierarchical structure, the present application designs an end-to-end map construction network. See Figure 2 , the map construction network of the present application includes:

[0033] The feature extraction network 101 is used to extract features from the road image data of the area to be mapped to obtain image features;

[0034] The semantic decoder 102 is used to perform semantic segmentation on the image features to obtain semantic segmentation features;

[0035] The query weighting module 103 is configured to weight the initial query based on the semantic segmentation features to obtain a weighted query.

[0036] The vector decoder 104 is configured to optimize the weighted query to obtain an optimized query, perform vector decoding on the optimized query and the image features based on the attention mechanism to obtain vector features, and construct a vectorized map of the area to be mapped based on the vector features.

[0037] First, the road image data of the area to be mapped can be obtained. The above-mentioned road image data may include one or more image frames collected by an in-vehicle camera. In some embodiments, the in-vehicle camera includes an in-vehicle surround-view camera, that is, a plurality of cameras deployed on the vehicle. These cameras provide a panoramic view of the surrounding environment of the vehicle through shooting at multiple angles, and the images collected by each camera can be stitched through image stitching technology to obtain the road image data.

[0038] After obtaining the road image data, the feature extraction network can be used to extract features from the road image data to obtain image features. In some embodiments, the feature extraction network may include a two-dimensional feature encoder and an aerial view feature encoder. Among them, the two-dimensional feature encoder is used to perform two-dimensional feature extraction on the road image data to obtain two-dimensional image features. The aerial view feature encoder is used to encode the two-dimensional image features to obtain aerial view image features. The two-dimensional features refer to the features directly extracted from the road image data, usually the pixel information of the planar image, which can represent the texture, color, shape, etc. of the road image data. The aerial view feature encoder can further perform deeper spatial information processing on the two-dimensional image features to obtain a feature representation with spatial perception ability (i.e., aerial view image features), so as to more comprehensively reflect the spatial layout and global information that can be extracted from looking down at the road image from the air. The aerial view image features can be used as the image features input to the semantic decoder and the vector decoder.

[0039] After obtaining the image features, the present application inputs the image features into two branches respectively. One branch is the semantic decoder, whose function corresponds to the pixel-level grid map layer of the above-mentioned hierarchical expression model, and obtains semantic segmentation features F by performing semantic segmentation on the image features. mask The other branch is the vector decoder, whose function corresponds to the vector instance map element layer of the above-mentioned hierarchical expression model. The vector decoder can perform vector decoding on the image features to obtain vectorized map elements M. vector

[0040] ​In the related art, there is a lack of connection between vectorized map elements and semantic segmentation maps, resulting in the inability to effectively utilize dense visual information, inconsistent semantic segmentation results and vector results, and low precision in map construction. To solve the above problems, this application fuses the semantic segmentation results and vector results. Specifically, according to the map element consistency between the pixel-level raster semantic segmentation map and the instance-level vector map in the hierarchical expression model of the autonomous driving map, this application designs a query weighting module and designs the structure of the vector decoder according to the vectorized key point expression method of map elements.

[0041] The query weighting module performs dimensionality conversion on the semantic segmentation features based on the dimension of the initial query, and uses the converted semantic segmentation features as attention weights to weight the initial query of the map element instance, obtaining a weighted query. In this way, it can effectively guide the initial query to focus on the shape and category of map element instances distributed at specific positions. By weighting the initial query with the semantic segmentation features, the finally weighted query constrained by the semantic segmentation shape can be obtained, and these weighted queries will be used as the input of the vector decoder. This method ensures that the vector decoder can more accurately identify and locate map elements, improving the accuracy and efficiency of the mapping process. The process of obtaining the weighted query can be expressed as:

[0042] Q masked =F mask *Q initial

[0043] where Q initial is the initial query, and Q masked is the weighted query.

[0044] When performing vector decoding, the vector decoder can first optimize the weighted query to obtain an optimized query, and then perform vector decoding on the optimized query and the image features based on the attention mechanism to obtain vector features, and construct a vectorized map of the area to be mapped based on the vector features. Since the query input to the vector decoder is no longer the randomly initialized value Q initial , but the result Q masked that has been constrained by the prior of the map element segmentation map. This improvement enables the vector decoder to focus more on the area that may contain map elements from the beginning, thereby improving the quality and efficiency of the overall mapping process. In this way, the map construction network of this application can more effectively improve the consistency between the semantic segmentation map and the vectorized map, and enhance the performance of online mapping.

[0045] When constructing a map, the acquired road image data usually includes multiple image frames. For any one of the multiple image frames (assumed to be the i-th image frame), the query weighting module can multiply the semantic segmentation features of the i-th image frame by the initial query of the i-th image frame to obtain the weighted query of the i-th image frame. The vector decoder can optimize the weighted query of the i-th image frame to obtain the optimized query of the i-th image frame, and perform vector decoding on the optimized query of the i-th image frame and the image features of the i-th image frame based on the attention mechanism to obtain the vector features of the i-th image frame. After obtaining the vector features of multiple image frames based on the above method, a vectorized map of the area to be built can be constructed based on the vector features of the multiple image frames.

[0046] In the related art, for each image frame, a number of initial queries are randomly obtained. The above method does not fully consider the consistency of the distribution of map elements. This neglect results in the query and reference point initialization strategies in each frame remaining unchanged, and each query must gradually approach the true value from scratch, which limits the convergence speed of the model.

[0047] To solve this problem, the present application proposes a novel query propagation mechanism that utilizes the cross-frame consistency of map elements in time and space. By this method, the present application can transfer the optimized query of the previous frame to the subsequent frames, thereby accelerating the approximation process of the optimized query of the subsequent frame. The above process can be implemented by a query compression and transmission module TSZT based on temporal consistency. The specific structure of the query compression and transmission module TSZT is as Figure 3 shown. Assume that the number of initial queries used for each image frame is N E , and each initial query is specifically used to detect a map element instance to ensure faster and more accurate feature alignment and map construction in all road image data.

[0048] For the first image frame, N E initial queries can be randomly generated. After the initial queries are fused with the semantic segmentation features to obtain the weighted queries, they are optimized by the vector decoder to obtain N E optimized queries of the first image frame, and these optimized queries are used to perform vector decoding on the image features of the first image frame.

[0049] For the (T - 1)-th image frame (T is an integer greater than 1), its optimized query is the result of the optimized weighted query of the (T - 1)-th image frame after being optimized by the vector decoder, denoted as:

[0050]

[0051] According to the temporal consistency of map element instances between different image frames, based on the pose relationship between the T-th image frame and the (T - 1)-th image frame, the query compression and transmission module TSZT can transfer the N E optimized queries from the (T - 1)-th image frame to the T-th image frame and compress them into K queries (denoted as compressed queries) as follows:

[0052]

[0053] These K queries obtained by compression can be used as the K initial queries for the T-th image frame. In addition, the remaining (N E - K) initial queries for the T-th image frame can be randomly generated. Thus, N E initial queries for the T-th image frame are obtained. The specific methods for obtaining the weighted queries and optimized queries for the T-th image frame are as described in the foregoing embodiments and will not be elaborated here.

[0054] In this way, the invariance of the geometric shape and position distribution of the same map element between adjacent frames can be fully considered, so as to effectively retain the detection information of the previous frame and improve the stability and accuracy of map element detection.

[0055] In addition, map elements are usually extended into a series of points for modeling, lacking a unified and effective representation method, resulting in possible differences in the mathematical models of the same map element between different frames. This application also designs a unified expression model to express map elements. This expression model can be represented in the form of a structure. Specifically, the i-th detected map element instance can be represented by the structure L i which i includes the element category class i , the instance number ID, and k vectorized key points (p i,1 , p i,2 , … p i,k ). The j-th vectorized key point p i,j of the i-th instance consists of its coordinate values (x j , y j ) and the attribute value attribute i,j .

[0056] L i = {class i , ID, (p i,1 , p i,2 , … p i,k )}

[0057] p i,j={(x j , y j ), attribute i,j}

[0058] Among them, the element category class i is used to represent the category of the i-th map element instance, such as lane lines, crosswalks, etc.; the instance number ID is used to uniquely identify each map element instance; k can be a preset positive integer; the coordinate values (x j , y j ) are used to represent the position of the j-th vectorized key point of the i-th instance; the attribute value attribute i,j is used to represent the key point category of the j-th vectorized key point of the i-th instance, and the key point category is used to characterize the importance level of the corresponding key point.

[0059] In some embodiments, the key point categories include:

[0060] The first category, which is used to indicate that the key point is the first key point that actually exists in the area to be mapped; the first key point is a map ground truth point, which depicts the inherent shape of the map, so it is a globally Figure 1 consistent key point. In the case of being converted to the global coordinate system, the map ground truth points of the same map element observed in different frames should be the same. Therefore, the points with this attribute are globally spatio-temporally consistent and are the most important.

[0061] The second category, which is used to indicate that the key point is the second key point at the field of view boundary of the image sensor. The second key point is usually the truncation point of the map element caused by the limited field of view of the image sensor. Since the image in the distance is prone to distortion, the prediction of this point is unstable, and its importance and priority are the lowest when obtaining the final result.

[0062] The third category, which is used to indicate that the key point is the third key point obtained by interpolation between the first key point and the second key point; the third key point is a map shape point, which is an interpolation point generated to maintain the shape of the map element and assist network prediction. Therefore, the map shape points may not be consistent between adjacent frames. Since the map shape is maintained, they also have great importance.

[0063] Different key point categories can be represented by different attribute values. For example, if the attribute value of a key point is 1, it means that the key point category of this key point is the first category; if the attribute value of a key point is 2, it means that the key point category of this key point is the second category; if the attribute value of a key point is 3, it means that the key point category of this key point is the third category.

[0064] In some embodiments, the vector decoder may include multiple classifiers, which are respectively used to obtain the position information of each key point on the map element instance, the element category of the map element instance, and the key point category of each key point on the map element instance.

[0065] Through unified representation, the present application can unify the key points of map element instances in different image frames, making the key points of the same map element instance in different image frames consistent. By distinguishing key points with different attributes, different importance can be assigned to key points with different attributes during map construction, improving the accuracy and reliability of the mapping result.

[0066] The solution of the present application has the following advantages:

[0067] (1) The map construction network of the present application adopts a hierarchical model, divides the autonomous driving map into a pixel-level grid map layer and a vector instance map element layer, establishes the correspondence between grid-level semantic segmentation features and vector features, promotes the information interaction between the two maps, and improves the mapping effect.

[0068] (2) The present application proposes a map element expression model that can unify the key points of map elements in the front and rear frames. This map model can distinguish shape key points that depict the shape of map elements, truncation points caused by field of view limitations, and interpolation points used to maintain the shape, effectively distinguishing the importance of different key points on the map element instance.

[0069] (3) Based on the general hierarchical model of the autonomous driving map and the map element expression model, the present invention also provides an online mapping network. Taking the single-frame image data collected by the camera sensor as input, the image features are learned through the deep feature extraction network, and the image features are input into the semantic decoder and the vector decoder to obtain two mapping results. In particular, for the vector decoder, vector decoding is performed on the image features and the weighted query, and then through three branches: a multi-layer perceptron branch (the first classifier) to obtain the map element key points that conform to the map element expression model, represented by the coordinate values of each key point; a multi-layer perceptron branch (the second classifier) to obtain the category corresponding to the map element instance, and the confidence of the map element instance is also the existence of the map element instance in the physical space; a multi-layer perceptron branch (the third classifier) to obtain the attribute value corresponding to the map element key point that conforms to the map element expression model, used to measure the importance of the map element key point.

[0070] (4) The present application proposes a map element temporal query transmission module, which transmits the optimized query of the previous image frame to the next image frame as the initial query of the next image frame, effectively improving the mapping accuracy.

[0071] See Figure 4, this application also provides a map construction method based on the map construction network in the foregoing embodiments. The method includes:

[0072] Step S1: Extract features from the road image data of the area to be mapped through a feature extraction network to obtain image features;

[0073] Step S2: Perform semantic segmentation on the image features through a semantic decoder to obtain semantic segmentation features; the semantic segmentation features are used to represent the semantic information of each pixel point in the road image data;

[0074] Step S3: Weight the initial query based on the semantic segmentation features through a query weighting module to obtain a weighted query;

[0075] Step S4: Optimize the weighted query through a vector decoder to obtain an optimized query, perform vector decoding on the optimized query and the image features based on the attention mechanism to obtain vector features, and construct a map of the area to be mapped based on the vector features.

[0076] An embodiment of this application also provides a computer device, which at least includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method described in any of the foregoing embodiments.

[0077] Figure 5 FIG. shows a more specific schematic diagram of the hardware structure of a computer device provided by an embodiment of this application. The device may include: a processor 301, a memory 302, an input / output interface 303, a communication interface 304, and a bus 305. Among them, the processor 301, the memory 302, the input / output interface 303, and the communication interface 304 are communicatively connected to each other inside the device through the bus 305.

[0078] The processor 301 may be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of this application. The processor 301 may further include a graphics card, and the graphics card may be an Nvidia titan X graphics card or a 1080Ti graphics card, etc.

[0079] The memory 302 may be implemented in the form of a Read Only Memory (ROM), a Random Access Memory (RAM), a static storage device, a dynamic storage device, etc. The memory 302 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of the present application through tools or firmware, the relevant program codes are stored in the memory 302 and are called and executed by the processor 301.

[0080] The input / output interface 303 is used to connect to the input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0081] The communication interface 304 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module may implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).

[0082] The bus 305 includes a path for transmitting information between various components of the device (such as the processor 301, the memory 302, the input / output interface 303, and the communication interface 304).

[0083] It should be noted that although the above device only shows the processor 301, the memory 302, the input / output interface 303, the communication interface 304, and the bus 305, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary for implementing the solutions of the embodiments of the present application and do not necessarily include all the components shown in the figure.

[0084] The embodiments of the present application provide a computer program product, including a computer program, which when executed by a processor implements the method described in any one of the embodiments of the present application.

[0085] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in any one of the foregoing embodiments.

[0086] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computer device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0087] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. When implementing the solutions of the embodiments of this application, the functions of the modules can be implemented in the same or multiple tools and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solutions of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0088] The above is only the specific implementation manner of the embodiments of this application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the embodiments of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of this application.

Claims

1. A map construction network, characterized in that: include: Feature extraction network, used to extract features from road image data in the area to be mapped to obtain image features; A semantic decoder, used for performing semantic segmentation on the image features to obtain semantic segmentation features; A query weighting module, used for weighting the initial query based on the semantic segmentation feature to obtain a weighted query; A vector decoder is used to optimize the weighted query to obtain an optimized query, perform vector decoding on the optimized query and the image features based on an attention mechanism to obtain vector features, and construct a vectorized map of the area to be mapped based on the vector features.

2. The map construction network according to claim 1, characterized in that: The vector decoder comprises: The first classifier is used to obtain the location information of each key point on the map element instance; A second classifier, used to obtain a feature category of a map feature instance; and The third classifier is used to obtain the key point category of each key point on the map element instance; the key point category is used to characterize the importance of the key point.

3. The map construction network according to claim 2, characterized in that: The key point categories include: The first category is used to indicate that the key point is a first key point that actually exists in the area to be mapped; A second category, used to indicate that the key point is a second key point at a boundary of a field of view of the image sensor; The third category is used to indicate that the key point is a third key point obtained by interpolation between the first key point and the second key point.

4. The map construction network according to claim 1, characterized in that: The road image data includes a plurality of image frames; The query weighting module optimizes the initial query based on the semantic segmentation feature to obtain a weighted query, including: Weighting the semantic segmentation features of the image frame and the initial query of the image frame to obtain a weighted query of the image frame; The process of constructing the map of the area to be mapped by the vector decoder includes: Optimizing the weighted query of the image frame to obtain an optimized query of the image frame; Performing vector decoding on the optimized query of the image frame and the image features of the image frame based on an attention mechanism to obtain the vector features of the image frame; A vectorized map of the area to be mapped is constructed based on the vector features of the multiple image frames.

5. The map construction network according to claim 4, characterized in that: The initial query of the first image frame among the plurality of image frames comprises a plurality of randomly generated initial queries; The initial query of the Tth image frame among the multiple image frames includes a plurality of randomly generated queries and a plurality of optimized queries of the T-1th image frame among the multiple image frames, where T is an integer greater than 1.

6. The map construction network according to claim 5, characterized in that: Before using the plurality of optimized queries of the T-1th image frame as the initial query of the Tth image frame, the query weighting module is further used for: Based on the relative position relationship between the T-th image frame and the T-1-th image frame, a plurality of optimization queries of the T-1-th image frame are transformed.

7. A map construction method based on the map construction network according to any one of claims 1 to 6, characterized in that: The method comprises: The feature extraction network is used to extract features from the road image data of the area to be mapped to obtain image features; Performing semantic segmentation on the image features through a semantic decoder to obtain semantic segmentation features; The initial query is optimized based on the semantic segmentation feature by a query weighting module to obtain a weighted query; The weighted query is optimized by a vector decoder to obtain an optimized query, the optimized query and the image features are vector decoded based on an attention mechanism to obtain vector features, and a vectorized map of the area to be mapped is constructed based on the vector features.

8. The method according to claim 7, characterized in that The vector decoder includes a first classifier, a second classifier and a third classifier; the vectorized map of the area to be mapped is constructed based on the vector features, including: Acquire the position information of each key point on the map element instance through the first classifier; Acquiring the feature category of the map feature instance by using the second classifier; and The key point category of each key point on the map element instance is obtained by the third classifier; the key point category is used to characterize the importance of the key point.

9. The method according to claim 8, characterized in that The key point categories include: The first category is used to indicate that the key point is a first key point that actually exists in the area to be mapped; A second category, used to indicate that the key point is a second key point at a boundary of a field of view of the image sensor; The third category is used to indicate that the key point is a third key point obtained by interpolation between the first key point and the second key point.

10. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 7 to 9 is implemented.

Citation Information

Cited By

  • Cross-region mowing method, device and equipment and storage medium

    CN120779968A