Map construction method and device, equipment, storage medium and program product
By inputting multi-view images into pre-trained map feature detection model for identification and modeling, the problems of slow construction speed and update lag of high-precision maps are solved, and more efficient high-precision map construction and update are achieved.
Patent Information
- Application Number
- CN202510097855.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-13
AI Technical Summary
Currently, the construction speed of high-precision maps is relatively slow, and there is lag in updates.
By inputting both the first and second viewing images acquired for the target area into the pre-trained map feature detection model for map feature recognition, the map feature recognition results are obtained, and closed and non-closed map elements are respectively modeled based on the recognition results to construct map data of the target area.
It realizes more efficient and high-precision map construction and update, improves processing efficiency, simplifies the conversion process from image to map representation, and solves the problems of limited perception range and occlusion of bicycles.
Smart Images

Figure CN120147567A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of traffic information processing, and particularly to a method, apparatus, device, storage medium and program product for map construction. Background Art
[0002] A high-precision map (also known as a high-accuracy map) provides a level of detail and accuracy far beyond that of traditional navigation maps. Such a map not only covers basic information of the road network, such as road shapes, intersection layouts, etc., but also includes centimeter-level details of traffic elements, vectorized topological structures, and rich navigation information. The high-precision map can inform the vehicle in advance about the conditions of the road ahead, such as slope, heading, curvature, etc., enabling the vehicle to better avoid potential risks.
[0003] The construction of a high-precision map is based on the Simultaneous Localization and Mapping (SLAM) technology. 3D point cloud data is obtained through sensors such as lidar, and the 3D point cloud data is projected into a Bird’s-Eye-View (BEV) space to obtain a ground grid. Then, the element information of the high-precision map is obtained on the ground grid through detection or manual annotation.
[0004] However, the current construction speed of high-precision maps is relatively slow, and there is a lag in updating. Summary of the Invention
[0005] Embodiments of the present application provide a method, apparatus, device, storage medium and program product for map construction, so as to achieve the effect of more efficient construction and updating of high-precision maps.
[0006] In a first aspect, an embodiment of the present application provides a method for map construction, including:
[0007] Input both the first perspective image and the second perspective image collected for a target area into a pre-trained map element detection model for map element recognition, to obtain a map element recognition result. The map element recognition result includes a vector point set corresponding to the recognized target map elements. The first perspective image is taken by a roadside device of the target area, the second perspective image is taken by a vehicle of the target area, and the target map elements include closed map elements and non-closed map elements;
[0008] Based on the map element recognition result, model the closed map elements and the non-closed map elements respectively, to generate a closed map element model and a non-closed map element model, so as to construct the map data of the target area based on the closed map element model and the non-closed map element model.
[0009] In a possible implementation, the map element recognition result includes the arrangement corresponding to the identified target map element, and the arrangement indicates the arrangement order of the vector points in the vector point set corresponding to the target map element. Based on the map element recognition result, modeling is respectively performed on the closed map element and the non-closed map element to generate a closed map element model and a non-closed map element model, including:
[0010] For any closed map element, modeling is performed according to a pre-configured first modeling strategy. During the modeling process, the connection order corresponding to the closed map element is determined based on the arrangement corresponding to the closed map element, and each vector point in the vector point set corresponding to the closed map element is connected in accordance with the connection order corresponding to the closed map element to form a closed figure, thereby constructing a closed map element model;
[0011] For any non-closed map element, modeling is performed according to a pre-configured second modeling strategy. During the modeling process, the connection order corresponding to the non-closed map element is determined based on the arrangement corresponding to the non-closed map element, and each vector point in the vector point set corresponding to the non-closed map element is connected to form a non-closed figure, thereby constructing a non-closed map element model.
[0012] In a possible implementation, inputting both the first perspective image and the second perspective image collected for the target area into a pre-trained map element detection model for map element recognition to obtain a map element recognition result includes:
[0013] Inputting both the first perspective image and the second perspective image collected for the target area into the encoding network of the pre-trained map element detection model for multi-scale feature extraction to obtain image domain features of different scales;
[0014] Inputting the image domain features into the feature fusion network of the map element detection model, where in the feature fusion network, the image domain features are fused to obtain fused image domain features, and the fused image domain features are mapped to the aerial view space to generate an aerial view space feature map;
[0015] Inputting the aerial view space feature map into the decoding network of the map element detection model to perform map element recognition on the aerial view space feature map to obtain a map element recognition result.
[0016] In a possible implementation, inputting the aerial view space feature map into the decoding network of the map element detection model to perform map element recognition on the aerial view space feature map to obtain a map element recognition result includes:
[0017] Input the aerial view space feature map into the decoding network of the map element detection model. In the decoding network, for any pixel in the aerial view space feature map, calculate the attention relationship between this pixel and other pixels in the same row to obtain the first attention feature of this pixel, and calculate the attention relationship between this pixel and other pixels in the same column to obtain the second attention feature of this pixel. Perform feature fusion on the first attention feature and the second attention feature to obtain the fused attention feature of this pixel;
[0018] Based on the fused attention features of each pixel, perform map element recognition to obtain the map element recognition result.
[0019] In a possible implementation manner, the training samples of the map element detection model include image sample pairs and the vector point set labels and arrangement mode labels corresponding to the image sample pairs. The image sample pairs include matching first perspective image samples and second perspective image samples. The first perspective image samples and the second perspective image samples are both labeled with corresponding map elements. The vector point set labels and arrangement mode labels corresponding to the image sample pairs are obtained through the following method:
[0020] Perform vectorization processing on the map elements in the first perspective image sample to obtain the first vector point set corresponding to the map elements in the first perspective image sample;
[0021] Perform vectorization processing on the map elements in the second perspective image sample to obtain the second vector point set corresponding to the map elements in the second perspective image sample;
[0022] According to the first vector point set and the second vector point set, obtain the vector point set labels and arrangement mode labels corresponding to the image sample pair.
[0023] In a possible implementation manner, the obtaining the vector point set labels and arrangement mode labels corresponding to the image sample pair according to the first vector point set and the second vector point set includes:
[0024] Fuse the first vector point set and the second vector point set to obtain a fused vector point set;
[0025] Uniformly sample the fused vector point set to obtain the vector point set labels corresponding to the image sample pair;
[0026] Determine the arrangement order of the vector points in the vector point set labels to generate the arrangement mode labels corresponding to the image sample pair.
[0027] In a possible implementation manner, the map element detection model is trained through the following method:
[0028] Input an image sample pair into a map element detection model to be trained, and obtain a map element recognition result of the image sample pair, where the map element recognition result includes a vector point set and an arrangement corresponding to the map element in the image sample pair;
[0029] Determine the loss of the arrangement corresponding to the map element in the image sample pair relative to the arrangement label corresponding to the image sample pair, and obtain first loss information;
[0030] Determine the loss of the vector point set corresponding to the map element in the image sample pair relative to the vector point set label corresponding to the image sample pair, and obtain second loss information;
[0031] Based on the first loss information and the second loss information, adjust the model parameters of the map element detection model.
[0032] In a second aspect, an embodiment of the present application provides a map construction device, including:
[0033] An identification module, configured to input both a first perspective image and a second perspective image collected for a target area into a pre-trained map element detection model for map element recognition, and obtain a map element recognition result, where the map element recognition result includes a vector point set corresponding to the identified target map element, the first perspective image is obtained by a roadside device photographing the target area, the second perspective image is obtained by a vehicle photographing the target area, and the target map element includes a closed map element and a non-closed map element;
[0034] A modeling module, configured to respectively model the closed map element and the non-closed map element based on the map element recognition result, generate a closed map element model and a non-closed map element model, and construct map data of the target area based on the closed map element model and the non-closed map element model.
[0035] In a possible implementation manner, the map element recognition result includes an arrangement corresponding to the identified target map element, and the arrangement indicates the arrangement order of vector points in the vector point set corresponding to the target map element. Specifically, the modeling module is configured to:
[0036] For any closed map element, perform modeling according to a pre-configured first modeling strategy. During the modeling process, determine the connection order corresponding to the closed map element based on the arrangement corresponding to the closed map element, and connect each vector point in the vector point set corresponding to the closed map element in the connection order corresponding to the closed map element to form a closed figure, and construct a closed map element model;
[0037] For any non-closed map element, model it according to a pre-configured second modeling strategy. During the modeling process, determine the connection order corresponding to the non-closed map element based on the arrangement mode corresponding to the non-closed map element, connect each vector point in the vector point set corresponding to the non-closed map element to form a non-closed figure, and construct a non-closed map element model.
[0038] In a possible implementation manner, the recognition module is specifically configured to:
[0039] Input both the first perspective image and the second perspective image collected for the target area into the encoding network of a pre-trained map element detection model for multi-scale feature extraction to obtain image domain features of different scales;
[0040] Input the image domain features into the feature fusion network of the map element detection model. In the feature fusion network, perform feature fusion on the image domain features to obtain fused image domain features, and map the fused image domain features to the aerial view space to generate an aerial view space feature map;
[0041] Input the aerial view space feature map into the decoding network of the map element detection model to perform map element recognition on the aerial view space feature map and obtain a map element recognition result.
[0042] In a possible implementation manner, the recognition module is specifically configured to:
[0043] Input the aerial view space feature map into the decoding network of the map element detection model. In the decoding network, for any pixel in the aerial view space feature map, calculate the attention relationship between the pixel and other pixels in the same row to obtain the first attention feature of the pixel, and calculate the attention relationship between the pixel and other pixels in the same column to obtain the second attention feature of the pixel. Perform feature fusion on the first attention feature and the second attention feature to obtain the fused attention feature of the pixel;
[0044] Perform map element recognition based on the fused attention features of each pixel to obtain the map element recognition result.
[0045] In a possible implementation manner, the training samples of the map element detection model include image sample pairs and vector point set labels and arrangement mode labels corresponding to the image sample pairs. The image sample pairs include matching first perspective image samples and second perspective image samples, and corresponding map elements are marked in both the first perspective image samples and the second perspective image samples. The vector point set labels and arrangement mode labels corresponding to the image sample pairs are obtained through the following method:
[0046] Vectorize the map elements in the first - perspective image sample to obtain a first vector point set corresponding to the map elements in the first - perspective image sample;
[0047] Vectorize the map elements in the second - perspective image sample to obtain a second vector point set corresponding to the map elements in the second - perspective image sample;
[0048] Based on the first vector point set and the second vector point set, obtain a vector point set label and an arrangement - style label corresponding to the image sample pair.
[0049] In a possible implementation manner, the step of obtaining a vector point set label and an arrangement - style label corresponding to the image sample pair according to the first vector point set and the second vector point set includes:
[0050] Fuse the first vector point set and the second vector point set to obtain a fused vector point set;
[0051] Uniformly sample the fused vector point set to obtain a vector point set label corresponding to the image sample pair;
[0052] Determine the arrangement order of the vector points in the vector point set label to generate an arrangement - style label corresponding to the image sample pair.
[0053] In a possible implementation manner, the map - element detection model is trained in the following way:
[0054] Input an image sample pair into a map - element detection model to be trained, and obtain a map - element recognition result of the image sample pair. The map - element recognition result includes a vector point set and an arrangement style corresponding to the map elements in the image sample pair;
[0055] Determine the loss of the arrangement style corresponding to the map elements in the image sample pair relative to the arrangement - style label corresponding to the image sample pair to obtain first loss information;
[0056] Determine the loss of the vector point set corresponding to the map elements in the image sample pair relative to the vector point set label corresponding to the image sample pair to obtain second loss information;
[0057] Based on the first loss information and the second loss information, adjust the model parameters of the map - element detection model.
[0058] In a third aspect, an embodiment of the present application provides a map construction device, including: a memory, a processor;
[0059] The memory stores computer - executable instructions;
[0060] The processor executes the computer-executable instructions stored in the memory, such that the processor executes the first aspect and / or various possible implementation manners of the first aspect as described above.
[0061] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the first aspect and / or various possible implementation manners of the first aspect as described above when being executed by a processor.
[0062] In a fifth aspect, an embodiment of the present application provides a computer program product including a computer program, which implements the first aspect and / or various possible implementation manners of the first aspect as described above when being executed by a processor.
[0063] The map construction method, device, equipment, storage medium and program product provided by the embodiments of the present application can input multi-view images collected for a target area into a pre-trained map element detection model to identify map elements, and obtain a map element identification result. The map element identification result includes a vector point set corresponding to the identified target map elements. The map element detection model of the present application can directly output the vector point set of the map elements. This representation method is convenient for subsequent map construction and update, because the vector point set can be directly used to generate a vector map or other forms of map representation. This direct output method improves the processing efficiency and simplifies the conversion process from an image to a map representation. In addition, by using multi-view images for map element identification, the multi-view images include first-view images captured by roadside devices. By introducing roadside devices at fixed positions, the limitations and occlusion problems of only using vehicles for environmental perception are solved. In the present application, according to the map element identification result, closed map elements and non-closed map elements are respectively modeled. This classification processing can improve the efficiency and accuracy of modeling. The present application provides a map construction method based on an end-to-end network architecture, achieving the effect of realizing more efficient high-precision map construction and update. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The accompanying drawings here are incorporated into the description and constitute a part of this description, showing embodiments consistent with the present application, and are used together with the description to explain the principles of the present application.
[0065] Figure 1 It is a schematic diagram of the scenario of the map construction method provided by the present application;
[0066] Figure 2 It is a schematic flowchart of the map construction method provided by the present application Figure 1 ;
[0067] Figure 3 It is a schematic flowchart of the map construction method provided by the present application Figure 2 ;
[0068] Figure 4 Schematic diagram of attention calculation based on direction decoupling provided for this application;
[0069] Figure 5 Schematic diagram of image sample pairs provided for this application;
[0070] Figure 6 Schematic diagram of annotation effect provided for this application;
[0071] Figure 7 Schematic diagram of generating vector point set labels provided for this application;
[0072] Figure 8 Schematic diagram of the structure of the map construction device provided for this application;
[0073] Figure 9 Schematic diagram of the structure of the map construction equipment provided for this application.
[0074] Through the above-mentioned drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0075] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.
[0076] For the high-precision map construction method based on simultaneous localization and mapping, there are limitations such as low efficiency of manual annotation and untimely update. The online high-precision map construction method based on vision and vehicle-mounted sensors is gradually replacing the traditional method based on simultaneous localization and mapping. The method based on vision and vehicle-mounted sensors uses multi-source sensor data, such as multi-view camera images and lidar point clouds, to construct a high-precision map, and is divided into two main methods: rasterization and vectorization.
[0077] The rasterization of high-precision map construction can be expressed as a semantic segmentation task in the BEV space. However, due to the difficulty in making instance-level distinctions and the lack of structural information of map elements, the rasterized high-precision map is not an ideal input for downstream tasks and requires cumbersome post-processing.
[0078] Vectorized high-precision map construction represents the map through map elements, which are usually represented by an ordered sequence of discrete points. Among them, through equivalent substitution modeling, map elements can be characterized as a set of point sets with equivalent arrangements, which can eliminate the ambiguity of the definition of the starting point and direction of map elements, enabling the model to more accurately understand and learn the shape characteristics of map elements.
[0079] Specifically, the rasterized high-precision map construction method relies on professional surveying vehicles and high-precision sampling data, with high annotation costs and untimely updates. In addition, the rasterized high-precision map construction method cannot distinguish map elements at the instance level. This means that different objects or feature points in the map may be treated as the same elements, thus losing the ability to identify and distinguish individual objects. Due to the lack of instance-level distinction and the structural information of map elements, a large amount of supplementation and correction is required in the subsequent processing, which is not friendly to downstream tasks.
[0080] The vectorized high-precision map construction method uses on-vehicle surround perception and point cloud data to construct a high-definition map. Although single-vehicle perception technology has made good progress, it is still affected by problems such as limited sensor field of view and changes in sensor data quality due to factors such as occlusion and rapid movement.
[0081] The map construction method provided in this application can input multi-view images collected for a target area into a pre-trained map element detection model for map element recognition to obtain map element recognition results. The map element recognition results contain the vector point set corresponding to the recognized target map elements. The map element detection model of this application can directly output the vector point set of map elements. This representation method is convenient for subsequent map construction and update because the vector point set can be directly used to generate a vector map or other forms of map representation. This direct output method improves the processing efficiency and simplifies the conversion process from images to map representation. In addition, by using multi-view images for map element recognition, the multi-view images include the first-view images captured by roadside devices. By introducing roadside devices at fixed positions, the limitations and occlusion problems of only using vehicles for environmental perception are solved. In this application, according to the map element recognition results, closed map elements and non-closed map elements are respectively modeled. This classification processing can improve the efficiency and accuracy of modeling. The map construction method provided based on the end-to-end network architecture solves the technical problems of slow high-precision map construction speed and untimely updates.
[0082] Figure 1 It is a schematic diagram of the scenario of the map construction method provided in this application, as Figure 1As shown in the figure, in this application scenario, the roadside device 101 captures the road and the surrounding environment to obtain a first - perspective image. The roadside device 101 sends the collected first - perspective image to the vehicle. The vehicle 102 performs map element recognition based on the obtained first - perspective image, combined with the second - perspective image captured by itself and the pre - deployed map element detection model, and constructs map data of the current area based on the obtained map element recognition result.
[0083] The following uses specific embodiments to elaborate in detail on the technical solution of the present application and how the technical solution of the present application solves the above - mentioned technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of the present application with reference to the accompanying drawings.
[0084] Figure 2 Flow schematic of the map construction method provided by the present application Figure 1 , as Figure 2 shown, the method includes:
[0085] S201. Input both the first - perspective image and the second - perspective image collected for the target area into the pre - trained map element detection model for map element recognition to obtain a map element recognition result. The map element recognition result includes a vector point set corresponding to the recognized target map elements. The first - perspective image is captured by the roadside device for the target area, and the second - perspective image is captured by the vehicle for the target area. The target map elements include closed map elements and non - closed map elements.
[0086] The first - perspective image is captured by a roadside device (such as a fixed camera), providing a top - down or wide - angle view. The roadside device is usually installed at a fixed position, capable of continuously monitoring a specific area and having the ability of continuous observation. Since the position of the roadside device is fixed, the internal and external parameters can be accurately calibrated. The roadside device is usually installed at a high place or a key position, capable of providing a wide - angle or top - down view, covering a larger area. This perspective can complement the limitations of in - vehicle sensors, especially in cases where the field of view is limited or blocked.
[0087] The second - perspective image is captured by the vehicle, providing a front - view or panoramic view.
[0088] In the embodiments of the present application, both the collected first - perspective image and the second - perspective image can be input into the pre - trained map element detection model. The model processes the input images and identifies the map elements therein, such as lane lines, crosswalks, and stop lines. Each map element in the map element recognition result is represented as a corresponding vector point set, and the vector point set describes the geometric shape and position of the corresponding map element. The map element recognition result includes two types of map elements: closed map elements and non - closed map elements.
[0089] In one embodiment, a map feature can be represented as a set of vector points with multiple equivalent arrangements, which means that all discrete points constituting a certain map feature are regarded as a set, and the arrangement order of these points in the set is equivalent for describing the shape, position and other characteristics of the map feature. In other words, no matter in what order these points are arranged, as long as their positions and attributes remain unchanged, the map feature they represent will not change.
[0090] S202. Based on the map feature recognition result, model the closed map feature and the non-closed map feature respectively to generate a closed map feature model and a non-closed map feature model, so as to construct the map data of the target area based on the closed map feature model and the non-closed map feature model.
[0091] In the embodiment of the present application, a closed map feature model and a non-closed map feature model are generated by using the set of vector points in the map feature recognition result, and then the generated closed map feature model and non-closed map feature model are used to construct the map data of the target area, realizing the rapid construction and dynamic update of the map.
[0092] The map construction method provided by the embodiment of the present application can input multi-view images collected for the target area into a pre-trained map feature detection model for map feature recognition to obtain a map feature recognition result. The map feature recognition result contains the set of vector points corresponding to the recognized target map feature. The map feature detection model of the present application can directly output the set of vector points of the map feature. This representation method is convenient for subsequent map construction and update, because the set of vector points can be directly used to generate a vector map or other forms of map representation. This direct output method improves the processing efficiency and simplifies the conversion process from an image to a map representation. In addition, by using multi-view images for map feature recognition, the multi-view images include the first-view images captured by roadside devices. By introducing roadside devices at fixed positions, the limitations and occlusion problems of only using vehicles for environmental perception are solved. In the present application, according to the map feature recognition result, the closed map feature and the non-closed map feature are modeled separately. This classification processing can improve the efficiency and accuracy of modeling. The present application provides a map construction method based on an end-to-end network architecture, achieving the effect of realizing more efficient high-precision map construction and update.
[0093] Figure 3 Flow schematic of the map construction method provided by the present application Figure 2 , such as Figure 3 shown. On the basis of the Figure 2 embodiment, the map construction method is described in detail. The method includes:
[0094] S301. Input both the first - perspective image and the second - perspective image collected for the target area into the pre - trained map element detection model for map element recognition to obtain a map element recognition result. The map element recognition result includes a vector point set corresponding to the recognized target map elements. The first - perspective image is obtained by a roadside device photographing the target area, and the second - perspective image is obtained by a vehicle photographing the target area. The target map elements include closed map elements and non - closed map elements. The map element recognition result includes the arrangement corresponding to the recognized target map elements, and the arrangement indicates the arrangement order of the vector points in the vector point set corresponding to the target map elements.
[0095] In the embodiment of the present application, the map element recognition result not only includes the vector point set corresponding to the recognized target map elements, but also may include the arrangement order of the vector points in the vector point set corresponding to the target map elements. The vector point set is the basic unit that composes the map element, usually represented in coordinate form. Each vector point represents a key position or inflection point of the map element. The arrangement order of the vector points determines how to connect these points to form a complete map element. Through the recognized vector point set and arrangement order, the map element model can be constructed more accurately. The map element model is a digital representation of various geographical elements in the map, and it is constructed based on the recognized vector point set and the arrangement order of these points. It should be noted that each map element in the map element recognition result may correspond to multiple arrangements, and these arrangements are equivalent. This means that in some cases, there can be multiple possible connection orders of the vector points without affecting the final representation of the map element.
[0096] In a possible implementation manner, inputting both the first - perspective image and the second - perspective image collected for the target area into the pre - trained map element detection model for map element recognition to obtain a map element recognition result may specifically include:
[0097] Input both the first - perspective image and the second - perspective image collected for the target area into the encoding network of the pre - trained map element detection model for multi - scale feature extraction to obtain image - domain features at different scales;
[0098] Input the image - domain features into the feature fusion network of the map element detection model. In the feature fusion network, perform feature fusion on the image - domain features to obtain fused image - domain features, and map the fused image - domain features to the aerial view space to generate an aerial view space feature map;
[0099] Input the aerial view space feature map into the decoding network of the map element detection model to perform map element recognition on the aerial view space feature map to obtain a map element recognition result.
[0100] The map element detection model includes an encoding network, a feature fusion network, and a decoding network. The encoding network is responsible for extracting multi-scale image domain features from the input image. The extracted image domain features are input into the feature fusion network. The role of the feature fusion network is to fuse features from different perspectives to generate more representative fused image domain features. The fused features are mapped to the bird's-eye view space to generate a bird's-eye view space feature map. The bird's-eye view space feature map is input into the decoding network. The decoding network is responsible for decoding the feature map into specific map element recognition results.
[0101] In this embodiment, performing multi-scale feature extraction helps to identify map elements of different scales. By using a Feature Pyramid Network (FPN), image domain features of different scales can be fused into a single-scale image domain feature, that is, the fused image domain feature. Mapping the fused image domain feature to the bird's-eye view space, a bird's-eye view space feature map is obtained, and the bird's-eye view space feature map can be expressed as X ∈ R H×W×C , where H is the spatial height, W is the spatial width, and C is the number of channels. In the decoding network, the bird's-eye view space feature map is used to identify specific map elements. Map elements are modeled as a point set structure.
[0102] In addition, in addition to effectively modeling the map elements in the image as a point set structure using the above map element recognition method based on multi-view images and deep learning models, point cloud data can also be used for map element recognition and the map elements can be modeled as a point set structure. In addition to deep learning models, other machine learning methods can also be considered to identify map elements. For example, classification algorithms such as support vector machines, random forests, and decision trees can be used to classify map elements. These methods usually need to first extract features in the image and then use the trained classifier for recognition.
[0103] In the embodiments of this application, a series of map elements are abstracted into a graph with a closed shape (such as a crosswalk) and a graph with an open shape (such as a lane line). Points are sequentially sampled through the graph boundary, and the closed map elements are abstracted into polygons, and the open map elements are abstracted into broken lines. Both polygons and broken lines can be represented as an ordered point set N v represents the number of points. Since the arrangement of the point set is not unique, there are equivalent arrangements for polygons and broken lines. For example, it is difficult to determine the direction of the lane separator in the middle of two opposite lanes. Both endpoints of the lane separator can be used as the starting point, resulting in the point set being organized in two directions. For an instance of a crosswalk abstracted into a closed polygon, any folding point can be selected as the starting point, so its equivalent arrangement will be more. Assuming that the point set is in a fixed arrangement is unreasonable. Therefore, in the embodiments of this application, represents the map element recognition result, where, A vector point set representing map elements, Γ = {γ k} represents the permutation corresponding to the vector point set V of map elements.
[0104] Specifically, for map elements of the polyline type, two equivalent permutations of Γ can be expressed as:
[0105]
[0106] For map elements of the polygon type, 2*N v equivalent permutations of Γ can be expressed as:
[0107]
[0108] The set of map elements identified from the top-down spatial feature map can be expressed as y i =(c i , V i , Γ i ), where c i , V i , Γ i respectively represent the instance category to which the map element belongs, the vector point set corresponding to the map element, and the permutation corresponding to the map element.
[0109] In a possible implementation, the top-down spatial feature map is input into the decoding network of the map element detection model to perform map element recognition on the top-down spatial feature map, and the map element recognition result can be obtained, specifically including:
[0110] The top-down spatial feature map is input into the decoding network of the map element detection model. In the decoding network, for any pixel in the top-down spatial feature map, calculate the attention relationship between the pixel and other pixels in the same row to obtain the first attention feature of the pixel, and calculate the attention relationship between the pixel and other pixels in the same column to obtain the second attention feature of the pixel, and perform feature fusion on the first attention feature and the second attention feature to obtain the fused attention feature of the pixel;
[0111] Based on the fused attention features of each pixel, perform map element recognition to obtain the map element recognition result.
[0112] In the decoding network, for any pixel in the bird's-eye view spatial feature map, the attention relationship between this pixel and other pixels in the same row is calculated. This horizontal attention relationship can be achieved through the dot product attention mechanism, and the first attention feature captures the contextual information of the pixel in the horizontal direction. Similarly, the attention relationship between this pixel and other pixels in the same column is calculated. This vertical attention relationship can also be achieved through the dot product attention mechanism, and the second attention feature captures the contextual information of the pixel in the vertical direction.
[0113] Referring to Figure 4 As shown, it is a schematic diagram of attention calculation based on direction decoupling provided by this application. As Figure 4 shown, for the input feature map x, the i-th row and the j-th column are represented by x i,j (i = 0, 1, …, H - 2, H - 1; j = 0, 1, …, W - 2, W - 1). For the pixel point x s,j , the set of pixel points in the same horizontal direction as it can be denoted as Ω H (s, j), and the set of pixel points in the same vertical direction as it can be denoted as Ω V (s, j). The attention relationship calculation formula can be:
[0114]
[0115] where θ() represents performing a convolution operation on the corresponding pixel point; φ() similarly represents performing a convolution operation on the corresponding pixel point; g() similarly represents performing a convolution operation on the corresponding pixel point; f() represents a similarity function. y s,j represents the attention feature of x s,j in the horizontal direction, and z i,j represents the attention feature of x s,j in the vertical direction.
[0116] In this embodiment, the attention calculation path is split into two directions, horizontal and vertical. This decoupling method allows the attention in the horizontal and vertical directions to be calculated separately, reducing the complexity when calculating both directions simultaneously. Although the calculation path is split into two directions, by calculating the attention in the horizontal and vertical directions sequentially, the model can still capture the global contextual information.
[0117] In a possible implementation manner, the training samples of the map element detection model include image sample pairs, vector point set labels corresponding to the image sample pairs, and arrangement mode labels. The image sample pairs include matching first perspective image samples and second perspective image samples. Both the first perspective image samples and the second perspective image samples are labeled with corresponding map elements. The vector point set labels and arrangement mode labels corresponding to the image sample pairs are obtained through the following method:
[0118] Performing vectorization processing on the map elements in the first-view image sample to obtain a first vector point set corresponding to the map elements in the first-view image sample;
[0119] Performing vectorization processing on the map elements in the second-view image sample to obtain a second vector point set corresponding to the map elements in the second-view image sample;
[0120] According to the first vector point set and the second vector point set, vector point set labels and arrangement mode labels corresponding to the image sample pairs are obtained.
[0121] The first-view image sample and the second-view image sample are labeled with the same map features. The map features in the first-view image sample are vectorized to obtain a first vector point set. Similarly, the map features in the second-view image sample are vectorized to obtain a second vector point set. The first vector point set and the second vector point set are compared and matched to generate a vector point set label for the image sample pair. The arrangement order of the vector points is determined to generate an arrangement label. The vector point set label represents the vector point set target or result that the map feature detection model needs to predict. The arrangement label represents the arrangement target or result that the map feature detection model needs to predict. Labels are a key part in supervised learning because labels provide standard answers for the model, enabling the model to adjust and optimize its parameters by comparing the predicted results with the labels.
[0122] Reference Figure 5 As shown in the figure, it is a schematic diagram of the image sample pair provided by the present application. Images captured by the vehicle and the roadside equipment at the same time can be extracted from the data set as image sample pairs for training. The first-view image samples and the second-view image samples are annotated with map elements to obtain the first-view image samples and the second-view image samples annotated with map elements. The annotation effect can be as follows Figure 6 The figure shows the pedestrian crossing, lane markings and stop line.
[0123] Since the high-precision map is constructed in the bird's-eye view space, the annotation results need to be converted from the original pixel coordinates to the world coordinates of the bird's-eye view space. The pixel coordinates can be converted to image coordinates, and then the image coordinates can be converted to camera coordinates using the camera's intrinsic parameters, and the camera coordinates can be converted to world coordinates using the camera's extrinsic parameters (rotation matrix and translation vector).
[0124] The relationship between pixel coordinates (usually expressed as u, v) and image coordinates (usually expressed as x, y) can be expressed by the following formula:
[0125]
[0126] Among them, (x 0 ,y 0) is the central coordinate of the imaging plane, and dx and dy respectively represent the lengths of physical pixels in the two coordinate axis directions. The above formula can be expressed in matrix form as:
[0127]
[0128] Its inverse transformation is:
[0129]
[0130] The three-dimensional point (x c , y c , z c ) in the camera coordinate system and the connection line between the camera center O c intersect the imaging plane at a point (x, y). According to the proportional relationship of similar triangles, we have:
[0131]
[0132] where f represents the focal length, and z c represents the depth. Through the above formula, the relationship between the image coordinates and the camera coordinates can be obtained:
[0133]
[0134] Finally, through rotation and translation, the annotation result in the overlooking space can be obtained. However, due to the lack of camera depth data, direct conversion is not possible. Therefore, the point cloud sample corresponding to the image sample pair can be obtained, the image depth information can be determined using the point cloud sample, and based on the image depth information, the map elements mapped to the overlooking space can be obtained. After that, the map elements mapped to the overlooking space can be converted into a vector point set, and the vector point set label and arrangement label of the image sample pair can be generated.
[0135] In a possible implementation manner, according to the first vector point set and the second vector point set, the vector point set label and arrangement label corresponding to the image sample pair are obtained, which may specifically include:
[0136] Fuse the first vector point set and the second vector point set to obtain a fused vector point set;
[0137] Uniformly sample the fused vector point set to obtain the vector point set label corresponding to the image sample pair;
[0138] Determine the arrangement order of the vector points in the vector point set label to generate the arrangement label corresponding to the image sample pair.
[0139] Represent the object to be modeled as a set of vector points. Then, using the principle of equivalent substitution, arrange or transform the points in the point set to generate multiple equivalent point set representations. These equivalent point sets are consistent with the original object in shape, but may differ in the arrangement and position of the points.
[0140] In this embodiment, the map elements annotated in the image sample pair are modeled as a vector point set containing equivalent arrangements. Specifically, the first vector point set and the second vector point set can be fused. The vector point sets extracted from the first perspective image sample and the second perspective image sample usually contain different perspective information of the same map element. Fusing these two vector point sets can combine the advantages of different perspectives. Uniformly sample the fused vector point set. This means selecting a set of representative points in the fused vector point set to form the final vector point set label. Determine the arrangement order of the vector points in the vector point set label. The arrangement order determines how to connect these points to form a complete map element. By fusing multi-perspective information and uniform sampling, the generated vector point set label and arrangement mode label can more accurately reflect the true shape and structure of the map element.
[0141] In one embodiment, the first vector point set and the second vector point set corresponding to a certain map element are fused to obtain the fused vector point set corresponding to the map element. Since each map element may be represented by a different number of vector points respectively, each map element can be standardized to be represented by a fixed number of vector points. Specifically, a linear interpolation method can be used for sampling. Based on the vector points in the fused vector point set, calculate the total length of the entire vector path. According to the required number of vector points, determine the sampling interval, and then start from the starting point and sample along the vector path at a fixed sampling interval to obtain a vector point set label containing the specified number of vector points. Refer to Figure 7 As shown, it is a schematic diagram of generating a vector point set label provided by this application. The vector point set generated by uniform sampling in the figure is the vector point set label.
[0142] In a possible embodiment, the map element detection model is trained in the following manner:
[0143] Input the image sample pair into the map element detection model to be trained, and obtain the map element recognition result of the image sample pair. The map element recognition result includes the vector point set and arrangement mode corresponding to the map element in the image sample pair;
[0144] Determine the loss of the arrangement mode corresponding to the map element in the image sample pair relative to the arrangement mode label corresponding to the image sample pair to obtain the first loss information;
[0145] Determine the loss of the vector point set corresponding to the map element in the image sample pair relative to the vector point set label of the image sample pair, and obtain the second loss information;
[0146] Based on the first loss information and the second loss information, adjust the model parameters of the map element detection model.
[0147] In model training, minimizing the loss is one of the key steps. The model loss usually includes classification loss and position matching loss. To minimize the matching cost between the ground truth label and the prediction result, a matching cost function can be defined and optimized. The prediction result (i.e., the map element recognition result) includes the instance prediction result and the point set prediction result. The instance prediction result includes the instance category to which the recognized map element belongs, and the point set prediction result includes the vector point set and the arrangement corresponding to the recognized map element. Correspondingly, the ground truth label includes the instance category label, the vector point set label, and the arrangement label.
[0148] Suppose represents the N predicted map elements, y i =(c i , V i , Γ i ), where c i , V i , Γ i represent the instance category to which the map element y i belongs, the vector point set corresponding to the map element y i , and the arrangement corresponding to the map element y i . The arrangement of map elements that minimizes the instance-level matching cost can be found using the following formula:
[0149]
[0150] where represents the i-th predicted map element under the map element arrangement π; y i represents the i-th ground truth map element; represents the instance matching cost between the predicted map element and the ground truth map element; π ∈ π N represents all possible map element arrangements; is the arrangement with the minimum instance-level matching cost among the arrangements of N map elements
[0151] In addition, the point-level loss also needs to be considered. The arrangement of vector points that minimizes the point-level matching cost can be found using the following formula:
[0152]
[0153] where represents the j-th prediction point in the prediction point set; v γ(j) represents the j-th true point under the vector point arrangement γ; represents the point-level matching cost between the prediction point and the true point; γ ∈ Γ represents all possible vector point arrangements; is N v the arrangement (γ ∈ Γ) with the minimum point-level matching cost among the arrangements of N vector points.
[0154] In this application, by abstracting map elements into vector point sets, constructing an end-to-end neural network through equivalent replacement vector modeling, and based on an improved attention calculation method, the operation efficiency on low-computing-power devices is improved. By introducing roadside device data, the problems of limited perception range of single vehicles and easy occlusion are solved. Through the end-to-end network architecture, online inference of high-precision maps is realized, avoiding the cumbersome post-processing process, and realizing the online construction of vehicle-road collaborative high-precision maps in low-cost and low-computing-power scenarios.
[0155] S302. For any closed map element, model it according to the pre-configured first modeling strategy. During the modeling process, determine the connection order corresponding to the closed map element based on the corresponding arrangement method of the closed map element, and connect each vector point in the vector point set corresponding to the closed map element in the connection order corresponding to the closed map element to form a closed figure, thereby constructing a closed map element model.
[0156] S303. For any non-closed map element, model it according to the pre-configured second modeling strategy. During the modeling process, determine the connection order corresponding to the non-closed map element based on the corresponding arrangement method of the non-closed map element, and connect each vector point in the vector point set corresponding to the non-closed map element to form a non-closed figure, thereby constructing a non-closed map element model.
[0157] In the embodiment of this application, different modeling strategies are adopted for the different characteristics of closed map elements and non-closed map elements. For closed map elements, each vector point in the vector point set is connected in sequence according to the determined connection order. Ensure that the last point is connected to the first point to form a complete closed figure.
[0158] For non-closed map elements, each vector point in the vector point set is connected in sequence according to the determined connection order. There is no need to connect the last point to the first point, keeping the path open.
[0159] In specific implementation, the map element recognition result output by the map element detection model also includes the instance category to which the recognized map element belongs. Based on the output instance category, it can be determined whether the recognized map element is a closed map element or a non-closed map element.
[0160] The map construction method provided by the embodiments of the present application can input multi-view images collected for a target area into a pre-trained map feature detection model for map feature recognition to obtain a map feature recognition result. The map feature recognition result includes a vector point set corresponding to the recognized target map features. The map feature detection model of the present application can directly output the vector point set of the map features. This representation method is convenient for subsequent map construction and update because the vector point set can be directly used to generate a vector map or other forms of map representation. This direct output method improves the processing efficiency and simplifies the conversion process from images to map representation. In addition, by using multi-view images for map feature recognition, the multi-view images include first-view images captured by roadside devices. By introducing roadside devices at fixed positions, the limitations and occlusion problems of only using vehicles for environmental perception are solved. In the present application, according to the map feature recognition result, the closed map features and non-closed map features are respectively modeled. This classification processing can improve the efficiency and accuracy of modeling. The present application provides a map construction method based on an end-to-end network architecture, achieving the effect of more efficient high-precision map construction and update.
[0161] Figure 8 FIG. is a schematic structural diagram of a map construction device provided by the present application, as Figure 8 shown. The map construction device 80 provided in this embodiment includes:
[0162] An identification module 801, configured to input both the first-view image and the second-view image collected for the target area into a pre-trained map feature detection model for map feature recognition, to obtain a map feature recognition result. The map feature recognition result includes a vector point set corresponding to the recognized target map features. The first-view image is captured by a roadside device for the target area, and the second-view image is captured by a vehicle for the target area. The target map features include closed map features and non-closed map features;
[0163] A modeling module 802, configured to respectively model the closed map features and the non-closed map features based on the map feature recognition result, to generate a closed map feature model and a non-closed map feature model, so as to construct map data of the target area based on the closed map feature model and the non-closed map feature model.
[0164] In a possible implementation manner, the map feature recognition result includes an arrangement method corresponding to the recognized target map features. The arrangement method indicates the arrangement order of the vector points in the vector point set corresponding to the target map features. The modeling module is specifically configured to:
[0165] For any closed map element, perform modeling according to a pre-configured first modeling strategy. During the modeling process, determine the connection order corresponding to the closed map element based on the arrangement mode corresponding to the closed map element, connect each vector point in the vector point set corresponding to the closed map element in the connection order corresponding to the closed map element to form a closed figure, and construct a closed map element model;
[0166] For any non-closed map element, perform modeling according to a pre-configured second modeling strategy. During the modeling process, determine the connection order corresponding to the non-closed map element based on the arrangement mode corresponding to the non-closed map element, connect each vector point in the vector point set corresponding to the non-closed map element to form a non-closed figure, and construct a non-closed map element model.
[0167] In a possible implementation manner, the recognition module is specifically configured to:
[0168] Input both the first perspective image and the second perspective image collected for the target area into the encoding network of a pre-trained map element detection model for multi-scale feature extraction to obtain image domain features of different scales;
[0169] Input the image domain features into the feature fusion network of the map element detection model. In the feature fusion network, perform feature fusion on the image domain features to obtain fused image domain features, and map the fused image domain features to the aerial view space to generate an aerial view space feature map;
[0170] Input the aerial view space feature map into the decoding network of the map element detection model to perform map element recognition on the aerial view space feature map and obtain a map element recognition result.
[0171] In a possible implementation manner, the recognition module is specifically configured to:
[0172] Input the aerial view space feature map into the decoding network of the map element detection model. In the decoding network, for any pixel in the aerial view space feature map, calculate the attention relationship between the pixel and other pixels in the same row to obtain the first attention feature of the pixel, and calculate the attention relationship between the pixel and other pixels in the same column to obtain the second attention feature of the pixel. Perform feature fusion on the first attention feature and the second attention feature to obtain the fused attention feature of the pixel;
[0173] Perform map element recognition based on the fused attention features of each pixel to obtain a map element recognition result.
[0174] In a possible implementation, the training samples of the map feature detection model include image sample pairs, vector point set labels corresponding to the image sample pairs, and arrangement mode labels. The image sample pairs include matching first-perspective image samples and second-perspective image samples. The corresponding map features are labeled in both the first-perspective image samples and the second-perspective image samples. The vector point set labels and arrangement mode labels corresponding to the image sample pairs are obtained through the following methods:
[0175] Perform vectorization processing on the map features in the first-perspective image samples to obtain a first vector point set corresponding to the map features in the first-perspective image samples;
[0176] Perform vectorization processing on the map features in the second-perspective image samples to obtain a second vector point set corresponding to the map features in the second-perspective image samples;
[0177] Based on the first vector point set and the second vector point set, obtain the vector point set labels and arrangement mode labels corresponding to the image sample pairs.
[0178] In a possible implementation, based on the first vector point set and the second vector point set, obtaining the vector point set labels and arrangement mode labels corresponding to the image sample pairs includes:
[0179] Fuse the first vector point set and the second vector point set to obtain a fused vector point set;
[0180] Uniformly sample the fused vector point set to obtain the vector point set labels corresponding to the image sample pairs;
[0181] Determine the arrangement order of the vector points in the vector point set labels to generate the arrangement mode labels corresponding to the image sample pairs.
[0182] In a possible implementation, the map feature detection model is trained through the following methods:
[0183] Input the image sample pairs into the map feature detection model to be trained to obtain the map feature recognition results of the image sample pairs. The map feature recognition results include the vector point set and arrangement mode corresponding to the map features in the image sample pairs;
[0184] Determine the loss of the arrangement mode corresponding to the map features in the image sample pairs relative to the arrangement mode labels corresponding to the image sample pairs to obtain the first loss information;
[0185] Determine the loss of the vector point set corresponding to the map features in the image sample pairs relative to the vector point set labels corresponding to the image sample pairs to obtain the second loss information;
[0186] Based on the first loss information and the second loss information, adjust the model parameters of the map feature detection model.
[0187] The map construction device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.
[0188] Figure 9 It is a schematic structural diagram of the map construction device provided in this application. As Figure 9 shown, the map construction device 90 provided in this embodiment includes: at least one processor 901 and a memory 902. Optionally, the device 90 further includes a communication component 903. Among them, the processor 901, the memory 902, and the communication component 903 are connected through a bus.
[0189] In the specific implementation process, at least one processor 901 executes the computer execution instructions stored in the memory 902, so that at least one processor 901 executes the above method.
[0190] For the specific implementation process of the processor 901, reference can be made to the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.
[0191] In the above embodiment, it should be understood that the processor may be a central processing unit (English: Central Processing Unit, abbreviated as: CPU), or other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0192] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0193] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.
[0194] This application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0195] This application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the above method is implemented.
[0196] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disk. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0197] An exemplary readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in the device.
[0198] The division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in an electrical, mechanical or other forms.
[0199] The unit described as a separating component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of these units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0200] In addition, in each embodiment of the present invention, each functional unit may be integrated into one processing unit, may exist physically separately for each unit, or two or more units may be integrated into one unit.
[0201] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0202] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the aforementioned storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0203] Finally, it should be noted that: after considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other implementation schemes of the present invention. The present invention aims to cover any variations, uses, or adaptive changes of the present invention. These variations, uses, or adaptive changes follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field of the present invention that are not disclosed in the present invention. It is not limited to the exact structure described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A map construction method, characterized in that: include: Inputting the first-perspective image and the second-perspective image collected for the target area into a pre-trained map element detection model for map element recognition to obtain a map element recognition result, wherein the map element recognition result includes a vector point set corresponding to the recognized target map element, the first-perspective image is obtained by photographing the target area with a roadside device, the second-perspective image is obtained by photographing the target area with a vehicle, and the target map element includes a closed map element and a non-closed map element; Based on the map element recognition result, the closed map elements and the non-closed map elements are modeled respectively to generate closed map element models and non-closed map element models, so as to construct map data of the target area based on the closed map element models and the non-closed map element models.
2. The method according to claim 1, characterized in that The map element recognition result includes an arrangement mode corresponding to the recognized target map element, and the arrangement mode indicates an arrangement order of vector points in a vector point set corresponding to the target map element. Based on the map element recognition result, the closed map element and the non-closed map element are modeled respectively to generate a closed map element model and a non-closed map element model, including: For any closed map element, modeling is performed according to a pre-configured first modeling strategy. During the modeling process, a connection order corresponding to the closed map element is determined based on an arrangement mode corresponding to the closed map element, and each vector point in a vector point set corresponding to the closed map element is connected according to the connection order corresponding to the closed map element to form a closed figure, thereby constructing a closed map element model; For any non-closed map element, modeling is performed according to a pre-configured second modeling strategy. During the modeling process, the connection order corresponding to the non-closed map element is determined based on the arrangement method corresponding to the non-closed map element, and each vector point in the vector point set corresponding to the non-closed map element is connected to form a non-closed figure, so as to construct a non-closed map element model.
3. The method according to claim 1, characterized in that The first-view image and the second-view image acquired for the target area are input into a pre-trained map element detection model to perform map element recognition, and obtain a map element recognition result, including: The first-view image and the second-view image collected for the target area are input into the encoding network of the pre-trained map feature detection model to extract multi-scale features, and obtain image domain features of different scales; Inputting the image domain features into the feature fusion network of the map element detection model, performing feature fusion on the image domain features in the feature fusion network to obtain fused image domain features, and mapping the fused image domain features to the bird's-eye view space to generate a bird's-eye view space feature map; The bird's-eye view spatial feature map is input into the decoding network of the map element detection model, and map element recognition is performed on the bird's-eye view spatial feature map to obtain a map element recognition result.
4. The method according to claim 3, characterized in that The step of inputting the bird's-eye view spatial feature map into the decoding network of the map element detection model, performing map element recognition on the bird's-eye view spatial feature map, and obtaining a map element recognition result includes: The bird's-eye view spatial feature map is input into a decoding network of the map element detection model. In the decoding network, for any pixel in the bird's-eye view spatial feature map, the attention relationship between the pixel and other pixels on the same row is calculated to obtain a first attention feature of the pixel, and the attention relationship between the pixel and other pixels on the same column is calculated to obtain a second attention feature of the pixel. The first attention feature and the second attention feature are feature fused to obtain a fused attention feature of the pixel. Map element recognition is performed based on the fused attention features of each pixel to obtain the map element recognition result.
5. The method according to any one of claims 1 to 4, characterized in that The training samples of the map feature detection model include an image sample pair and a vector point set label and an arrangement mode label corresponding to the image sample pair, the image sample pair includes a matched first-view image sample and a second-view image sample, the first-view image sample and the second-view image sample are both annotated with corresponding map elements, and the vector point set label and the arrangement mode label corresponding to the image sample pair are obtained in the following manner: Performing vectorization processing on the map elements in the first-view image samples to obtain a first vector point set corresponding to the map elements in the first-view image samples; Performing vectorization processing on the map elements in the second-view image samples to obtain a second vector point set corresponding to the map elements in the second-view image samples; According to the first vector point set and the second vector point set, vector point set labels and arrangement mode labels corresponding to the image sample pairs are obtained.
6. The method according to claim 5, characterized in that The step of obtaining the vector point set labels and arrangement mode labels corresponding to the image sample pairs according to the first vector point set and the second vector point set includes: Fusing the first vector point set and the second vector point set to obtain a fused vector point set; Uniformly collecting the fused vector point set to obtain vector point set labels corresponding to the image sample pairs; Determine the arrangement order of the vector points in the vector point set label, and generate an arrangement mode label corresponding to the image sample pair.
7. The method according to any one of claims 1 to 4, characterized in that The map feature detection model is trained in the following way: Inputting the image sample pair into the map element detection model to be trained to obtain the map element recognition result of the image sample pair, wherein the map element recognition result includes the vector point set and arrangement method corresponding to the map element in the image sample pair; Determine the loss of the arrangement mode corresponding to the map elements in the image sample pair relative to the arrangement mode label corresponding to the image sample pair to obtain first loss information; Determine the loss of the vector point set corresponding to the map element in the image sample pair relative to the label of the vector point set corresponding to the image sample pair to obtain second loss information; Based on the first loss information and the second loss information, model parameters of a map feature detection model are adjusted.
8. A map construction device, characterized in that: include: a recognition module, used for inputting the first-perspective image and the second-perspective image acquired for the target area into a pre-trained map element detection model for map element recognition, and obtaining a map element recognition result, wherein the map element recognition result includes a vector point set corresponding to the recognized target map element, the first-perspective image is obtained by photographing the target area by a roadside device, the second-perspective image is obtained by photographing the target area by a vehicle, and the target map element includes a closed map element and a non-closed map element; A generation module is used to model the closed map elements and the non-closed map elements respectively based on the map element recognition results, generate closed map element models and non-closed map element models, and construct map data of the target area based on the closed map element models and the non-closed map element models.
9. A map construction device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 7.
10. A computer-readable storage medium / computer program product, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor; and / or, The computer program product comprises a computer program, which implements the method according to any one of claims 1 to 7 when executed by a processor.
Citation Information
Cited By
Method for processing map element detection model and method for generating vectorized map
CN120766026A