Vector map construction method and system based on prior map and multi-stage diffusion reasoning

Through a vector map construction method based on prior maps and multi-stage diffusion inference, the robustness and accuracy of online map construction in complex environments is solved, high-quality image recovery and enhancement effects are achieved, and stable navigation of the autonomous driving perception system under harsh conditions is supported.

CN120429459AActive Publication Date: 2025-08-05BEIHANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510529068.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-05
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing online map construction technology has insufficient robustness and accuracy in complex environments (such as camera occlusion, bad weather, etc.), and it is difficult to deal with multiple image degradation problems in a unified manner, affecting the stability and accuracy of the autonomous driving perception system.

Method used

A vector map construction method based on prior maps and multi-stage diffusion inference is adopted, and a variety of prior maps are encoded through a unified vector encoder, and multi-stage diffusion inference and feature fusion are combined with vehicle-mounted sensor data to generate high-quality online vector maps.

Benefits of technology

It significantly improves the robustness and accuracy of online map construction, can provide high-quality image recovery and enhancement effects in complex environments, and improves the stability and accuracy of autonomous driving perception systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429459A_ABST
    Figure CN120429459A_ABST
Patent Text Reader

Abstract

The invention provides a priori map and multi-stage diffusion reasoning-based vector map construction method and system, and the method comprises the steps: carrying out the coding of a plurality of priori maps through employing a unified vector encoder, and generating priori features; and converting the obtained real-time data into bird's-eye view features, and executing multi-stage diffusion reasoning to map the bird's-eye view features of a degradation domain to diffusion features of a normal domain. And performing fusion processing on the prior features and the diffusion features to obtain fusion features. And inputting the fusion features into a constructed map decoder to generate an online vector map. According to the scheme, multiple prior maps are introduced, and different prior maps are efficiently coded by using the unified vector encoder, so that the robustness and accuracy of online map construction are enhanced. Through a multi-stage diffusion reasoning technology, image details are gradually optimized, various image degradation problems can be processed in a unified manner, high-quality image restoration and enhancement effects can be provided in a complex environment, and powerful support is provided for an automatic driving perception system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and more specifically, to a method and system for constructing a vector map based on a priori maps and multi-stage diffusion reasoning. Background Art

[0002] The field of autonomous driving perception relies heavily on high-precision maps and real-time environmental perception technologies to ensure vehicles can navigate safely and accurately in complex environments. However, existing online map-building technologies have significant limitations when dealing with complex environments (such as camera occlusion and inclement weather). In particular, traditional methods for building and updating high-precision maps often rely on a single source of perception data (such as onboard cameras or LiDAR), which is susceptible to environmental interference, resulting in incomplete map data or large errors.

[0003] In addition, existing image restoration and enhancement methods are usually designed for a single task (such as denoising, deblurring, or low-light enhancement), and it is difficult to handle multiple degradation problems in a unified manner, which limits their application in complex scenes. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a vector map construction method and system based on a priori maps and multi-stage diffusion reasoning, so as to improve image restoration and enhancement effects and enhance the robustness and accuracy of online map construction.

[0005] In a first aspect, the present invention provides a method for constructing a vector map based on a priori maps and multi-stage diffusion reasoning, the method comprising:

[0006] Obtaining a plurality of pre-stored prior maps, encoding each of the prior maps using a unified vector encoder to generate prior features;

[0007] obtaining real-time data detected by an on-board sensor, converting the real-time data into bird's-eye view features, and performing multi-stage diffusion inference on the bird's-eye view features to map the bird's-eye view features in a degraded domain to diffusion features in a normal domain;

[0008] Fusing the priori feature and the diffusion feature to obtain a fused feature;

[0009] The fused features are input into the constructed map decoder to generate an online vector map.

[0010] In an optional embodiment, the a priori map includes a plurality of data vectors, each of the data vectors being composed of a plurality of points;

[0011] The step of encoding each of the prior maps using a unified vector encoder to generate prior features includes:

[0012] Processing the data vectors and points in the data vectors in each of the prior maps to obtain various types of feature information;

[0013] Inputting multiple types of feature information into the unified vector encoder, and using the intra-vector attention encoder in the unified vector encoder to process the interaction between points within the data vector;

[0014] The inter-vector attention encoder in the unified vector encoder is used to process the interactions between different data vectors to generate prior features.

[0015] In an optional embodiment, the step of processing the data vectors and points in the data vectors in each of the prior maps to obtain multiple types of feature information includes:

[0016] Encoding the position information and direction information of each point of each data vector in the prior map to obtain local features of each point;

[0017] Setting an instance-level tag for each of the data vectors, and encoding the instance-level tag using a multi-layer perceptron to generate instance-level embedded features;

[0018] assigning a corresponding type code based on the element type of each data vector, and encoding the type code using an embedding matrix to generate a map element feature;

[0019] A position index is assigned to each of the points, and the position index is encoded using a position encoding function to generate a position information feature.

[0020] In an optional embodiment, the step of performing multi-stage diffusion inference on the bird's-eye view feature to map the bird's-eye view feature in the degraded domain to a diffusion feature in the normal domain includes:

[0021] Performing a forward diffusion process on the bird's-eye view feature in the degraded domain to add noise data and convert the image into a noise image;

[0022] Performing a reverse reasoning operation on the noisy image to restore the original data, generating a preliminary image of the normal domain, and learning degradation parameters;

[0023] Continuing to perform a reverse reasoning operation based on the preliminary image and the degradation parameter to generate a normal domain image of a normal domain;

[0024] Calibration and optimization are performed on the normal domain image to generate diffusion features of the normal domain.

[0025] In an optional embodiment, the step of continuing to perform a reverse reasoning operation based on the preliminary image and the degradation parameter to generate a normal domain image of the normal domain includes:

[0026] Continuing to perform a reverse reasoning operation on the preliminary image based on the degradation parameter to gradually remove noise in the preliminary image, wherein the preliminary image has a text prompt;

[0027] Encoding the preliminary image based on the constructed image encoder, and encoding the text prompt based on the constructed text encoder;

[0028] The semantic similarity between the encoded preliminary image and the text prompt is calculated based on the constructed loss function, and the preliminary image is optimized by iterating the loss function to generate a normal domain image of the normal domain.

[0029] In an optional embodiment, the step of performing calibration and optimization processing on the normal domain image to generate the diffusion feature of the normal domain includes:

[0030] Decomposing the information of the normal domain image into low-frequency information and high-frequency information;

[0031] Calibrate the low-frequency information, and perform diffusion inference on the calibrated low-frequency information based on the learned degradation prior knowledge to refine the low-frequency information;

[0032] Using a feature gain model to optimize the high-frequency information to remove redundant features in the high-frequency information;

[0033] The processed low-frequency information and high-frequency information are merged to obtain the diffusion characteristics of the normal domain.

[0034] In an optional embodiment, the step of optimizing the high-frequency information using a feature gain model to remove redundant features in the high-frequency information includes:

[0035] Using the convolutional layer in the feature gain model to extract features from the high-frequency information to obtain shallow features;

[0036] Extracting refined features of the shallow features using a residual layer and an activation function;

[0037] Fusing the shallow features with the refined features;

[0038] The convolution layer and activation function are used to reconstruct the noise map of the fused features to obtain high-frequency information after removing redundant features.

[0039] In an optional embodiment, the prior features include instance-level prior features and point-level prior features;

[0040] The step of fusing the priori feature and the diffusion feature to obtain a fused feature includes:

[0041] Converting the diffusion features into instance-level diffusion features and point-level diffusion features;

[0042] Addition operations, replacement operations, and connection operations are performed on the instance-level prior features and instance-level diffusion features, and the point-level prior features and point-level diffusion features, respectively, to obtain fused features.

[0043] In an optional embodiment, the step of inputting the fused features into a constructed map decoder to generate an online vector map includes:

[0044] Inputting the fused features into a constructed map decoder to generate vectorized map elements;

[0045] An online vector map is generated based on the point set, direction and type contained in each of the map elements.

[0046] In a second aspect, the present invention provides a vector map construction system based on a priori maps and multi-stage diffusion reasoning, the system comprising:

[0047] An encoding module, configured to obtain a plurality of pre-stored prior maps, and encode each of the prior maps using a unified vector encoder to generate prior features;

[0048] an inference module, configured to obtain real-time data detected by an on-board sensor, convert the real-time data into bird's-eye view features, and perform multi-stage diffusion inference on the bird's-eye view features to map the bird's-eye view features in a degraded domain to diffusion features in a normal domain;

[0049] A fusion module, configured to fuse the prior features and the diffusion features to obtain fused features;

[0050] The generation module is used to input the fusion features into the constructed map decoder to generate an online vector map.

[0051] The vector map construction method and system based on prior maps and multi-stage diffusion reasoning provided in the embodiments of this application significantly enhance the robustness and accuracy of online map construction by introducing multiple prior maps and efficiently encoding different prior map data using a unified vector encoder. In terms of image processing, multi-stage diffusion reasoning technology is used to gradually optimize image details. This not only uniformly handles various image degradation issues, but also provides high-quality image restoration and enhancement in complex environments, providing strong support for autonomous driving perception systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0053] Figure 1 A flowchart of a vector map construction method based on a priori maps and multi-stage diffusion reasoning provided in an embodiment of the present application;

[0054] Figure 2 for Figure 1 Flowchart of the sub-steps included in S11;

[0055] Figure 3 for Figure 1 Flowchart of the sub-steps included in S12;

[0056] Figure 4 for Figure 1 Flowchart of the sub-steps included in S13;

[0057] Figure 5 for Figure 1 Flowchart of the sub-steps included in S14;

[0058] Figure 6 A functional module block diagram of a vector map construction system based on a priori maps and multi-stage diffusion reasoning provided in an embodiment of the present application;

[0059] Figure 7 This is a structural block diagram of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0060] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0061] To facilitate understanding of the embodiments of the present application, further explanation will be given below with reference to specific embodiments in conjunction with the accompanying drawings. The embodiments do not constitute a limitation on the embodiments of the present application.

[0062] See also Figure 1The following is a flowchart of a vector map construction method based on a priori maps and multi-stage diffusion reasoning, according to an embodiment of the present invention. This vector map construction method can be executed by a vector map construction system. This vector map construction system can be implemented using software and / or hardware and can be configured in an electronic device, such as a computer or server, for example, a server in a backend control platform. The detailed steps of this vector map construction method are described below.

[0063] S11, obtaining a plurality of pre-stored priori maps, and encoding each of the priori maps using a unified vector encoder to generate priori features.

[0064] S12, obtaining real-time data detected by the vehicle-mounted sensor, converting the real-time data into bird's-eye view features, and performing multi-stage diffusion inference on the bird's-eye view features to map the bird's-eye view features in the degraded domain to diffusion features in the normal domain.

[0065] S13, fusing the priori feature and the diffusion feature to obtain a fused feature.

[0066] S14: Input the fused features into the constructed map decoder to generate an online vector map.

[0067] The vector map construction method provided in this embodiment is applied in an autonomous driving scenario and can be used to generate an online vector map for autonomous driving control guidance.

[0068] The various pre-existing prior maps can be obtained from an open prior knowledge base. These prior maps can include standard definition maps, existing but outdated high-precision maps, and historical prediction maps. Standard definition maps provide road skeleton information, while high-precision maps are highly accurate but less frequently updated. Historical prediction maps reflect the latest observation data from multiple vehicle passes through the same location.

[0069] In this embodiment, by combining the three types of prior maps mentioned above and complementing each other, key road information such as road skeletons, lane lines, intersections, etc. can be provided more completely.

[0070] A unified vector encoder is used to encode each prior map. This involves performing multi-feature representation on the vector data (such as lane lines and intersections) in the prior map, processing the internal information of the vector data, and processing the associations between the vector data. Ultimately, prior features are generated based on the prior map. These prior features represent the static, existing map information.

[0071] Furthermore, the vehicle is equipped with multiple onboard sensors, including cameras and lidar. During operation, these sensors generate real-time data, including multi-view image data and lidar point cloud data. To reflect the global state of the vehicle's environment, this real-time data can be converted into bird's-eye-view features.

[0072] Real-time data detected by on-board sensors may be captured in scenarios such as camera obstructions or inclement weather. This type of real-time data is of poor quality and is referred to as degraded domain data. This degraded domain data needs to be converted to the normal domain. Therefore, the terms "normal domain" and "degraded domain" are relative terms: data represented in the degraded domain is of poor quality, while data represented in the normal domain is of higher quality.

[0073] In this embodiment, multi-stage diffusion inference is performed on the bird's-eye view features of the degraded domain, and image details are gradually optimized through processing such as denoising, deblurring, and low-light enhancement, ultimately obtaining the diffusion features of the normal domain.

[0074] On this basis, the diffusion features are fused with the prior features obtained based on the prior map to obtain the fused features, which are finally decoded by the map decoder to generate an online vector map.

[0075] This solution significantly enhances the robustness and stability of online map construction by introducing multiple prior maps. This approach, particularly in complex environments (such as occlusion and inclement weather), effectively compensates for the shortcomings of single-source perceptual data and improves map accuracy and completeness. Furthermore, by uniformly encoding each prior map using a unified vector encoder, it efficiently processes diverse prior map data, improving the model's generalization and adapting it to a variety of map types and environmental variations.

[0076] In addition, through the multi-stage diffusion inference process, image details can be gradually optimized to ensure high-quality mapping from the degraded domain to the normal domain. It can uniformly handle multiple image degradation problems, including denoising, deblurring, low-light enhancement, etc., to improve the effects of image restoration and enhancement.

[0077] This solution supports real-time map updates and rapid construction, can quickly adapt to changes in road structure, and significantly improve the stability of the autonomous driving perception system in complex environments, especially in conditions such as severe weather and occlusion, ensuring that the vehicle can accurately perceive the surrounding environment and achieve safe navigation.

[0078] The specific implementation of each of the above steps is described below. In this embodiment, each prior map includes multiple data vectors, each data vector is composed of multiple points, where the data vectors are lane lines, intersections, etc. Figure 2, the above steps of encoding each prior map using a unified vector encoder to generate prior features can be implemented as follows:

[0079] S111 , processing the data vectors and points in the data vectors in each of the prior maps to obtain various types of feature information.

[0080] S112, inputting multiple types of feature information into the unified vector encoder, and utilizing the intra-vector attention encoder in the unified vector encoder to process the interactions between points within the data vector.

[0081] S113, using the inter-vector attention encoder in the unified vector encoder to process the interaction between different data vectors to generate prior features.

[0082] In this embodiment, when processing each prior map, the vector data in the prior map can be processed using a hybrid prior embedding and dual encoding mechanism, encoding it into a fixed-length feature representation. Specifically, the prior map is first processed using hybrid prior embedding. In this embodiment, the data vectors in the prior map and the points in the data vectors are processed. The various types of feature information obtained include local features, instance-level embedding features, map element features, and location information features. Specifically, they can be obtained through the following methods:

[0083] The position and direction information of each point in each data vector in the prior map are encoded to obtain the local features of each point; instance-level tags are set for each data vector, and the instance-level tags are encoded using a multi-layer perceptron to generate instance-level embedded features; corresponding type codes are assigned based on the element type of each data vector, and the type codes are encoded using an embedding matrix to generate map element features; a position index is assigned to each point, and the position encoding function is used to encode the position index to generate position information features.

[0084] In this embodiment, each point in each data vector is embedded at the point level, and the position information (x, y) and direction information (v) of the point are encoded by the multi-layer perceptron. x ,v y ), to capture the local features E of each point point (p i ), can be characterized as follows:

[0085]

[0086] Among them, p i Represents point p i , x i and y i Represents point p i location information, and Represents point p i , MLP represents the processing of multi-layer perceptron, and Concat represents the concatenation function.

[0087] Since each data vector is composed of multiple ordered points, the points in the same data vector are included in the same vector instance, and an instance-level tag [VEC] is set for each vector instance. The tag can be a number, for example. The instance-level tag can uniformly represent the characteristics of the entire vector instance, capture the relationship between points within the instance vector, and use a multi-layer perceptron to encode the instance-level tag to generate the instance-level embedded feature E. instance , characterized as follows:

[0088] E instance =MLP([VEC])

[0089] In addition, this embodiment introduces type embedding, that is, the data vector has different element types, that is, intersections, road lines, etc. are different element types. A corresponding type code t is set for each data vector with different element types. The form of the type code is not limited, as long as it can distinguish different element types. For example, it can be a label or a value. Then, the type code is encoded using a learnable embedding matrix Embedding to generate a feature E that can distinguish different types of map elements. type :

[0090] Et ype =Embedding(t)

[0091] In addition, in this embodiment, a position index pos is assigned to each point based on the position information of each point. i , the position index can be indexed in a certain order according to the position of the point, so as to maintain the order information of the point. Then use the sine and cosine position encoding function PositionalEncoding to encode the position index and obtain the position information feature E pos (p i ):

[0092] E pos (p i )=PositionalEncoding(pos i )

[0093] On the basis of the above, multiple types of feature information are spliced together to form a hybrid prior embedding E mixed :

[0094] E mixed =Concat(E point ,E instance ,Etype ,E pos )

[0095] The mixed prior embedding is used as the input of the unified vector encoder, providing the encoder with multi-level feature representation.

[0096] In the unified vector encoder, interaction processing is performed between points and between data vectors.

[0097] Specifically, the unified vector encoder includes M layers of intra-vector attention encoders, which use the intra-vector attention encoders to process the interactions between points within the data vector. mixed , first split it into the embedding representation E of each data vector vector , for each data vector, the interaction between points is calculated using the following formula:

[0098]

[0099] Among them, Q1, K1, V1 are respectively obtained by linear transformation from E vector The resulting query, key, and value matrix, d k The multi-layer self-attention mechanism enhances the ability of each point to perceive other points in the same data vector and captures the local features within the data vector.

[0100] In addition, the unified vector encoder also includes an N-layer inter-vector attention encoder, which uses the inter-vector attention encoder to process the interaction between different data vectors and capture the global context information between data vectors. First, the embedding representation E of all data vectors is vector Then splice them together to form the global context representation E global , and then calculate the interaction between different data vectors according to the following formula:

[0101]

[0102] Among them, Q2, K2, V2 are respectively obtained by linear transformation from E global The resulting query, key, and value matrices capture the global contextual information between data vectors through a multi-layer self-attention mechanism.

[0103] During the dual encoding process, an attention mask mechanism is used to control the attention interactions within and between data vectors, ensuring that points within a data vector only interact with other points within the same data vector, while interactions between data vectors span different vector instances.

[0104] After the above two stages, the unified vector encoder can output the prior feature f prior .

[0105] On this basis, after obtaining the real-time data detected by the vehicle-mounted sensor, the real-time data is first converted into bird's-eye view features. For example, the features of the real-time data can be extracted by a perspective view feature extractor, and the extracted features are converted into bird's-eye view features.

[0106] On this basis, through multi-stage diffusion reasoning, preliminary normal domain map features are gradually generated from the degraded domain, and then further optimized to generate refined normal domain map features.

[0107] See also Figure 3 , the step of performing multi-stage diffusion inference on the bird's-eye view features to map the bird's-eye view features of the degraded domain to the diffusion features of the normal domain can be achieved by:

[0108] S121 , performing forward diffusion processing on the bird's-eye view feature in the degraded domain to add noise data and convert the image into a noise image.

[0109] S122 , performing a reverse reasoning operation on the noise image to restore the original data, generating a preliminary image of the normal domain, and learning degradation parameters.

[0110] S123 , continuing to perform a reverse reasoning operation based on the preliminary image and the degradation parameter to generate a normal domain image of the normal domain.

[0111] S124 , performing calibration and optimization processing on the normal domain image to generate diffusion features of the normal domain.

[0112] First, in the first stage of multi-stage diffusion inference, the bird's-eye view features of the degraded domain are gradually added with noise through the forward diffusion process. The forward diffusion process can be expressed as follows:

[0113]

[0114] Among them, x0 represents the bird's-eye view feature of the degradation domain of the original input, x t is the corrupted noise data at time step t, β t is the predefined noise variance, ∈ represents Gaussian noise, which obeys the standard normal distribution N(0,I), α t =1-β t , Through the forward diffusion process, the bird's-eye view feature LQ of the degradation domain is gradually transformed into the noise image x T .

[0115] Then, through the reverse reasoning process, the original data is restored from the Gaussian noise to generate a preliminary image of the rough normal domain. The reverse reasoning process can be expressed as follows:

[0116]

[0117] Among them, ∈ θ μ is the noise predictor of the diffusion model, and the noise is gradually removed by optimizing the noise predictor. θ is the mean of the Gaussian distribution in the backward inference, which is given by the noise predictor ∈ θ Calculated, Represents the variance of the Gaussian distribution in reverse inference, usually a fixed value or the same as β t related. Represents the predicted data in the reverse reasoning process, which is at time step t, the model passes the noise predictor ∈ θ From the noisy data x t Estimated denoised data.

[0118] In this stage, the model learns the degradation parameters and stores some degradation prior knowledge.

[0119] On this basis, in the second stage of the multi-stage diffusion inference, processing is performed based on the information obtained in the first stage. Specifically, the reverse inference operation is continued based on the preliminary image and the degradation parameter to generate the normal domain image of the normal domain. This can be achieved by:

[0120] A reverse inference operation is continued to be performed on the preliminary image based on the degradation parameter to gradually remove noise in the preliminary image, wherein the preliminary image has a text prompt; the preliminary image is encoded based on the constructed image encoder, and the text prompt is encoded based on the constructed text encoder; the semantic similarity between the encoded preliminary image and the text prompt is calculated based on the constructed loss function, and the preliminary image is optimized by iterating the loss function to generate a normal domain image of the normal domain.

[0121] In the second stage of multi-stage diffusion inference, the preliminary image generated in the first stage is input Using the degradation parameters learned in the first stage, we continue to perform reverse reasoning through the diffusion model to generate more refined normal domain images. The reverse reasoning process can be expressed as follows:

[0122]

[0123] Where y is the input condition (i.e. ),∈ θ (x t ,t,y) is the noise predictor, which is used to predict the noise and remove it step by step.

[0124] At this stage, we also need to consider image-related textual cues. Textual cues are text used to represent image representation information. These cues can include positive and negative textual cues. Positive textual cues are textual cues where the image representation information is consistent with the textual cues’ representation information, while negative textual cues are textual cues where the image representation information is inconsistent with the textual cues’ representation information.

[0125] In this stage, the contrastive language-image pre-training (CLIP) model is used to encode the text prompts and calculate the semantic similarity between the preliminary image and the text prompts to ensure that the generated image results are closer to normal images. Among them, the loss function constructed is used to measure the semantic similarity between the preliminary image and the text prompt. The loss function L clip as follows:

[0126]

[0127] Among them, T p is a positive text prompt, T n is a negative text prompt, φ text and φ image They are text encoder and image encoder respectively.

[0128] A loss function is used to measure the semantic similarity between the generated preliminary image and the text prompt. The generation of the preliminary image is guided by the loss function. When certain conditions are met, such as the convergence of the loss function, a normal domain image that meets the requirements can be obtained.

[0129] In the third stage of multi-stage diffusion inference, the normal domain image generated in the second stage is calibrated and optimized to generate the diffusion features of the normal domain. Specifically, this can be achieved by:

[0130] The information of the normal domain image is decomposed into low-frequency information and high-frequency information; the low-frequency information is calibrated, and diffusion reasoning is performed on the calibrated low-frequency information based on the learned degradation prior knowledge to refine the low-frequency information; the high-frequency information is optimized using a feature gain model to remove redundant features in the high-frequency information; the processed low-frequency information and high-frequency information are merged to obtain the diffusion features of the normal domain.

[0131] In this embodiment, the normal domain image obtained Perform discrete wavelet transform processing to decompose it into low-frequency information L (containing the main structural information of the image content) and high-frequency information H (containing high-frequency sub-bands in the vertical, horizontal and diagonal directions), which can be represented as follows:

[0132]

[0133] Finely calibrate the low-frequency information in the wavelet low-frequency domain, remove redundant information, and generate calibrated low-frequency information In this embodiment, the degradation prior knowledge learned in the previous stages can be used to perform 10 steps of diffusion reasoning to further refine the low-frequency information.

[0134] For high-frequency information, redundant features are further removed through the feature gain model to achieve the purpose of optimizing high-frequency information. Specifically, this step can be achieved by:

[0135] The convolution layer in the feature gain model is used to extract features of the high-frequency information to obtain shallow features; the residual layer and activation function are used to extract refined features of the shallow features; the shallow features and the refined features are fused; the convolution layer and activation function are used to reconstruct the noise map of the fused features to obtain high-frequency information after removing redundant features.

[0136] In this embodiment, the feature gain model includes a convolution layer, a residual layer, an activation function, etc. First, the features of the high-frequency information are extracted through the convolution layer as shallow features. Multiple layers of residual layers (for example, four layers) and activation functions (such as ReLU activation function) are used to extract further features of the shallow features as refined features. Then, the shallow features and the refined features are fused through the residual operation of the residual layer to enhance the memory ability of the features. Finally, the noise map is reconstructed through the convolution layer and the activation function to remove redundant features such as noise to generate clean high-frequency information.

[0137] On the basis of the above, the calibrated low-frequency information and optimized high-frequency information Merge to generate high-quality diffusion features HQ of the normal domain:

[0138]

[0139] After obtaining the prior features and diffusion features through the above processing, please refer to Figure 4 , the step of fusing the prior features and the diffusion features to obtain the fusion features can be achieved by the following methods:

[0140] S131, converting the diffusion features into instance-level diffusion features and point-level diffusion features.

[0141] S132 , performing addition operations, replacement operations, and concatenation operations on the instance-level prior features and instance-level diffusion features, and the point-level prior features and point-level diffusion features, respectively, to obtain fused features.

[0142] As can be seen from the above, prior features include instance-level prior features and point-level prior features. To align the diffusion features with the prior features, the diffusion features must first be converted into instance-level diffusion features and point-level diffusion features. The conversion method is similar to the conversion method for the prior features described above and will not be repeated here. In this way, instance-level or point-level features can perform more fine-grained interactive processing when faced with complex processing scenarios.

[0143] On this basis, for the prior feature f prior (including instance-level prior features f ins and point-level prior features f pt ) and diffusion features Q (including instance-level diffusion features q ins and point-level diffusion feature q pt ) Through addition, replacement and connection operations, the fusion feature is obtained. The specific process can be characterized as follows:

[0144]

[0145] Addition operation:

[0146] Replace operation:

[0147] Connection operation:

[0148] in, It is the fusion feature after fusion.

[0149] Experimental results show that the concatenation operation is the most effective integration method and can significantly improve the performance of the model.

[0150] After obtaining the fusion features through the above method, please refer to Figure 5 , the fused features are input into the constructed map decoder to generate an online vector map. This step can be achieved by:

[0151] S141: Input the fused features into the constructed map decoder to generate vectorized map elements.

[0152] S142: Generate an online vector map based on the point set, direction, and type included in each of the map elements.

[0153] In this embodiment, the map decoder can be a Transformer decoder, for example. By inputting the fused features into the map decoder, vectorized map elements, such as lane lines, intersections, and lane segments, can be directly generated. Each map element is composed of a corresponding point set, and each map element has direction information and its type information. Combining the point set, direction, and type of each map element ultimately generates an online vector map.

[0154] In summary, the vector map construction solution based on prior maps and multi-stage diffusion reasoning provided in this embodiment significantly enhances the robustness and accuracy of online map construction by introducing multiple prior maps and using a unified vector encoder to efficiently encode different prior map data. In terms of image processing, through the multi-stage diffusion reasoning process, the image details are gradually optimized, which can not only uniformly handle multiple image degradation problems (such as denoising, deblurring, low-light enhancement, etc.), but also provide high-quality image restoration and enhancement effects in complex environments, providing strong support for autonomous driving perception systems. It can help autonomous driving vehicles achieve safer and more reliable navigation and perception in complex environments, especially under conditions such as bad weather and camera occlusion, significantly improving the robustness and accuracy of the system. It can help autonomous driving vehicles achieve safer and more reliable navigation and perception in complex environments, and provide a good theoretical basis for navigation in bad weather.

[0155] Based on the same inventive concept, please refer to Figure 6 , an embodiment of the present invention also provides a functional module diagram of a vector map construction system based on a priori maps and multi-stage diffusion reasoning. This embodiment can divide the functional modules of the vector map construction system according to the above method embodiment. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present invention is schematic and is only a logical functional division. There may be other division methods in actual implementation.

[0156] For example, when each functional module is divided into corresponding functional modules, Figure 6 The vector map construction system shown is only a schematic diagram of a device. The vector map construction system may include an encoding module, an inference module, a fusion module, and a generation module. The functions of each functional module of the vector map construction system are described in detail below.

[0157] An encoding module, configured to obtain a plurality of pre-stored prior maps, and encode each of the prior maps using a unified vector encoder to generate prior features;

[0158] an inference module, configured to obtain real-time data detected by an on-board sensor, convert the real-time data into bird's-eye view features, and perform multi-stage diffusion inference on the bird's-eye view features to map the bird's-eye view features in a degraded domain to diffusion features in a normal domain;

[0159] A fusion module, configured to fuse the prior features and the diffusion features to obtain fused features;

[0160] The generation module is used to input the fusion features into the constructed map decoder to generate an online vector map.

[0161] The vector map construction system provided in this embodiment can be used to execute the vector map construction method under any implementation method of the above embodiments. For details not provided in this embodiment, please refer to the corresponding description of the above embodiments, and this embodiment will not be repeated here.

[0162] See also Figure 7 , is a block diagram of the structure of an electronic device provided in an embodiment of the present invention. This electronic device may be a computer device, server, or the like in an autonomous driving control platform. The electronic device includes a memory, a processor, and a communication module. The memory, processor, and communication module components are electrically connected to each other, directly or indirectly, to enable data transmission or interaction. For example, these components may be electrically connected to each other via one or more communication buses or signal lines.

[0163] Memory is used to store computer programs or data. Memory can be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM).

[0164] The processor is used to read / write data or programs stored in the memory and execute the vector map construction method based on prior map and multi-stage diffusion reasoning provided by any embodiment of the present invention.

[0165] The communication module is used to establish a communication connection between the electronic device and other communication terminals through the network, and is used to send and receive data through the network.

[0166] It should be understood that Figure 7 The structure shown is only a schematic diagram of the structure of the electronic device. The electronic device may also include Figure 7 More or fewer components than shown, or with Figure 7 Different configurations shown.

[0167] Furthermore, an embodiment of the present invention also provides a computer-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are executed, the vector map construction method based on prior maps and multi-stage diffusion reasoning provided in the above embodiment is implemented.

[0168] Specifically, the computer-readable storage medium can be a general-purpose storage medium, such as a removable disk or hard disk. When the computer program on the computer-readable storage medium is executed, it can execute the aforementioned vector map construction method based on a priori maps and multi-stage diffusion reasoning. Regarding the processes involved in executing the computer-readable storage medium and its executable instructions, please refer to the relevant description in the aforementioned method embodiment and will not be detailed here.

[0169] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0170] In addition, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0171] Furthermore, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0172] It should be noted that if the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0173] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any actual relationship or order between these entities or operations.

[0174] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A vector map construction method based on prior maps and multi-stage diffusion reasoning, characterized in that: The method comprises: Obtaining a plurality of pre-stored prior maps, encoding each of the prior maps using a unified vector encoder to generate prior features; obtaining real-time data detected by an on-board sensor, converting the real-time data into bird's-eye view features, and performing multi-stage diffusion inference on the bird's-eye view features to map the bird's-eye view features in a degraded domain to diffusion features in a normal domain; Fusing the priori feature and the diffusion feature to obtain a fused feature; The fused features are input into the constructed map decoder to generate an online vector map.

2. The vector map construction method based on prior map and multi-stage diffusion reasoning according to claim 1 is characterized in that: The prior map includes a plurality of data vectors, each of the data vectors being composed of a plurality of points; The step of encoding each of the prior maps using a unified vector encoder to generate prior features includes: Processing the data vectors and points in the data vectors in each of the prior maps to obtain various types of feature information; Inputting multiple types of feature information into the unified vector encoder, and using the intra-vector attention encoder in the unified vector encoder to process the interaction between points within the data vector; The inter-vector attention encoder in the unified vector encoder is used to process the interactions between different data vectors to generate prior features.

3. The vector map construction method based on prior map and multi-stage diffusion reasoning according to claim 2 is characterized in that: The step of processing the data vectors and points in the data vectors in each of the prior maps to obtain multiple types of feature information includes: Encoding the position information and direction information of each point of each data vector in the prior map to obtain local features of each point; Setting an instance-level tag for each of the data vectors, and encoding the instance-level tag using a multi-layer perceptron to generate instance-level embedded features; assigning a corresponding type code based on the element type of each data vector, and encoding the type code using an embedding matrix to generate a map element feature; A position index is assigned to each of the points, and the position index is encoded using a position encoding function to generate a position information feature.

4. The vector map construction method based on prior map and multi-stage diffusion reasoning according to claim 1 is characterized in that: The step of performing multi-stage diffusion inference on the bird's-eye view feature to map the bird's-eye view feature in the degraded domain to the diffusion feature in the normal domain includes: Performing a forward diffusion process on the bird's-eye view feature in the degraded domain to add noise data and convert the image into a noise image; Performing a reverse reasoning operation on the noisy image to restore the original data, generating a preliminary image of the normal domain, and learning degradation parameters; Continuing to perform a reverse reasoning operation based on the preliminary image and the degradation parameter to generate a normal domain image of a normal domain; Calibration and optimization are performed on the normal domain image to generate diffusion features of the normal domain.

5. The vector map construction method based on prior map and multi-stage diffusion reasoning according to claim 4 is characterized in that: The step of continuing to perform a reverse reasoning operation based on the preliminary image and the degradation parameter to generate a normal domain image of the normal domain includes: Continuing to perform a reverse reasoning operation on the preliminary image based on the degradation parameter to gradually remove noise in the preliminary image, wherein the preliminary image has a text prompt; Encoding the preliminary image based on the constructed image encoder, and encoding the text prompt based on the constructed text encoder; The semantic similarity between the encoded preliminary image and the text prompt is calculated based on the constructed loss function, and the preliminary image is optimized by iterating the loss function to generate a normal domain image of the normal domain.

6. The vector map construction method based on prior map and multi-stage diffusion reasoning according to claim 4 is characterized in that: The step of calibrating and optimizing the normal domain image to generate the diffusion feature of the normal domain includes: Decomposing the information of the normal domain image into low-frequency information and high-frequency information; Calibrate the low-frequency information, and perform diffusion inference on the calibrated low-frequency information based on the learned degradation prior knowledge to refine the low-frequency information; Using a feature gain model to optimize the high-frequency information to remove redundant features in the high-frequency information; The processed low-frequency information and high-frequency information are merged to obtain the diffusion characteristics of the normal domain.

7. The vector map construction method based on prior map and multi-stage diffusion reasoning according to claim 6 is characterized in that: The step of optimizing the high-frequency information using a feature gain model to remove redundant features in the high-frequency information includes: Using the convolutional layer in the feature gain model to extract features from the high-frequency information to obtain shallow features; Extracting refined features of the shallow features using a residual layer and an activation function; Fusing the shallow features with the refined features; The convolution layer and activation function are used to reconstruct the noise map of the fused features to obtain high-frequency information after removing redundant features.

8. The vector map construction method based on prior map and multi-stage diffusion reasoning according to claim 1 is characterized in that: The prior features include instance-level prior features and point-level prior features; The step of fusing the priori feature and the diffusion feature to obtain a fused feature includes: Converting the diffusion features into instance-level diffusion features and point-level diffusion features; Addition operations, replacement operations, and connection operations are performed on the instance-level prior features and instance-level diffusion features, and the point-level prior features and point-level diffusion features, respectively, to obtain fused features.

9. The vector map construction method based on prior map and multi-stage diffusion reasoning according to claim 1 is characterized in that: The step of inputting the fusion features into the constructed map decoder to generate an online vector map includes: Inputting the fused features into a constructed map decoder to generate vectorized map elements; An online vector map is generated based on the point set, direction and type contained in each of the map elements.

10. A vector map construction system based on prior maps and multi-stage diffusion reasoning, characterized in that: The system comprises: An encoding module, configured to obtain a plurality of pre-stored prior maps, and encode each of the prior maps using a unified vector encoder to generate prior features; an inference module, configured to obtain real-time data detected by an on-board sensor, convert the real-time data into bird's-eye view features, and perform multi-stage diffusion inference on the bird's-eye view features to map the bird's-eye view features in a degraded domain to diffusion features in a normal domain; A fusion module, configured to fuse the prior features and the diffusion features to obtain fused features; The generation module is used to input the fusion features into the constructed map decoder to generate an online vector map.

Citation Information

Patent Citations

  • Online vectorization high-precision map construction method based on multi-modal instance fusion

    CN118864651A

  • Vector map generation method and device, equipment and storage medium

    CN119437198A

  • Map generation method and apparatus, electronic device, and storage medium

    WO2023123837A1