Low-light road perception and topology reasoning method and system based on trajectory prior

By enhancing the low-light image and encoding and decoding methods combining trajectory prior data, the problem of insufficient image quality under low-light conditions is solved, the road perception and topological inference accuracy of autonomous driving is improved, and the cost of obtaining prior information is reduced.

CN120430971AActive Publication Date: 2025-08-05BEIHANG UNIV

Patent Information

Application Number
CN202510527056.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-05
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

Under low light, night or special lighting conditions, the images collected by the on-board cameras are insufficient brightness and noise, making it difficult to accurately extract key road features such as lane lines, affecting the positioning, path planning and decision-making of autonomous driving. At the same time, the construction and update of high-precision maps are expensive.

Method used

The image pre-training model is used to enhance the low-light images, and rasterize and vectorize the trajectory prior data for rasterization and vectorize encoding, generate rasterized heat maps and vectorized trajectory information, and decode them through the cross attention mechanism to generate road perception and topological inference results.

Benefits of technology

It improves image quality and map construction accuracy under low-light conditions, reduces the cost of a priori information acquisition, and enhances the road perception and topological inference accuracy of autonomous driving under complex lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430971A_ABST
    Figure CN120430971A_ABST
Patent Text Reader

Abstract

The invention provides a low-light road perception and topological reasoning method and system based on track prior, and the method carries out the enhancement processing of a low-light image through an image pre-training model, thereby restoring the details of the image, improving the brightness of the image, and reducing the distortion phenomenon of the image in the enhancement process. Moreover, crowdsourcing trajectory data is effectively utilized as prior information, and the rasterized thermodynamic diagram and vectorized trajectory information are constructed in combination with the trajectory prior data, so that the accuracy and robustness of the map construction process are enhanced, and the difficulty and cost of obtaining the prior information are reduced. According to the scheme, the capability of dealing with an automatic driving scene under the conditions of low illumination and complex illumination can be effectively improved, and the road perception and topological reasoning precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to a low-light road perception and topology reasoning method and system based on trajectory prior. Background Art

[0002] In recent years, with the rapid development of artificial intelligence and automation, autonomous driving technology has become a key focus of national scientific and technological development. This rapid development has also driven the demand for road perception and mapping technologies. However, traditional mapping methods typically rely on offline mapping using lidar point clouds and camera-collected image data. This inherently suffers from high costs, delayed updates, and an inability to dynamically adapt to environmental changes, making it insufficient for autonomous driving. Online mapping methods, onboard sensors, achieve real-time environmental perception, dynamically reflecting environmental changes and enabling real-time map adjustments.

[0003] However, in low light, at night or under special lighting conditions, the images collected by on-board cameras often have problems such as insufficient brightness, high noise, and loss of details, making it difficult to accurately extract key road features such as lane lines, thereby affecting vehicle positioning, path planning and decision-making.

[0004] At the same time, current online mapping technologies often use high-precision maps as prior information, but these maps are expensive to build and update. With the widespread adoption of crowdsourcing, a large amount of historical trajectory data has been collected, which can reflect actual vehicle trajectories and road structure information. One area of current technological development is how to use this trajectory data as prior information and integrate it with real-time perception data to improve the accuracy of road perception and map construction while reducing the cost of acquiring prior information. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a low-light road perception and topology reasoning method and system based on trajectory priors to enhance the quality of low-light images and combine trajectory prior data to improve the accuracy and robustness of map construction.

[0006] In a first aspect, the present invention provides a low-light road perception and topology reasoning method based on trajectory prior, the method comprising:

[0007] Obtaining a surround view image captured by a vehicle-mounted camera, wherein the surround view image is a low-light image, and enhancing the surround view image using an image pre-training model to obtain an enhanced image;

[0008] Obtaining trajectory priori data, and performing rasterization encoding and vectorization encoding on the trajectory priori data to generate a rasterized heat map and vectorized trajectory information;

[0009] Encoding the enhanced image to obtain a bird's-eye view feature, and fusing and aligning the bird's-eye view feature with the rasterized heat map to obtain an aligned fused feature;

[0010] The vectorized trajectory information and the fused features are input into a decoder, decoded using a cross-attention mechanism, and the decoded results are transformed through a linear layer to generate road perception and topological reasoning results, which include the geometric positions and topological relationships of lane segments.

[0011] In an optional embodiment, the step of enhancing the surround view image using the image pre-training model to obtain an enhanced image includes:

[0012] Inputting the surround view image into an image pre-training model to obtain an average brightness of the surround view image;

[0013] Obtaining an edge map based on the surround view image according to an average brightness of the surround view image and a set first brightness threshold;

[0014] Obtaining an initialized potential representation based on the surround view image according to an average brightness of the surround view image and a set second brightness threshold;

[0015] An enhanced image is generated according to the edge map and the initialized latent representation.

[0016] In an optional embodiment, the step of obtaining an initialized potential representation based on the surround view image according to the average brightness of the surround view image and a set second brightness threshold includes:

[0017] comparing the average brightness of the surround view image with a set second brightness threshold, and obtaining a brightness-enhanced image based on the surround view image according to the comparison result;

[0018] Converting the brightness-enhanced image into a latent representation through a denoising diffusion implicit model and generating random noise;

[0019] The latent representation and the random noise are fused to generate an initialized latent representation.

[0020] In an optional embodiment, the step of generating an enhanced image based on the edge map and the initialized latent representation includes:

[0021] Denoising the initialized latent representation and extracting self-attention features at each time step;

[0022] generating prediction noise based on the self-attention features, the edge map, and the initialized latent representation;

[0023] The initialized latent representation and the predicted noise are used to iteratively optimize the initialized latent representation according to time steps to obtain a final latent representation, and the final latent representation is decoded to generate an enhanced image.

[0024] In an optional embodiment, the trajectory priori data includes multiple trajectory data, each trajectory data consists of multiple two-dimensional points;

[0025] The rasterized heat map is generated in the following way:

[0026] For each piece of trajectory data in the trajectory priori data, obtaining the direction angle of each two adjacent two-dimensional points in the trajectory data;

[0027] Create a grid containing multiple grid cells and calculate the number and average direction angle of trajectory data passing through each grid cell;

[0028] Normalizing the number of trajectory data items of all grid cells, and transforming the average direction angle to a set interval using an inverse tangent function to generate a rasterized heat map containing density and direction information;

[0029] The vectorized trajectory information is generated in the following manner:

[0030] Determine a representative trajectory data sample from the plurality of trajectory data using a clustering algorithm or a farthest point sampling method;

[0031] Vectorization processing is performed on the representative trajectory data samples to obtain vectorized trajectory information.

[0032] In an optional embodiment, the vehicle-mounted camera includes a plurality of;

[0033] The step of encoding the enhanced image to obtain a bird's-eye view feature comprises:

[0034] Performing feature extraction on the enhanced images of each of the vehicle-mounted cameras to obtain a corresponding feature map;

[0035] Inputting the obtained multiple feature maps into an encoder comprising multiple encoding layers, performing bird's-eye view feature extraction for each of the feature maps for multiple time steps, and in each encoding layer, querying the bird's-eye view features obtained at the previous time step through a temporal self-attention mechanism based on set query information to obtain the temporal information of the current time step, and querying the spatial information of the current time step from the multiple feature maps through a spatial cross-attention mechanism;

[0036] Based on the temporal and spatial information output by the last encoding layer, bird's-eye view features are generated.

[0037] In an optional embodiment, the step of fusing and aligning the bird's-eye view feature with the rasterized heat map to obtain an aligned fused feature includes:

[0038] Obtaining a trajectory priori features based on the rasterized heat map, and splicing the trajectory priori features with the bird's-eye view features to obtain spliced features;

[0039] Processing the splicing features using multiple convolutional layers to predict the coordinate offset of each pixel in the trajectory prior features;

[0040] Mapping the position information of each pixel point to a new position according to the coordinate offset of the pixel point;

[0041] Upsampling the trajectory prior features to align with the bird's-eye view features;

[0042] The aligned trajectory prior features and bird's-eye view features are weightedly fused according to the learned weights to obtain the fused features.

[0043] In an optional embodiment, the step of inputting the vectorized trajectory information and the fused features into a decoder and performing decoding processing through a cross-attention mechanism includes:

[0044] Converting the vectorized trajectory information into high-dimensional vectorized trajectory information through a multi-layer perceptron, and performing a linear transformation on the high-dimensional vectorized trajectory information to generate an initial query;

[0045] Combining the initial query with a preset learnable query vector to generate a query embedding;

[0046] The query embedding and the fused features are input into a decoder, and a decoding result is obtained by combining the query embedding and the fused features using a cross attention mechanism.

[0047] In an optional embodiment, the step of converting the decoding result obtained after decoding through a linear layer to generate a road perception and topology reasoning result includes:

[0048] The decoded results are converted into coordinate information through linear transformation to obtain the geometric position of the lane segment;

[0049] The connection relationship between lane segments is calculated through graph neural network to obtain the topological adjacency matrix of the lane segments.

[0050] In a second aspect, the present invention provides a low-light road perception and topology reasoning system based on trajectory prior, the system comprising:

[0051] an enhancement processing module, configured to obtain a surround view image captured by an on-board camera, the surround view image being a low-light image, and perform enhancement processing on the surround view image using an image pre-training model to obtain an enhanced image;

[0052] An encoding module is used to obtain trajectory prior data, perform raster encoding and vector encoding on the trajectory prior data, and generate a rasterized heat map and vectorized trajectory information;

[0053] a fusion module, configured to encode the enhanced image to obtain a bird's-eye view feature, and fuse and align the bird's-eye view feature with the rasterized heat map to obtain an aligned fusion feature;

[0054] A decoding module is configured to input the vectorized trajectory information and the fused features into a decoder, perform decoding processing through a cross-attention mechanism, and transform the decoding results obtained after decoding through a linear layer to generate road perception and topological reasoning results, wherein the road perception and topological reasoning results include the geometric positions and topological relationships of lane segments.

[0055] The low-light road perception and topology reasoning method and system based on trajectory priors provided by the embodiments of the present invention enhance low-light images through an image pre-training model, restoring image details and improving image brightness to a certain extent, while reducing image distortion during the enhancement process. Furthermore, the method effectively utilizes crowdsourced trajectory data as prior information, combines it with trajectory prior data to construct rasterized heat maps and vectorized trajectory information, enhancing the accuracy and robustness of the map-building process and reducing the difficulty and cost of obtaining prior information. This solution can effectively improve the ability to cope with autonomous driving scenarios in low-light and complex lighting conditions, and enhance the accuracy of road perception and topology reasoning. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0057] Figure 1 A flowchart of a low-light road perception and topology reasoning method based on trajectory prior provided by an embodiment of the present invention;

[0058] Figure 2 for Figure 1 Flowchart of the sub-steps included in S11;

[0059] Figure 3 for Figure 1Flowchart of the sub-steps included in S12;

[0060] Figure 4 for Figure 1 Flowchart of the sub-steps included in S13;

[0061] Figure 5 for Figure 1 Another flow chart of the sub-steps included in S13;

[0062] Figure 6 for Figure 1 Flowchart of the sub-steps included in S14;

[0063] Figure 7 for Figure 1 Another flow chart of the sub-steps included in S14;

[0064] Figure 8 A functional module block diagram of a low-light road perception and topology reasoning system based on trajectory priors provided by an embodiment of the present invention;

[0065] Figure 9 This is a structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0066] The technical solutions in the embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention.

[0067] See also Figure 1 This is a flowchart of a trajectory prior-based low-light road perception and topology reasoning method provided by an embodiment of the present invention. This low-light road perception and topology reasoning method can be performed by a low-light road perception and topology reasoning system. This low-light road perception and topology reasoning system can be implemented in software and / or hardware and configured in an electronic device, such as a computer or server, for example, a server in a backend control platform. The detailed steps of this low-light road perception and topology reasoning method are described below.

[0068] S11, obtaining a surround view image captured by a vehicle-mounted camera, wherein the surround view image is a low-light image, and enhancing the surround view image using an image pre-training model to obtain an enhanced image.

[0069] S12 , obtaining trajectory priori data, performing rasterization encoding and vectorization encoding on the trajectory priori data, and generating a rasterized heat map and vectorized trajectory information.

[0070] S13, encoding the enhanced image to obtain a bird's-eye view feature, and fusing and aligning the bird's-eye view feature with the rasterized heat map to obtain an aligned fusion feature.

[0071] S14: Input the vectorized trajectory information and the fused features into a decoder, perform decoding processing through a cross-attention mechanism, and transform the decoding results obtained after decoding through a linear layer to generate road perception and topological reasoning results. The road perception and topological reasoning results include the geometric position and topological relationship of the lane segments.

[0072] A vehicle is equipped with multiple onboard cameras, each of which can capture images within its shooting range. In this embodiment, the images captured by multiple onboard cameras are combined to obtain a surround view image, which can be understood as images involving multiple directions of the vehicle. The low-light road perception and topology reasoning method provided in this embodiment can be used for perception and reasoning processing in low light, nighttime, or special lighting conditions. Therefore, the obtained surround view image can be a low-light image, that is, an image captured in low light, nighttime, or special lighting conditions.

[0073] First, a preset image pre-training model can be used to enhance the surround view image with poor quality, including processes such as improving image brightness and removing image noise to obtain an enhanced image.

[0074] Furthermore, trajectory prior data can be extracted from a crowdsourced trajectory database. This prior data includes multiple trajectory data. This prior data can first be filtered, for example, to remove short or abnormal trajectories. Then, an averaging filter can be used to filter the prior data to reduce random fluctuations, allowing the processed data to better reflect the main motion trends.

[0075] On this basis, the trajectory prior data is rasterized and vectorized to obtain two forms of prior trajectory information, including rasterized heat map and vectorized trajectory information.

[0076] After obtaining the enhanced image of the surround view image, in order to perceive the surrounding environment from a global perspective, the enhanced image can be encoded to obtain bird's-eye view features, which are image features from a bird's-eye view perspective.

[0077] The rasterized heat map obtained based on trajectory prior data is a raster map constructed from a global perspective. Therefore, the bird's-eye view features from a global perspective can be aligned and fused with the rasterized heat map to obtain the fused features after integrating the bird's-eye view features into the rasterized heat map.

[0078] The vectorized trajectory information obtained from the trajectory prior data provides vector information on the map. Therefore, the fused features and vectorized trajectory information can be combined and cross-attention encoded to obtain road perception and topological reasoning results, including the geometric positions and topological relationships of lane segments in the map. This can be used for autonomous driving control guidance in autonomous driving scenarios.

[0079] The trajectory prior-based low-light road perception and topology reasoning method provided in this embodiment utilizes a pre-trained image model to enhance low-light images, resulting in enhanced images. This method enhances the feature extraction capabilities of map construction in low-light environments, effectively improving the ability to cope with autonomous driving scenarios in low and complex lighting conditions, and enhancing the accuracy of road perception and topology reasoning. Furthermore, extracting trajectory prior data as prior information reduces the cost of using high-precision maps as prior information, while also enhancing the accuracy and robustness of the map construction process.

[0080] The following is a detailed description of the specific implementation of each of the above steps. Figure 2 The above-mentioned step of enhancing the surround view image using the image pre-training model to obtain the enhanced image can be achieved by the following methods:

[0081] S111: Input the surround view image into an image pre-training model to obtain an average brightness of the surround view image.

[0082] S112 : Obtain an edge map based on the surround view image according to the average brightness of the surround view image and a set first brightness threshold.

[0083] S113 : Obtain an initialized latent representation based on the surround view image according to the average brightness of the surround view image and a set second brightness threshold.

[0084] S114: Generate an enhanced image according to the edge map and the initialized latent representation.

[0085] In this embodiment, the surround view image is input into the image pre-training model, and the average brightness x of the surround view image is first calculated. c .

[0086] In this embodiment, a first brightness threshold and a second brightness threshold are provided. The first brightness threshold is greater than the second brightness threshold. For example, the first brightness threshold may be 70, and the second brightness threshold may be 40. The first brightness threshold may be used to extract edge information of the surround view image, while the second brightness threshold may be used to determine the potential representation of the surround view image.

[0087] The average brightness of the surround view image can be compared with a first brightness threshold. If the average brightness of the surround view image is greater than or equal to the first brightness threshold, the original surround view image is output. If the average brightness is less than the first brightness threshold, the brightness of the surround view image is proportionally increased to the first brightness threshold. The representation is as follows:

[0088]

[0089] Where x′ cIndicates the brightness of the surround image after brightness enhancement, x c represents the original brightness of the surround image, τ avg represents the first brightness threshold, Indicates the average brightness of the surround image.

[0090] In addition, the average brightness of the surround view image is compared with the second brightness threshold. The comparison method and the image output method are similar to the above and are not described in detail here.

[0091] In this embodiment, the output image obtained by processing the surround view image based on the first brightness threshold is recorded as x′ c1 , the output image obtained by processing the surround view image based on the second brightness threshold is recorded as x′ c2 .

[0092] The output image x′ is detected using a fully nested edge detection method. c1 Perform edge extraction to obtain edge map e.

[0093] For the output image x′ c2 An inverse transformation is performed using a denoising diffusion implicit model to obtain an initialized latent representation of the surround view image. Specifically, the step of obtaining the initialized latent representation of the surround view image based on the average brightness of the surround view image and a set second brightness threshold can be achieved by:

[0094] The average brightness of the surround view image is compared with a set second brightness threshold, and a brightness-enhanced image is obtained based on the surround view image according to the comparison result; the brightness-enhanced image is converted into a latent representation through a denoising diffusion implicit model, and random noise is generated; the latent representation and the random noise are fused to generate an initialized latent representation.

[0095] In this embodiment, the average brightness of the surround view image is compared with the second brightness threshold, and the brightness-enhanced image obtained is the output image x′. c2 The brightness-enhanced image is inversely converted to the latent representation by the denoising diffusion implicit model And, generate random noise in Represents a normal distribution with zero mean, unit variance, and independent and identical distribution in each dimension, and I represents the unit covariance matrix. Through adaptive instance normalization, the potential representation obtained is With random noise Perform fusion to generate initialized potential representation The fusion methods are as follows:

[0096]

[0097] Among them, σ() represents the standard deviation function and μ() represents the mean function.

[0098] On this basis, the enhanced image is generated according to the edge map and the initialized potential representation. Specifically, this step can be achieved by:

[0099] The initialized latent representation is denoised and self-attention features are extracted at each time step; prediction noise is generated based on the self-attention features, edge map and initialized latent representation; the initialized latent representation and prediction noise are used to iteratively optimize the initialized latent representation according to the time step to obtain the final latent representation, and the final latent representation is decoded to generate an enhanced image.

[0100] In this embodiment, the denoising implicit diffusion model ∈ θ Fine-tuning can be performed by low-rank adaptation. Specifically, low-rank matrices A and B can be added to the denoising implicit diffusion model so that these low-rank matrices can be fine-tuned according to the low-light image data without changing the weights of the original model, thereby obtaining a fine-tuned stable denoising implicit diffusion model ∈′ θ .

[0101] On this basis, the initial potential representation of the image Using a stable denoising implicit diffusion model ∈′ θ Denoise the image and extract the self-attention feature A at each time step (t=T,…,1,0) in the process t , characterized as follows:

[0102]

[0103] Among them, q t 、k t 、v t are the embedding vectors of query, key, and value respectively, and d is q t 、k t 、v t At the same time, q t 、k t 、v t The potential representation of the current time step can be utilized It is generated by three different linear transformations, the specific formula is:

[0104]

[0105] Among them, W q 、W k 、W v are the weight matrices of query, key, and value, which are learnable linear parameters. At the same time, the self-attention feature A is taken out t Attention features of the lth layer and the current potential representation The stable diffusion model ∈ that controls the network expansion with time step t and edge graph e as conditional input θcn , generating prediction noise The formula is as follows:

[0106]

[0107] Then the potential representation is initialized by the denoising diffusion implicit model And the predicted noise obtained according to the above formula in time step T Generate the latent representation for the next time step Then at time step T-1, using the potential representation and the prediction noise at the current time step Get potential representation The latent representation is then iterated and optimized in this way based on the time step.

[0108] Finally, when the denoising process is completed (i.e., time step t = 0), the final potential representation is obtained And decode it to get the enhanced image I with low noise and normal brightness e , as the output of the pre-trained model for low-light images.

[0109] In addition, in this embodiment, the obtained trajectory prior data is rasterized and vectorized to generate a rasterized heat map and vectorized trajectory information, wherein, please refer to Figure 3 The above steps of rasterizing and encoding the trajectory priori data and generating a rasterized heat map can be achieved by the following methods:

[0110] S121 : For each piece of trajectory data in the trajectory priori data, obtain the direction angle of each two adjacent two-dimensional points in the trajectory data.

[0111] S122: Create a grid including multiple grid cells, and calculate the number and average direction angle of trajectory data passing through each grid cell.

[0112] S123 , normalizing the number of trajectory data of all grid cells, and transforming the average direction angle to a set interval using an inverse tangent function to generate a rasterized heat map containing density information and direction information.

[0113] In this embodiment, the trajectory priori data is composed of multiple trajectory data, and the trajectory data can be expressed as T={P (1) ,...,P (m)}, where if each trajectory data P (i) It is composed of n two-dimensional points and can be expressed as For every two adjacent two-dimensional points in the trajectory data, for example, point p i and p i+1 , set the direction angle between the two to be θ i,i+1 =arctan(y i+1 –y i ,x i+1 –x i ).

[0114] Then, we can create a grid with H×W grid cells, where each grid cell represents a location point (Δx, Δy) in the real world. We can then calculate the number of trajectory data N and the average direction angle θ that pass through each grid cell, and obtain the maximum number of trajectories N that pass through all grid cells. max Then, for the number of trajectory data in all grid cells, the number of trajectory data N is normalized by the Sigmoid function, and the formula is as follows:

[0115]

[0116] Then the average direction angle θ is transformed into a set interval through the inverse tangent function, such as the interval Finally, a rasterized heat map containing density and direction information is obtained.

[0117] In addition, the step of vectorizing and encoding the trajectory prior data to generate vectorized trajectory information can be achieved by:

[0118] A clustering algorithm or a farthest point sampling method is used to determine representative trajectory data samples from multiple trajectory data; the representative trajectory data samples are vectorized to obtain vectorized trajectory information.

[0119] Specifically, the most representative trajectory data samples are extracted through the K-Means clustering method or the farthest point sampling method.

[0120] For the K-Means clustering method, first select m trajectory data from all trajectory data as the initial cluster center, then calculate each trajectory data P (i) The distance to m cluster centers is calculated and assigned to the nearest cluster center, and then the position of each cluster center is updated. The specific formula is as follows:

[0121]

[0122] Among them, μ (k) represents the kth cluster center, S kThis represents all trajectory data assigned to the kth cluster. The algorithm then iterates in this manner until the maximum number of iterations is reached or the change in the cluster center falls below a preset threshold, such as 0.0001. Finally, the cluster center is designated as the representative trajectory, forming the most representative representative trajectory data sample.

[0123] For the farthest point sampling method, first, a random trajectory data is selected as the initial point and added to the trajectory set S. Then, a trajectory data with the largest minimum distance from all trajectory data in the currently selected trajectory set S is selected and added to S. This process is repeated until a sufficient number of trajectory data are obtained as the most representative representative trajectory data sample.

[0124] The most representative trajectory data samples are vectorized to obtain vectorized trajectory information, which is expressed as Where m represents the number of selected trajectory data, and n represents the feature dimension of the trajectory data.

[0125] After obtaining the enhanced image, rasterized heat map and vectorized trajectory information, the enhanced image is encoded to obtain the bird's-eye view feature. Figure 4 , this step can be achieved by:

[0126] S131, performing feature extraction on the enhanced image of each vehicle-mounted camera to obtain a corresponding feature map.

[0127] S132, input the obtained multiple feature maps into an encoder comprising multiple coding layers, perform bird's-eye view feature extraction for multiple time steps on each of the feature maps, and in each coding layer, query the bird's-eye view features obtained in the previous time step through a temporal self-attention mechanism based on the set query information to obtain the time information of the current time step, and query the multiple feature maps through a spatial cross-attention mechanism to obtain the spatial information of the current time step.

[0128] S133: Generate a bird's-eye view feature based on the temporal information and spatial information output by the last coding layer.

[0129] First, the enhanced images corresponding to each vehicle-mounted camera are extracted through the backbone network to obtain the feature map of each vehicle-mounted camera. in represents the feature map of the i-th vehicle camera, N view is the total number of onboard cameras.

[0130] In addition, through a set of predefined grid-like learnable parameters As query information, H and W represent the spatial shape of the Bird's Eye View (BEV) plane, and C represents the feature dimension. The query information is understood as a processing unit. That is, the processing for the entire space is divided into multiple processing units, and processing is performed on a single processing unit at a time, which can improve processing speed and precision.

[0131] Multiple feature maps are input into an encoder, which includes multiple coding layers, for example, six. The feature maps are processed sequentially by each coding layer. The processing of each coding layer can be understood as processing at each time step. Each coding layer is processed based on the bird's-eye view features obtained at the previous coding layer (i.e., the previous time step) and the set query information.

[0132] Each encoding layer processes from the temporal and spatial perspectives. At each encoder layer, the BEV query information Q is first used to query the temporal information from the bird’s-eye view features obtained in the previous time step through the temporal self-attention mechanism. Then, the BEV query information Q is used to query the temporal information from the multi-camera features F through the spatial cross-attention mechanism. t After the feedforward network, the encoding layer outputs the refined bird's-eye view features and serves as the input of the next encoding layer. Through the iteration of multiple encoding layers, the bird's-eye view features are gradually refined, and finally the unified bird's-eye view features of the current time step are generated.

[0133] On this basis, see Figure 5 , the step of fusing and aligning the bird's-eye view features with the rasterized heat map to obtain the aligned fusion features can be achieved by the following methods:

[0134] S134 , obtaining a trajectory priori features based on the rasterized heat map, and concatenating the trajectory priori features with the bird's-eye view features to obtain concatenated features.

[0135] S135 , using multiple convolutional layers to process the splicing features, and predicting the coordinate offset of each pixel point in the trajectory prior features.

[0136] S136 , mapping the position information of each pixel point to a new position according to the coordinate offset of each pixel point.

[0137] S137 , upsampling the trajectory priori feature to align it with the bird's-eye view feature.

[0138] S138: Perform weighted fusion on the aligned trajectory prior features and the bird's-eye view features according to the learned weights to obtain fused features.

[0139] In this embodiment, the trajectory prior features are first obtained based on the rasterized heat map. The trajectory prior features With bird's-eye view features The splicing features obtained by processing multiple convolutional layers are used to predict the coordinate offset of each pixel point in the spatial range, which is expressed as Since the plane range is mainly considered, C here is 2, which is the two-dimensional space range.

[0140] Among them, the coordinate offset includes horizontal offset and vertical offset, which are expressed as Δ hw1 and Δ hw2 .

[0141] The optical flow model can be used to perform bilinear interpolation on the trajectory prior features, and the position of each pixel (h, w) is calculated based on the predicted coordinate offset (Δ hw1 ,Δ hw2 ), mapped to the new position (h′, w′), the formula is as follows:

[0142]

[0143] Where (h′, w′) represents the position (h+Δ hw1 ,w+Δ hw2 ) surrounding neighborhood pixels (upper left, upper right, lower left, lower right), w p is the bilinear interpolation weight.

[0144] Next, run the confidence fusion module, input the trajectory prior features and bird's-eye view features into the confidence fusion module, and perform the above-processed trajectory prior features. Upsampling is performed to align it with the bird's-eye view features to match. Then two spatial importance weight matrices α are introduced l and β l , use 1×1 convolution kernel to perform bird’s-eye view features and the upsampled trajectory prior features Calculate weights:

[0145]

[0146] The unnormalized weights obtained from the above formula are normalized using the Softmax function to calculate the final weights while ensuring that α l +β l =1, and α l ,β l ∈[0,1]:

[0147]

[0148] Finally, according to the learned weights, the two trajectory prior features and bird's-eye view features are weighted fused to obtain the weighted fusion aligned fusion feature y l , the specific formula is as follows:

[0149]

[0150] where ⊙ represents the element-wise multiplication of matrices.

[0151] On this basis, the vectorized trajectory information and fusion features are input into the decoder and decoded through the cross attention mechanism. Figure 6 , this step can be achieved by:

[0152] S141, converting the vectorized trajectory information into high-dimensional vectorized trajectory information through a multi-layer perceptron, and performing a linear transformation on the high-dimensional vectorized trajectory information to generate an initial query.

[0153] S142: Generate a query embedding by combining the initial query and a preset learnable query vector.

[0154] S143: Input the query embedding and the fused features into a decoder, and adopt a cross-attention mechanism to combine the query embedding and the fused features to obtain a decoding result.

[0155] Vectorize the trajectory prior information T v The high-dimensional vectorized trajectory information h is converted into high-dimensional vectorized trajectory information h through a multi-layer perceptron (MLP). Then, the high-dimensional vectorized trajectory information h is transformed into the initial query Q through a learnable linear transformation. init The specific formula is as follows:

[0156] Q init =W q h+b q

[0157] Among them, h is the high-dimensional vectorized trajectory information, W q is the learnable weight matrix, b q is the bias term.

[0158] The coordinates of the trajectory data are used as the initial reference point for the initial query. Finally, the initial query and the pre-set learnable query vector are combined to form a complete query embedding. The Transformer decoder of the Transformer-based object detection algorithm (DETR) then employs a cross-attention mechanism to combine the query embedding with the enhanced fusion features to produce the decoding result.

[0159] See also Figure 7The decoding results are converted through linear layers to generate road perception and topology reasoning results. This can be achieved by:

[0160] S144 , converting the decoded result into coordinate information through linear transformation to obtain the geometric position of the lane segment.

[0161] S145: Calculate the connection relationship between lane segments through the graph neural network to obtain the topological adjacency matrix of the lane segments.

[0162] In this embodiment, the decoded results of the Transformer decoder are converted into coordinate information through linear transformation to obtain the geometric position of the lane segment. The connectivity between lane segments is calculated through a graph neural network to obtain the lane topology adjacency matrix.

[0163] The Sigmoid function is used to normalize the topological adjacency matrix so that the probability value of the connection relationship is between 0 and 1. This is used to obtain the topological reasoning of the lane segment and finally output the final road perception and topological reasoning results.

[0164] The trajectory prior-based low-light road perception and topology reasoning method provided in this embodiment uses a pre-trained image model to enhance low-light images using a stepwise diffusion approach. This method can restore image details and improve image brightness while reducing image distortion during the enhancement process. Furthermore, by combining trajectory prior data with a multimodal fusion system to enhance lane perception and topology reasoning capabilities, it can be an effective way to improve the accuracy of road perception and topology reasoning in low-light environments.

[0165] This solution effectively improves the ability to cope with autonomous driving scenarios in low-light and complex lighting conditions, enhances the accuracy of road perception and topology reasoning, and effectively improves the anti-interference performance of map construction in complex situations. It also effectively utilizes crowdsourced trajectory data as prior information, enhancing the accuracy and robustness of the map construction process while reducing the difficulty and cost of obtaining prior information.

[0166] Based on the same inventive concept, please refer to Figure 8 , an embodiment of the present invention also provides a functional module diagram of a low-light road perception and topology reasoning system based on trajectory prior. This embodiment can divide the functional modules of the low-light road perception and topology reasoning system according to the above-mentioned method embodiment. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present invention is schematic and is only a logical functional division. There may be other division methods in actual implementation.

[0167] For example, when each functional module is divided into corresponding functional modules, Figure 8 The low-light road perception and topology inference system shown is only a schematic diagram. It can include an enhancement processing module, an encoding module, a fusion module, and a decoding module. The functions of each module of the low-light road perception and topology inference system are described in detail below.

[0168] an enhancement processing module, configured to obtain a surround view image captured by an on-board camera, the surround view image being a low-light image, and perform enhancement processing on the surround view image using an image pre-training model to obtain an enhanced image;

[0169] An encoding module is used to obtain trajectory prior data, perform raster encoding and vector encoding on the trajectory prior data, and generate a rasterized heat map and vectorized trajectory information;

[0170] a fusion module, configured to encode the enhanced image to obtain a bird's-eye view feature, and fuse and align the bird's-eye view feature with the rasterized heat map to obtain an aligned fusion feature;

[0171] A decoding module is configured to input the vectorized trajectory information and the fused features into a decoder, perform decoding processing through a cross-attention mechanism, and transform the decoding results obtained after decoding through a linear layer to generate road perception and topological reasoning results, wherein the road perception and topological reasoning results include the geometric positions and topological relationships of lane segments.

[0172] The low-light road perception and topology reasoning system provided in this embodiment can be used to execute the low-light road perception and topology reasoning method under any implementation method of the above embodiments. For matters not detailed in this embodiment, please refer to the corresponding description of the above embodiments, and this embodiment will not be repeated here.

[0173] See also Figure 9 , is a block diagram of the structure of an electronic device provided in an embodiment of the present invention. This electronic device may be a computer device, server, or the like in an autonomous driving control platform. The electronic device includes a memory, a processor, and a communication module. The memory, processor, and communication module components are electrically connected to each other, directly or indirectly, to enable data transmission or interaction. For example, these components may be electrically connected to each other via one or more communication buses or signal lines.

[0174] Memory is used to store computer programs or data. Memory can be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM).

[0175] The processor is used to read / write data or programs stored in the memory and execute the low-light road perception and topology reasoning method provided by any embodiment of the present invention.

[0176] The communication module is used to establish a communication connection between the electronic device and other communication terminals through the network, and is used to send and receive data through the network.

[0177] It should be understood that Figure 9 The structure shown is only a schematic diagram of the structure of the electronic device. The electronic device may also include Figure 9 More or fewer components than shown, or with Figure 9 Different configurations shown.

[0178] Furthermore, an embodiment of the present invention also provides a computer-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are executed, the low-light road perception and topology reasoning method provided in the above embodiment is implemented.

[0179] Specifically, the computer-readable storage medium can be a general-purpose storage medium, such as a removable disk or hard disk. When the computer program on the computer-readable storage medium is executed, it can perform the aforementioned low-light road perception and topology reasoning method. The processes involved in executing the computer-readable storage medium and its executable instructions can be found in the description of the aforementioned method embodiments and will not be further elaborated here.

[0180] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some communication interface, the indirect coupling or communication connection of the device or unit may be electrical, mechanical or other forms.

[0181] In addition, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0182] Furthermore, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.

[0183] It should be noted that if the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0184] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any actual relationship or order between these entities or operations.

[0185] The foregoing description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A low-light road perception and topology reasoning method based on trajectory prior, characterized by: The method comprises: Obtaining a surround view image captured by a vehicle-mounted camera, wherein the surround view image is a low-light image, and enhancing the surround view image using an image pre-training model to obtain an enhanced image; Obtaining trajectory priori data, and performing rasterization encoding and vectorization encoding on the trajectory priori data to generate a rasterized heat map and vectorized trajectory information; Encoding the enhanced image to obtain a bird's-eye view feature, and fusing and aligning the bird's-eye view feature with the rasterized heat map to obtain an aligned fused feature; The vectorized trajectory information and the fused features are input into a decoder, decoded using a cross-attention mechanism, and the decoded results are transformed through a linear layer to generate road perception and topological reasoning results, which include the geometric positions and topological relationships of lane segments.

2. The low-light road perception and topology reasoning method based on trajectory prior according to claim 1, characterized in that: The step of enhancing the surround view image using the image pre-training model to obtain an enhanced image includes: Inputting the surround view image into an image pre-training model to obtain an average brightness of the surround view image; Obtaining an edge map based on the surround view image according to an average brightness of the surround view image and a set first brightness threshold; Obtaining an initialized potential representation based on the surround view image according to an average brightness of the surround view image and a set second brightness threshold; An enhanced image is generated according to the edge map and the initialized latent representation.

3. The low-light road perception and topology reasoning method based on trajectory prior according to claim 2, characterized in that: The step of obtaining an initialized potential representation based on the surround view image according to the average brightness of the surround view image and a set second brightness threshold comprises: comparing the average brightness of the surround view image with a set second brightness threshold, and obtaining a brightness-enhanced image based on the surround view image according to the comparison result; Converting the brightness-enhanced image into a latent representation through a denoising diffusion implicit model and generating random noise; The latent representation and the random noise are fused to generate an initialized latent representation.

4. The low-light road perception and topology reasoning method based on trajectory prior according to claim 3, characterized in that: The step of generating an enhanced image according to the edge map and the initialized latent representation comprises: Denoising the initialized latent representation and extracting self-attention features at each time step; generating prediction noise based on the self-attention features, the edge map, and the initialized latent representation; The initialized latent representation and the predicted noise are used to iteratively optimize the initialized latent representation according to time steps to obtain a final latent representation, and the final latent representation is decoded to generate an enhanced image.

5. The low-light road perception and topology reasoning method based on trajectory prior according to claim 1, characterized in that: The trajectory priori data includes a plurality of trajectory data, each trajectory data is composed of a plurality of two-dimensional points; The rasterized heat map is generated in the following way: For each piece of trajectory data in the trajectory priori data, obtaining the direction angle of each two adjacent two-dimensional points in the trajectory data; Create a grid containing multiple grid cells and calculate the number and average direction angle of trajectory data passing through each grid cell; Normalizing the number of trajectory data items of all grid cells, and transforming the average direction angle to a set interval using an inverse tangent function to generate a rasterized heat map containing density and direction information; The vectorized trajectory information is generated in the following manner: Determine a representative trajectory data sample from the plurality of trajectory data using a clustering algorithm or a farthest point sampling method; Vectorization processing is performed on the representative trajectory data samples to obtain vectorized trajectory information.

6. The low-light road perception and topology reasoning method based on trajectory prior according to claim 1, characterized in that: The vehicle-mounted camera includes a plurality of; The step of encoding the enhanced image to obtain a bird's-eye view feature comprises: Performing feature extraction on the enhanced images of each of the vehicle-mounted cameras to obtain a corresponding feature map; Inputting the obtained multiple feature maps into an encoder comprising multiple encoding layers, performing bird's-eye view feature extraction for each of the feature maps for multiple time steps, and in each encoding layer, querying the bird's-eye view features obtained at the previous time step through a temporal self-attention mechanism based on set query information to obtain the temporal information of the current time step, and querying the spatial information of the current time step from the multiple feature maps through a spatial cross-attention mechanism; Based on the temporal and spatial information output by the last encoding layer, bird's-eye view features are generated.

7. The low-light road perception and topology reasoning method based on trajectory prior according to claim 1, characterized in that: The step of fusing and aligning the bird's-eye view feature with the rasterized heat map to obtain an aligned fused feature includes: Obtaining a trajectory priori features based on the rasterized heat map, and splicing the trajectory priori features with the bird's-eye view features to obtain spliced features; Processing the splicing features using multiple convolutional layers to predict the coordinate offset of each pixel in the trajectory prior features; Mapping the position information of each pixel point to a new position according to the coordinate offset of the pixel point; Upsampling the trajectory prior features to align with the bird's-eye view features; The aligned trajectory prior features and bird's-eye view features are weightedly fused according to the learned weights to obtain the fused features.

8. The low-light road perception and topology reasoning method based on trajectory prior according to claim 1, characterized in that: The step of inputting the vectorized trajectory information and the fused features into a decoder and performing decoding processing through a cross-attention mechanism includes: Converting the vectorized trajectory information into high-dimensional vectorized trajectory information through a multi-layer perceptron, and performing a linear transformation on the high-dimensional vectorized trajectory information to generate an initial query; Combining the initial query with a preset learnable query vector to generate a query embedding; The query embedding and the fused features are input into a decoder, and a decoding result is obtained by combining the query embedding and the fused features using a cross attention mechanism.

9. The low-light road perception and topology reasoning method based on trajectory prior according to claim 1, characterized in that: The step of converting the decoding result obtained after decoding through a linear layer to generate a road perception and topology reasoning result includes: The decoded results are converted into coordinate information through linear transformation to obtain the geometric position of the lane segment; The connection relationship between lane segments is calculated through graph neural network to obtain the topological adjacency matrix of the lane segments.

10. A low-light road perception and topology reasoning system based on trajectory prior, characterized by: The system comprises: an enhancement processing module, configured to obtain a surround view image captured by an on-board camera, the surround view image being a low-light image, and perform enhancement processing on the surround view image using an image pre-training model to obtain an enhanced image; An encoding module is used to obtain trajectory prior data, perform raster encoding and vector encoding on the trajectory prior data, and generate a rasterized heat map and vectorized trajectory information; a fusion module, configured to encode the enhanced image to obtain a bird's-eye view feature, and fuse and align the bird's-eye view feature with the rasterized heat map to obtain an aligned fusion feature; A decoding module is configured to input the vectorized trajectory information and the fused features into a decoder, perform decoding processing through a cross-attention mechanism, and transform the decoding results obtained after decoding through a linear layer to generate road perception and topological reasoning results, wherein the road perception and topological reasoning results include the geometric positions and topological relationships of lane segments.

Citation Information

Patent Citations

  • Method and system for establishing lane line map and extracting topological structure

    CN117671494A

  • Three-dimensional sensing method based on millimeter wave radar and camera aerial view fusion

    CN118038396A

  • Vehicle trajectory determination method and device based on thermodynamic diagram, vehicle and medium

    CN118429377A

  • Online vectorization high-precision map construction method based on multi-modal instance fusion

    CN118864651A

  • High-altitude aerial image road identification method based on deep learning

    CN119131577A

Cited By

  • Unmanned aerial vehicle aerial image road detection method based on deep learning

    CN121170654A