Online vector map construction method for low-light and rainfall scene

Through multi-scale foreground enhancement and bird's-eye view feature fusion technology, the problem of insufficient perception accuracy of autonomous driving systems in low light and rain conditions is solved, efficient and accurate vector map construction is achieved, and the adaptability and robustness of the system are enhanced.

CN120707693APending Publication Date: 2025-09-26BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510792610.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The perception accuracy of existing autonomous driving systems decreases in low-light and rainy conditions, affecting safety and reliability. Existing technologies make it difficult to accurately construct vector maps in real time in low-light and rainy environments.

Method used

Multi-scale foreground enhancement processing and bird's-eye view feature dynamic fusion technology are adopted. Low-light and rainfall images are distinguished through image classifiers, and targeted processing is performed. Combined with feature encoders, multi-layer perceptrons and deformable attention modules, high-precision vector maps are generated.

Benefits of technology

It improves the perception capability of the autonomous driving system in low-light and rainy conditions, can accurately identify road features, enhances the adaptability and robustness of the system, optimizes the vector map construction process, reduces the amount of calculation, and improves real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707693A_ABST
    Figure CN120707693A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of automatic driving perception, and particularly relates to an online vector map construction method for a low-light and rainfall scene, and the method comprises the steps: obtaining a real-time image, and carrying out the classification of the real-time image based on an image classifier, and the real-time image comprises a multi-view image; when the classification type of the real-time image is a low-light image, performing first processing; when the classification type of the real-time image is a rain image, performing second processing; when the classification type of the real-time image is a normal image, the real-time image is not processed; establishing a vector map based on the input image, including performing multi-scale foreground enhancement processing based on the input image to obtain masked multi-scale perspective features; based on the masked multi-scale perspective view features, integration of height priori knowledge and multi-scale features is carried out, and multi-scale aerial view features are obtained; and performing feature dynamic fusion processing on the aerial view based on the multi-scale aerial view features to obtain a vector map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of autonomous driving perception technology, and specifically relates to an online vector map construction method for low-light and rainfall scenes. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, the safety and accuracy standards for autonomous driving are becoming increasingly stringent. In this context, autonomous driving systems must demonstrate superior environmental recognition and responsiveness in changing traffic conditions. Therefore, a deep understanding of the road environment, particularly the ability to construct accurate semantic vector maps, is crucial for the practical deployment and operation of autonomous driving technology. To meet these high standards, the design of autonomous driving systems must focus on improving their perception accuracy and responsiveness to complex road conditions. This includes detailed analysis of the surrounding environment and rapid identification of potential hazards. Deepening semantic understanding through the construction of accurate vector maps provides a solid foundation for autonomous driving technology, ensuring its effectiveness and reliability in the real world.

[0003] Existing technologies typically rely on multiple sensors, such as cameras, radar, and lidar, to gather information about the surrounding environment. However, the performance of these systems is often limited in low-light and rainy conditions, resulting in reduced perception accuracy and impacting the safety and reliability of autonomous driving. Nighttime or low-light environments reduce camera visibility, making it difficult for vision-based perception systems to accurately identify road signs, lane markings, and other key elements. Raindrops block the camera lens, affecting image quality, and the specific frequency components they create in the image increase the complexity of image processing.

[0004] Research on advanced visual tasks in low-light environments is facing many challenges, such as insufficient lighting, shadows, reflections, and noise. To address these challenges, researchers have developed a number of technologies. For example, image enhancement technology improves image quality by adjusting the brightness, contrast, saturation, etc. of the image, but it may not be able to restore details lost in extremely low light, and may introduce noise or distortion. Multi-exposure image fusion technology combines multiple images with different exposure settings to create a more balanced image, but requires additional hardware support and cannot be processed in real time in dynamic scenes. Domain adaptation and transfer learning technologies adjust the model to adapt to different lighting conditions, but require additional labeled data to train models in new domains, and performance may degrade under unseen lighting conditions.

[0005] Research on advanced visual tasks in rainy environments faces numerous challenges, such as raindrop occlusion, light scattering, and image blur. To address these challenges, researchers have developed a number of technologies. For example, raindrop removal technology uses image processing algorithms, such as physical-based raindrop simulation, deep learning networks, or frequency-domain filtering transforms, to identify and remove raindrops. This may not fully restore image details obscured by raindrops, and the effect may decrease under heavy rainfall conditions. Multi-sensor fusion technology combines sensor data from cameras, radars, lidars, and other sensors to improve perception capabilities in rainy conditions, but this increases system complexity and cost, and has high data processing requirements. Attention mechanisms and Transformer models are used to enhance the model's focus on key features, but the computational cost is high and may not be suitable for real-time applications.

[0006] It can be seen that there is an urgent need to provide a vector map construction method for low-light and rainfall environments that can meet both accuracy and timeliness. Summary of the Invention

[0007] The present invention is proposed based on the above-mentioned requirements of the prior art. The technical problem to be solved by the present invention is to provide an online vector map construction method for low-light and rainfall scenes to improve the real-time performance and accuracy of vector map generation.

[0008] In order to solve the above problems, the technical solutions provided by the present invention include:

[0009] Provided is an online vector map construction method for low-light and rainy scenes, comprising: acquiring a real-time image and classifying the real-time image based on an image classifier, wherein the real-time image includes a multi-view image; performing different processing according to different classification types to obtain an input image; performing a first processing when the classification type of the real-time image is a low-light image; performing a second processing when the classification type of the real-time image is a rainy image; and performing no processing when the classification type of the real-time image is a normal image; establishing a vector map based on the input image, comprising: extracting the input image through a feature encoder to obtain multi-scale perspective features; performing multi-scale foreground enhancement processing on the multi-scale perspective features to obtain masked multi-scale perspective features, comprising: generating a foreground mask through a multi-layer perceptron based on the multi-scale perspective features; connecting the foreground mask and the multi-scale perspective features through a residual, and Element-wise multiplication is used to adjust the feature weights at different positions of the image to obtain masked multi-scale perspective features; based on the masked multi-scale perspective features, the height prior knowledge and multi-scale features are integrated to obtain multi-scale bird's-eye view features; based on the height distribution probability and multi-scale height distribution features, multi-scale bird's-eye view features are obtained; based on the multi-scale bird's-eye view features, dynamic fusion processing of bird's-eye view features is performed, including: splicing bird's-eye view features of different scales to form a coarse integrated feature; after the coarse integrated feature passes through the convolutional neural network, deformable attention is applied to dynamically adjust the feature focus, selectively enhance the key feature area, refine the features through a multi-layer perceptron, and perform residual connection with the output of the variable attention module to achieve feature fusion, to obtain the fused bird's-eye view features; the fused bird's-eye view features are converted into a real-time vectorized scene through a feature decoder.

[0010] Preferably, the masked multi-scale perspective feature is used to integrate the height prior knowledge and the multi-scale features to obtain the multi-scale bird's-eye view feature, including: position encoding and average pooling the masked multi-scale perspective feature to generate a global feature; integrating the global feature into a predefined query and calculating the height distribution probability through a multi-layer perceptron; obtaining a predefined reference point close to the road under the bird's-eye view perspective, projecting the reference point into the multi-scale perspective space, and obtaining the multi-scale height distribution feature.

[0011] Preferably, the foreground mask is generated by a multi-layer perceptron based on the multi-scale perspective image features, which is expressed as: in, represents the foreground mask, F i Represents multi-scale perspective features, Conv() is a convolutional layer, ReLU() is an activation function, and Sigmoid() is an activation function.

[0012] Preferably, the foreground mask and the multi-scale perspective feature are connected through the residual, and the feature weights of different positions of the image are adjusted by element-wise multiplication to obtain the masked multi-scale perspective feature, which is expressed as: in is the masked multi-scale perspective feature, is the foreground mask, F i is the multi-scale perspective feature, and ⊙ is the element-wise multiplication.

[0013] Preferably, the global features are integrated into a predefined query, and the height distribution probability is calculated by a multi-layer perceptron, including, expressed as: Among them, Height pro is the height distribution probability, MLP() is the multi-layer perceptron, AP() is the average pooling, PE is the position encoding, and Query is the predefined query vector.

[0014] Preferably, the method obtains a predefined reference point close to the road from a bird's-eye view, projects the reference point into a multi-scale perspective image space, and obtains a multi-scale height distribution feature, which is expressed as: in, Sampling is a multi-scale height distribution feature. i () is the feature sampling from the projected image features at the i-th scale, I is the intrinsic parameter matrix of the camera, K is the extrinsic parameter matrix of the camera, is a predefined reference point close to the road in the bird's-eye view perspective; the multi-scale bird's-eye view feature is obtained based on the height distribution probability and the multi-scale height distribution feature, which is expressed as: AP() is average pooling, is the i-th bird's-eye view feature.

[0015] Preferably, after the coarse integrated features pass through the neural network, the deformable attention is applied to dynamically adjust the feature focus, selectively enhance the key feature area, refine the features through the multi-layer perceptron, and perform feature fusion with the output residual connection of the deformable attention. The fused bird's-eye view feature is represented as follows: Among them, F bev′ is the fused bird's-eye view feature, Conv() is the convolution layer, MLP() is the multi-layer perceptron, DA() is the deformable attention, and Concat() is the feature integration.

[0016] Preferably, the first processing includes inputting a real-time image into a trained first model and outputting an image with normal illumination as an input image; the training process of the first model includes: acquiring a low-light image and a sufficient-light image; converting the low-light image to a first intermediate-state illumination image through visual enhancement; converting the sufficient-light image to a second intermediate-state illumination image through a preset network; aligning the low-light image and the sufficient-light image through adversarial learning under cycle consistency to update the first intermediate-state illumination image and the second intermediate-state illumination image; performing feature-level alignment of image semantics based on the updated first intermediate-state illumination image and the second intermediate-state illumination image to obtain an intermediate-state illumination image; and The over-spectral residual generates a saliency map of the sufficient illumination image, and the saliency map is mapped back to the spatial domain through inverse fast Fourier transform, and the salient region is proposed to obtain a static attention map; the sufficient illumination image and the intermediate state illumination image are divided into multiple image blocks, and the sufficient illumination image is combined with the static attention map to obtain a first image; the intermediate state illumination image is combined with the static attention map to obtain a second image; the target area and non-target area of ​​the first and second images are distinguished, the image blocks in the target area are removed or replaced, and some image blocks in the non-target area are retained or occluded to obtain a first updated image and a second updated image; based on the first and second update images, a bootstrapped non-negative contrastive learning loss is used, and the loss is adjusted to minimize the loss.

[0017] Preferably, the second processing includes inputting the real-time image into the trained second model and outputting an image without raindrop traces as the input image; the training process of the second model includes: step one, obtaining a rainy image; step two, performing the following processing to obtain a first output, the processing including: normalizing the rainy image, and after passing through the visual state model, performing element-level multiplication operation, combining it with a learnable scale factor, and inputting it into a multi-scale convolution gating unit, and obtaining the first output through multiple convolution layers in series and parallel forms; step three, normalizing the first output, passing through the visual state model and fast Fourier transform respectively, combining the output of the visual state model with the output after fast Fourier transform through element-level multiplication, and then combining it with a learnable scale factor, and inputting it into a multi-scale convolution gating unit, and obtaining the second output through multiple convolution layers in series and parallel forms; step four, repeating steps two to three seven times, each time with a different resolution and number of channels; step five, calculating the loss function L total =L L1 +μ f L Freq , where L L1 is the L1 norm, μ f To adjust the loss ratio coefficient, L Freq is the loss for each frequency component.

[0018] Preferably, the input is input into the multi-scale convolution gate control unit, including normalizing the input value, expanding the number of channels through 1×1 convolution, and then entering three parallel branches: a gate branch, a 3×3 depth convolution branch and a 5×5 depth convolution branch. In the gate branch, a 3×3 depth convolution is performed first, and then a GeLU activation function is performed to obtain the first data. In the 3×3 and 5×5 depth convolution branches, 3×3 and 5×5 depth convolutions are performed respectively, and the second data and the third data are output respectively. The first data, the second data and the third data are integrated, and then a 1×1 convolution is performed to restore the number of channels, and a jump connection with a learnable scale factor is added and output.

[0019] An online vector map construction system for low-light and rainy scenes is also provided, including: an image acquisition and classification module, which acquires real-time images and classifies the real-time images based on an image classifier, wherein the real-time images include multi-view input images; a low-light image processing module, which performs a first processing to form an input image when the classification type of the real-time image is a low-light image; a rainy image processing module, which performs a second processing to form an input image when the classification type of the real-time image is a rainy image; a normal image processing module, which does not perform any processing to form an input image when the classification type of the real-time image is a normal image; a vector map establishment module, which establishes a vector map based on the input image, including extracting the input image through a feature encoder to obtain multi-scale perspective features; generating a foreground mask through a multi-layer perceptron based on the multi-scale perspective features; connecting the foreground mask and the multi-scale perspective features through residuals, and adjusting the image at different positions using element-level multiplication. The feature weights are integrated to obtain masked multi-scale perspective features; the masked multi-scale perspective features are positionally encoded and average-pooled to generate global features; the global features are integrated into predefined queries, and the height distribution probability is calculated through a multi-layer perceptron; pre-defined reference points close to the road under the bird's-eye view perspective are obtained, and the reference points are projected into the multi-scale perspective space to obtain multi-scale height distribution features; multi-scale bird's-eye view features are obtained based on the height distribution probability and the multi-scale height distribution features; bird's-eye view features of different scales are spliced ​​to form coarse integrated features; after the coarse integrated features pass through the neural network, deformable attention is applied to dynamically adjust the feature focus, and key feature areas are selectively enhanced. The features are refined through a multi-layer perceptron and fused with the output residual connection of the deformable attention to obtain the fused bird's-eye view features; the fused bird's-eye view features are converted into real-time vectorized scenes through a feature decoder.

[0020] Compared to existing technologies, this invention significantly improves the perception capabilities of autonomous driving systems in low-light and rainy conditions through multi-scale foreground enhancement processing and dynamic fusion of bird's-eye view features. These technologies enable the system to more accurately identify and process key road features, such as lane markings and road signs, maintaining high performance even in adverse weather conditions, thereby enhancing the adaptability and robustness of the aforementioned method. The online vector map construction process is optimized by integrating highly prior knowledge and multi-scale features. This not only enhances the model's ability to understand and interpret the road environment, but also improves real-time performance by reducing unnecessary computation, enabling the system to respond quickly and efficiently utilize computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0022] Figure 1 Schematic diagram of the steps of a method for constructing an online vector map for low-light and rainfall scenes according to an embodiment of the present invention;

[0023] Figure 2 Schematic diagram of data transmission of an online vector map construction method for low-light and rainfall scenes according to an embodiment of the present invention;

[0024] Figure 3 A schematic diagram of data transmission during the training process of the first model of the present invention;

[0025] Figure 4 A schematic diagram of data transmission during the training process of the second model of the present invention;

[0026] Figure 5 Schematic diagram of data transmission in the vector map construction process of the present invention. DETAILED DESCRIPTION

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0028] In the description of the embodiments of the present invention, it should be noted that, unless otherwise specified or limited, the term "connected" should be understood in a broad sense. For example, it can mean a fixed connection, a detachable connection, or an integral connection. It can be a mechanical connection, an electrical connection, a direct connection, or an indirect connection through an intermediate medium. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0029] The terms "top," "bottom," "above," "below," and "on" used throughout the description refer to relative positions of components of a device, such as the relative positions of top and bottom substrates within a device. It will be understood that devices are multifunctional regardless of their orientation in space.

[0030] To facilitate understanding of the embodiments of the present invention, specific embodiments will be further explained below with reference to the accompanying drawings. The embodiments do not limit the embodiments of the present invention.

[0031] Example 1

[0032] This embodiment provides an online vector map construction method for low light and rainy scenes. Figure 1-Figure 5 shown.

[0033] The online vector map construction method for low-light and rainfall scenes includes:

[0034] A real-time image is acquired and classified based on an image classifier, wherein the real-time image includes a multi-view input image.

[0035] The acquired images form an image data set, which includes a landscape image, a low-light image, and a normal image.

[0036] Annotated images include rainy images, low-light images, and normal images.

[0037] Preprocess the images in the image dataset; preprocessing includes methods such as rotation, scaling, and cropping to enhance the data, increase data diversity, and improve the generalization ability of the model; and perform normalization to normalize the image pixel values ​​to the same range to obtain a third image.

[0038] The third image is extracted through ResNet50 to output a feature map; global average pooling is performed on the feature map to convert the feature map into a feature vector of fixed length.

[0039] The feature vectors are mapped to the category space through multiple fully connected layers, and the unnormalized category scores are output. The category scores are normalized to obtain the predicted probability distribution of the three categories, which sum to 1. The categories in the category space are divided into rain, low light, and normal.

[0040] Multiple fully connected layers (MLPs) map the feature vectors to three categories: rainy images, low-light images, and normal images. After the fully connected layers, a softmax layer is used to convert the output into a probability distribution corresponding to the three categories.

[0041] Build a training network so that the category corresponding to the highest probability in the predicted probability distribution is close to the true label; use cross entropy as the loss function; update the network parameters according to the loss gradient using the Adam optimizer; calculate accuracy, precision, recall, and F1 score for performance evaluation.

[0042] Input a single original image for image enhancement and normalization to generate the corresponding new first image, input the new first image into the updated and loaded training network to obtain the predicted probability distribution, and select the category with the highest probability as the final classification result.

[0043] Through the above process, an image classifier capable of distinguishing between rainy, low-light, and normal images is constructed. The data preprocessing module performs data collection, labeling, augmentation, and normalization. The feature extraction and processing module extracts image features and performs global average pooling. The classification module applies fully connected layers and a softmax layer to output image class probabilities. The training module uses the cross-entropy loss function and the Adam optimizer to train and optimize the model, and evaluates model performance based on evaluation metrics. The inference module inputs the preprocessed image into the loaded model and outputs the image class.

[0044] Different processing is performed according to different classification types to obtain the input image.

[0045] When the classification type of the real-time image is a low-light image, a first process is performed to form an input image.

[0046] The first processing includes inputting a real-time image into a trained first model and outputting an image with normal lighting as an input image.

[0047] The training process of the first model is as follows Figure 3 Shown, including:

[0048] Obtain a low-light image and a sufficient-light image. The low-light image is represented by I l , the image with sufficient light is represented as I h .

[0049] Image-level alignment is to find an image I with an intermediate illumination state m , ensuring their consistency in image appearance.

[0050] The low-light image is converted into a first intermediate-state light image through visual enhancement.

[0051] That is, there is a set of low-light image enhancement parameters The visual enhancement function can be expressed as in For I m The approximate value of , E(·) is the enhancement function, that is, the adjustment curve derived under the prior algorithm Φ.

[0052] The sufficient illumination image is converted into a second intermediate state illumination image through a preset network.

[0053] Training a network To achieve a well-lit image I h To the intermediate state illumination image I m The conversion of F θ is an encoder-decoder network with parameters θ.

[0054] The low-light image and the sufficient-light image are aligned through adversarial learning under cycle consistency to update the first intermediate-state illumination image and the second intermediate-state illumination image.

[0055] Right now:

[0056] L adv represents the adversarial loss function, Represents the data distribution from dark domain images (low light images) Medium Sampling I l The expected value of Represents the data distribution from bright area images (high light images) Medium Sampling I h The expected value of , D(·) is the discriminator.

[0057] Feature-level alignment of image semantics is performed based on the updated first intermediate state illumination image and the second intermediate state illumination image to obtain an intermediate state illumination image.

[0058] Consider aligning them at the level of image semantic features, that is, based on the aligned intermediate state image and Perform feature-level alignment, expressed as:

[0059]

[0060] Among them F ψ It is a pre-trained network for extracting high-dimensional features. Represents a mapping function The square of the maximum average difference loss in the reproduction kernel Hilbert space of , which is usually a set of radial basis functions. On the right side of the equation, the first term calculates the square of the maximum average difference loss in the reproduction kernel Hilbert space from N l Sample The second term computes the average value of the mapped features from N in the Hilbert space of the reproduction kernel h Sample By minimizing L MMD , reducing the intermediate state I in the new feature space m The average difference between the two approximations of I m ′, i.e., the intermediate state illumination image, thus promoting the alignment of features.

[0061] The saliency map of the well-illuminated image is generated by spectral residual, and the saliency map is mapped back to the spatial domain through inverse fast Fourier transform, and the salient regions are proposed to obtain a static attention map.

[0062] The target-aware masked representation learning scheme includes a target highlighting strategy, a masking strategy, and a contrast learning scheme.

[0063] The target highlighting strategy should be in the most attractive area corresponding to the human visual attention mechanism. h It has a high spectral residual, so the spectral residual is used to generate the saliency map. The spectral residual can be calculated as R(I h )=L(I h )-L(I h )*H, where R(I h ) represents the input image I h The spectral residual of I h Unique statistical singularity, L(I h ) represents the logarithmic amplitude spectrum, calculated as L(I h )=log(|F(I h )|), where F(·) is I h The Fast Fourier Transform (FFT) of , H is the average spectrum that can be approximated by convolution. The spectral residual highlights the significantly different areas in the image and is therefore more likely to attract visual attention. It is then necessary to use the inverse Fast Fourier Transform to map it back to the spatial domain, resulting in a saliency map of Among them, S(I h ) represents the saliency map obtained after mapping back to the spatial domain, i.e., the saliency map, F -1 represents the inverse fast Fourier transform process, P(F(I h )) is the phase spectrum of the image, The static attention map can be obtained by extracting the salient regions.

[0064] The sufficient illumination image and the intermediate state illumination image are divided into multiple image blocks, and the sufficient illumination image is combined with the static attention map to obtain a first image; the intermediate state illumination image is combined with the static attention map to obtain a second image; the target area and non-target area of ​​the first image and the second image are distinguished, the image blocks in the target area are removed or replaced, and some image blocks in the non-target area are retained or blocked to obtain a first updated image and a second updated image.

[0065] The intermediate state illumination image I m ′ and sufficient illumination image I h The image is divided into multiple blocks; combined with a static attention map, it distinguishes which small blocks are target areas and which are non-target areas. The image is processed using a masking strategy, and the masking strategy is determined based on the importance of different areas in the image. For image blocks within the target area, an erasure strategy is applied, including filling with noise or a fixed value, completely removing or replacing the content of these image blocks, so that the model can learn how to recognize or process images without this information. For image blocks in non-target areas, a missing block strategy is adopted, randomly retaining or lightly occluding parts of the non-target area, retaining some background information, so that the model can capture and learn as many features as possible.

[0066] Based on the first update image and the second update image, a bootstrapped non-negative contrastive learning loss is used and adjusted to minimize the loss.

[0067] The loss function is defined as:

[0068]

[0069] where v and v + Represent the anchor point and positive sample respectively, z is the prediction head, p and p′ are the projection heads, F and F′ are feature extractors, and the gradient is only passed in F and p. In order to track the dynamic changes of the model during training, an exponential moving average is used to update the corresponding points. In the feature representation learning scheme, there are two branches: one is the original model trained on data with sufficient daylight, and the other is I obtained by two-level bidirectional domain alignment. m Since only the anchor branch has gradient, a dual-branch optimization strategy is adopted to obtain the optimal feature representation. The final contrastive learning loss L CL It can be expressed as:

[0070] L CL =L BYOL (I h1 ,I m1 )+L BYOL (I m1 ,I h1 )

[0071] Among them, I h1 is the first updated image, I m1 Update image for the second.

[0072] The above process builds a target perception representation learning framework for high-level visual tasks in low-light environments, which is used to recognize and process low-light images. A two-level bidirectional domain alignment scheme is designed to align image information in bright and dark domains from the two aspects of image appearance and semantic features, integrating data captured under different lighting conditions. A target highlighting strategy is designed to use the saliency mechanism to emphasize task-related targets. A masking strategy is designed to display task-related targets and hide non-targets to filter out interference from task-irrelevant factors during representation learning and focus on the target itself. A contrastive learning scheme is designed to learn how to recognize targets from different angles and scenarios, generalize the model so that it can work in any scenario.

[0073] When the classification type of the real-time image is a rainy image, a second process is performed to form an input image.

[0074] The second processing includes inputting the real-time image into the trained second model and outputting an image without raindrop traces as the input image. The training process of the second model is as follows: Figure 4 Shown, including:

[0075] Step 1: Obtain rainy images;

[0076] Step 2: Perform the following processing to obtain a first output. The processing includes: normalizing the rainy image, passing it through the visual state model, performing element-level multiplication, combining it with a learnable scale factor, and inputting it into a multi-scale convolutional gated unit to obtain the first output through multiple convolutional layers in series and parallel.

[0077] Specifically, it is expressed as:

[0078] X out =VSSM(LN(X))+sX

[0079] Where LN represents layer normalization, which is used to improve the stability and performance of training, s represents a learnable scaling factor, which is used to control the information of the skip connection, X is a rainy image, and VSSM() is a visual state space module (Visual State Space Module), a neural network module for capturing global or long-range dependencies. In the visual state space module, the 2D selective scanning module scans the input two-dimensional feature map from four basic directions (from upper left to lower right, from lower right to upper left, from lower left to upper right, and from upper right to lower left) to capture the features of the image in different directions, and then flattens the image into a one-dimensional sequence. This is because the state space model usually processes sequence data rather than two-dimensional spatial data. Then, the long-range dependencies of each sequence are captured according to the discrete state space equation. Finally, summation is used to merge all sequences, and the input image is restored to a two-dimensional feature map through the reconstruction operation, i.e., X out .

[0080] Data X out It is normalized, and then the number of channels is expanded by 1×1 convolution, and then enters three parallel branches: a gate branch, a 3×3 depth convolution branch and a 5×5 depth convolution branch. In the gate branch, a 3×3 depth convolution is performed first, and then the GELU activation function is performed to obtain X gate In the 3×3 and 5×5 depth convolution branches, 3×3 and 5×5 depth convolutions are performed respectively, and the outputs are X 3×3 and X 5×5 The output results of these branches are integrated and then a 1×1 convolution is performed to restore the number of channels. In addition, we use channel attention to consider the global context information in this block. Finally, a jump connection with a learnable scale factor is added to allow the model to reuse features at different levels, enhancing the feature fusion ability of the model, and obtaining the fused feature map X out ′, which is the first output. In summary, the multi-scale convolutional gated unit can be expressed as:

[0081] X gate =GELU(3×3DW-Conv(LN(X out ))),

[0082] X 3×3 =3×3DW-Conv(LN(X out )),

[0083] X 5×5 =5×5DW-Conv(LN(X out )),

[0084] X out ′=CA(X gate ⊙[X 3×3 ,X 5×5])+s·X out ,

[0085] Where DW-Conv represents depthwise convolution, CA represents channel attention, and [·] represents channel concatenation.

[0086] The above process uses convolution kernels of different sizes to identify raindrop traces of different sizes in the image and remove rain from the perspective of local details.

[0087] In step three, the first output is normalized and passed through the visual state model and fast Fourier transform respectively. The output of the visual state model is combined with the output after the fast Fourier transform through element-wise multiplication, and then combined with the learnable scale factor and input into the multi-scale convolutional gated unit. The second output is obtained through multiple convolutional layers in series and parallel.

[0088] Specifically, it is expressed as follows: Converting the first output information into frequency domain information can separate high-frequency signals representing details and textures from low-frequency signals corresponding to flat areas. Processing these different features separately in the frequency domain can enhance the expressiveness of the model and help remove raindrop traces. The frequency domain state space block uses the visual state space module and the fast Fourier transform module:

[0089] X′ out =VSSM(LN(X out ′))+FFTM(LN(X out ′))+sX out '

[0090] In the Fast Fourier Transform (FFT) module, a 1×1 convolution is initially used to halve the number of channels, and a SiLU activation function (LN()) is used to obtain a feature map suitable for frequency domain processing. Subsequently, a Fast Fourier Transform (FFTM()) is used to transform the feature map from the spatial domain to the frequency domain. Given that image features are represented as real tensors, the symmetry of the two-dimensional real FFT is exploited to reduce computational complexity.

[0091] To further extract features in the frequency domain, filter out the frequency components of raindrop traces, and emphasize the frequency components without rain, a 1×1 convolution with a SiLU activation function is used. Then, a 2D real inverse fast Fourier transform is used to convert the image from the frequency domain back to the spatial domain. Finally, a 1×1 convolution expands the channels to the original number. Given the input feature Z, the fast Fourier transform module is defined as:

[0092]

[0093] Z out =1×1Conv(Z f ),

[0094] Where F represents the two-dimensional real fast Fourier transform, F -1 Stands for two-dimensional real inverse fast Fourier transform.

[0095] The above process performs global rain removal based on the overall structure of the image.

[0096] Afterwards, the image is processed through the same multi-scale convolution gating process as in step 2 to output local rain removal features. Convolution kernels of different sizes are used to identify raindrop traces of different sizes in the image, removing rain from the perspective of local details.

[0097] Step 4: Repeat steps 2 to 3 seven times, each time with different resolution and number of channels.

[0098] After multiple stages, each with a different resolution and number of channels, the final stage yields a high-level feature map. Deconvolution or upsampling restores the original resolution to output an image with raindrop traces removed. For example, this embodiment includes eight stages, from stage 0 to stage 7.

[0099] Step 5: Calculate the loss function L total =L L1 +μ f L Freq , where L L1 is the L1 norm, μ f To adjust the loss ratio coefficient, L Freq is the loss for each frequency component.

[0100] The expression of L1 loss is:

[0101]

[0102] Where ‖·‖1 is the L1 norm. L1 constrains pixel-level consistency, It represents the output result after step 4, and G represents the real image, that is, the ideal target image without rain and noise.

[0103] The frequency reconstruction loss calculates the loss for each frequency component:

[0104]

[0105] where F(·) represents the two-dimensional fast Fourier transform, and the frequency loss ensures the consistency of the frequency domain components.

[0106] The total loss is:

[0107] L total =L L1 +μ f L Freq

[0108] where μ fTo adjust the coefficient of loss ratio, the total loss jointly optimizes the model parameters.

[0109] Through the above process, raindrop traces are gradually removed from the global to the local, and from the spatial domain to the frequency domain, and finally a high-quality rain-free image is generated.

[0110] A frequency-domain state-space model was developed for high-level visual tasks in rainy environments to remove raindrop traces from rainy images. Raindrop traces are often represented by specific frequency components in images. To effectively identify and remove raindrop traces of these high-intensity frequency components, a state-space model and frequency-domain processing were employed simultaneously. Furthermore, a multi-scale convolutional gating unit was developed. This unit uses convolutional kernels of varying sizes to effectively capture raindrop traces of varying sizes and integrates a gating mechanism to manage information flow.

[0111] When the classification type of the real-time image is a normal image, no processing is performed to form the input image.

[0112] Create a vector map based on the input image, such as Figure 5 Shown, including:

[0113] The input image is extracted through a feature encoder to obtain multi-scale perspective features.

[0114] Perform multi-scale foreground enhancement processing to obtain the masked multi-scale perspective image features. Specifically including:

[0115] Generate foreground mask based on multi-scale perspective features through multi-layer perceptron.

[0116] Specifically, convolution and activation operations are performed on the multi-scale perspective feature map to generate a foreground mask, which is expressed as:

[0117]

[0118] in, represents the foreground mask, F i Represents multi-scale perspective features, Conv() is a convolutional layer, ReLU() is an activation function, and Sigmoid() is an activation function. The generated foreground mask corresponds to the input multi-scale perspective features and is used to indicate whether each position belongs to the foreground.

[0119] The foreground mask and multi-scale perspective features are connected through residuals, and the feature weights of different image positions are adjusted by element-wise multiplication to obtain the masked multi-scale perspective features.

[0120] Specifically expressed as:

[0121]

[0122] in is the masked multi-scale perspective feature, is the foreground mask, F i The multi-scale perspective features are obtained by element-wise multiplication of the foreground mask with the original multi-scale perspective features, enhancing foreground features and suppressing background features. The resulting mask is added to the original multi-scale perspective features and combined with confidence information to enrich the feature set, improving the model's foreground recognition and robustness. This process classifies the features into road features and non-road features.

[0123] Based on the masked multi-scale perspective image features, we integrate high-level prior knowledge and multi-scale features to obtain multi-scale bird's-eye view features. Specifically, we include:

[0124] The masked multi-scale perspective features are positionally encoded and average-pooled to generate global features.

[0125] Perform global abstraction of image features to establish a foundational layer for subsequent height distribution probability modeling.

[0126] Global features are integrated into predefined queries and the highly distributed probabilities are calculated through a multi-layer perceptron.

[0127] Specifically expressed as:

[0128]

[0129] Among them, Height pro is the height distribution probability, MLP() is the multi-layer perceptron, AP() is the average pooling, PE is the position encoding, and Query is the predefined query vector.

[0130] The predefined reference points close to the road are obtained from the bird's-eye view, and the reference points are projected into the multi-scale perspective space to obtain the multi-scale height distribution features.

[0131] Use predefined reference points closer to the road from a bird's eye view These reference points are projected into a multi-scale perspective space using the camera’s intrinsic and extrinsic parameters to effectively capture features at multiple scales.

[0132]

[0133] in, Sampling is a multi-scale height distribution feature. i () is the feature sampling from the projected image features at the i-th scale, I is the intrinsic parameter matrix of the camera, K is the extrinsic parameter matrix of the camera, These are predefined reference points close to the road from a bird's-eye view perspective.

[0134] Multi-scale bird's-eye view features are obtained based on height distribution probability and multi-scale height distribution features.

[0135] These image features with known height distribution probabilities are weighted pooled along the z-axis to obtain the final bird's-eye view features

[0136]

[0137] Among them, AP() is average pooling, through which the multi-scale perspective features of road elements are converted into bird's-eye view features, laying the foundation for the subsequent bird's-eye view feature dynamic fusion unit.

[0138] Based on the multi-scale bird's-eye view features, the features of the bird's-eye view are dynamically fused to obtain a vector map. Specifically, it includes:

[0139] Splice bird's-eye view features of different scales to form a coarse integrated feature.

[0140] This step provides a multi-scale feature input for subsequent feature processing.

[0141] After the coarsely integrated features pass through the neural network, deformable attention is applied to dynamically adjust the feature focus, selectively enhance the key feature areas, refine the features through a multi-layer perceptron, and perform feature fusion with the output residual connection of the deformable attention to obtain the fused bird's-eye view features.

[0142] A specific convolutional neural network is applied to process the features, enabling the model to gain a deeper understanding of spatial dynamics, thereby providing a more detailed interpretation of the road environment. A deformable attention mechanism is then applied to enable the model to selectively enhance key feature areas and dynamically adjust the model's focus between different feature areas to ensure the pertinence of the feature fusion process. After the fusion process, a multi-layer perceptron is applied to fine-tune the features. The multi-layer perceptron integrates the features using a series of linear layers and enhances them by strategically placing residual connections, ensuring that the resulting bird's-eye view feature F bev′ Both rich and stable.

[0143]

[0144] Among them, F bev′ is the fused bird's-eye view feature, Conv() is the convolution layer, MLP() is the multi-layer perceptron, DA() is the deformable attention, Concat() is the feature concatenation, and AP() is the average pooling.

[0145] A feature decoder is used to convert the fused bird’s-eye view features into a vectorized scene representation, i.e., the final prediction result, which should contain key static road elements such as dividers, boundaries, and crosswalks.

[0146] This process integrates bird's-eye view features captured at different scales, improving the online vector map construction model's ability to understand and interpret the road environment.

[0147] The fused bird's-eye view features are converted into real-time vectorized scenes through a feature decoder.

[0148] The fused bird's-eye view features are processed, including restoring the resolution using deconvolution or upsampling, to generate a vectorized scene representation to output a vector map.

[0149] By the classification loss L cls , point-to-point loss L pos , edge direction loss L dir , mask loss L mask The four key losses are composed of λ, α, β and γ, which represent the weight coefficients corresponding to the above losses, as shown below:

[0150] L=λL cls +αL pos +βL dir +γL mask

[0151] The above process significantly improves the perception capability of the autonomous driving system in low-light and rainy conditions. These technologies enable the system to more accurately identify and process key road features, such as lane lines and road signs, and maintain high performance even in adverse weather conditions. The model's ability to understand and interpret the road environment is improved, and by reducing unnecessary calculations, real-time performance is improved, allowing the system to respond quickly and effectively utilize computing resources. Accurate environmental perception is crucial to avoiding accidents and ensuring driving safety. The present invention enhances the safety and reliability of the autonomous driving system by improving the perception accuracy under these conditions. The line vector map construction method significantly improves the perception and response capabilities of the autonomous driving system in complex environments such as low light and rain.

[0152] Example 2

[0153] This embodiment provides an online vector map construction system for low-light and rainy scenes.

[0154] The online vector map construction system for low-light and rainfall scenes includes:

[0155] The image acquisition and classification module acquires real-time images and classifies the real-time images based on an image classifier, wherein the real-time images include multi-view input images.

[0156] The low-light image processing module performs a first processing to form an input image when the classification type of the real-time image is a low-light image.

[0157] The rain image processing module performs a second processing to form an input image when the classification type of the real-time image is a rain image.

[0158] Normal image processing module, when the classification type of the real-time image is a normal image, no processing is performed to form an input image.

[0159] The vector map establishment module establishes a vector map based on the input image, including extracting the input image through a feature encoder to obtain multi-scale perspective features; generating a foreground mask through a multi-layer perceptron based on the multi-scale perspective features; connecting the foreground mask and the multi-scale perspective features through residuals, and adjusting the feature weights of different positions of the image by element-level multiplication to integrate them to obtain masked multi-scale perspective features; position encoding and average pooling of the masked multi-scale perspective features to generate global features; integrating the global features into a predefined query, and calculating the height distribution probability through a multi-layer perceptron; obtaining a predefined bird's-eye view perspective. A reference point close to the road is projected into the multi-scale perspective image space to obtain multi-scale height distribution features; multi-scale bird's-eye view features are obtained based on the height distribution probability and the multi-scale height distribution features; bird's-eye view features of different scales are spliced ​​to form a coarse integrated feature; after the coarse integrated feature passes through the neural network, deformable attention is applied to dynamically adjust the feature focus, and key feature areas are selectively enhanced. The features are refined through a multi-layer perceptron and fused with the residual connection of the output of the deformable attention to obtain the fused bird's-eye view features; the fused bird's-eye view features are converted into a real-time vectorized scene through a feature decoder.

[0160] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for constructing an online vector map for low-light and rainfall scenes, characterized by: include: Acquire a real-time image and classify the real-time image based on an image classifier, wherein the real-time image includes a multi-view image; Perform different processing according to different classification types to obtain the input image; When the classification type of the real-time image is a low-light image, performing a first process; When the classification type of the real-time image is a rainy image, performing the second processing; When the classification type of the real-time image is normal image, no processing is performed; Build a vector map based on the input image, including, The input image is extracted through a feature encoder to obtain multi-scale perspective features; Performing multi-scale foreground enhancement processing on the multi-scale perspective features to obtain masked multi-scale perspective features includes: generating a foreground mask based on the multi-scale perspective features through a multi-layer perceptron; connecting the foreground mask and the multi-scale perspective features through residuals, and adjusting feature weights at different positions of the image through element-wise multiplication to obtain the masked multi-scale perspective features; Based on the masked multi-scale perspective image features, the high-level prior knowledge and multi-scale features are integrated to obtain the multi-scale bird's-eye view features; Based on multi-scale bird's-eye view features, dynamic fusion processing of bird's-eye view features is performed, including: splicing bird's-eye view features of different scales to form a coarse integrated feature; after the coarse integrated feature passes through the convolutional neural network, deformable attention is used to dynamically adjust the feature focus, selectively enhance key feature areas, and refine the features through a multi-layer perceptron. The residual connection is then made with the output of the deformable attention module to achieve feature fusion, resulting in the fused bird's-eye view feature; The fused bird's-eye view features are converted into real-time vectorized scenes through a feature decoder.

2. The online vector map construction method for low-light and rainfall scenes according to claim 1 is characterized in that: The masked multi-scale perspective view feature is based on integrating height prior knowledge and multi-scale features to obtain multi-scale bird's-eye view features, including: position encoding and average pooling of the masked multi-scale perspective view features to generate global features; integrating the global features into a predefined query and calculating the height distribution probability through a multi-layer perceptron; obtaining pre-defined reference points close to the road from the bird's-eye view perspective, projecting the reference points into the multi-scale perspective view space to obtain multi-scale height distribution features; and obtaining the multi-scale bird's-eye view features based on the height distribution probability and the multi-scale height distribution features.

3. The online vector map construction method for low-light and rainfall scenes according to claim 1 is characterized in that: The foreground mask is generated by a multi-layer perceptron based on the multi-scale perspective feature, which is expressed as: in, represents the foreground mask, F i Represents multi-scale perspective features, Conv() is a convolutional layer, ReLU() is an activation function, and Sigmoid() is an activation function.

4. The online vector map construction method for low-light and rainfall scenes according to claim 1 is characterized in that: The foreground mask and the multi-scale perspective feature are connected by residuals, and the feature weights of different positions of the image are adjusted by element-wise multiplication to obtain the masked multi-scale perspective feature, which is expressed as: in is the masked multi-scale perspective feature, is the foreground mask, F i is the multi-scale perspective feature, and ⊙ is the element-wise multiplication.

5. The online vector map construction method for low-light and rainfall scenes according to claim 2 is characterized in that: The global features are integrated into the predefined query and the height distribution probability is calculated by the multi-layer perceptron, which is expressed as: Among them, Height pro is the height distribution probability, MLP() is the multi-layer perceptron, AP() is the average pooling, PE is the position encoding, and Query is the predefined query vector.

6. The online vector map construction method for low-light and rainfall scenes according to claim 1, characterized in that: The method obtains a predefined reference point close to the road from a bird's-eye view, projects the reference point into a multi-scale perspective image space, and obtains a multi-scale height distribution feature, which is expressed as: in, Sampling is a multi-scale height distribution feature. i () is the feature sampling from the projected image features at the i-th scale, I is the intrinsic parameter matrix of the camera, K is the extrinsic parameter matrix of the camera, A predefined reference point close to the road from a bird's-eye view perspective; The multi-scale bird's-eye view feature is obtained based on the height distribution probability and the multi-scale height distribution feature, which is expressed as: AP() is average pooling, is the i-th bird's-eye view feature.

7. The online vector map construction method for low-light and rainfall scenes according to claim 1 is characterized in that: After the coarse integrated features pass through the neural network, the deformable attention is applied to dynamically adjust the feature focus, selectively enhance the key feature areas, and refine the features through the multi-layer perceptron. The features are then fused with the output residual of the deformable attention. The fused bird's-eye view feature is represented as follows: Among them, F bev′ is the fused bird's-eye view feature, Conv() is the convolution layer, MLP() is the multi-layer perceptron, DA() is the deformable attention, and Concat() is the feature integration.

8. The online vector map construction method for low-light and rainfall scenes according to claim 1, characterized in that: The first processing includes inputting a real-time image into a trained first model and outputting an image with normal lighting as an input image; The training process of the first model includes: Acquire low-light and bright-light images; converting the low-light image into a first intermediate-state light image by visual enhancement; Converting the sufficient illumination image into a second intermediate state illumination image through a preset network; Aligning the low-light image and the sufficient-light image through adversarial learning under cycle consistency to update the first intermediate-state illumination image and the second intermediate-state illumination image; Performing feature-level alignment of image semantics based on the updated first intermediate state illumination image and the second intermediate state illumination image to obtain an intermediate state illumination image; The saliency map of the well-illuminated image is generated by spectral residual, and the saliency map is mapped back to the spatial domain through inverse fast Fourier transform, and the salient region is proposed to obtain a static attention map; The sufficient illumination image and the intermediate state illumination image are divided into multiple image blocks, and the sufficient illumination image is combined with the static attention map to obtain a first image; the intermediate state illumination image is combined with the static attention map to obtain a second image; the target area and non-target area of ​​the first image and the second image are distinguished, the image blocks in the target area are removed or replaced, and the image blocks in the non-target area are retained or blocked to obtain a first updated image and a second updated image; Based on the first update image and the second update image, a bootstrapped non-negative contrastive learning loss is used and adjusted to minimize the loss.

9. The online vector map construction method for low-light and rainfall scenes according to claim 1, characterized in that: The second processing includes inputting the real-time image into the trained second model and outputting an image without raindrop traces as the input image; The training process of the second model includes: Step 1: Obtain rainy images; Step 2: Perform the following processing to obtain a first output, the processing including: normalizing the rainy image, passing it through the visual state model, performing element-wise multiplication, combining it with a learnable scale factor, and inputting it into a multi-scale convolutional gating unit to obtain the first output through multiple convolutional layers in series and parallel. Step 3: The first output is normalized and passed through the visual state model and fast Fourier transform respectively. The output of the visual state model is combined with the output after fast Fourier transform through element-wise multiplication, and then combined with the learnable scale factor and input into the multi-scale convolutional gated unit. The second output is obtained through multiple convolutional layers in series and parallel. Step 4: Repeat steps 2 to 3 seven times, each time with different resolution and number of channels; Step 5: Calculate the loss function L total =L L1 +μ f L Freq , where L L1 is the L1 norm, μ f To adjust the loss ratio coefficient, L Freq is the loss for each frequency component.

10. The online vector map construction method for low-light and rainfall scenes according to claim 9, characterized in that: The input is input into the multi-scale convolution gate unit, including normalizing the input value, expanding the number of channels through 1×1 convolution, and then entering three parallel branches: a gate branch, a 3×3 depth convolution branch and a 5×5 depth convolution branch. In the gate branch, a 3×3 depth convolution is performed first, and then the GeLU activation function is performed to obtain the first data. In the 3×3 and 5×5 depth convolution branches, 3×3 and 5×5 depth convolutions are performed respectively, and the second and third data are output respectively. The first data, second data and third data are integrated, and then a 1×1 convolution is performed to restore the number of channels. A jump connection with a learnable scale factor is added and then output.