Model and method for automatic driving image segmentation under challenging weather conditions
By adopting the encoder-decoder network architecture and multiple attention modules in the autonomous driving system, the problem of insufficient image segmentation reliability under severe weather conditions is solved, and efficient identification and segmentation of objects in driving scenarios is achieved.
Patent Information
- Application Number
- CN202510106192.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
AI Technical Summary
In the severe weather conditions, different lighting scenarios and complex situations, the existing autonomous driving system has insufficient reliability in image segmentation, making it difficult to effectively identify and segment objects in driving scenarios.
The encoder-decoder network architecture is adopted, combining coordinate attention, triple attention and prospective attention modules to enhance the encoding and decoding capabilities of image features, capture complex interactions between space and channel dimensions, and improve the robustness of image segmentation.
Under extremely harsh weather conditions, the model can effectively capture subtle features such as fog, snow and rain, improve object recognition performance, and achieve clear segmentation and identification of buildings, cars and sidewalks.
Smart Images

Figure CN119992094A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving image processing technology, and in particular to a model and method for autonomous driving image segmentation under challenging weather conditions. Background Art
[0002] Typically, autonomous driving systems consist of two key stages. The first stage is related to perception and localization, which involves extracting salient features that represent the driving environment, allowing the autonomous car to understand and interpret its surroundings. The subsequent stages are centered around decision-making and are further divided into two parts: motion planner and trajectory controller. The motion planner is tasked with developing appropriate action plans, while the trajectory controller is responsible for overseeing the execution of these plans, ensuring that the vehicle drives smoothly and efficiently.
[0003] To successfully complete the scene understanding task, autonomous vehicles must be able to discern the semantic segmentation of each pixel in the received image signal. This complex process involves classifying each pixel according to the corresponding segmentation label, and in the field of computer vision and autonomous driving systems, this challenge is often referred to as "semantic segmentation". Semantic segmentation plays a pivotal role in promoting the comprehensive understanding of the scene by robotic systems. With the maturity of image segmentation technology and its wide application in various sectors of the transportation field, it has become an important tool for detecting and identifying the number of motor vehicles in driving scenes.
[0004] However, the reliability of these systems is still unsatisfactory in complex situations such as adverse weather conditions, varying lighting scenarios, and the presence of a large number of overlapping objects. To improve the reliability of the systems, especially in challenging driving environments, it is essential to exploit the spectral reflectance properties of different objects in the driving scene (beyond the visible spectrum). This additional information has the potential to greatly enhance the robustness of these systems. Furthermore, this spectral data is essential for the development of advanced vision systems that can more deeply understand and interpret the entire driving scene. Summary of the invention
[0005] In view of the problems existing in the prior art, the present invention provides a model for autonomous driving image segmentation under challenging weather conditions, wherein the model includes an encoder and a decoder corresponding to the encoder;
[0006] The encoder comprises an input module and four encoder layers in sequence, wherein a coordinate attention module is connected between the input module and the encoder layer connected thereto and between adjacent encoder layers;
[0007] The decoder comprises four decoder layers and an output module in sequence, wherein a prospective attention module is connected between the output module and the decoder layer connected thereto and between adjacent decoder layers;
[0008] A triple attention module is connected between the encoder and the encoder.
[0009] Based on the above scheme, there are four triple attention modules, which are connected before each encoder layer and after the decoder layer respectively.
[0010] Based on the above scheme, in the coordinate attention module, the input is H-pooled and W-pooled respectively, and then the feature maps H and W obtained after pooling are spliced to obtain HW; after splicing, the feature maps are successively convolved, batch normalized and activated to obtain HWC, HWB and HWA respectively; then segmented to obtain SH and SW; the two results are respectively subjected to convolution and nonlinear activation functions to obtain SHS and SWS; finally, SHS and SWS are vector multiplied.
[0011] Based on the above scheme, the triple attention module adopts triple parallel processing branches, where:
[0012] In the first branch, the input tensor is rotated 90 degrees counterclockwise along the height axis to obtain a rotated tensor; the rotated tensor passes through the Z-pooling layer, convolution layer, batch normalization layer, and activation function layer in sequence to obtain the attention weight; then the attention weight and the rotation tensor are vector-multiplied and the result is rotated exactly 90 degrees clockwise along the height axis;
[0013] In the second branch, the input tensor is rotated 90 degrees counterclockwise along the W axis to obtain a rotated tensor; the rotated tensor passes through the Z-pooling layer, convolution layer, batch normalization layer, and activation function layer in sequence to obtain the attention weight; then the attention weight and the rotation tensor are vector-multiplied, and the result is rotated exactly 90 degrees clockwise along the W axis;
[0014] In the third branch, the input tensor is first reduced in dimension by channel pooling; then, the tensor is sent to the convolution layer, batch normalization layer, and activation function layer in sequence to generate attention weights, which are then multiplied by the original input tensor;
[0015] The results in the three branches are added and averaged to combine the information from the three branches into a single comprehensive feature.
[0016] Based on the above scheme, in the prospective attention module, the input X is projected through two different linear layers, one is a linear projection, and the other is a linear + anti-folding projection; finally, the results of the two linear layer projections are multiplied to obtain the folded weighted average.
[0017] Based on the above scheme, the results of the two linear layer projections are multiplied to obtain the formula for the folded weighted average:
[0018]
[0019] The present invention also provides a method for image segmentation for autonomous driving under challenging weather conditions, using the above-mentioned model.
[0020] The present invention also provides a server, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the above method steps are implemented when the processor executes the computer program.
[0021] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and is characterized in that the computer program implements the above method steps when executed by a processor.
[0022] Beneficial effects of the present invention:
[0023] The model of the present invention is specifically targeted at autonomous driving scene graph segmentation in extreme weather conditions. It uses "coordinate attention" to carefully encode spatial information, combines "triple attention" to capture the intricate cross-dimensional interactions between spatial and channel dimensions in the input, and uses "lookahead attention" to encode fine-grained features and contextual nuances as tags. This comprehensive encoding strategy is very beneficial to improving recognition performance. In scenes affected by weather, the model of the present invention is able to cleverly capture subtle features such as fog, snow, and rain, which helps to recognize objects in these challenging environments.
[0024] Through actual tests, it is proved that the method of the present invention can clearly segment the input image and identify buildings, cars and sidewalks. In terms of prediction, it is very close to the ground truth. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is the overall architecture diagram of the model of the present invention;
[0026] Figure 2 This is the architecture diagram of the coordinate attention module in the model of the present invention;
[0027] Figure 3 The architecture diagram of the triple attention module in the model of the present invention;
[0028] Figure 4 Schematic diagram of the prospective attention module in the model of the present invention;
[0029] Figure 5 The training loss effect of the model in Example 2 of the present invention using the urban landscape set;
[0030] Figure 6 This is a rendering of the change of mIoU of the model on the urban landscape set with Epoch in Example 2 of the present invention;
[0031] Figure 7 This is a visualization diagram of the model in Example 2 of the present invention on the urban landscape dataset;
[0032] Figure 8 The training loss effect of the model in Example 2 of the present invention using the Shandong highway landscape data set;
[0033] Fig. 9 This is a diagram showing the effect of the model in Example 2 of the present invention on the Shandong Province highway landscape dataset on the change of mIoU with Epoch;
[0034] Fig.10 This is a visualization diagram of the model in Example 2 of the present invention on the Shandong highway landscape dataset. DETAILED DESCRIPTION
[0035] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation mode, structure, characteristics and effects of the present invention are described in detail below in combination with the accompanying drawings and preferred embodiments.
[0036] Example 1
[0037] The proposed model for autonomous driving image segmentation in challenging weather conditions uses an encoder-decoder network architecture, named SceneSegAttentionNet, which is tailored for autonomous driving scene segmentation in extreme weather conditions. The proposed network aims to improve the accuracy and robustness of scene perception in complex environments, thereby improving the overall safety and performance of autonomous vehicles.
[0038] SceneSegAttentionNet consists of an encoder and a decoder. There is an attention module between each encoder layer and decoder layer, which is located after the encoder and connected to the decoder. The encoder is composed of a pre-trained ResNet18 (the ResNet18 pre-training model is based on the ImageNet picture set, which contains more than 1 million annotated pictures covering thousands of categories. It uses a large number of computing resources, such as GPU or TPU clusters, for training. During the training process, a mini-batch stochastic gradient descent (Mini-batch SGD) is used to update the network weight training until the performance of the model on the validation set is no longer significantly improved, or a predetermined number of iterations is reached. After the training is completed, the weights of the model are saved and released as a pre-trained model.). Each encoder layer in this structure is paired with the corresponding decoder layer to ensure coordinated and orderly data flow. Finally, the output of the decoder is imported into an output module equipped with a deconvolution function to perform accurate pixel prediction.
[0039] The overall architecture of SceneSegAttentionNet is as follows Figure 1 As shown in the figure. First, the ResNet18 backbone network serves as the foundation to accurately extract complex image feature representations. Subsequently, an encoder-decoder network architecture is integrated to facilitate the efficient transformation and refinement of these features. Finally, a set of output modules completes the process, providing accurate detection predictions that cover the essence of the visual scene.
[0040] The core of our proposed model is the encoder-decoder architecture network, which is carefully designed as two independent components: encoder and decoder. In particular, a triple attention module is inserted between the encoder and decoder, which is placed after the encoder and seamlessly connected to the decoder. The encoder uses a pre-trained ResNet1 8-layer as its main encoding block, ensuring strong feature extraction capabilities.
[0041] In the encoder hierarchy, a coordinate attention layer is carefully integrated after the initial input module and each encoder layer (layers 1, 2, and 3) to enhance the spatial awareness in the encoder representation. Conversely, a look-ahead attention mechanism is introduced before each decoder layer (decoder in layers 1, 2, and 3) of the decoder to provide a forward-looking perspective to the encoding process.
[0042] Crucially, in addition to the coordinate attention, triple attention is also extracted after each encoder layer. The attention extracted from the encoder is then concatenated with the output of each decoder as the key input of their respective prospective attention modules. This innovative integration of multiple attention mechanisms highlights the complex design of SceneSegAttentionNet, enabling it to achieve superior performance in scene segmentation tasks, even under extreme weather conditions.
[0043] Specifically, the input image vector first passes through the input module, producing an intermediate representation value, denoted as X. Subsequently, in the encoder path, X is input into the first coordinate attention module 1 to generate CA1, which is then processed by encoder layer 1 to generate output E1. This process is repeated in each subsequent stage, and E1 is used as the input of coordinate attention 2 to generate CA2, which is then processed by encoder layer 2 to generate E2. Similarly, E2 is passed to coordinate attention 3 to generate CA3, which is then passed to encoder layer 3 to generate output E3. Finally, E3 undergoes the same process as coordinate attention 4 to generate CA4, which is then input to encoder layer 4 to generate the final encoder layer output E4.
[0044] The final output of the encoder, E4, is directly passed to the decoder layer 4 to generate the preliminary decoder layer output D4. At the same time, E3 is processed by the triple attention module 4, and the processing result is concatenated with D4 to obtain TD4. Subsequently, TD4 is input to the prospective attention module 4 to generate OA4, and is passed as input to the decoder layer 3 to produce D3. This process is repeated in a cascade manner. For E2 and E1, E2 is processed by the triple attention module 3 and concatenated with D3 to form TD3, and then passes through the prospective attention module 3 to generate OA3, and finally obtains D2. Similarly, E1 is processed by the triple attention module 2, concatenated with D2 to generate TD2, and passes through the prospective attention module 2 to obtain OA2, and finally generates D1. Finally, the initial input X is processed by the triple attention module 1, and the output is concatenated with D1 to generate TD1. TD1 is then processed by the prospective attention module 1 to obtain OA1, and is finally passed to the output module to generate the prediction map.
[0045] After each encoder, the "Coordinate Attention" function is combined to capture both local features and surrounding features, and further enhance them with global contextual information. This mechanism ensures that the encoded representation is rich in details and contextual features. Between the encoder and decoder, the network uses a triple attention technique to connect their respective feature maps to each other. This helps to exploit the spatial details of the encoder and give the decoder a deeper understanding of the spatial nuances of the encoded features. The output of each decoder is carefully connected with the corresponding triple attention, which promotes seamless fusion of features. This fusion process allows the two branches to promote each other and complement each other's strengths, resulting in a stronger and more comprehensive feature set. Finally, these fused features are fed into the "Look Ahead Attention" module, where they are further refined to generate richer and more discriminative feature representations. This multi-faceted attention mechanism highlights the complexity of "SceneSegAttentionNet", enabling it to achieve excellent performance in scene segmentation tasks, especially under challenging conditions.
[0046] As a specific implementation, the coordinate attention is implemented after the input module and each encoder layer (encoder layer 1, encoder layer 2, and encoder layer 3). Conceptually, the coordinate attention module can be viewed as a complex computational unit whose task is to enhance the expressiveness of the learned features. It accepts an intermediate feature tensor (denoted as X = [x1, x2, ..., xC] ∈ R C×H×W as input) and produces a transformed tensor (Y = [y1, y2, ..., yC]) with the same dimension as X but richer features.
[0047] The coordinate attention module combines channel relations and long-range dependencies, enriched by carefully designed position information. This integration is achieved through a carefully designed two-stage process: first, coordinate information embedding ensures that spatial nuances are captured; then, the generated coordinate attention further refines and enhances these representations, promoting a comprehensive understanding of the visual context.
[0048] Specifically, given an input tensor X, two different spatial poolings, H-pooling and W-pooling, are used, with sizes (H, 1) and (1, W) respectively. Spatial pooling can encode the features of each channel along the horizontal and vertical coordinates, thus gaining a more detailed understanding of the spatial layout. Therefore, for the c-th channel with a height of h, the output of the pooling operation can be mathematically expressed as follows:
[0049]
[0050] Similarly, the output of the pooling operation corresponding to the cth channel with a width of w can also be expressed by a similar formula as follows:
[0051]
[0052] The above-mentioned pooling operation accurately aggregates features along two orthogonal spatial dimensions through dual transformations to generate a pair of direction-sensitive feature maps. This approach is in sharp contrast to the compression operation in traditional channel attention methods, which usually compresses spatial information into a single feature vector. It is worth noting that the attention module constructed by the present invention, with the support of these two transformations, can capture long-range dependencies along one spatial axis while accurately retaining detailed position clues on the orthogonal axis. This unique ability enables the network to accurately locate objects of interest in complex visual scenes, thereby improving overall performance in complex scene analysis.
[0053] Formulas (1) and (2) help to understand and preserve the global receptive field and encode subtle location information. On this basis, we introduce a key second transformation: the generation of coordinate attention.
[0054] Specifically, after obtaining the feature maps H and W obtained by equations (1) and (2), we concatenate them to fuse the encoding information in two directions, and the result is HW. Then, the fused feature map HW is subjected to shared 1x1 convolution (Conv) and batch normalization (BatchNorm), as shown below:
[0055] f=δ(F1([z h , z w ])) (3)
[0056] As mentioned above, the symbol [-, -] indicates the concatenation operation along the spatial dimension; F1 indicates the operation of Conv and BatchNorm. δ indicates the nonlinear activation function.
[0057] Here, the nonlinear activation function (Activation, represented by δ) introduces nonlinear calculations and enhances the representation ability of the network. The resulting intermediate feature map ( Figure 2 is HWA), f∈R C / r×(H+W) The horizontal and vertical spatial information is fully encoded, where r represents the reduction ratio, which is a parameter that controls the dimension of the block to improve computational efficiency.
[0058] Subsequently, HWA is carefully split along the spatial dimension into two different tensors, namely SH and f h ∈R C / r×H and SW, f w ∈R C / r×W , each focusing on horizontal or vertical position clues. In order to make the channel dimensions of these tensors consistent with the input X, two additional 1×1 convolution (Conv) transformations are implemented, namely F h and F w , and the result is Figure 2 Then, SHC and SWC are respectively subjected to a nonlinear activation function (Sigmoid, represented by δ) to obtain Figure 2 SHS and SWS in.
[0059] In order to complete SHS and SWS and ensure that they are compatible with the input channel X, thereby promoting seamless integration within the network architecture, vector multiplication is performed on X and SHS and SWS to obtain the output.
[0060] The results of these transformations are as follows:
[0061] g h =σ(F h (f h )) (4)
[0062] g w =σ(F w (f w )) (5)
[0063]
[0064] In the above formula, f h indicates cues that focus on horizontal position, f w Indicates cues that focus on vertical position. h and F wrepresents an additional 1×1 convolution (Conv) transformation. δ represents a nonlinear activation function (Sigmoid). The output result is Y, and c represents a channel. g h express Figure 2 SHS in g w express Figure 2 SWS in y c express Figure 2 In the output, i and j represent the elements in the i-th row and j-th column respectively.
[0065] The coordinate attention module incorporates the encoding of spatial information into its framework. The attention mechanism acts on both horizontal and vertical dimensions, acting on the input tensor. In particular, each component of the dual attention map acts as a beacon, indicating whether the queried object lies within its corresponding row or column. This complex encoding process not only enhances the ability of coordinate attention in accurately locating the target object, but also improves the recognition ability of the overall model.
[0066] As a specific implementation, the triple attention module demonstrates the ability of cross-dimensional interaction through a three-branch structure. When facing the input tensor, the first and second branches of the triple attention first perform a rotation (Permute) operation and then perform a pooling operation, thereby cleverly constructing a cross-dimensional dependency relationship. This design not only promotes interaction between different dimensions, but also effectively encodes inter-channel information and spatial information.
[0067] like Figure 3 As shown in Figure 1, the triple attention module adopts a three-pronged parallel processing branch approach. Among them, two different branches play a key role in capturing the intricate cross-dimensional interactions between the channel dimension (denoted as C) and the spatial dimension (H or W). This two-pronged approach ensures that the interactions between these different but interdependent fields are comprehensively studied. As a complement to these two branches, the remaining last branch carefully constructs spatial attention to provide a focusing lens for the intricate spatial relationships in the input data. The harmonious integration of the three branches is further promoted by a simple aggregation strategy, in which the output of each branch is integrated through a simple average.
[0068] The task of the three branches of the triple attention module is to capture the underlying dependencies between (C, H), (C, W), and (H, W) in the input tensor that span different dimensions but are interrelated.
[0069] The Z-pooling layer compresses the zero dimension of a tensor into a two-dimensional representation. This is achieved by carefully merging the results of average and max pooling operations along this dimension, thereby promoting a more nuanced and comprehensive understanding of the data. This strategy not only ensures that the depth and subtle representation of the original tensor are preserved, but also cleverly reduces its depth, thereby reducing the computational burden of subsequent processing stages. This complex process can be summarized by the following equation:
[0070] Z-pool(χ)=[MaxPool 0d (χ), AvgPool 0d (χ)] (7)
[0071] In the Z pooling layer, the symbol "0d" in the above formula is a key marker, indicating the zero dimension on which operations such as the maximum pooling (MaxPool) and the average pooling (AvgPool) are performed. This symbol emphasizes the importance of this dimension in the operation of this layer and highlights the accuracy and meticulousness of the execution of the pooling function. A tensor with a shape of (C×H×W) undergoes the Z pooling layer transformation process and finally generates a transformed tensor with a shape of (2×H×W).
[0072] Triple Attention is a modular architecture consisting of three different branches. It receives a shape of χ∈R C ×H×W The input tensor is processed by these three branches and outputs a refined tensor of the same shape. Essentially, the input tensor is first processed by the three branches in the proposed triple attention module.
[0073] Specifically, Figure 3 As shown in Figure 1, the first branch of the triple attention module is to establish an interaction between the height and channel dimensions of the input tensor. Specifically, the input tensor χ is rotated (Permute) 90 degrees counterclockwise along the height axis to obtain a transformed tensor χ1 with a size of (W×H×C). The rotated tensor then passes through a Z-pooling layer to reduce its size to the shape of χ1 (2×H×C). Subsequently, χ1 is processed by a convolutional layer (Conv) with a kernel size of k×k, and then by a batch normalization layer (BatchNorm). This series of operations produces an intermediate output tensor with a dimension of (1×H×C). In order to generate attention weights, the intermediate output tensor passes through a Sigmoid activation function layer (denoted as σ) to obtain the attention weights.
[0074] Subsequently, these attention weights are vector-multiplied with the original rotation tensor χ1, and the result is precisely rotated (Permute) 90 degrees clockwise along the height axis. This transformation restores the input tensor χ to its original shape, ensuring that salient features are accurately preserved and enhanced.
[0075] In the second branch, the tensor χ is rotated (Permute) 90 degrees counterclockwise along the W axis to become a tensor χ2. This rotated tensor of size (H×C×W) then enters the Z-pooling layer to restore it to a tensor χ2 of shape (2×C×W). Subsequently, χ2 passes through a convolutional (Conv) layer with a kernel size of k×k. Subsequently, the batch normalization layer (BatchNorm) further processes the tensor to produce an output of shape (1×C×W). To obtain the attention weights, this refined tensor is strategically guided through a Sigmoid activation function layer (σ) to ensure a smooth and accurate transformation of the underlying features, thus obtaining the attention weights. Subsequently, the attention weights and χ2 are vector-multiplied, and the result is precisely rotated (Permute) 90 degrees clockwise along the W axis to ensure that its shape is consistent with the original input tensor, thereby maintaining the integrity of the data χ.
[0076] In the third branch, the input tensor χ is subjected to dimensionality reduction by channel pooling, which reduces its channels to two. Channel pooling is an operation that aggregates or selects channels of feature maps in deep learning models. Its purpose is to reduce the number of channels of feature maps while retaining the most important information. This results in a reduced tensor χ3 of shape (2×H×W). This tensor is then fed into a convolutional (Conv) layer with a kernel size of k to extract meaningful features. This is followed by a batch normalization layer (BatchNorm), which normalizes the input data of each mini-batch to make the distribution of inputs to each layer more stable, which helps in higher learning rates without causing training instability. The refined output of this normalization layer is then passed through a sigmoid activation function layer (σ) to produce attention weights (1×H×W), thereby emphasizing salient features and promoting accurate predictions. These attention weights are then multiplied with the original input tensor χ to produce a refined tensor of shape (C×H×W).
[0077] For each of the three branches, a refinement tensor of shape (C×H×W) is generated. These refinement tensors are then summarized through a simple additive (+) averaging process to combine the information from all three branches into a single comprehensive feature.
[0078] In summary, driven by the “triple attention” mechanism, for the input tensor χ∈R C×H×W The process of obtaining the refined attention application tensor y can be summarized as the following equation:
[0079]
[0080] Here, σ represents the sigmoid activation function; ψ1, ψ2, and ψ3 represent the convolutional (Conv) layers, which are determined by the kernel size k in the three branches of triple attention. χ represents the input tensor, A tensor representing the first branch, A tensor representing the second branch, A tensor representing the third branch.
[0081] The present invention uses a triple attention module, which can extract complex and discriminative feature representations with minimal computational cost. Especially in scenes affected by weather, it can cleverly capture subtle features such as fog, snow and rain, which can help to identify objects in these challenging environments. This approach highlights the importance of promoting cross-dimensional interactions without resorting to dimensionality reduction. By avoiding such dimensionality reduction, it eliminates any potential indirect connection between channels and weights, promoting a more direct and efficient feature extraction process.
[0082] The look-ahead attention module is used to encode subtle features and contextual information into tags, an ability that is critical to recognition performance but is often overlooked by traditional self-attention models. This approach effectively enhances the tag representation with fine-grained information, thereby enriching the semantic content of the tag. Look-ahead attention revolutionizes the approach to generating attention weights during tag aggregation. It enables the model to effectively encode fine-grained details, thereby enhancing its overall representational ability. Specifically, it cleverly exploits efficient linear projections to infer the aggregation of neighboring tags directly from the feature representation of the anchor tag, thereby avoiding the computationally intensive dot-product attention calculation.
[0083] For each spatial position (i, j), it effectively evaluates the similarity between that position and all its neighboring positions within a local window of size K×K, centered at (i, j). Compared to self-attention, which requires computationally intensive Query-Key matrix multiplications to calculate attention weights, look-ahead attention greatly simplifies this process through a simple reshape operation.
[0084] The specific definitions are as follows: Figure 4 , given an input X, it will be projected through two different linear layers, one is a linear projection and the other is a linear + anti-folding projection.
[0085] In the linear projection operation, the weight WA∈R C×K4 , generating the outlook weight A∈R H×W×K4. Then reshape the weight to A i,j ∈R K2×K2 , and then normalize it using the Softmax function.
[0086] In the linear + anti-folding operation, the weight WV∈R C×C , after linear + anti-folding, the generated value is expressed as V∈R H ×W×C . Use V Δij ∈R C×K2 represents the set of values residing within a localization window centered at spatial position (i, j), that is:
[0087]
[0088] Finally, the results of the two linear layer projections are multiplied to obtain the folded weighted average. The formula is:
[0089]
[0090] A i,j Represents the result of weighted deformation, V Δi,j represents the set of values that reside within a localization window centered at spatial position (i, j).
[0091] The look-ahead attention aggregates the predicted value representations in a dense manner. By adding weighted values from different local windows corresponding to the same spatial position, the final output is constructed as follows:
[0092]
[0093] i, j represent the elements in the i-th row and j-th column, m, n represent the offsets of the horizontal and vertical coordinates respectively, Y represents the input vector, and K represents the size of the K×K local window.
[0094] The insights of lookahead attention lie in two key aspects: First, the feature representation at each spatial location is robust enough to generate attention weights that promote local aggregation of neighboring features. This highlights the intrinsic value of these localized representations, especially in challenging scenarios such as severe weather conditions, where different objects are often intertwined. In such scenarios, the ability of lookahead attention to model features of neighboring objects becomes critical, enabling accurate object segmentation. Second, the adoption of a dense and localized spatial aggregation strategy can effectively encode fine-grained information, capturing subtle differences that are often overlooked by coarse-grained methods. This meticulous encoding process ensures that the model can discern intricate details, thereby improving recognition and segmentation performance, especially in complex environments.
[0095] Example 2
[0096] Based on the model in Example 1, the present invention uses two different datasets to rigorously evaluate the performance of the proposed model: one is a public urban landscape dataset Cityscapes known for its comprehensiveness, and the other is a Shandong highway scene dataset specially customized for the nuances of the present invention.
[0097] The Cityscape dataset is a high-resolution image taken during the day from the driver's perspective, which is characterized by large crowds of people and vehicles. The target objects are small, such as street lights and traffic signs. The Cityscape dataset is widely used in the fields of urban scene understanding and autonomous driving. It contains 19 categories and 5,000 images, of which 2,975 are used for training, 500 for validation, and 1,525 for testing.
[0098] In order to expand the applicability of the proposed model, we collect and annotate a customized dataset specifically for the highway landscape in Shandong Province. This dataset is enriched with images describing adverse weather conditions such as rain, snow, and fog on Shandong Expressway, providing a powerful platform for evaluating the model's adaptability to adverse environmental factors. The dataset includes 9 object categories such as cars, buses, bicycles, pedestrians, buildings, signs, street lights, and vegetation, and contains a total of 2,000 images, of which 1,000 images are used for training, 500 images are used for validation, and 500 images are used for testing.
[0099] The model of the present invention is carried out on a dedicated computer system based on Ubuntu, which is equipped with a hardware configuration including Intel Core i7 8700 CPU, NVIDIA TM A GeForce GTX3080 GPU and an ample 16GB of RAM. To realize the potential of the model, we also used PyTorch for deep learning, OpenCV for image processing, and Python for scripting and programming.
[0100] To systematically train all variants of the model, we adopted stochastic gradient descent (SGD) as the optimization algorithm and calibrated a learning rate of 0.1 and a momentum factor of 0.9 to accelerate convergence. The training process continued until the training loss reached a steady state (indicating convergence). Before each training, we thoroughly shuffled the training set to ensure randomness and fairness, and then sequentially sampled mini-batches of 8 images to ensure that each image was used exactly once in each training cycle. To determine the best model, we evaluated the performance of each variant on a dedicated validation dataset and selected the variant with the highest accuracy. In addition, we also conducted a comprehensive training of 100 times for each model variant to ensure that the capabilities of the model were fully explored.
[0101] Cross entropy loss is used for training, as shown in the following figure:
[0102]
[0103] Here, i represents the index of the pixel, n*n represents the size of the output image, p represents the true label of the sample, 1 represents positive, 0 represents negative, and represents the probability that the sample is predicted as positive. The loss is the sum of all pixels in a mini-batch.
[0104] The present invention uses Mean Intersection over Union (MeanIoU) as the main evaluation metric. In this framework, each pixel in the scene is used as a positive sample with its unique color in the image. In order to carefully evaluate the performance of the model, we classify the pixels into four different types: true positive (TP), false positive (FP), true negative (TN), and false negative (FN) based on the interaction between the labeled results and the predicted results. This strict classification scheme ensures a comprehensive evaluation of the model's segmentation ability.
[0105] Mean Intersection over Union (MIoU) is a standard metric for semantic segmentation. MIoU represents the average of the "Intersection over Union" (IoU) values calculated for each category in the dataset. MIoU takes the IoU values of all categories and calculates their arithmetic mean. If the dataset contains multiple categories, each category will have its own IoU value, and MIoU will aggregate these values into a single metric to evaluate the performance of the model across all categories as a whole.
[0106] This metric is defined as the overlap between the predicted segmentation and the ground truth divided by the total area covered by the two combined. This metric can be calculated for each class separately or for all classes simultaneously. The best value for this metric is 1 and the worst value is 0.
[0107]
[0108] Here, k represents the number of categories.
[0109] MIoU is an important metric in semantic segmentation, which can quantitatively evaluate the prediction accuracy of the model. The higher the MIoU value, the closer the predicted segmentation is to the ground truth, and the better the model performance.
[0110] We trained the model on the categories provided by the official evaluation script and, as shown in Table 1, it outperforms existing methods.
[0111] Table 1 Results of different models on the urban landscape dataset
[0112]
[0113] As can be seen from Table 1, the proposed method SceneSegAttentionNet establishes a novel state-of-the-art balance between speed and ultra-high accuracy. In particular, it achieves the highest MIoU of 80.2% while maintaining an excellent inference performance of 75.3 FPS. Compared with STDC-2-Seg75, RegSeg, and SFNet, SceneSegAttentionNet surpasses them with 3.2%, 2.1%, and 1.2% higher MIoU, respectively, while also achieving faster inference speed. Even compared with other leading models, including SegNet, ENet, DeepLabv3+, and BiSeNet, DSNet maintains its dominance in MIoU and maintains commendable frame rate performance.
[0114] Figure 5 It shows that the loss value is more significant in the early stage of model training. However, as the training process progresses, the loss value shows a clear and consistent downward trend. This downward curve is very smooth, and there are no sudden peaks in all iterations. In particular, when the training reaches 6000 steps, the loss value tends to stabilize. After completing 1000 iterations, the final result is an optimized model.
[0115] like Figure 6 As shown in the figure, the initial training phase of the established model shows a significantly lower mIoU value. However, as the training progresses, this indicator gradually increases, although there will be slight fluctuations, but it is generally increasing. In particular, after about 1000 iterations, mIoU reaches a stable high point and eventually reaches the best value of 0.802.
[0116] In order to clearly demonstrate the performance of the established model, we visualize the segmentation results of SceneSegAttentionNet, as shown in Figure 7 As shown in Figure 2. We can see that for the selected input image of a city landscape, SceneSegAttentionNet can clearly segment and identify buildings, cars, and sidewalks. In terms of prediction, it is very close to the ground truth.
[0117] Evaluation of Shandong Highway Landscape Dataset
[0118] Table 2 Results of Shandong Province Highway Scene Dataset on Different Models
[0119]
[0120]
[0121] As shown in Table 2, SceneSegAttentionNet sets a new benchmark in the speed-accuracy trade-off for semantic segmentation. Specifically, it achieves an MIoU of 75.3%, surpassing several state-of-the-art models such as STDC-2-Seg75, RegSeg, and SFNet by a significant margin of 5.1%, 4.1%, and 7.3%, respectively. Notably, while improving the accuracy, the inference speed is maintained at an impressive 106 frames per second, which is a high-efficiency performance of the model for real-time applications. Compared with other leading models such as SegNet, ENet, DeepLabv3+, and BiSeNet, SceneSegAttentionNet continues to shine, not only with the highest MIoU but also competitive FPS performance. This comprehensive evaluation highlights the effectiveness of the proposed method in balancing the dual objectives of speed and accuracy.
[0122] Figure 8 A key observation is presented that highlights the effectiveness of the adopted training regimen. Initially, the loss values are significantly higher, marking the beginning of model learning. This initial surge in loss is a common feature that stems from the random initialization of the model. However, as the training process unfolds and progresses through subsequent iterations, a phenomenon emerges: the loss values systematically and consistently decrease. This downward trend is represented by a smoothly descending curve, indicating that the model is gradually adapting to the complexity of the data and adjusting its internal parameters to better approximate the objective function. A key milestone is that the loss values gradually stabilize after about 200 iterations. This indicates that the model has converted to a local minimum and further training will not significantly reduce the loss.
[0123] like Fig. 9 As shown in the figure, the mean intersection over union (mIoU) value is significantly low at the beginning of the training process. As the number of iterations increases, this indicator gradually increases, although the increase is small. In particular, after the 300th iteration, mIoU shows a significant increase, accompanied by small fluctuations in the middle of the training stage. Finally, after 800 iterations, mIoU reaches a stable plateau, reaching an optimal value of 0.753, and maintains this peak performance in the next 200 iterations (until the 1000th iteration).
[0124] The first row shows a cloudy day with overcast sky. In the corresponding predictions, the main focus was on achieving accurate segmentation and buildings are clearly outlined. The second row depicts a foggy day with limited visibility, even at close range. However, when applying SceneSegAttentionNet for segmentation, even white areas in the distance are effectively captured and segmented. In particular, buildings remain clearly discernible in the fog. Finally, the third row presents a snowy day with the ground covered with snow. Here, cars and bicycles are accurately segmented, and although a minibus is misclassified as a bus, this slight error is acceptable in this case.
[0125] The analysis shows that SceneSegAttentionNet can selectively process input images in the Shandong Highway Landscape Dataset, excelling in providing accurate segmentation and proficiently distinguishing cars, buses, bicycles, pedestrians, and buildings, even under adverse weather conditions such as rain and snow. In particular, the predictions generated by our model are very close to the ground truth, demonstrating its excellent accuracy and robustness.
[0126] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A model for autonomous driving image segmentation in challenging weather conditions, characterized in that The model includes an encoder and a decoder corresponding to the encoder; The encoder comprises an input module and four encoder layers in sequence, wherein a coordinate attention module is connected between the input module and the encoder layer connected thereto and between adjacent encoder layers; The decoder comprises four decoder layers and an output module in sequence, wherein a prospective attention module is connected between the output module and the decoder layer connected thereto and between adjacent decoder layers; A triple attention module is connected between the encoder and the encoder.
2. The model for autonomous driving image segmentation under challenging weather conditions according to claim 1, characterized in that There are four triple attention modules, connected before each encoder layer and after the decoder layer.
3. The model for autonomous driving image segmentation under challenging weather conditions according to claim 1, characterized in that In the coordinate attention module, the input is H-pooled and W-pooled respectively, and then the feature maps H and W obtained after pooling are spliced to obtain HW; after splicing, the feature maps are sequentially convolved, batch normalized and activated to obtain HWC, HWB and HWA respectively; then segmented to obtain SH and SW; the two results are respectively subjected to convolution and nonlinear activation functions to obtain SHS and SWS; finally, SHS and SWS are vector multiplied.
4. The model for autonomous driving image segmentation under challenging weather conditions according to claim 1, characterized in that The triple attention module adopts triple parallel processing branches, where: In the first branch, the input tensor is rotated 90 degrees counterclockwise along the height axis to obtain a rotated tensor; the rotated tensor passes through the Z-pooling layer, convolution layer, batch normalization layer, and activation function layer in sequence to obtain the attention weight; then the attention weight and the rotation tensor are vector-multiplied and the result is rotated exactly 90 degrees clockwise along the height axis; In the second branch, the input tensor is rotated 90 degrees counterclockwise along the W axis to obtain a rotated tensor; the rotated tensor passes through the Z-pooling layer, convolution layer, batch normalization layer, and activation function layer in sequence to obtain the attention weight; then the attention weight and the rotation tensor are vector-multiplied, and the result is rotated exactly 90 degrees clockwise along the W axis; In the third branch, the input tensor is first reduced in dimension by channel pooling; then, the tensor is sent to the convolution layer, batch normalization layer, and activation function layer in sequence to generate attention weights, which are then multiplied by the original input tensor; The results in the three branches are added and averaged to combine the information from the three branches into a single comprehensive feature.
5. The model for autonomous driving image segmentation under challenging weather conditions according to claim 4, characterized in that The calculation formula of triple attention is as follows: Where: σ represents the sigmoid activation function; ψ1, ψ2 and ψ3 represent convolution; χ represents the input tensor, A tensor representing the first branch; A tensor representing the second branch; A tensor representing the third branch.
6. The model for autonomous driving image segmentation under challenging weather conditions according to claim 1, characterized in that In the prospective attention module, the input X is projected through two different linear layers, one is a linear projection and the other is a linear + anti-folding projection; finally, the results of the two linear layer projections are multiplied to obtain the folded weighted average.
7. The model for autonomous driving image segmentation under challenging weather conditions according to claim 6, characterized in that Multiplying the results of the two linear layer projections, the formula for the folded weighted average is: Where: A i,j Represents the result of weighted deformation, V Δi,j represents the set of values that reside within a localization window centered at spatial position (i, j).
8. A method for image segmentation for autonomous driving under challenging weather conditions, characterized in that Use the model described in any one of claims 1 to 7.
9. A server comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor implements the method steps of claim 8 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: The computer program implements the method steps of claim 8 when executed by a processor.