Inverse residual cross-modal automatic driving scene perception data set construction method

Through the inverse residual cross-modal structure and adaptive convolution kernel combined with the cross attention mechanism, the problem of image generation quality and feature information loss in the construction of the autonomous driving scene perception data set is solved, high-quality and diverse image generation is achieved, and the accuracy and authenticity of autonomous driving scene perception is improved.

CN120339980APending Publication Date: 2025-07-18XIAN TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510320962.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing method of building autonomous driving scene perception data sets is insufficient in terms of image generation quality and feature information diversity, and there is a risk of feature information loss when processing the original residual compression input feature map.

Method used

The inverse residual cross-modal structure is adopted, and through multi-scale inverse residual structure, cross-modal multi-scale convolution and channel transformation technology, combined with the adaptive convolution kernel and the cross-attention mechanism, the expansion and fusion of feature map information is enhanced to form an inverse residual cross-modal text generation image model.

Benefits of technology

The details and clarity of image generation are improved, the capture ability and cross-modal understanding of multi-resolution features are enhanced, and the accuracy and authenticity of scene images of autonomous driving are significantly improved. The IS and FID indicators are increased by 11.9% and 17.9% respectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339980A_ABST
    Figure CN120339980A_ABST
Patent Text Reader

Abstract

The invention discloses an inverse residual cross-modal automatic driving scene perception data set construction method. The method comprises the following implementation steps of: firstly, preprocessing a text encoder; secondly, dimension raising is conducted on the preprocessed feature map, the feature map is divided into three groups, and it is ensured that features are effectively reserved in the image generation process; then, introducing adaptive convolution kernels, capturing features of different resolutions, introducing a cross attention mechanism, enhancing feature interaction, and performing channel shuffling on mapping features obtained by each convolution kernel; then, obtained features are fused and spliced, and channel expansion is carried out in point convolution; and finally, forming an inverse residual cross-modal text generation image model, and evaluating model performance and image quality. According to the method, feature information can be effectively reserved, the diversity of image generation is enhanced, and a rich test environment can be provided. Experiments prove that the method can effectively construct an automatic driving perception data set, and compared with an original model ATTNGAN, the method provided by the invention has the advantages that the IS and FID are improved by 11.9% and 17.9% respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and image processing, and particularly relates to the generation of an inverse residual cross-modal automatic driving scene perception dataset. Background Art

[0002] As a cutting-edge technology such as artificial intelligence, Internet of Things, and high-performance computing, autonomous driving technology is increasingly becoming the focus and hotspot in global strategic competition. The construction of an autonomous driving scene perception dataset is crucial for the testing of autonomous driving technology. It is the prerequisite for virtual simulation testing. The document with the application number "CN202011095861.4" discloses "a method for generating a road scene and related devices". The construction of an autonomous driving perception dataset depends on the corresponding point cloud data of the target road. By extracting key features from the point cloud and performing clustering, a target road scene is generated. This technology shows the practical application potential in autonomous driving perception testing by overcoming the problems restricted by environmental factors in the traditional point cloud data acquisition process. However, the acquisition of point cloud data is greatly affected by environmental factors such as weather and light, and complex operations such as extracting key features, clustering, and spatial establishment need to be performed on a large amount of point cloud data. The text-to-image technology can effectively avoid these limitations with its low cost, high adaptability, and flexibility, and can generate specific scenarios and simulate dangerous scenarios according to requirements.

[0003] In recent years, the text-to-image technology based on the GAN network has shown great potential in many fields, and thus has also been introduced into the generation of autonomous driving scene images. Models such as ControlGAN, MirrorGAN, and AttnGAN can all be used in the generation of autonomous driving perception scene images. Compared with ControlGAN and MirrorGAN, AttnGAN can generate high-quality images consistent with the text content according to the text description by introducing an attention mechanism. The patent with the application number "CN202111109265.1" discloses "a method for generating an image from text", which uses a Transformer module and an AttnGAN network to generate a rough image. Subsequently, by combining the improved feature vector with the word feature and performing neural network upsampling, higher-resolution images are gradually generated, and images with more detailed outlines and clearer details are generated. This method proves the feasibility of constructing a dataset by the text-to-image method, but it is slightly inferior in terms of image feature diversity, and there is a risk of feature information attenuation or loss during the processing of the original input feature map compressed by residuals.

[0004] For virtual testing of autonomous driving perception scenarios, it is necessary to construct a high-quality image dataset for autonomous driving scenario perception, which poses higher requirements for the quality and feature information of the generated images. Although the current image generation quality and semantic consistency have improved, there are still deficiencies in meeting the diversity requirements of images for autonomous driving perception testing, and there is a risk of loss of extracted feature information when processing the original residual compressed input feature map. Summary of the Invention

[0005] In response to the above analysis, the present invention proposes a method for constructing an autonomous driving scenario perception dataset with inverse residual cross-modal. Based on AttnGAN, an integrated multi-scale inverse residual structure is adopted to expand the information content of the feature map, and cross-modal multi-scale convolution and channel transformation technology are introduced to effectively solve the problem of information loss in image generation.

[0006] To achieve the above object, the technical solution steps of the present invention are as follows: A method for constructing an autonomous driving scenario perception dataset with inverse residual cross-modal includes the following steps:

[0007] Step 1: Construct a text-to-image model with inverse residual cross-modal: Follow the multi-stage structure of AttnGAN, including text processing and feature extraction, attention-guided image generation, and multi-scale feature fusion and discrimination. Preprocess the text image encoder, including converting text into a digital sequence, sequence padding or truncation, encoding global sentence vectors and word vectors, and training the text encoder (RNN) and image encoder (CNN) for feature extraction;

[0008] Step 2: In the model construction stage, convert the residual of the original network in AttnGAN into an inverse residual structure, increase the dimension of the input feature map and perform batch normalization, and then divide the number of channels into three groups to enrich the number of amplified information;

[0009] Step 3: Capture features of different resolutions through an adaptive convolution kernel and a cross-attention mechanism, and then perform channel shuffling on the features obtained by each convolution kernel to achieve efficient fusion of multi-resolution features;

[0010] Step 4: Fuse and splice the feature maps processed by channel shuffling, then expand the channels in pointwise convolution, and finally perform dimensionality reduction through batch normalization, activation, and convolution layers to form a text-to-image model with inverse residual cross-modal;

[0011] Step 5: Train the constructed text-to-image model with inverse residual cross-modal. After training, use text descriptions and expectations to generate an autonomous driving scenario perception image dataset, and evaluate the quality of the generated images and the performance of the model using IS and FID as indicators.

[0012] Further, in the above step 2, an inverse residual structure is constructed and used as one of the components of the text-to-image inverse residual cross-modal model to augment the image information. The specific operation process is as follows:

[0013] Suppose \(x\in R\) H×W×C represents the input feature map, where \(H\) and \(W\) are the height and width of the feature map respectively, \(C\) is the number of channels, and \(h\) represents the output feature map

[0014] First, the number of channels of the input feature is increased through a group of 1×1 convolutional layers to perform dimensionality increase on the input feature map. The formula is:

[0015] \(h1 = W1 * x + b1(1)\)

[0016] where \(h1\in R\) H×W×C' is the output feature map after dimensionality increase, \(x\) is the input feature map, \(W1\) and \(b1\) are the weights and biases of the 1×1 convolutional layer respectively, \(*\) represents the convolution operation, and the convolution kernel \(W1\in R\) 1×1×C×C′ Then, batch normalization is performed on the feature map after dimensionality increase to accelerate training and improve the generalization ability of the model:

[0017]

[0018] where \(\mu\) is the mean, \(\sigma\) is the standard deviation, \(\gamma\) and \(\beta\) are learnable parameters, and \(h2\) is the output feature map after dimensionality increase;

[0019] Finally, the ReLU6 activation function is used for the feature map after batch normalization to introduce non-linearity and limit the range of output values. The relevant formula is:

[0020] \(h3 = f\) ReLU6 (x)=\(\min(\max(0, h2), 6)(3)\)

[0021] The input channels are divided into three groups, and an adaptive convolution kernel is introduced to dynamically adjust the shape and size of the convolution kernel to adapt to the input features. Each convolution kernel only performs convolution operations on its corresponding channel group to extract local features.

[0022] Further, in the above step 3, an adaptive convolution kernel is introduced into each of the three groups of input channels. The specific operation process is as follows:

[0023] \(K = KernelPredictor(h3)(4)\)

[0024] where \(K\) is the dynamically generated convolution kernel, and \(h3\) is the output feature map after being processed by the activation function after dimensionality increase;

[0025] The channels of the input feature map are divided into three groups, and adaptive convolution is applied to each group respectively:

[0026]

[0027] where K( g ) is the convolution kernel generated for the g-th group, where 0 < g ≤ 3;

[0028] Perform cross-attention mechanism processing on the features obtained by each convolution kernel to enhance the interaction between features:

[0029]

[0030] where h4 is the output after cross-attention mechanism processing, which are the outputs of 3 groups of channels respectively. The cross-attention mechanism further enhances feature fusion by dynamically allocating attention weights;

[0031] The feature maps obtained by each convolution kernel are shuffled in the channel dimension. The shuffling operation is performed within the group to ensure the integrity of information.

[0032] Furthermore, in the above step four, the features are fused and concatenated, and channel expansion is performed. The specific operation process is as follows:

[0033] The shuffled feature maps are concatenated together to form a new feature map:

[0034]

[0035] h4 is the output after cross-attention mechanism processing, which are the outputs of 3 groups of channels respectively, and a1 is the fused output feature map;

[0036] Then, a 1×1 convolution kernel is used to perform channel expansion on the fused feature map to further enhance the expression ability of the features:

[0037] a2 = w3 * a1 + b3(8)

[0038] where w3 and b3 are the weights and biases of the 1×1 convolution kernel, a1 is the fused feature map after concatenation, a2 is the feature map after channel expansion, the convolution kernel W3 ∈ R 1×1×(C′+C″)×C″ , and the output feature map a2 ∈ R H×W×C″ ;

[0039] Apply batch normalization and activation function activation to the expanded feature map again:

[0040]

[0041] h6 = f ReLU6 (x) = min(max(0, x), 6) (10)

[0042] Finally, perform dimensionality reduction processing through a group of 1×1 convolutional layers:

[0043] h7 = w4 * h6 + b4 (11)

[0044] where w4 and b4 are the weights and biases of the 1×1 convolutional kernel, h6 is the fused feature map after concatenation and batch normalization and activation processing, and the convolutional kernel W4 ∈ R 1×1×C″×C , h7 is the feature map after dimensionality reduction processing, and C is the dimension of the model input feature map.

[0045] Compared with the existing technologies, the beneficial effects of the present invention are as follows:

[0046] (1) Aiming at the problem of limited information processing ability in cross-modal feature interaction of the traditional residual structure, the present invention proposes a solution method based on the inverse residual structure. By first increasing the dimension of the input feature map and performing grouped processing, the number of amplified information is enriched, and the model's ability to capture features of different scales is enhanced. Compared with other methods, the details and clarity of image generation are improved.

[0047] (2) Aiming at the problem of insufficient scale adaptability caused by fixed convolutional kernels in multi-scale feature fusion, the present invention proposes a method combining an adaptive convolutional kernel and cross-attention. The adaptive convolutional kernel can dynamically adjust its shape and size to adapt to the input features, and the cross-attention mechanism enhances the interaction between features, enabling efficient fusion of multi-resolution features. Compared with other methods, the accuracy and authenticity of the generated images in the autonomous driving perception scenario are significantly improved.

[0048] (3) Aiming at the problem of insufficient cross-modal information interaction, the present invention uses the method of channel shuffle. By fusing the features obtained by concatenating the adaptive convolutional kernels and performing channel shuffle, cross-modal information fusion and interaction are achieved, further mixing different channel features and improving the model's cross-modal understanding ability. Compared with the original model ATTNGAN, the method of the present invention improves by 11.9% and 17.9% respectively in terms of IS and FID. Description of the Drawings

[0049] Figure 1 It is the overall flowchart of the method for constructing an inverse residual cross-modal autonomous driving scenario perception dataset;

[0050] Figure 2 It is the framework diagram of the image generation model for the inverse residual cross-modal autonomous driving perception scenario constructed;

[0051] Figure 3 It is the structural difference diagram between the original residual and the inverse residual;

[0052] Figure 4 It is the structural diagram showing the capture of features of different resolutions by depthwise separable convolution;

[0053] Figure 5Network structure diagram for channel shuffling;

[0054] Figure 6 Result diagram of the autonomous driving perception scenario generated by the method of the present invention. Specific implementation manner

[0055] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0056] See Figure 1 , to construct an image generation model for autonomous driving perception scenarios with cross-modal inverse residuals, the present invention first preprocesses the text image encoder, and then in the image generation stage, converts the residuals in the original network into inverse residuals, and divides the input channels into multiple groups to form cross-modal multi-scale channels, so as to enhance the non-linear mapping ability of the network. Secondly, depthwise separable convolutions are used to expand the dimension of the input feature map, capture features at different resolutions, and then perform channel shuffling on the fused features, so as to achieve the purpose of enriching the image information content. Finally, the obtained features are fused and spliced, and then the fused feature map is dimension-reduced by another convolutional kernel. This method effectively enhances the diversity of image generation, thus realizing the generation of autonomous driving perception scenario images. Use datasets such as COCO to train the constructed inverse residual cross-modal text generation image model to obtain an improved inverse residual cross-modal text-to-image (MSRR-T2I) model, and use text descriptions and expectations to generate an autonomous driving scene perception image dataset, and evaluate the quality of the generated images and the performance of the model with IS and FID as indicators.

[0057] Based on the above basic idea, the present invention provides a method for constructing an autonomous driving scenario perception dataset with cross-modal inverse residuals, and the specific implementation steps are as follows:

[0058] Step 1: Construct a cross-modal inverse residual model from text to image, following the multi-stage structure of AttnGAN. The multi-stage structure of AttnGAN includes text processing and feature extraction, attention-guided image generation, and multi-scale feature fusion and discrimination.

[0059] The pre-trained text processing and feature extraction are used to extract the semantic features of the text description. The text encoder uses a pre-trained recurrent neural network (RNN) model, and the image encoder is pre-trained using a convolutional neural network (CNN).

[0060] The attention-guided image generation uses an attention mechanism to combine text features with the generated image features. The image is gradually constructed by a multi-stage generator, and each stage includes upsampling and convolutional operations, as well as an attention module to refine the image content.

[0061] The multi-scale feature fusion and discrimination adopts a multi-scale incremental generation strategy (MSRR) to generate images at different resolutions, from rough to detailed. The local encoder extracts the local features of the image and aligns them with the text features. The discriminator evaluates the authenticity of the generated image and forms an adversarial process with the generator to improve the image quality.

[0062] Step 2: In the image generation part, the residual in the original network is transformed into an inverted residual structure, which first increases the dimension of the input feature map, divides the number of channels into three groups, and enriches the quantity of amplified information.

[0063] As Figure 3 shown in the structural comparison diagram of the inverted residual and the original residual, assume x ∈ R H×W×C represents the input feature map, where H and W are the height and width of the feature map respectively, C is the number of channels, and h represents the output feature map.

[0064] The input feature is passed through a group of 1×1 convolutional layers to increase the number of channels, the input feature map is dimensionally upsampled, and batch normalization and an activation function are used for activation. The formula is:

[0065] h1 = W1 * x + b1(1)

[0066] where h1 is the output feature map after dimensional upsampling, x is the input feature map, W1 and b1 are the weights and biases of the 1×1 convolutional layer respectively, and * represents the convolution operation.

[0067] Batch normalization is performed on the dimensionally upsampled feature map to accelerate training and improve the generalization ability of the model:

[0068]

[0069] where μ is the mean, σ is the standard deviation, γ and β are learnable parameters, and h2 is the output feature map after dimensional upsampling;

[0070] Finally, the ReLU6 activation function is used for the batch-normalized feature map to introduce non-linearity and limit the range of output values. The relevant formula is:

[0071] h 3= f ReLU6 (x) = min(max(0, h2), 6)(3)

[0072] The main function of the activation function is to introduce non-linearity, alleviate the vanishing gradient, enable the network to learn complex patterns, and contribute to the training of deep networks.

[0073] Step 3: Adaptive convolutional kernels and cross-attention mechanisms are used to capture features at different resolutions, and then channel shuffling is performed on the features obtained by each convolutional kernel to achieve efficient fusion of multi-resolution features.

[0074] First, an adaptive convolution kernel is introduced to capture features of different resolutions, and the specific description is as follows:

[0075] As Figure 4 shown in the structural diagram of depthwise separable convolution, the input channels are divided into three groups, and an adaptive convolution kernel is introduced to dynamically adjust the shape and size of the convolution kernel to adapt to the input features. Each convolution kernel only performs convolution operations on its corresponding channel group to extract local features. The adaptive convolution kernel dynamically generates convolution kernel weights according to the input features through a convolution kernel generator, enhancing the ability to capture features of different scales:

[0076] K = KernelPredictor(h3) (4)

[0077] where K is the dynamically generated convolution kernel, and h3 is the output feature map after being processed by the activation function after dimensionality increase.

[0078] The channels of the input feature map are divided into three groups, and adaptive convolution is applied to each group respectively:

[0079]

[0080] where K( g ) is the convolution kernel generated for the g-th group, 0 < g ≤ 3.

[0081] Then, a cross-attention mechanism is introduced to perform cross-attention mechanism processing on the features obtained by each convolution kernel, enhancing the interaction between features:

[0082]

[0083] where h4 is the output after being processed by the cross-attention mechanism, are the outputs of the three groups of channels respectively. The cross-attention mechanism further enhances feature fusion by dynamically allocating attention weights.

[0084] Finally, after depth convolution, the feature maps obtained by each convolution kernel are shuffled in the channel dimension. The shuffling operation is performed within the group to ensure the integrity of information.

[0085] Step 4: The features obtained by convolution kernels of different sizes are fused and concatenated. By adopting channel transformation technology, information fusion and interaction between different groups are realized. Finally, the obtained feature map is dimensionally reduced again to obtain the inverse residual cross-modal text-to-image generation model.

[0086] As Figure 5 shown in the structural diagram of the channel shuffling network.

[0087] Step 4.1 The shuffled feature maps are concatenated together to form a new feature map:

[0088]

[0089] h4 is the output after the cross-attention mechanism processes, which are the outputs of 3 groups of channels respectively, and a1 is the output feature map after fusion.

[0090] Step 4.2 uses a 1×1 convolutional kernel to expand the channels of the fused feature map:

[0091] a2 = w3 * a1 + b3 (8)

[0092] where W3 and b3 are the weights and biases of the 1×1 convolutional kernel, a1 is the fused feature map after splicing, and a2 is the feature map after channel expansion.

[0093] Apply batch normalization and activation function activation to the expanded feature map again:

[0094]

[0095] h6 = f ReLU6 (x) = min(max(0, x), 6) (10)

[0096] Finally, perform dimensionality reduction through a group of 1×1 convolutional layers:

[0097] h7 = w4 * h6 + b4 (11)

[0098] where w4 and b4 are the weights and biases of the 1×1 convolutional kernel, h6 is the fused feature map after splicing and batch normalization and activation processing, and h7 is the feature map after dimensionality reduction processing.

[0099] Step Five: Use datasets such as COCO to train the constructed text-to-image model to obtain an improved inverse residual cross-modal text-to-image (MSRR-T2I) model, use text descriptions and expectations to generate an autonomous driving scene perception image dataset, and evaluate the quality of the generated images and the performance of the model with IS and FID as indicators.

[0100] Inception Score (IS) and Fréchet Inception Distance (FID) are two commonly used indicators for evaluating the performance of a generation model, and they are mainly used to measure the quality and diversity of the generated images.

[0101]

[0102] p data is the distribution of real data, p(y|x) is the conditional probability distribution given by the Inception model on a specific image x, and p(y) is the average of the conditional probability distributions on all images.

[0103] The higher the IS value, the more diverse and realistic the generated image categories are.

[0104] FID = ||μ1 - μ2|| 2 + Tr(Σ1 + Σ2 - 2(Σ1Σ2) 1 / 2 ) (13)

[0105] μ1 and Σ1 are the mean and covariance of real images in the feature space of the Inception model, μ2 and Σ2 are the mean and covariance of generated images in the same feature space, and Tr represents the trace (the sum of the diagonal elements of the matrix).

[0106] FID evaluates the quality of generated images by comparing the distributions of generated images and real images in the feature space. It calculates the Fréchet distance between the feature means and covariance matrices of two image sets. The lower the FID value, the closer the generated images are to the real images in terms of features, and the higher the image quality.

[0107] As shown in Table 1, the specific index comparison between this invention patent and other algorithms is as follows. Compared with the basic network ATTNGAN, the IS score of the COCO dataset has increased from 25.89 to 28.98, an increase of 11.9%. The FID score has decreased from 35.49 to 29.12, an improvement of 17.9%, indicating that the model in this paper performs better in terms of image quality and diversity.

[0108] Table 1

[0109]

[0110] Figure 6 are the self-driving perception scene images of the original model and the model of this invention. By comparison, it can be seen that the improved model of this invention can effectively generate target objects such as roads, traffic signs, and signals in the self-driving perception scene, further verifying the feasibility of the method for constructing the inverse residual cross-modal self-driving scene perception dataset.

[0111] The above is an explanation of the specific implementation of the present invention, rather than a limitation thereof. Those skilled in the relevant technical field can also make various equivalent technical solutions without departing from the scope of the present invention. Therefore, all equivalent technical solutions should be included in the protection scope of the present invention.

Claims

1. A method for constructing an inverse residual cross-modal dataset for autonomous driving scene perception, characterized by the following steps: Step 1: Construct an inverse residual cross-modal text-to-image generation model: Follow the multi-stage structure of AttnGAN, including text processing and feature extraction, attention-guided image generation, multi-scale feature fusion and discrimination. Preprocess the text image encoder, including converting text into digital sequences, sequence padding or truncation, encoding global sentence vectors and word vectors, and training the text encoder (RNN) and image encoder (CNN) for feature extraction; Step 2: In the model construction stage, convert the residual of the original network in AttnGAN into an inverse residual structure, increase the dimension of the input feature map and perform batch normalization, and then divide the number of channels into three groups to enrich the number of amplified information; Step 3: Capture features of different resolutions through an adaptive convolution kernel and a cross-attention mechanism, and then perform channel shuffling on the features obtained by each convolution kernel to achieve efficient fusion of multi-resolution features; Step 4: Fuse and splice the feature maps processed by channel shuffling, then perform channel expansion in pointwise convolution, and finally perform dimensionality reduction through batch normalization, activation, and convolutional layers to form an inverse residual cross-modal text-to-image generation model; Step 5: Train the constructed inverse residual cross-modal text-to-image generation model. After training, use text descriptions and expectations to generate an autonomous driving scene perception image dataset, and evaluate the quality of the generated images and the performance of the model using IS and FID as metrics.

2. The method for constructing an inverse residual cross-modal automatic driving scene perception data set according to claim 1, characterized in that: In the second step, an inverse residual structure is constructed and used as one of the components of the text-to-image inverse residual cross-modal model to amplify image information. The specific operation process is as follows: Assume \(x\in R\) H×W×C represents the input feature map, where \(H\) and \(W\) are the height and width of the feature map respectively, \(C\) is the number of channels, and \(h\) represents the output feature map First, increase the number of channels of the input feature through a group of 1×1 convolutional layers to increase the dimension of the input feature map. The formula is: h1 = W1 * x + b1 (1) where h1 ∈ R H×W×C' is the output feature map after dimension elevation, x is the input feature map, W1 and b1 are the weights and biases of the 1×1 convolutional layer respectively, * represents the convolution operation, and the convolution kernel W1 ∈ R 1×1×C×C′ , batch normalization is performed on the feature map after dimension elevation to accelerate training and improve the generalization ability of the model: where μ is the mean, σ is the standard deviation, γ and β are learnable parameters, and h2 is the output feature map after dimensionality increase; Finally, use the ReLU6 activation function for the feature map after batch normalization to introduce non-linearity and limit the range of output values. The relevant formula is: h3 = f ReLU6 (x) = min(max(0, h2), 6) (3) Divide the input channels into three groups, introduce an adaptive convolution kernel, dynamically adjust the shape and size of the convolution kernel to adapt to the input features, and each convolution kernel only performs convolution operations on its corresponding channel group to extract local features.

3. A method for constructing an inverse residual cross-modal automatic driving scene perception data set according to claim 2, characterized in that: In the third step, an adaptive convolution kernel is introduced into each of the three input channels respectively. The specific operation process is as follows: K = KernelPredictor(h3) (4) where K is the dynamically generated convolution kernel, and h3 is the output feature map after being processed by the activation function after dimensionality increase; Divide the channels of the input feature map into three groups, and apply adaptive convolution to each group respectively: where K (g) is the convolution kernel generated for the g-th group, where 0 < g ≤ 3; Perform cross-attention mechanism processing on the features obtained by each convolution kernel to enhance the interaction between features: where h4 is the output after the cross-attention mechanism processes, which are the outputs of 3 groups of channels respectively. The cross-attention mechanism further enhances feature fusion by dynamically allocating attention weights; The feature maps obtained by each convolution kernel are shuffled in the channel dimension, and the shuffling operation is performed within the group to ensure the integrity of information.

4. A method for constructing an inverse residual cross-modal automatic driving scene perception data set according to claim 3, characterized in that: In the fourth step, fuse and splice the features and perform channel expansion. The specific operation process is as follows: The shuffled feature maps are concatenated together to form a new feature map: h4 is the output after the cross-attention mechanism processes, which are the outputs of 3 groups of channels respectively, and a1 is the output feature map after fusion; Then, a 1×1 convolutional kernel is used to expand the channels of the fused feature map, further enhancing the feature expression ability: a2 = w3 * a1 + b3 (8) where w3 and b3 are the weights and biases of the 1×1 convolutional kernel, a1 is the fused feature map after concatenation, a2 is the feature map after channel expansion, the convolutional kernel W3 ∈ R 1×1×(C′+C″×C″ , and the output feature map a2 ∈ R H×W×C ; Batch normalization and an activation function are applied again to the expanded feature map for activation: h6 = f ReLU6 (x) = min(max(0, x), 6) (10) Finally, dimensionality reduction is performed through a set of 1×1 convolutional layers: h7 = w4 * h6 + b4 (11) where w4 and b4 are the weights and biases of the 1×1 convolutional kernel, h6 is the fused feature map after concatenation and batch normalization and activation processing, and the convolutional kernel W4 ∈ R 1×1×C″×C , h7 is the feature map after dimensionality reduction processing, and C is the dimension of the input feature map of the model.

Citation Information

Patent Citations

  • Road scene generation method and related device

    CN112435333A

  • Method for generating image through text

    CN114022582A