Vehicle acquisition image processing method and device, vehicle, medium and program product

The vehicle image processing method enhances feature extraction in smart driving vehicles by using SMSB, CAPEB, and ESCB modules to address the challenge of capturing fine-grained differences, reducing complexity and resource use.

CN120318792APending Publication Date: 2025-07-15CHERY INTELLIGENT VEHICLE TECH (HEFEI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510478625.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing feature extraction methods are difficult to capture subtle differences, resulting in insufficient feature distinction capabilities, high computational complexity, and large resource consumption.

Method used

The fine-grained image recognition model is adopted, and dimensionality reduction and multi-head attention calculation are performed through the SMSB module, the CAPEB module performs channel attention weighting and patch embedding, and the ESCB module performs deep convolution and spatial attention processing to generate feature representations of global perception capabilities.

Benefits of technology

It improves feature distinction ability, reduces computing complexity, reduces resource consumption, and improves computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318792A_ABST
    Figure CN120318792A_ABST
Patent Text Reader

Abstract

The invention relates to a vehicle acquisition image processing method and device, a vehicle, a medium and a program product. Comprising the following steps: acquiring a to-be-processed image acquired by at least one sensor of a vehicle; inputting an image to be processed into the fine-grained image recognition model, and performing dimension reduction processing and multi-head attention calculation on the image to be processed through an SMSB module of the model to obtain feature representation after dimension reduction; channel attention weighting and patch embedding processing are carried out on the feature representation after dimension reduction through a CAPEB module of the model to obtain enhanced local feature representation, and deep convolution and space attention processing are carried out on the enhanced local feature representation through an ESCB module of the model to obtain feature representation of global perception ability; therefore, a final output image is generated. Therefore, the problem that the feature distinguishing ability is insufficient due to the fact that tiny differences are difficult to capture in an existing feature extraction method is solved, the calculation complexity is reduced, resource consumption is reduced, and the calculation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent driving technology, and particularly to a method and apparatus for processing images collected by a vehicle, a vehicle, a medium, and a program product. Background Art

[0002] The goal in the field of intelligent driving is to achieve fully autonomous vehicle operation, which requires the vehicle to be able to identify and understand complex road environments in real time and accurately. Image recognition algorithms play a core role in this process. It needs to identify traffic signs, signal lights, pedestrians, vehicles, obstacles, and road boundaries on the road, and this information is crucial for the safe driving of the vehicle. In addition, the intelligent driving system also needs to be able to process visual information under various lighting conditions, weather conditions, and different time periods, which further increases the challenge of image recognition.

[0003] In related technologies, using traditional convolutional neural networks, through multi-level feature extraction, complex visual features are automatically learned from a large amount of image data, providing a powerful tool for image recognition. However, this method is difficult to capture subtle differences, resulting in insufficient feature discrimination ability, which urgently needs to be solved. Summary of the Invention

[0004] This application provides a method and apparatus for processing images collected by a vehicle, a vehicle, a medium, and a program product to solve the problem that the existing feature extraction method is difficult to capture subtle differences, resulting in insufficient feature discrimination ability, reduce the computational complexity while reducing resource consumption, and improve the computational efficiency.

[0005] To achieve the above object, a first aspect embodiment of this application proposes a method for processing an image collected by a vehicle, including the following steps:

[0006] Obtain an image to be processed, where the image to be processed is collected by at least one sensor of the vehicle;

[0007] Input the image to be processed into the fine-grained image recognition model. Use the SMSB (Subspace Multi-Head Self-Attention Block) module of the fine-grained image recognition model to perform dimensionality reduction processing and multi-head attention calculation on the image to be processed to obtain a feature representation after dimensionality reduction. Then, use the CAPEB (Channel Attention Patch Embedding Block) module of the fine-grained image recognition model to perform channel attention weighting and patch embedding processing on the feature representation after dimensionality reduction to obtain an enhanced local feature representation. Moreover, use the ESCB (Efficient Semantic Convolution Block) module of the fine-grained image recognition model to perform depth convolution and spatial attention processing on the enhanced local feature representation to obtain a feature representation with global perception ability;

[0008] Generate the final output image according to the feature representation with global perception ability.

[0009] According to an embodiment of the present application, the step of using the SMSB module of the fine-grained image recognition model to perform dimensionality reduction processing and multi-head attention calculation on the image to be processed to obtain a feature representation after dimensionality reduction includes:

[0010] Map the feature representation of the image to be processed to a subspace of a preset dimension through a linear transformation to generate the query, key, and value after dimensionality reduction;

[0011] Based on a preset multi-head attention mechanism, decompose the query, key, and value after dimensionality reduction into multiple independent attention heads, where each attention head focuses on a different subspace;

[0012] Concatenate the outputs of the multiple independent attention heads, and perform a linear transformation on the concatenation result through a preset output weight matrix to obtain the feature representation after dimensionality reduction.

[0013] According to an embodiment of the present application, the step of using the CAPEB module of the fine-grained image recognition model to perform channel attention weighting and patch embedding processing on the feature representation after dimensionality reduction to obtain an enhanced local feature representation includes:

[0014] Segment the feature representation after dimensionality reduction into multiple patches of the same size, and convert the multiple patches of the same size into vector representations to obtain a set of feature vectors after dimensionality reduction;

[0015] Based on a preset channel attention mechanism, calculate the weights of each channel, and perform weighted processing on the downsampled feature vectors based on the calculation results. Then, segment the feature vectors weighted by the preset channel attention mechanism into multiple new patches of the same size, and convert the multiple new patches of the same size into vector representations to obtain a vector matrix;

[0016] Based on a preset normalization layer, normalize each vector in the vector matrix to obtain the enhanced local feature representation.

[0017] According to an embodiment of the present application, the process of obtaining the feature representation with global perception ability by performing depth convolution and spatial attention processing on the enhanced local feature representation through the ESCB module of the fine-grained image recognition model includes:

[0018] Perform depth convolution processing on the enhanced local feature representation to obtain a feature map after depth convolution;

[0019] Perform spatial attention processing on the feature map after depth convolution, use the self-attention mechanism to learn the interaction and dependency relationships between different positions in the image, and obtain the feature representation with global perception ability according to the learning results.

[0020] According to the vehicle-acquired image processing method proposed in the embodiment of the present application, by inputting the to-be-processed image collected by at least one sensor of the vehicle into the fine-grained image recognition model, the to-be-processed image can be downsampled and multi-head attention calculated through the SMSB module of the model to obtain a downsampled feature representation, and the downsampled feature representation can be subjected to channel attention weighting and patch embedding processing through the CAPEB module of the model to obtain an enhanced local feature representation, and the enhanced local feature representation can be subjected to depth convolution and spatial attention processing through the ESCB module of the model to obtain a feature representation with global perception ability, thereby generating a final output image. Thus, the problem that existing feature extraction methods are difficult to capture subtle differences, resulting in insufficient feature discrimination ability, is solved, the computational complexity is reduced, resource consumption is reduced, and computational efficiency is improved.

[0021] To achieve the above object, an embodiment of the second aspect of the present application proposes a vehicle-acquired image processing device, including:

[0022] An acquisition module, configured to acquire a to-be-processed image, where the to-be-processed image is collected by at least one sensor of the vehicle;

[0023] A processing module, configured to input the image to be processed into a fine-grained image recognition model, perform dimensionality reduction processing and multi-head attention calculation on the image to be processed through the SMSB module of the fine-grained image recognition model to obtain a feature representation after dimensionality reduction, and perform channel attention weighting and patch embedding processing on the feature representation after dimensionality reduction through the CAPEB module of the fine-grained image recognition model to obtain an enhanced local feature representation, and perform depth convolution and spatial attention processing on the enhanced local feature representation through the ESCB module of the fine-grained image recognition model to obtain a feature representation with global perception ability;

[0024] A generation module, configured to generate a final output image according to the feature representation with global perception ability.

[0025] According to an embodiment of the present application, the processing module is specifically configured to:

[0026] Map the feature representation of the image to be processed to a subspace of a preset dimension through a linear transformation to generate a query, a key, and a value after dimensionality reduction;

[0027] Based on a preset multi-head attention mechanism, decompose the query, key, and value after dimensionality reduction into multiple independent attention heads, where each attention head focuses on a different subspace;

[0028] Concatenate the outputs of the multiple independent attention heads, and perform a linear transformation on the concatenation result through a preset output weight matrix to obtain the feature representation after dimensionality reduction.

[0029] According to an embodiment of the present application, the processing module is specifically configured to:

[0030] Segment the feature representation after dimensionality reduction into multiple patches of the same size, and convert the multiple patches of the same size into vector representations to obtain a set of feature vectors after dimensionality reduction;

[0031] Based on a preset channel attention mechanism, calculate the weight of each channel, and perform weighted processing on the feature vectors after dimensionality reduction based on the calculation result, and segment the features weighted by the preset channel attention mechanism into multiple new patches of the same size, and convert the multiple new patches of the same size into vector representations to obtain a vector matrix;

[0032] Based on a preset normalization layer, normalize each vector in the vector matrix to obtain the enhanced local feature representation.

[0033] According to an embodiment of the present application, the processing module is specifically configured to:

[0034] Perform depth convolution processing on the enhanced local feature representation to obtain a feature map after depth convolution;

[0035] Perform spatial attention processing on the feature map after depth convolution, use the self-attention mechanism to learn the interaction and dependence relationships between different positions in the image, and obtain the feature representation of the global perception ability according to the learning results.

[0036] According to the vehicle-acquired image processing device proposed in the embodiments of the present application, by inputting the to-be-processed image collected by at least one sensor of the vehicle into the fine-grained image recognition model, the SMSB module of the model can perform dimensionality reduction processing and multi-head attention calculation on the to-be-processed image to obtain the dimensionality-reduced feature representation, and the CAPEB module of the model can perform channel attention weighting and patch embedding processing on the dimensionality-reduced feature representation to obtain the enhanced local feature representation, and the ESCB module of the model can perform depth convolution and spatial attention processing on the enhanced local feature representation to obtain the feature representation of the global perception ability, thereby generating the final output image. Thereby, the problem that existing feature extraction methods are difficult to capture subtle differences, resulting in insufficient feature discrimination ability, is solved, the computational complexity is reduced, resource consumption is reduced, and the computational efficiency is improved.

[0037] To achieve the above object, an embodiment of the third aspect of the present application proposes a vehicle, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the vehicle-acquired image processing method as described in the above embodiments.

[0038] To achieve the above object, an embodiment of the fourth aspect of the present application proposes a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to be used to implement the vehicle-acquired image processing method as described in the above embodiments.

[0039] To achieve the above object, an embodiment of the fifth aspect of the present application proposes a computer program product, which includes a computer program, and when the computer program is executed by a processor, it is used to implement the vehicle-acquired image processing method as described in the above embodiments.

[0040] The additional aspects and advantages of the present application will be partly given in the following description, partly become obvious from the following description, or be understood through the practice of the present application. Description of the Drawings

[0041] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0042] Figure 1 It is a flowchart of a vehicle-acquired image processing method provided according to an embodiment of the present application;

[0043] Figure 2 Schematic structural diagram of a fine-grained image recognition model according to an embodiment of the present application;

[0044] Figure 3 Block diagram of a processing device for vehicle-acquired images provided according to an embodiment of the present application;

[0045] Figure 4 Schematic structural diagram of a vehicle provided according to an embodiment of the present application. Detailed implementation manners

[0046] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.

[0047] A method, device, vehicle, medium, and program product for processing vehicle-acquired images according to embodiments of the present application will be described below with reference to the accompanying drawings. First, a method for processing vehicle-acquired images according to embodiments of the present application will be described with reference to the accompanying drawings.

[0048] Figure 1 Flowchart of a method for processing vehicle-acquired images according to an embodiment of the present application.

[0049] Before introducing the method for processing vehicle-acquired images proposed in the embodiments of the present application, the relevant technical background will be briefly introduced.

[0050] In the field of traditional deep learning, through a multi-level feature extraction process, a convolutional neural network can automatically extract complex visual features from a large amount of image data, providing strong support for image recognition technology. However, traditional convolutional neural networks have limitations in terms of recognition accuracy and computing speed. The introduction of the attention mechanism algorithm has significantly improved this situation. With the help of the attention mechanism, the network can focus more on the learning and application of key information, thereby improving the accuracy of image recognition.

[0051] In the field of fine-grained image recognition, the key differentiations often lie in the local regions of the images. Using the attention mechanism, the model can automatically filter and focus on those crucial local regions. In this way, the model can identify which regions are crucial for the image recognition task, thereby improving the recognition accuracy. A commonly used method is to weight different parts of the feature map through spatial attention or channel attention.

[0052] However, the field of fine-grained image recognition also faces a series of challenges. Fine-grained image recognition tasks usually involve numerous categories, while the number of training samples contained in each category is limited. This phenomenon leads to the problems of data sparsity and class imbalance, which in turn makes it difficult for the model to learn sufficient feature discrimination ability from limited data. Fine-grained image recognition requires the model to be able to learn feature representations that are sensitive to subtle differences. Although traditional hand-designed feature extraction methods and deep learning-based convolutional neural networks have certain limitations in capturing these subtle differences, they are still widely used in this field. To improve the recognition accuracy, there is an urgent need to develop more effective feature representation learning methods. Although ViT (Vision Transformer) has made some progress in computer vision tasks, it also has some defects. The number of parameters of ViT is usually much larger than that of traditional convolutional neural network models, which not only increases the computational resources and storage space requirements for training and deployment, but also enhances the complexity of the model and reduces the interpretability and explainability of the model. The ViT model mainly relies on the self-attention mechanism to model the global relationships in the image. In contrast, convolutional neural networks are more proficient in capturing the spatial information of local regions. Due to the lack of an explicit local receptive field in the self-attention mechanism, ViT may not perform as well as convolutional neural network models when dealing with tasks that rely on local context.

[0053] Based on the above problems, the embodiments of this application propose a method for processing vehicle-acquired images. By inputting the to-be-processed images collected by at least one sensor of the vehicle into a fine-grained image recognition model, the SMSB module of the model can perform dimensionality reduction processing and multi-head attention calculation on the to-be-processed images to obtain the dimensionality-reduced feature representations, and the CAPEB module of the model can perform channel attention weighting and patch embedding processing on the dimensionality-reduced feature representations to obtain enhanced local feature representations, and the ESCB module of the model can perform depth convolution and spatial attention processing on the enhanced local feature representations to obtain feature representations with global perception ability, thereby generating the final output images. Thus, the problem that existing feature extraction methods are difficult to capture subtle differences, resulting in insufficient feature discrimination ability, is solved, the computational complexity is reduced, the resource consumption is reduced, and the computational efficiency is improved.

[0054] Exemplarily, as Figure 1 shown, the method for processing vehicle-acquired images includes the following steps:

[0055] In step S101, obtain the to-be-processed images, where the to-be-processed images are collected by at least one sensor of the vehicle.

[0056] It can be understood that the image to be processed refers to image data that needs to be further processed or analyzed. A sensor refers to a device on a vehicle used to detect and respond to the external or internal environment of the vehicle, such as a camera, radar, etc.

[0057] That is to say, at least one sensor on the vehicle can obtain the image to be processed for analyzing the vehicle's surrounding environment, performing driving assistance, or safety detection, etc.

[0058] In step S102, the image to be processed is input into the fine-grained image recognition model. The SMSB module of the fine-grained image recognition model performs dimensionality reduction processing and multi-head attention calculation on the image to be processed to obtain a dimensionality-reduced feature representation. The CAPEB module of the fine-grained image recognition model performs channel attention weighting and patch embedding processing on the dimensionality-reduced feature representation to obtain an enhanced local feature representation. And the ESCB module of the fine-grained image recognition model performs depth convolution and spatial attention processing on the enhanced local feature representation to obtain a feature representation with global perception ability.

[0059] It should be noted that in order to more efficiently capture image features, an embodiment of the present application proposes a fine-grained image recognition model, such as Figure 2As shown in Figure 1, the model consists of three main modules, namely the SMAB module, the CAPEB module and the ESCB module. The SMAB module builds an efficient and powerful global feature extraction mechanism. The SMAB module can reduce the dimension of QKV, reduce the high-dimensional QKV to a low-dimensional subspace, and remove irrelevant information while ensuring that key information is retained. In addition, the SMAB module decomposes the traditional attention mechanism into several independent attention heads, each of which is responsible for a different subspace and performs calculations at different scales, thereby capturing multi-level feature representations and more effectively processing dependencies between different positions and scales. The CAPEB module builds an efficient and effective local feature extraction mechanism. The CAPEB module integrates the channel attention mechanism. By assigning different attention weights to channels, the module can more accurately identify key features in the image, reduce dependence on irrelevant information, and thus improve classification accuracy. At the same time, the patch embedding mechanism of the CAPEB module converts image patches into vector form, allowing the module to map them to a higher-dimensional feature space while maintaining the local structure of the image. This process enhances the expressiveness of the image and enables the module to more effectively capture the details and texture information in the image. The ESCB module can enhance the global and contextual information of the image. The ESCB module uses deep convolution technology to further reduce the number of parameters and computational requirements, while ensuring independence between channels and improving the expressiveness of features within the channel. In addition, the ESCB module uses the spatial attention mechanism to learn the interactions and dependencies between different positions in the image, thereby enhancing the global and contextual information of the image. The spatial attention mechanism can dynamically adjust the weights of each position to ensure that key areas receive more concentrated attention, while secondary areas receive less attention.

[0060] For ease of understanding, the SMSB module, CAPEB module and ESCB module of the fine-grained image recognition model are explained in detail below.

[0061] As a possible implementation method, in some embodiments, the SMSB module of the fine-grained image recognition model performs dimensionality reduction processing and multi-head attention calculation on the image to be processed to obtain the reduced-dimensional feature representation, including: mapping the feature representation of the image to be processed to a subspace of a preset dimension through a linear transformation, generating reduced-dimensional queries, keys, and values; based on a preset multi-head attention mechanism, decomposing the reduced-dimensional queries, keys, and values into multiple independent attention heads, wherein each attention head focuses on a different subspace; splicing the outputs of multiple independent attention heads, and performing a linear transformation on the splicing result through a preset output weight matrix to obtain the reduced-dimensional feature representation.

[0062] It is understandable that the SMSB module is a Transformer-based module that can capture dependencies at any distance between data and has strong generalization ability and interpretability. However, there are also some problems with Transformer, such as excessive computational complexity and memory consumption, as well as insufficient utilization of position information and multi-scale information. To solve these problems, the embodiments of this application have improved Transformer using this module.

[0063] Since the computational bottleneck of Transformer is the dimension of QKV (Query-Key-Value), the larger the dimension, the greater the computational complexity. By designing a new module to reduce the dimension of QKV (this method is logically feasible because there are redundant items in the feature information and an appropriate reduction factor can be found), this module provides an efficient and flexible implementation for the self-attention mechanism through advantages such as multi-head attention, position encoding, scaled dot-product attention, feature fusion, and downsampling processing.

[0064] Specifically, first, the SMSB module performs dimensionality reduction on QKV (the feature representation of the image to be processed), reducing the high-dimensional QKV to a low-dimensional subspace while ensuring that the reduced QKV (the reduced query, key, and value) can still retain sufficient information, as shown in the following formula:

[0065]

[0066] where i is the number of attention heads, is the dimension of the subspace, X ∈ R n×d ,

[0067] Second, the SMSB module introduces the multi-head attention mechanism, further improving the efficiency and expressiveness of Transformer. By decomposing the original attention mechanism into multiple independent attention heads, each head can focus on different subspaces and perform calculations at different scales. This multi-head mechanism enables Transformer to capture feature representations at different levels simultaneously and handle dependencies between different positions and scales better. This step can be represented by the following formula:

[0068]

[0069] where Q i , K i , V i are the query, key, and value vectors of the i-th head respectively.

[0070] Finally, the SMSB module can concatenate the outputs of all heads and multiply by an output weight matrix W o , to obtain the final output vector. The purpose of this is to fuse the information from different subspaces and restore it to the dimension of the original input data. This step can be expressed by the following formula:

[0071] Output=[head1·head2·…·head h W o ;

[0072] where, [·] represents the concatenation operation,

[0073] As a possible implementation, in some other embodiments, the CAPEB module of the fine-grained image recognition model performs channel attention weighting and patch embedding processing on the dimension-reduced feature representation to obtain an enhanced local feature representation, including: dividing the dimension-reduced feature representation into multiple patches of the same size, and converting the multiple patches of the same size into vector representations to obtain a set of dimension-reduced feature vectors; based on a preset channel attention mechanism, calculating the weight of each channel, and performing weighting processing on the dimension-reduced feature vectors based on the calculation result, and dividing the feature weighted by the preset channel attention mechanism into multiple new patches of the same size, and converting the multiple new patches of the same size into vector representations to obtain a vector matrix; based on a preset normalization layer, normalizing each vector in the vector matrix to obtain an enhanced local feature representation.

[0074] It can be understood that the CAPEB module is a module that divides the input image into multiple patches of the same size and converts each patch into a vector representation. It provides an effective local feature extractor for image classification tasks through advantages such as efficient image transformation, spatial perception ability, normalization processing, and flexible input adaptability. It can reduce the computational complexity while still capturing the important features of the image, improving the performance and efficiency of the model. This module mainly includes two parts: the channel attention mechanism and the patch embedding mechanism.

[0075] Among them, the channel attention mechanism is a method for adaptively adjusting the importance between different channels, enabling the module to pay more attention to channels with higher importance, improving the feature expression ability, and weakening the dependence on irrelevant information. By weighting the attention of channels, the module can more accurately capture the key features in the image, thereby improving the classification accuracy.

[0076] Assume the input image is:

[0077] I ∈ ℝ C×H×w ;

[0078] Where C is the number of channels, H is the height, and W is the width.

[0079] The formula for the channel attention mechanism is:

[0080] S = softmax(W2δ(W1I)) ∈ ℝ C×1×1 ;

[0081]

[0082] Where W1, r is the reduction ratio, δ(·) is the activation function, and ⊙ is the Hadamard product (element-wise multiplication).

[0083] This step calculates the attention score S for each channel through a global average pooling layer and two fully connected layers, and then weights the input image I with these scores to obtain the image after channel attention weighting. The purpose of doing this is to adaptively adjust the importance between different channels, so that the module can pay more attention to the channels with higher importance.

[0084] The patch embedding mechanism is a method that divides the input image into multiple patches of the same size and converts each patch into a vector representation. It can reduce the spatial dimension of the image while retaining the global information of the image. This conversion enables the module to better process the local information of the image. By converting the patches into vector representations, the module can map them to a higher-dimensional feature space while maintaining the local structure of the image. This can enhance the expressive power of the image, enabling the module to better capture the details and texture information in the image. At the same time, the patch embedding mechanism can also improve the stability and robustness of the module through normalization processing, enabling the module to perform accurate feature extraction on images of different scales and resolutions.

[0085] The formula for the patch embedding mechanism is:

[0086]

[0087] E = PW3 + B3 ∈ ℝ N×D ;

[0088]

[0089] Where P h 、P w are the height and width of the patch, N = H / P h W / P w is the number of patches, W3, D is the embedding dimension, and LN(·) is layer normalization.

[0090] In this step, the image I weighted by channel attention is segmented into multiple patches of the same size, and each patch is converted into a vector representation to obtain a matrix P. Then, a linear transformation layer is used to map each vector into a higher-dimensional feature space to obtain a matrix E. Finally, a layer normalization layer is used to normalize each vector to obtain a matrix The purpose of doing this is to reduce the spatial dimension of the image, while retaining the global information of the image and enhancing the perception ability of the image.

[0091] As a possible implementation method, in some other embodiments, the enhanced local feature representation is subjected to depth convolution and spatial attention processing through the ESCB module of the fine-grained image recognition model to obtain a feature representation with global perception ability, including: performing depth convolution processing on the enhanced local feature representation to obtain a feature map after depth convolution; performing spatial attention processing on the feature map after depth convolution, using the self-attention mechanism to learn the interaction and dependency relationships between different positions in the image, and obtaining a feature representation with global perception ability according to the learning results.

[0092] It can be understood that the ESCB module is a convolution-based image feature extractor, which can effectively process image inputs with high resolution and complex scenes, and provides an efficient feature representation for the image semantic segmentation task. This module mainly includes two parts: depth convolution (Dwconv) and spatial attention (SpatialAttention).

[0093] Among them, depth convolution is a method that uses depth-wise convolution to reduce the number of parameters and computational amount. It can enhance the feature expression within the channel while maintaining the independence between channels. Depth convolution can effectively reduce the spatial dimension of the feature map while retaining the channel dimension of the feature map, providing more information for subsequent spatial attention.

[0094] Spatial attention is a method that uses the self-attention mechanism to learn the interaction and dependency relationships between different positions in the image. It can enhance the global information and context information of the image. Spatial attention can dynamically adjust the weight of each position, so that important positions receive more attention while unimportant positions are suppressed. The calculation formula is as follows:

[0095] Assume the input image is:

[0096] I∈R C×H×W ;

[0097] The depth convolution formula is as follows:

[0098] F = Dwconv(I) ∈ R C×H′×W′ ;

[0099] where H' = H / S h , W' = W / S w , S h , S w is the convolution stride.

[0100] This step performs a convolution operation on the input image I through channel-wise convolution to obtain a feature map F.

[0101] The spatial attention formula is as follows:

[0102] A = softmax(FF T )R N′×N′ ;

[0103] O = AF × R N′×C ;

[0104] where N' = H' × W', F T is the transpose of the matrix.

[0105] The spatial attention module can establish connections between different spatial positions to form a spatial attention matrix A, which reflects the influence degree of each position in the feature map F on other positions, that is, the attention weight. Then, by using this matrix to perform weighted summation on the feature map F, an output feature vector O can be obtained, which contains the global information and context information of the image. Thus, important positions can be highlighted and unimportant positions can be suppressed according to the attention weight, so as to extract more semantic features.

[0106] In step S103, the final output image is generated according to the feature representation of the global perception ability.

[0107] That is to say, after being processed by the SMSB module, CAPEB module, and ESCB module of the fine-grained image recognition model in sequence, the feature representation of the global perception ability can be obtained, and the final output image can be generated based on this feature representation of the global perception ability.

[0108] Furthermore, by using the CUB-200-2011 dataset, Keras-Birds dataset, and NABirds dataset, the performance of the fine-grained image recognition model (LWT) proposed in the embodiments of the present application can be compared with that of other different models (such as ResNet-152, MobileNetV2, ViT) in the fine-grained image recognition task, and the comparison results are shown in Tables 1 to 3.

[0109] Table 1

[0110]

[0111] Table 2

[0112]

[0113] Table 3

[0114]

[0115]

[0116] Thus, by comparing the performance metrics of different models on different datasets, it can be clearly found that the fine-grained image recognition model (LWT) proposed in the embodiments of the present application has superiority in fine-grained image recognition.

[0117] According to the method for processing vehicle-acquired images proposed in the embodiments of the present application, by inputting the to-be-processed image collected by at least one sensor of the vehicle into the fine-grained image recognition model, the SMSB module of the model can perform dimensionality reduction processing and multi-head attention calculation on the to-be-processed image to obtain a dimensionality-reduced feature representation, and the CAPEB module of the model can perform channel attention weighting and patch embedding processing on the dimensionality-reduced feature representation to obtain an enhanced local feature representation, and the ESCB module of the model can perform depth convolution and spatial attention processing on the enhanced local feature representation to obtain a feature representation with global perception ability, thereby generating a final output image. Thus, the problem that existing feature extraction methods are difficult to capture subtle differences, resulting in insufficient feature discrimination ability, is solved, the computational complexity is reduced, resource consumption is reduced, and the computational efficiency is improved.

[0118] Next, a device for processing vehicle-acquired images proposed in the embodiments of the present application will be described with reference to the accompanying drawings.

[0119] Figure 3 It is a block diagram of a device for processing vehicle-acquired images according to an embodiment of the present application.

[0120] As Figure 3 shown, the device 10 for processing vehicle-acquired images includes: an acquisition module 100, a processing module 200, and a generation module 300.

[0121] Among them, the acquisition module 100 is configured to acquire a to-be-processed image, where the to-be-processed image is collected by at least one sensor of the vehicle;

[0122] A processing module 200 is configured to input an image to be processed into a fine-grained image recognition model, perform dimensionality reduction processing and multi-head attention calculation on the image to be processed through the SMSB module of the fine-grained image recognition model to obtain a dimensionally reduced feature representation, and perform channel attention weighting and patch embedding processing on the dimensionally reduced feature representation through the CAPEB module of the fine-grained image recognition model to obtain an enhanced local feature representation, and perform depth convolution and spatial attention processing on the enhanced local feature representation through the ESCB module of the fine-grained image recognition model to obtain a feature representation with global perception ability;

[0123] A generation module 300 is configured to generate a final output image according to the feature representation with global perception ability.

[0124] Further, in some embodiments, the processing module 200 is specifically configured to:

[0125] Map the feature representation of the image to be processed to a subspace of a preset dimension through a linear transformation to generate a dimensionally reduced query, key, and value;

[0126] Based on a preset multi-head attention mechanism, decompose the dimensionally reduced query, key, and value into multiple independent attention heads, where each attention head focuses on a different subspace;

[0127] Concatenate the outputs of the multiple independent attention heads, and perform a linear transformation on the concatenation result through a preset output weight matrix to obtain a dimensionally reduced feature representation.

[0128] Further, in some embodiments, the processing module 200 is specifically configured to:

[0129] Segment the dimensionally reduced feature representation into multiple patches of the same size, and convert the multiple patches of the same size into vector representations to obtain a set of dimensionally reduced feature vectors;

[0130] Based on a preset channel attention mechanism, calculate the weight of each channel, and perform a weighting process on the dimensionally reduced feature vectors based on the calculation result, and segment the features weighted by the preset channel attention mechanism into multiple new patches of the same size, and convert the multiple new patches of the same size into vector representations to obtain a vector matrix;

[0131] Based on a preset normalization layer, normalize each vector in the vector matrix to obtain an enhanced local feature representation.

[0132] Further, in some embodiments, the processing module 200 is specifically configured to:

[0133] Perform depth convolution processing on the enhanced local feature representation to obtain a depth-convolved feature map;

[0134] Perform spatial attention processing on the feature map after depth convolution, use the self-attention mechanism to learn the interactions and dependencies between different positions in the image, and obtain a feature representation with global perception ability according to the learning results.

[0135] It should be noted that the above explanation of the embodiments of the method for processing vehicle-acquired images also applies to the device for processing vehicle-acquired images in this embodiment, and will not be elaborated here.

[0136] According to the device for processing vehicle-acquired images proposed in the embodiments of the present application, by inputting the to-be-processed images collected by at least one sensor of the vehicle into the fine-grained image recognition model, the SMSB module of the model can perform dimensionality reduction processing and multi-head attention calculation on the to-be-processed images to obtain a dimensionality-reduced feature representation, and the CAPEB module of the model can perform channel attention weighting and patch embedding processing on the dimensionality-reduced feature representation to obtain an enhanced local feature representation, and the ESCB module of the model can perform depth convolution and spatial attention processing on the enhanced local feature representation to obtain a feature representation with global perception ability, thereby generating the final output image. Thus, the problem that existing feature extraction methods are difficult to capture subtle differences, resulting in insufficient feature discrimination ability, is solved, the computational complexity is reduced, resource consumption is reduced, and the computational efficiency is improved.

[0137] Figure 4 This is a schematic structural diagram of a vehicle provided by the embodiments of the present application. The vehicle may include:

[0138] A memory 401, a processor 402, and a computer program stored on the memory 401 and executable on the processor 402.

[0139] When the processor 402 executes the program, it implements the method for processing vehicle-acquired images provided in the above embodiments.

[0140] Furthermore, the vehicle further includes:

[0141] A communication interface 403 for communication between the memory 401 and the processor 402.

[0142] The memory 401 is used to store a computer program executable on the processor 402.

[0143] The memory 401 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.

[0144] If the memory 401, the processor 402, and the communication interface 403 are implemented independently, the communication interface 403, the memory 401, and the processor 402 can be interconnected via a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 only a thick line is used to represent it in Figure 4 , but it does not mean that there is only one bus or one type of bus.

[0145] Optionally, in a specific implementation, if the memory 401, the processor 402, and the communication interface 403 are integrated on a single chip, the memory 401, the processor 402, and the communication interface 403 can communicate with each other through an internal interface.

[0146] The processor 402 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.

[0147] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned method for processing vehicle-acquired images is implemented.

[0148] The embodiments of the present application also provide a computer program product, which includes a computer program, and when the computer program is executed by a processor, the above-mentioned method for processing vehicle-acquired images is implemented.

[0149] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0150] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0151] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A method for processing vehicle-acquired images, characterized in that, Including the following steps: Obtain an image to be processed, where the image to be processed is collected by at least one sensor of a vehicle; Input the image to be processed into a fine-grained image recognition model. Through the SMSB module of the fine-grained image recognition model, perform dimensionality reduction processing and multi-head attention calculation on the image to be processed to obtain a feature representation after dimensionality reduction, and through the CAPEB module of the fine-grained image recognition model, perform channel attention weighting and patch embedding processing on the feature representation after dimensionality reduction to obtain an enhanced local feature representation, and through the ESCB module of the fine-grained image recognition model, perform depth convolution and spatial attention processing on the enhanced local feature representation to obtain a feature representation with global perception ability; Generate a final output image according to the feature representation with global perception ability.

2. The method according to claim 1, characterized in that, The step of performing dimensionality reduction processing and multi-head attention calculation on the image to be processed through the SMSB module of the fine-grained image recognition model to obtain a feature representation after dimensionality reduction includes: Map the feature representation of the image to be processed to a subspace of a preset dimension through a linear transformation to generate a query, a key, and a value after dimensionality reduction; Based on a preset multi-head attention mechanism, decompose the query, key, and value after dimensionality reduction into multiple independent attention heads, where each attention head focuses on a different subspace; Concatenate the outputs of the multiple independent attention heads, and perform a linear transformation on the concatenation result through a preset output weight matrix to obtain the feature representation after dimensionality reduction.

3. The method according to claim 2, characterized in that, The step of performing channel attention weighting and patch embedding processing on the feature representation after dimensionality reduction through the CAPEB module of the fine-grained image recognition model to obtain an enhanced local feature representation includes: Segment the feature representation after dimensionality reduction into multiple patches of the same size, and convert the multiple patches of the same size into vector representations to obtain a set of feature vectors after dimensionality reduction; Based on a preset channel attention mechanism, calculate the weight of each channel, and perform weighting processing on the feature vectors after dimensionality reduction based on the calculation result. Segment the features weighted by the preset channel attention mechanism into multiple new patches of the same size, and convert the multiple new patches of the same size into vector representations to obtain a vector matrix; Based on a preset normalization layer, normalize each vector in the vector matrix to obtain the enhanced local feature representation.

4. The method according to claim 3, wherein The step of performing depth convolution and spatial attention processing on the enhanced local feature representation through the ESCB module of the fine-grained image recognition model to obtain a feature representation with global perception ability includes: Perform depth convolution processing on the enhanced local feature representation to obtain a feature map after depth convolution; Perform spatial attention processing on the feature map after depth convolution, use the self-attention mechanism to learn the interaction and dependence relationships between different positions in the image, and obtain the feature representation with global perception ability according to the learning result.

5. A processing device for vehicle-acquired images, characterized in that, Including: An acquisition module for acquiring an image to be processed, where the image to be processed is collected by at least one sensor of a vehicle; A processing module, configured to input the image to be processed into a fine-grained image recognition model, perform dimensionality reduction processing and multi-head attention calculation on the image to be processed through the SMSB module of the fine-grained image recognition model to obtain a feature representation after dimensionality reduction, and perform channel attention weighting and patch embedding processing on the feature representation after dimensionality reduction through the CAPEB module of the fine-grained image recognition model to obtain an enhanced local feature representation, and perform depth convolution and spatial attention processing on the enhanced local feature representation through the ESCB module of the fine-grained image recognition model to obtain a feature representation with global perception ability; A generation module, configured to generate a final output image according to the feature representation with global perception ability.

6. The device according to claim 5, characterized in that, The processing module is specifically configured to: Map the feature representation of the image to be processed to a subspace of a preset dimension through a linear transformation to generate a query, a key, and a value after dimensionality reduction; Based on a preset multi-head attention mechanism, decompose the query, the key, and the value after dimensionality reduction into multiple independent attention heads, where each attention head focuses on a different subspace; Concatenate the outputs of the multiple independent attention heads, and perform a linear transformation on the concatenation result through a preset output weight matrix to obtain the feature representation after dimensionality reduction.

7. The device according to claim 6, wherein The processing module is specifically configured to: Segment the feature representation after dimensionality reduction into multiple patches of the same size, and convert the multiple patches of the same size into vector representations to obtain a set of feature vectors after dimensionality reduction; Based on a preset channel attention mechanism, calculate the weight of each channel, and perform weighting processing on the feature vectors after dimensionality reduction based on the calculation result, and segment the features weighted by the preset channel attention mechanism into multiple new patches of the same size, and convert the multiple new patches of the same size into vector representations to obtain a vector matrix; Based on a preset normalization layer, normalize each vector in the vector matrix to obtain the enhanced local feature representation.

8. A vehicle, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the method for processing a vehicle-acquired image according to any one of claims 1-4.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to be used to implement the method for processing a vehicle-acquired image according to any one of claims 1-4.

10. A computer program product, characterized in that, It includes a computer program, where when the computer program is executed by the processor, it is used to implement the method for processing a vehicle-acquired image according to any one of claims 1-4.