Vehicle identification method, electronic equipment and readable storage medium
By preprocessing and feature separation of vehicle images, combined with the vehicle recognition model of multi-head self-attention unit and key information selection module, the problem of insufficient subtle feature capture in vehicle segmentation and charging type recognition is solved, and the recognition accuracy and robustness are improved.
Patent Information
- Application Number
- CN202510394917.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art lacks the ability to capture subtle features in the identification of vehicle segmentation charging types, especially in complex environments (such as light changes, occlusions and viewing angle changes), and classification performance deteriorates.
By preprocessing the vehicle image, separating the image based on the preset sliding step length and extracting features, the vehicle category recognition model of the multi-head self-attention unit and the key information selection module is used to identify the vehicle category, including a multi-layer convolutional layer and a self-attention layer, to enhance the attention to local areas.
提高了车辆细分类型的识别精度、鲁棒性和实用性,增强了对复杂环境的适应能力。
Smart Images

Figure CN120298788A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, and more particularly, to a vehicle recognition method, an electronic device, and a readable storage medium. Background Art
[0002] With the rapid development of intelligent transportation systems, the automatic recognition of vehicle sub-classification charging types has become an important part of traffic management. In practical applications, the recognition of vehicle sub-classification charging types not only needs to distinguish between major vehicle categories, such as trucks and buses, but also needs to further identify sub-categories, such as different load types of trucks, in order to achieve accurate charging.
[0003] Currently, the following methods are mainly used for the recognition of vehicle sub-classification charging types: Based on local feature extraction and traditional machine learning models, first identify the major vehicle category, and then perform sub-classification based on local features such as the number of wheels.
[0004] However, the existing technology has insufficient ability to capture subtle features, and the classification performance decreases in the face of complex environments, such as changes in lighting, occlusion, and viewing angles. Summary of the Invention
[0005] The purpose of this application is to provide a vehicle recognition method, an electronic device, and a readable storage medium to solve the problem of poor vehicle classification performance in the existing technology for the deficiencies in the above-mentioned existing technology.
[0006] To achieve the above purpose, the technical solutions adopted in this application are as follows:
[0007] In a first aspect, this application provides a vehicle recognition method, the method includes:
[0008] Obtain a vehicle image of the current vehicle;
[0009] Preprocess the vehicle image of the current vehicle to obtain an image to be segmented;
[0010] Based on a preset sliding step, segment the image to be segmented and perform feature extraction to obtain a plurality of initial image features, where there is an overlapping area between the image blocks corresponding to adjacent initial image features;
[0011] Input the plurality of initial image features into a pre-trained vehicle recognition model to obtain the vehicle category of the current vehicle. The vehicle recognition model includes a plurality of encoding modules and a key information selection module. The key information selection module is located before the last encoding module. The key information selection module is used to screen the received image features and output them to the last encoding model. Each encoding module sequentially includes a multi-head self-attention unit and a multi-layer perceptron unit. The multi-head self-attention unit includes: a first convolutional layer, a multi-head self-attention layer, and a second convolutional layer.
[0012] Optionally, preprocessing the vehicle image of the current vehicle to obtain an image to be segmented, including:
[0013] Performing parameter prediction, coordinate mapping, and resampling on the vehicle image of the current vehicle in sequence based on a spatial transformation network to obtain an image to be segmented.
[0014] Optionally, segmenting the image to be segmented based on a preset sliding step and performing feature extraction to obtain a plurality of initial image features, including:
[0015] Segmenting the image to be segmented based on a preset sliding step to obtain a plurality of image patches;
[0016] Projecting the plurality of image patches into a preset embedding space to generate a plurality of embedding vectors;
[0017] Determining a plurality of initial image features according to the plurality of embedding vectors.
[0018] Optionally, the determining a plurality of initial image features according to the plurality of embedding vectors includes:
[0019] Determining the position encoding corresponding to each image patch;
[0020] Taking the sum of each embedding vector and the corresponding position encoding as an initial image feature.
[0021] Optionally, the determining the position encoding corresponding to each image patch includes:
[0022] Determining the position encoding corresponding to each image patch based on a sine-cosine position encoding algorithm.
[0023] Optionally, the key information selection module is specifically configured to:
[0024] Receive the first image feature sent by the previous encoding module;
[0025] Generating a plurality of target attention weight matrices based on the attention weight matrices of each encoding module, where each target attention weight matrix corresponds to one attention head in each encoding module;
[0026] Determining a target position index according to each target attention weight matrix, where the target position index is used to represent the position index of the image patch corresponding to vehicle classification;
[0027] Determining a second image feature according to the target position index and the first image feature, and inputting the second image feature into the last encoding module.
[0028] Optionally, generating a plurality of target attention weight matrices based on the attention weight matrices of the respective encoding modules includes:
[0029] Performing a recursive multiplication operation on the attention weight matrices of the respective encoding modules to obtain a target attention weight matrix.
[0030] Optionally, determining a target position index according to the respective target attention weight matrices includes:
[0031] Determining a preset number of attention heads with the largest weights according to the respective target attention weight matrices, and using the position indices of the preset number of attention heads as the target position index.
[0032] In a second aspect, the present application provides a vehicle recognition device, and the device includes:
[0033] An acquisition module, configured to acquire a vehicle image of a current vehicle;
[0034] A preprocessing module, configured to preprocess the vehicle image of the current vehicle to obtain an image to be segmented;
[0035] A segmentation module, configured to segment the image to be segmented based on a preset sliding step length and perform feature extraction to obtain a plurality of initial image features, wherein there is an overlapping area between the image blocks corresponding to adjacent initial image features;
[0036] An identification module, configured to input the plurality of initial image features into a pre-trained vehicle recognition model to obtain the vehicle category of the current vehicle. The vehicle recognition model includes a plurality of encoding modules and a key information selection module. The key information selection module is located before the last encoding module. The key information selection module is configured to screen the received image features and output them to the last encoding model. Each of the encoding modules sequentially includes a multi-head self-attention unit and a multi-layer perceptron unit. The multi-head self-attention unit includes: a first convolutional layer, a multi-head self-attention layer, and a second convolutional layer.
[0037] In a third aspect, the present application provides an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the vehicle recognition method described in the first aspect above.
[0038] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it performs the steps of the vehicle recognition method described in the first aspect above.
[0039] The beneficial effects of this application are as follows: The vehicle image of the current vehicle obtained is preprocessed to improve fine-grained classification, so as to better focus on the details related to the classification task. Then, based on a preset sliding step size, the image is segmented and features are extracted to obtain multiple initial image features, so as to better retain local area information. Then, the multiple initial image features are input into the vehicle recognition model to obtain the vehicle category of the current vehicle. In the vehicle recognition model, a first convolutional layer and a second convolutional layer are included before and after multiple self-attention layers in the multi-head self-attention unit, so as to capture shallow features. The key information selection module in the vehicle recognition model reduces the misleading of the single-layer attention weight and enhances the attention to the discriminative area. This embodiment improves the recognition accuracy, robustness and practicability of the vehicle sub-types. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0041] Figure 1 is a flowchart of a vehicle recognition method provided by an embodiment of the present application;
[0042] Figure 2 is a schematic structural diagram of a vehicle recognition model provided by an embodiment of the present application;
[0043] Figure 3 is a schematic structural diagram of a spatial transformation network provided by an embodiment of the present application;
[0044] Figure 4 is a schematic flowchart of obtaining multiple initial image features provided by an embodiment of the present application;
[0045] Figure 5 is a schematic working flowchart of a key information selection module provided by an embodiment of the present application;
[0046] Figure 6 is a schematic structural diagram of a vehicle recognition device provided by an embodiment of the present application;
[0047] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. It should be understood that the accompanying drawings in this application are only for the purposes of illustration and description, and are not used to limit the protection scope of this application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of this application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without a logical context relationship may be reversed or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of this application.
[0049] In addition, the described embodiments are only some embodiments of this application, rather than all embodiments. The components of the embodiments of this application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of this application that is claimed, but merely represents the selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative efforts belong to the scope of protection of this application.
[0050] It should be noted that the term "including" will be used in the embodiments of this application to indicate the existence of the features stated thereafter, but does not exclude the addition of other features.
[0051] In scenarios such as highway toll stations or parking lots, there are different toll standards for different vehicle types. Vehicle types can include mini and small passenger cars with less than 9 seats (including 9 seats) and a vehicle length less than 6 meters, medium-sized passenger cars with 10 - 19 seats and a vehicle length less than 6 meters, large passenger cars with 20 - 39 seats (including 39 seats) and a vehicle length not less than 6 meters, large passenger cars with 40 seats or more (including 40 seats) and a vehicle length not less than 9 meters, and trucks of categories 1 to 6. The toll station needs to identify the vehicle type and charge the vehicle.
[0052] In the prior art, the method for identifying vehicle sub - classification toll types is based on local feature extraction and traditional machine learning models. First, the major vehicle category is identified, and then fine - classification is performed according to local features such as the number of wheels. However, the prior art has insufficient ability to capture subtle features, and when facing complex environments such as light changes, occlusion, and perspective changes, the classification performance deteriorates.
[0053] Based on this, the present application provides a vehicle recognition method. After preprocessing the acquired vehicle image of the current vehicle, the image is segmented based on a preset sliding step size and features are extracted to obtain multiple initial image features. Then, the multiple initial image features are input into a vehicle recognition model to obtain the vehicle category of the current vehicle, thereby improving the recognition accuracy, robustness, and practicality of the vehicle sub-type.
[0054] Next, refer to Figure 1 to introduce the specific steps of the vehicle recognition method. Among them, Figure 1 is the method flowchart of a vehicle recognition method provided by an embodiment of the present application.
[0055] S101. Obtain the vehicle image of the current vehicle.
[0056] Optionally, the vehicle image can be an image of the current vehicle passing by a highway toll station or a parking lot.
[0057] S102. Preprocess the vehicle image of the current vehicle to obtain an image to be segmented.
[0058] Optionally, the preprocessing can be to perform an affine transformation on the vehicle image to obtain the image to be segmented. By preprocessing the vehicle image, the key information in the vehicle image can be extracted, and the region of interest can be found, such as deleting the background region in the picture and performing operations such as rotating and magnifying the vehicle region.
[0059] S103. Based on the preset sliding step size, segment the image to be segmented and perform feature extraction to obtain multiple initial image features, where there is an overlapping region between the image blocks corresponding to adjacent initial image features.
[0060] Among them, the preset sliding step size is the window length for segmenting the image to be segmented, and the width of the overlapping region between the image blocks corresponding to two adjacent initial image features is the preset sliding step size.
[0061] Optionally, the smaller the preset sliding step size, the more accurate the vehicle recognition, but the computational complexity increases. The larger the preset sliding step size, the coarser the vehicle recognition, but the computational complexity is smaller.
[0062] S104. Input the multiple initial image features into a pre-trained vehicle recognition model to obtain the vehicle category of the current vehicle. The vehicle recognition model includes multiple encoding modules and a key information selection module. The key information selection module is located before the last encoding module. The key information selection module is used to screen the received image features and output them to the last encoding model. Each encoding module sequentially includes a multi-head self-attention unit and a multi-layer perceptron unit. The multi-head self-attention unit includes: a first convolutional layer, a multi-head self-attention layer, and a second convolutional layer.
[0063] Optionally,Figure 2 This is a schematic structural diagram of a vehicle recognition model provided by an embodiment of the present application. As Figure 2 shown, the vehicle recognition model includes multiple encoding modules and a key information selection module, and the key information selection module is located before the last encoding module.
[0064] Among them, each encoding module sequentially includes a multi-head self-attention unit and a multi-layer perceptron unit. The multi-head self-attention unit includes: a first convolutional layer, a multi-head self-attention layer, and a second convolutional layer. Each encoding module also includes a first normalization layer, and the multi-layer perceptron unit includes a second normalization layer and a multi-layer perceptron layer. Specifically, after the encoding module receives the input features, the first normalization layer normalizes them to improve the training stability and accelerate convergence. The first convolutional layer extracts shallow features through convolutional operations to enhance the model's ability to capture local structures. The multi-head self-attention layer captures the relationships at different positions in the sequence through the multi-head attention mechanism to enhance the understanding of global context. The second convolutional layer further extracts features to refine the features extracted by the multi-head sub-attention layer. A residual connection is made between the extracted features and the input features to alleviate the vanishing gradient and further train the deep network. The second normalization layer further normalizes the features to improve the training stability. The multi-layer perceptron layer performs a non-linear transformation through a fully connected layer to enhance the model's expressive ability to obtain output features. The features obtained from the first residual connection and the output features are residually connected again to further alleviate the vanishing gradient and improve the model performance.
[0065] Specifically, the single-head attention in the feature extraction process of the multi-head self-attention layer can be represented by formula (1):
[0066]
[0067] where Q, K, and V are the query matrix, key matrix, and value matrix respectively, and Q, K, and V are generated from the input feature sequence. d k is the dimension of the query and the key, and softmax is the normalization. After obtaining the single-head attention result, a concatenation operation is performed on it to obtain the output results of multiple heads.
[0068] The multi-layer perceptron layer can be represented by the following formula (2):
[0069] FFN(x) = ReLU(xW1 + b1)W2 + b2 (2)
[0070] where x is the input feature, ReLU is the activation function, W1 is the weight matrix of the first linear transformation, b1 is the bias vector of the first linear transformation, W2 is the weight matrix of the second linear transformation, and b2 is the bias vector of the second linear transformation.
[0071] Each residual connection and normalization layer can be represented by the following formula (3):
[0072] xout = LayerNorm(x + Submodule(x)) (3)
[0073] Among them, LayerNorm is a layer normalization operation, x is the input feature, and Submodule represents adding the input and output of a module. Among them, the module can be, for example, the above-mentioned multi-head self-attention unit or multi-layer perceptron unit.
[0074] Optionally, the key information selection module is used to filter the received image features, so as to retain the features corresponding to the information-rich regions, and then output these features to the last encoding model for encoding. As an alternative implementation, the attention weights of all previous layers can be integrated to obtain the target attention weights, and then the image features are filtered based on the target attention weights. Thus, the misleading caused by the single-layer attention weights can be reduced.
[0075] In this embodiment, the vehicle image of the current vehicle obtained is preprocessed to improve fine-grained classification, so as to better focus on the details related to the classification task. Then, based on a preset sliding step, the image is segmented and features are extracted to obtain multiple initial image features, so as to better retain the local region information. Then, the multiple initial image features are input into the vehicle recognition model to obtain the vehicle category of the current vehicle. In the vehicle recognition model, a first convolutional layer and a second convolutional layer are included before and after the multi-layer self-attention layer in the multi-head self-attention unit, so as to capture shallow features. The key information selection module in the vehicle recognition model reduces the misleading of the single-layer attention weights and enhances the attention to the discriminative regions. This embodiment improves the recognition accuracy, robustness, and practicality of the vehicle sub-types.
[0076] As an alternative implementation, the specific method for preprocessing the vehicle image of the current vehicle in step S102 above to obtain the image to be segmented can be: based on the spatial transformation network, the vehicle image of the current vehicle is sequentially subjected to parameter prediction, coordinate mapping, and resampling to obtain the image to be segmented.
[0077] Among them, the spatial transformation network (Spatial Transformer Networks, STN) can extract the key information of the vehicle image through spatial transformation in the spatial domain. Specifically, STN can autonomously find the region of interest and perform rotation, scaling, and affine transformation on it, so as to discard unimportant information.
[0078] Figure 3 It is a schematic structural diagram of a spatial transformation network provided by an embodiment of the present application. As Figure 3As shown in the figure, the STN includes an input module, a parameter prediction network module, a coordinate mapping module, a resampling module, and an output module. Among them, the input module is used to receive vehicle images, the parameter prediction network module is used to analyze the input data through a neural network, predict the parameters required for spatial transformation, such as multiple parameters of affine transformation, and send the parameters to the coordinate mapping module. The coordinate mapping module is used to establish the mapping relationship between the positions of each pixel in the output image and the input image by using the predicted parameters. The resampling module is used to sample pixel values from the input image according to the coordinate mapping result to generate a transformed image. The output module is used to output the transformed image, that is, the image to be segmented.
[0079] Optionally, after testing, using the spatial transformation network to process vehicle images, compared with directly obtaining the image to be segmented without processing the vehicle images, the vehicle classification accuracy is increased by 1%.
[0080] In this embodiment, the parameter prediction, coordinate mapping, and resampling of the vehicle image are performed through the spatial transformation network to obtain the image to be segmented, thereby magnifying or focusing on the local area in the vehicle image, which is beneficial for fine-grained classification and enables the model to better focus on the details related to the classification task.
[0081] Next, with reference to Figure 4 The specific process of segmenting the image to be segmented based on a preset sliding step length and extracting features in step S103 above to obtain multiple initial image features will be introduced. Among them, Figure 4 is a schematic flowchart of a process for obtaining multiple initial image features provided by an embodiment of the present application.
[0082] S401. Segment the image to be segmented based on a preset sliding step length to obtain a plurality of image patches.
[0083] Optionally, if no preset sliding step length is set and the image to be segmented is directly segmented to obtain non-overlapping image patches, it may damage the local adjacent structure. Especially when the discriminative region is segmented, it will affect the vehicle type judgment result. Therefore, segmenting the image to be segmented based on a preset sliding step length to obtain a plurality of image patches, thereby retaining the local area information.
[0084] Exemplarily, if the resolution of the image to be segmented is H*W, the size of the image patch is P, and the preset sliding step length is S, then N image patches can be obtained, as shown in the following formula (4):
[0085]
[0086] There is an overlapping region of size (P - S)*P between two adjacent image patches.
[0087] S402. Project the plurality of image patches into a preset embedding space to generate a plurality of embedding vectors.
[0088] Specifically, after vectorizing multiple image patches, they are mapped into a trainable embedding space, which can be a D-dimensional projection matrix E. This matrix is a linear layer, and then a linear transformation is performed on each image patch vector, and finally the embedding vectors of each image patch are obtained. Among them, the process of mapping each image patch to a fixed embedding vector using a linear layer or a convolution operation is shown in the following formula (5):
[0089]
[0090] S403. Determine multiple initial image features according to multiple embedding vectors.
[0091] As an alternative implementation, position encodings corresponding to multiple embedding vectors can be determined, and then multiple initial image features are determined according to the multiple embedding vectors and the position encodings.
[0092] In this embodiment, by segmenting the image to be segmented based on a preset sliding step, the local adjacent structure between each image patch is protected. Then, multiple image patches are projected into a preset embedding space to generate multiple embedding vectors, and multiple initial image features are determined according to the multiple embedding vectors, thereby improving the dependence relationship between features.
[0093] As an alternative implementation, the specific steps of determining multiple initial image features according to multiple embedding vectors in the above step S403 are as follows:
[0094] Optionally, determine the position encoding corresponding to each image patch.
[0095] Specifically, according to the position of each image patch in the image to be segmented, the position encoding corresponding to each image patch is determined.
[0096] Optionally, the sum of each embedding vector and the corresponding position encoding is used as an initial image feature.
[0097] Specifically, the sum of each embedding vector and the corresponding position encoding can be as shown in the following formula (6):
[0098]
[0099] where E pos is the position encoding.
[0100] In this embodiment, by using the sum of each embedding vector and the corresponding position encoding as an initial image feature, the image patch is mapped from the pixel space to the high-dimensional semantic space to capture local features.
[0101] Among them, the method for determining the position encoding corresponding to each image patch can be: based on the sine-cosine position encoding algorithm, determine the position encoding corresponding to each image patch.
[0102] Specifically, the position encoding can adopt the sine-cosine position encoding method to preserve the sequence information, as shown in the following formulas (7) and (8):
[0103]
[0104]
[0105] Among them, pos represents the position index in the image patch, d represents the total number of dimensions, and i represents the dimension index.
[0106] Optionally, if the embedding vector corresponding to the image patch is d-dimensional, then the position encoding corresponding to the image patch is also d-dimensional.
[0107] In this embodiment, by using the sine-cosine position encoding algorithm to determine the position encoding corresponding to each image patch, the position information of each image patch can be retained, which is beneficial to capturing sequence dependencies.
[0108] Next, refer to Figure 5 to introduce the specific working process of the key information selection module in the vehicle recognition model. Among them, Figure 5 is a schematic diagram of the working process of a key information selection module provided by an embodiment of the present application.
[0109] S501. Receive the first image feature sent by the previous encoding module.
[0110] Optionally, if there are 10 encoding modules, for the 3rd encoding module, the 3rd encoding module receives the first image feature sent by the 2nd encoding module.
[0111] S502. Generate a plurality of target attention weight matrices based on the attention weight matrices of each encoding module, and each target attention weight matrix corresponds to one attention head in each encoding module.
[0112] Specifically, the attention weights of each layer in the vehicle recognition module only reflect the local attention pattern of the current layer. As the network deepens, the high-level attention may overly focus on regions irrelevant to the category, such as the background or common features, resulting in the neglect of fine-grained features. Exemplarily, in the vehicle category recognition task, the high-level attention may focus on common regions such as the wheel size and the head size of the vehicle, while ignoring key details such as the wheel position.
[0113] Therefore, a plurality of target attention weight matrices are generated according to the attention weight matrices of each encoding module. Among them, the number of target attention weight matrices is the same as the number of attention heads in each encoding module.
[0114] S503. Determine a target position index according to each target attention weight matrix, where the target position index is used to represent the position index of the image patch corresponding to the vehicle classification.
[0115] Specifically, use the position indices corresponding to the multiple attention heads with the largest weights in each target attention weight matrix as the target position index. Wherein, one attention head corresponds to one image patch, and the image patch corresponding to the target position index focuses on the local details that need to be concerned for vehicle classification.
[0116] Exemplarily, if there are attention head 1, attention head 2, attention head 3, and attention head 4, and the weights of attention head 2 and attention head 4 are greater than those of attention head 1 and attention head 3, then use the position indices of attention head 2 and attention head 4 as the target position index.
[0117] S504. Determine a second image feature according to each target position index and the first image feature, and input the second image feature into the last encoding module.
[0118] It should be noted that when initializing the vehicle recognition model, a classification vector can be randomly generated, for example, sampled from a normal distribution, and its dimension is the same as that of the embedding vector. Before inputting the initial image feature into the encoding module in the vehicle recognition model, splice the classification vector to the first position of the image feature. Exemplarily, assume there are N image patch embeddings with a shape of N×D, and the spliced image feature is (N + 1)×D.
[0119] Optionally, splice the vectors corresponding to each target position index in the first image feature with the classification vector. Exemplarily, the second image feature can be as shown in the following formula (9):
[0120]
[0121] Wherein, is the classification vector, K is the target position index, and L is the total number of layers. In this embodiment, the total number of layers is the total number of layers included in all encoding modules except the last encoding module.
[0122] In this embodiment, by generating multiple target attention weight matrices based on the attention weight matrices of each encoding module, determining the target position index according to each target attention weight matrix, determining the second image feature according to each target position index and the first image feature, and inputting the second image feature into the last encoding module, while ensuring the robustness of the key area of vehicle classification, only focusing on local details and ignoring non-discriminative areas, thereby improving the fine-grained classification performance.
[0123] As an alternative implementation, the specific steps of generating multiple target attention weight matrices based on the attention weight matrices of each encoding module in step S502 above may be: performing a recursive multiplication operation on the attention weight matrices of each encoding module to obtain the target attention weight matrix.
[0124] Optionally, the target attention weight sequence accumulates the attention weights of each layer and can more accurately reflect the relative importance of each image patch in the entire image.
[0125] Exemplarily, the target attention weight matrix can be represented by the following formula (10):
[0126]
[0127] where L represents the total number of layers. a final captures how information propagates from the input to the higher-level embedding representations. Compared with directly using the first image features for image feature extraction in the last layer, it provides a selection of discriminative regions.
[0128] In this embodiment, the target attention weight matrix is determined through a recursive multiplication operation, so that the last encoding module focuses on the details that the vehicle category focuses on, improving the vehicle classification accuracy.
[0129] As an alternative implementation, the specific process of determining the target position index according to each target attention weight matrix in step S503 above may be as follows: according to each target attention weight matrix, determine a preset number of attention heads with the largest weights, and use the position indices of the preset number of attention heads as the target position index.
[0130] Specifically, the maximum value or average value of the weights of each attention head in the target attention weight matrix can be taken to obtain the weights of the image patches corresponding to each attention head after summarization. Then, the image patches are sorted from largest to smallest based on the weights, and the position indices of the first preset number of attention heads are used as the target position index.
[0131] Exemplarily, if the weights corresponding to each image patch after summarization are: 0.15 for image patch 1, 0.25 for image patch 2, 0.18 for image patch 3, and 0.20 for image patch 4. Then, the position indices of the attention heads corresponding to image patch 2 and image patch 4 are used as the target position index.
[0132] In this embodiment, the position indices of a preset number of attention heads with the largest weights are selected as the target position index, so as to focus on the key local details and improve the fine-grained classification performance.
[0133] Based on the same inventive concept, an embodiment of the present application further provides a vehicle recognition device corresponding to the vehicle recognition method. Since the principle of problem-solving of the device in the embodiment of the present application is similar to that of the above-mentioned vehicle recognition method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0134] Refer to Figure 6 As shown, it is a schematic structural diagram of a vehicle recognition device provided by an embodiment of the present application. The device includes: an acquisition module 601, a preprocessing module 602, a segmentation module 603, and a recognition module 604; wherein:
[0135] The acquisition module 601 is used to acquire a vehicle image of the current vehicle;
[0136] The preprocessing module 602 is used to preprocess the vehicle image of the current vehicle to obtain an image to be segmented;
[0137] The segmentation module 603 is used to segment the image to be segmented based on a preset sliding step length and perform feature extraction to obtain a plurality of initial image features. Among them, there is an overlapping area between the image blocks corresponding to adjacent initial image features;
[0138] The recognition module 604 is used to input a plurality of the initial image features into a pre-trained vehicle recognition model to obtain the vehicle category of the current vehicle. The vehicle recognition model includes a plurality of encoding modules and a key information selection module. The key information selection module is located before the last encoding module. The key information selection module is used to screen the received image features and output them to the last encoding model. Each of the encoding modules sequentially includes a multi-head self-attention unit and a multi-layer perceptron unit. The multi-head self-attention unit includes: a first convolutional layer, a multi-head self-attention layer, and a second convolutional layer.
[0139] The preprocessing module 602 is specifically used to sequentially perform parameter prediction, coordinate mapping, and resampling on the vehicle image of the current vehicle based on a spatial transformation network to obtain an image to be segmented.
[0140] The segmentation module 603 is specifically used for:
[0141] Segment the image to be segmented based on a preset sliding step length to obtain a plurality of image blocks;
[0142] Project the plurality of image blocks into a preset embedding space to generate a plurality of embedding vectors;
[0143] Determine a plurality of initial image features according to the plurality of embedding vectors.
[0144] The segmentation module 603 is specifically used for:
[0145] Determine the position encoding corresponding to each image patch;
[0146] Use the sum of each embedding vector and the corresponding position encoding as an initial image feature.
[0147] The splitting module 603 is specifically configured to:
[0148] Based on the sine-cosine position encoding algorithm, determine the position encoding corresponding to each image patch.
[0149] The recognition module 604 is specifically configured to:
[0150] Receive the first image feature sent by the previous encoding module;
[0151] Based on the attention weight matrices of each of the encoding modules, generate multiple target attention weight matrices, where each of the target attention weight matrices corresponds to one attention head in each encoding module;
[0152] According to each of the target attention weight matrices, determine a target position index, where the target position index is used to represent the position index of the image patch corresponding to vehicle classification;
[0153] According to the target position index and the first image feature, determine a second image feature, and input the second image feature into the last encoding module.
[0154] The recognition module 604 is specifically configured to: perform a recursive multiplication operation on the attention weight matrices of each of the encoding modules to obtain a target attention weight matrix.
[0155] The recognition module 604 is specifically configured to: according to each of the target attention weight matrices, determine a preset number of attention heads with the largest weights, and use the position indices of the preset number of attention heads as the target position index.
[0156] The description of the processing flow of each module in the device and the interaction flow between modules can refer to the relevant description in the above method embodiments, and will not be elaborated here.
[0157] This application embodiment also provides an electronic device, as Figure 7 shown, which is a schematic structural diagram of an electronic device provided by this application embodiment, including: a processor 701, a memory 702, and a bus. The memory 702 stores machine-readable instructions executable by the processor 701 (for example, Figure 6The execution instructions corresponding to the acquisition module 601, the preprocessing module 602, the segmentation module 603, and the recognition module 604 in the device in , etc. When the computer device runs, the processor 701 communicates with the memory 702 through a bus, and when the machine-readable instructions are executed by the processor 701, the processing of the above vehicle recognition method is executed.
[0158] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, the steps of the above vehicle recognition method are executed.
[0159] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the method embodiments, and will not be repeated in this application. In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical, mechanical or other forms.
[0160] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. And the foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, RandomAccess Memory), a magnetic disk, or an optical disc that can store program codes.
[0161] The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. A vehicle recognition method, characterized in that, The method includes: Obtaining a vehicle image of the current vehicle; Preprocessing the vehicle image of the current vehicle to obtain a to-be-segmented image; Based on a preset sliding step, segmenting the to-be-segmented image and performing feature extraction to obtain a plurality of initial image features, wherein there is an overlapping area between the image blocks corresponding to adjacent initial image features; Inputting the plurality of initial image features into a pre-trained vehicle recognition model to obtain the vehicle category of the current vehicle, the vehicle recognition model includes a plurality of encoding modules and a key information selection module, the key information selection module is located before the last encoding module, the key information selection module is used to screen the received image features and output them to the last encoding model, each of the encoding modules sequentially includes a multi-head self-attention unit and a multi-layer perceptron unit, and the multi-head self-attention unit includes: a first convolutional layer, a multi-head self-attention layer, and a second convolutional layer.
2. The vehicle identification method according to claim 1, characterized in that The preprocessing the vehicle image of the current vehicle to obtain a to-be-segmented image includes: Based on a spatial transformation network, sequentially performing parameter prediction, coordinate mapping, and resampling on the vehicle image of the current vehicle to obtain a to-be-segmented image.
3. The vehicle recognition method according to claim 1, wherein The segmenting the to-be-segmented image and performing feature extraction based on a preset sliding step to obtain a plurality of initial image features includes: Based on a preset sliding step, segmenting the to-be-segmented image to obtain a plurality of image blocks; Projecting the plurality of image blocks into a preset embedding space to generate a plurality of embedding vectors; Determining a plurality of initial image features according to the plurality of embedding vectors.
4. The vehicle recognition method according to claim 3, characterized in that, The determining a plurality of initial image features according to the plurality of embedding vectors includes: Determining the position encoding corresponding to each image block; Using the sum of each embedding vector and the corresponding position encoding as an initial image feature.
5. The vehicle identification method according to claim 4, wherein, The determining the position encoding corresponding to each image block includes: Based on the sine-cosine position encoding algorithm, determining the position encoding corresponding to each image block.
6. The vehicle identification method according to claim 1, wherein The key information selection module is specifically used for: Receiving a first image feature sent by the previous encoding module; Generating a plurality of target attention weight matrices based on the attention weight matrices of each of the encoding modules, and each of the target attention weight matrices corresponds to one attention head in each of the encoding modules; Determining a target position index according to each of the target attention weight matrices, and the target position index is used to represent the position index of the image block corresponding to vehicle classification; Determining a second image feature according to the target position index and the first image feature, and inputting the second image feature into the last encoding module.
7. The vehicle recognition method according to claim 6, wherein, The generating a plurality of target attention weight matrices based on the attention weight matrices of each of the encoding modules includes: Performing a recursive multiplication operation on the attention weight matrices of each of the encoding modules to obtain a target attention weight matrix.
8. The vehicle identification method according to claim 6, characterized in that The determining a target position index according to each of the target attention weight matrices includes: Determining a preset number of attention heads with the largest weights according to each of the target attention weight matrices, and using the position indices of the preset number of attention heads as the target position index.
9. An electronic device, characterized in that, Includes: A processor and a memory, the memory storing machine-readable instructions executable by the processor, and when the electronic device runs, the processor executes the machine-readable instructions to perform the steps of the vehicle recognition method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the steps of the vehicle recognition method according to any one of claims 1 to 8 are performed.