Vehicle type classification method with capsule packet hybrid network

By building a capsule packet hybrid network, combining dual-domain feature extraction, 3D convolution and dynamic routing processes, the problem of low accuracy in vehicle type classification is solved, and higher accuracy and stronger robustness is achieved, which is suitable for vehicle type identification in intelligent transportation systems.

CN120339689APending Publication Date: 2025-07-18JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510397283.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has the problem of low accuracy in vehicle type classification, especially in complex contexts, where the misclassification rate is high, and traditional methods are difficult to effectively utilize the learning ability of deep convolutional neural networks and capsule networks.

Method used

A method with a capsule packet hybrid network is adopted, by dimensional adjustment and Conv+BN+ReLU operations on the input image, the dual-domain feature extraction subnet is used to extract features from the spatial domain and frequency domain, combined with the basic capsule layer, the capsule type group layer and the 3D convolution capsule layer, and dynamic routing process is carried out, and vehicle type classification is finally realized through the classification capsule layer, and the objective function composed of margin loss and classification loss is trained.

Benefits of technology

It improves the accuracy of vehicle type classification, alleviates the problem of category imbalance, enhances the robustness of the model, and can better accurately classify vehicle types in intelligent traffic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339689A_ABST
    Figure CN120339689A_ABST
Patent Text Reader

Abstract

The invention relates to a vehicle type classification method with a capsule packet hybrid network, and belongs to the field of intelligent traffic systems. The method comprises the following steps: adjusting the size of an input image to obtain a feature map, carrying out double-domain feature extraction subnet processing, mapping the feature map into a low-layer capsule through a basic capsule layer, grouping layers through capsule types, carrying out 3D convolution operation on each group of middle capsules, and reasoning the low-layer feature map into a high-layer feature map through a dynamic routing process. And final vehicle type classification is realized. The method has the advantages that a hybrid network with capsule type grouping is constructed, so that the hybrid network has higher accuracy and stronger robustness in a vehicle type classification task, and particularly, global and local information of a target is effectively represented by using a hybrid spatial domain and frequency domain feature processing technology, so that more accurate vehicle type classification learning is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent transportation systems, and particularly relates to traffic flow statistics, ETC intelligent toll collection systems, vehicle type detection, etc., and in particular refers to a method for classifying vehicle types in a massive vehicle image database captured by a camera. Background Art

[0002] With the development of technology and the continuous improvement of people's living standards, cars have become an indispensable part of modern people's lives, and their numbers are increasing year by year. The use of a large number of cars reflects the mobility of the population, economic capabilities, and the degree of intimacy among people, which is not only of great significance to urban development and government decision-making, but also brings new research topics for researchers. Vehicle type classification is one of the basic topics in many studies, and it is also an important part of intelligent transportation systems. It is widely used in traffic scenarios such as assisted driving systems, traffic flow statistics, and ETC intelligent toll collection systems, providing data support for intelligent transportation systems and having important research significance and value for solving traffic problems.

[0003] In recent years, the popularization of visual surveillance networks has made it convenient to collect a large amount of vehicle-related information. At the same time, with the improvement of computer performance and the rapid development of image processing technology, computer vision technology, and pattern recognition technology, vehicle type classification technology based on images has been realized. Compared with the contact method based on induction coil method and the non-contact method based on ultrasonic and infrared detection methods, the automatic vehicle type classification technology based on images has the advantages of simple sensor installation and low cost.

[0004] Early image-based vehicle type classification methods mainly relied on low-level features or handcrafted features, such as color, shape, texture, etc., focusing on local or global designs. Typical features commonly used for vehicle classification include Scale-Invariant Feature Transform (SIFT), color or texture histograms, Local Binary Pattern (LBP), Histogram of Oriented Gradients (HOG), etc., or combinations thereof. Although using low-level features of images can describe the image content, it cannot represent human understanding of the image. Compared with methods based on low-level features, mid-level features attempt to calculate the overall image representation formed by local features. The Bag of Words (BOW) model is the most popular mid-level representation method. This method improves the representation of meaningless and disordered low-level local features into an ordered visual word representation with abstract semantic information, and has achieved good experimental results in image classification. However, the expressive ability of the BOW model is limited, resulting in no further breakthrough in image classification tasks.

[0005] In recent years, after AlexNet achieved excellent performance on the ImageNet dataset, deep learning-based methods have become popular and achieved remarkable results in many tasks such as image classification, object recognition, and semantic segmentation, marking a new era in the feature representation of images. Different from low-level and mid-level features, deep learning models can automatically learn more discriminative and abstract features without the need for a large amount of domain knowledge. Nowadays, these deep learning models, especially convolutional neural networks (CNNs), are considered the most widely used deep learning technologies, especially in applications related to image classification. Currently, CNN-based methods have greatly improved the accuracy of vehicle type classification, but still prone to misclassification in some scenarios.

[0006] Compared with CNNs, Capsule Network (CapsNet) is a new structure that encodes the nature and spatial relationship of image features. Capsule Network is composed of capsules instead of neurons. It can encode spatial information and calculate the probability of the existence of an object at the same time. Capsule Network alleviates the problems of invariance caused by the pooling operation in CNNs and the inability to understand the spatial relationship between features. It performs excellently on simple datasets such as MNIST. However, the performance of the original Capsule Network on complex background datasets is not satisfactory, and it has many parameters and low model training efficiency. In the vehicle type classification task, in order to improve the classification accuracy, it is necessary to make full use of the powerful learning capabilities of deep convolutional neural networks and Capsule Network. Summary of the Invention

[0007] The present invention provides a vehicle type classification method with a capsule grouping hybrid network to solve the problem of low accuracy in current vehicle type classification.

[0008] The technical solution adopted by the present invention includes the following steps:

[0009] S1: Take the vehicle image I captured in the real monitoring scenario in the city as the input;

[0010] S2: Adjust the size of the input image i, and perform Conv+BN+ReLU operations on the adjusted image, where Conv represents the convolutional layer, BN is batch normalization, and ReLU is the activation function, to obtain the feature map F, which is used as the input of the dual-domain feature extraction subnet;

[0011] S3: The feature map F is processed by the dual-domain feature extraction subnet to obtain the feature map F';

[0012] S4: The feature map F' passes through the basic capsule layer to map F' into the low-level capsule F D_Cap ;

[0013] S5: The low-level capsule FD_Cap The capsules are further divided into multiple groups of intermediate capsules through the capsule type grouping layer, so that each group contains multiple types of capsules;

[0014] S6: Perform 3D convolution operations on each group of intermediate capsules to obtain a three-dimensional low-level feature map F 3D , after the 3D convolution capsule layer, the spatial relationship similar to that of objects in three-dimensional space is obtained;

[0015] S7: 3D low-level feature map F 3D The low-level feature map is inferred to the high-level feature map through the dynamic routing process, and then the final vehicle type classification is achieved through the classification capsule layer.

[0016] S8: After building the network model, use the objective function composed of margin loss and classification loss to train the network.

[0017] In step S1 of the present invention, the input is a large number of vehicle images captured in real monitoring scenes in the city. Where W and H represent the width and height of the image respectively, and each vehicle image is annotated with the vehicle type.

[0018] In step S3 of the present invention, the feature map F extracts global features and multi-scale local features from two dimensions, frequency domain and spatial domain, respectively, through a dual-branch structure, and the output F of the spatial domain branch S The output of the frequency domain branch F F The frequency-space domain features are initially integrated through element-wise addition operations, and the feature map F′ is generated through two layers of 2D convolution operations.

[0019] The specific operation of the dual-branch structure of the present invention to simultaneously process the features of the frequency and spatial domains is as follows: one branch captures the spatial domain features by selecting attention based on Top-K Token, and its structure is a self-attention mechanism consisting of query Q, key K, and value V; first, the size of Q and K is adjusted by reshape operation, and the attention matrix M is generated by dot product operation. S ; Then, the elements with low attention values are shielded through the adaptive selection strategy, and only the first 50% of the elements are retained (the remaining elements are set to 0) for Sigmoid activation; finally, the dot product is performed with V to obtain the output F of the spatial domain branch S ; The other branch first obtains the frequency domain Q through a fast Fourier transform f =FFT(F), K f =FFT(F), V f =FFT(F), where FFT(·) stands for Fast Fourier Transform; then Q f and K fPerform a reshape operation to generate a frequency-domain attention map M through a dot product operation F , and extract the real part and the imaginary part Then, perform Softmax activation operations on the real part and the imaginary part respectively, and merge them to obtain the activated attention map Subsequently, use to optimize the weights on V f , and convert it to the spatial domain using the inverse fast Fourier transform IFFT(·); finally, obtain the output F of the frequency-domain branch through modulo operation F .

[0020] In step S4 of the present invention, the basic capsule layer converts the feature map F' extracted by the dual-domain feature extraction subnet into different low-level capsules containing vectorized feature information Among them, N represents the number of capsules, k represents the number of neurons in each capsule, and each capsule is a vector composed of a group of scalars, reflecting the features of specific types of entities, such as position, direction, size, texture, etc. The basic capsule layer is implemented by multiple filters with a fixed stride

[0021] In step S5 of the present invention, the capsule type grouping layer further divides the N capsules obtained by the basic capsule layer into m groups from top to bottom according to the serpentine sampling strategy. Each group contains N / m capsules, and the m groups of different capsules containing N / m capsules are respectively concatenated along the channel dimension, so as to obtain m groups of intermediate capsules

[0022] In step S6 of the present invention, the 3D convolutional capsule layer performs 3D convolutional operations on the m groups of intermediate capsules respectively. In this way, the dynamic routing object changes from the vector of the original capsule network to a three-dimensional feature map. The process of one 3D convolutional operation is expressed as:

[0023]

[0024] Among them, the dimension of the output feature map of layer l is w l represents the width and height of layer l, c l represents the depth of the feature map obtained after a convolutional kernel performs a 3D convolutional operation, n l represents the stride of the convolutional kernel in the depth direction, that is, the depth of the convolutional kernel, and F 3D is a series of feature maps obtained after 3D convolution

[0025] In step S7 of the present invention, the dynamic routing process updates the coupling coefficient through the similarity between capsule vectors. The specific operation is: reshape the feature map obtained after 3D convolution into a new feature map, expressed as where n l+1Indicates the number of neurons contained in each capsule, and then uses the capsules obtained after reshape to route the classification capsules. The routing process is completed through the following formula:

[0026] C [pqrs] = softmax3D(b [pqrs] )

[0027]

[0028] V [pqrs] = squask3D(S [pqrs] )

[0029]

[0030] Among them, C [pqrs] is the coupling coefficient, which is a four-dimensional matrix representing the contribution weight of the lower-layer capsules to the higher-layer capsules; after several iterations, its value can be determined. b [pqrs] represents a four-dimensional matrix, p ∈ w l+1 , q ∈ w l+1 , r ∈ c l + 1, s ∈ c l , it is initialized to all 0. Squash3D means compressing the norm lengths of all capsules in S [pqrs] to between 0 and 1.

[0031] In step S8 of the present invention, the objective function is expressed as L total , and is composed of the margin loss L k and the classification loss L c . Specifically, it is expressed as:

[0032] L k = T k max(0, m + - ||v k ||) 2 + λ(1 - T k )max(0, ||v k || - m - ) 2

[0033]

[0034] L total = L k + αL c

[0035] Among them, if the input picture belongs to the kth class, then T k is equal to 1, otherwise it is equal to 0. m +and m - are the upper bound for penalizing false positives and the lower bound for penalizing false negatives, where m + = 0.9 and m - = 0.1. λ is used to control the effect of gradient backpropagation in the initial stage of training, and λ = 0.5. In the classification loss, γ is set to 2. In the objective function, the margin loss dominates, and α is set to 0.3.

[0036] The beneficial effects of the present invention are as follows:

[0037] The present invention extracts initial features of vehicle images from two different perspectives, the spatial domain and the frequency domain. At the same time, taking advantage of the fact that the capsule network can extract more features, a vehicle type classification method with a capsule grouping hybrid network is proposed. In the present invention, not only the initial input features are extracted and optimized from two complementary dimensions, the spatial domain and the frequency domain, but also a 3D capsule network is integrated to learn different variants of the features. These components effectively alleviate the problem of insufficient ability of traditional convolutional neural networks and the original capsule network to extract vehicle type features, and can better complete the accurate classification of vehicle types in the intelligent transportation scenario. In addition, there is a phenomenon of class imbalance both in real life and in the currently existing public vehicle datasets. On the one hand, the present invention adopts a data augmentation method, and on the other hand, a hybrid loss composed of margin loss and classification loss is used to guide the network, which can alleviate the problem of reduced accuracy caused by class imbalance and also increase the class probability of the real class. Description of the Drawings

[0038] Figure 1 is a schematic diagram of the vehicle type classification task on the MIO-TCD dataset, where the image marked with a red box in the rightmost column is suspected of being misclassified;

[0039] Figure 2 is an image display of the vehicle type classification datasets BIT-Vehicle and MIO-TCD. Among them, (a) is an example of vehicle image data in the BIT-Vehicle dataset; (b) is an example of vehicle image data in the MIO-TCD dataset;

[0040] Figure 3 is a framework diagram of a vehicle type classification method with a capsule grouping hybrid network provided by an experimental example of the present invention;

[0041] Figure 4 is a schematic diagram of the capsule type grouping layer on the MIO-TCD dataset in an experimental example of the present invention;

[0042] Figure 5 is the confusion matrix on the MIO-TCD dataset in an experimental example of the present invention. Detailed implementation manners

[0043] The following will describe the detailed implementation manners of the present invention in more detail with reference to the accompanying drawings, so that the technical solutions of the present invention are easier to understand and master. The following content is only a better specific implementation method of the present invention, but the protection scope of the present invention is not limited thereto. The present invention can also have other various specific implementation manners. Any equivalent replacement or equivalent transformation disclosed by those skilled in the art in this technology field is covered within the scope of protection required by the present invention. In addition, the accompanying drawings are only for illustrative purposes, and are only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0044] It includes the following steps:

[0045] S1: Use the vehicle image I captured in the real monitoring scenario in the city as the input;

[0046] S2: Adjust the size of the input image i, and perform Conv+BN+ReLU operations on the adjusted image, where Conv represents the convolutional layer, BN is batch normalization, and ReLU is the activation function, to obtain the feature map F as the input of the dual-domain feature extraction subnet;

[0047] S3: The feature map F is processed by the dual-domain feature extraction subnet to obtain the feature map F';

[0048] S4: The feature map F' passes through the basic capsule layer to map F' into the low-level capsule F D_Cap ;

[0049] S5: The low-level capsule F D_Cap is further divided into multiple groups of intermediate capsules through the capsule type grouping layer, so that each group contains multiple types of capsules;

[0050] S6: Perform 3D convolution operations on each group of intermediate capsules respectively to obtain the three-dimensional low-level feature map F 3D , and obtain the spatial relationship similar to the object in the three-dimensional space through the 3D convolutional capsule layer;

[0051] S7: The three-dimensional low-level feature map F 3D realizes the inference from the low-level feature map to the high-level feature map through the dynamic routing process, and then realizes the final vehicle type classification through the classification capsule layer.

[0052] S8: After the network model is constructed, use the objective function composed of the margin loss and the classification loss to train the network.

[0053] In step S1 of the present invention, the input is a large number of vehicle images captured in a real monitoring scenario in the city. Where W and H respectively represent the width and height of the image, and each vehicle image has a vehicle type annotation.

[0054] In step S3 of the present invention, the processing operation of the dual-domain feature extraction subnet is as follows: The feature map F extracts global features and multi-scale local features from two dimensions of the frequency domain and the spatial domain through a dual-branch structure.

[0055] The dual-branch structure can process features in both the frequency and spatial domains simultaneously. The specific operation is as follows: One branch captures spatial domain features based on Top-K Token selection attention. Its structure is a self-attention mechanism composed of query Q, key K, and value V. First, the sizes of Q and K are adjusted through a reshape operation, and a dot product operation is performed to generate an attention matrix M. S ; Then, low-attention-value elements are masked out through an adaptive selection strategy, and only the top 50% of the elements (the remaining elements are set to 0) are retained for Sigmoid activation; finally, a dot product is performed with V to obtain the output F of the spatial domain branch. S ; The other branch first obtains Q in the frequency domain through a fast Fourier transform. f = FFT(F), K f = FFT(F), V f = FFT(F), where FFT(·) represents the fast Fourier transform; then Q f and K f are subjected to a reshape operation, and a dot product operation is performed to generate a frequency domain attention map M. F , and the real part is extracted. and the imaginary part. Then, Softmax activation operations are respectively performed on the real part and the imaginary part, and they are combined to obtain an activated attention map. Subsequently, use to optimize the weights on V. f Convert it to the spatial domain using the inverse fast Fourier transform IFFT(·); finally, an analog operation is used to obtain the output F of the frequency domain branch. F ;

[0056] Finally, the output F of the spatial domain branch. S and the output F of the frequency domain branch. F are initially integrated with the frequency-spatial domain features through an element-wise addition operation (denoted as ), and then two layers of 2D convolution operations are performed to generate the feature map F'.

[0057] As a further optimization solution of the present invention, in step S4, the basic capsule layer converts the feature map F' extracted by the dual-domain feature extraction subnet into different low-level capsules containing vectorized feature information. Among them, N represents the number of capsules, and k represents the number of neurons in each capsule. Each capsule is a vector composed of a group of scalars, reflecting the characteristics of specific types of entities, such as position, direction, size, texture, etc. The basic capsule layer can be implemented by multiple filters with a fixed stride.

[0058] As a further optimization solution of the present invention, in step S5, the capsule type grouping layer further divides the N capsules obtained from the basic capsule layer into m groups from top to bottom according to the serpentine sampling strategy. Each group contains N / m capsules, and the m groups of different capsules containing N / m capsules are respectively spliced along the channel dimension, so that m groups of intermediate capsules are obtained.

[0059] As a further optimization solution of the present invention, in step S6, the 3D convolutional capsule layer performs 3D convolutional operations on each of the m groups of intermediate capsules, so that the dynamic routing object changes from the vector of the original capsule network to a three-dimensional feature map. The process of one 3D convolutional operation is expressed as:

[0060]

[0061] Among them, the dimension of the output feature map of the l layer is w l represents the width and height of the l layer, c l represents the depth of the feature map obtained after a convolutional kernel performs a 3D convolutional operation, n l represents the stride of the convolutional kernel in the depth direction, that is, the depth of the convolutional kernel, F 3D is a series of feature maps obtained after 3D convolution.

[0062] As a further optimization solution of the present invention, in step S7, the dynamic routing process updates the coupling coefficient through the similarity between capsule vectors. The specific operation is as follows: Reshape the feature map obtained after 3D convolution into a new feature map, expressed as where n l+1 represents the number of neurons contained in each capsule. Then use the capsules obtained after reshape to route the classification capsules, and the routing process is completed through the following formula:

[0063] C [pqrs] = softmax3D(b [pqrs] )

[0064]

[0065] V [pqrs]= squash3D(S [pqrs] )

[0066]

[0067] where C [pqrs] is the coupling coefficient, which is a four - dimensional matrix representing the contribution weight of the lower - layer capsules to the higher - layer capsules; its value can be determined after several iterations. b [pqrs] represents a four - dimensional matrix, p ∈ w l+1 , q ∈ w l+1 , r ∈ c l+1 , s ∈ c l , and it is initialized to all 0. squash3D means compressing the norms of all capsules in S [pqrs] to be between 0 and 1.

[0068] As a further optimization scheme of the present invention, the objective function in step S8 is expressed as L total , which is composed of the margin loss L k and the classification loss L c , and is specifically expressed as:

[0069] L k = T k max(0, m + - ||v k ||) 2 + λ(1 - T k )max(0, ||v k || - m - ) 2

[0070]

[0071] L total = L k + αL c

[0072] where, if the input image belongs to the k - th class, then T k equals 1, otherwise it equals 0. m + and m - are respectively the upper bound for punishing false positives and the lower bound for punishing false negatives. Here, m + = 0.9 and m - = 0.1. λ is used to control the effect of gradient backpropagation in the initial stage of training, and λ = 0.5. In the classification loss, γ is set to 2. In the objective function, the margin loss is dominant, and α is set to 0.3.

[0073] The effects of the present invention will be further described below through specific experimental examples.

[0074] The present invention is a hybrid model with 2D convolution, 3D convolution, frequency-domain convolution, capsule type grouping, and dynamic routing. It consists of a Dual-domain Feature Extraction Sub-network, a PrimaryCaps Layer, a CapsType GroupingLayer, a 3D ConvCaps Layer, and a ClassCaps Layer. The proposed Dual-domain Feature Extraction Sub-network fuses spatial features with local attributes and frequency features generated by Fourier transform to optimize the initial input features and enhance their discriminative ability; the PrimaryCaps Layer maps the output of the Dual-domain Feature Extraction Sub-network into low-level capsules; the CapsType GroupingLayer groups various types of low-level capsules obtained from the PrimaryCaps Layer according to the concept of "neutralization"; the 3D ConvCaps Layer is used to obtain the spatial relationship similar to an object in three-dimensional space, which can synthesize capsule features on the one hand and reduce the computational complexity of the model on the other hand. Finally, through the dynamic routing process based on 3D convolution, the capsules obtained from the 3D ConvCaps Layer are routed to the ClassCaps Layer, and are continuously iteratively updated to finally achieve the classification of vehicle types. The present invention enables the model to have higher accuracy in vehicle type classification.

[0075] Experimental Example

[0076] As shown in the appendix Figure 1 intuitively shows the meaning of the vehicle type classification task. The essence of vehicle type classification is to perform coarse-grained classification of categories on all image samples in a given dataset. That is, vehicle type classification is to label each sample in the dataset. Taking the MIO-TCD dataset as an example, each picture is labeled as vehicle types such as bus, car, bicycle, SUV, etc.

[0077] The present invention performs vehicle type classification on two datasets, BIT-Vehicle and MIO-TCD. For BIT-Vehicle, the sample size in this example is adjusted to 512×512 pixels; for MIO-TCD, the sample size in this example is adjusted to 256×256 pixels.

[0078] Dataset: As shown in the appendix Figure 2 shows two publicly available vehicle datasets specifically constructed for the vehicle type classification task. As Figure 2 (a) shows examples of different vehicle models in BIT-Vehicle. Figure 2(b) shows an example of vehicle types in MIO-TCD. The vehicle images used for vehicle type classification in the present invention are from the above two datasets. BIT-Vehicle contains 9,850 images with sizes of 1600×1200 and 1920×1080 pixels respectively. The vehicles in this dataset are divided into 6 categories, as shown in Figure 2 (a). MIO-TCD consists of 648,959 images, among which 519,164 images are in the training set and 129,795 images are in the test set. This dataset contains a total of 11 different categories, as shown in Figure 2 (b). Due to the imbalance in the number of each category, the above two datasets are both complex and challenging.

[0079] The framework structure of the vehicle type classification method is as shown in Appendix Figure 3 and includes six parts: an input layer, a dual-domain feature extraction subnet, a basic capsule layer, a capsule type grouping layer, a 3D convolutional capsule layer, and a classification capsule layer. The framework model of the vehicle type classification method based on the capsule grouping hybrid network takes vehicle images captured in real monitoring scenarios in the city as input. For the BIT-Vehicle dataset, through a convolutional layer with a convolutional kernel size of 3×3, a BN batch normalization layer, and a ReLU activation function, a feature map F with a size of 256×256×256 is obtained; for the MIO-TCD dataset, a feature map F with a size of 128×128×256 is obtained; F is used as the input of the dual-domain feature extraction subnet.

[0080] The dual-domain feature extraction subnet includes two parts: a self-attention module that focuses on spatial domain features and a self-attention module that focuses on frequency domain features. Among them, the spatial domain part adaptively selects the most important Tokens based on the Top-K selection mechanism. Taking the MIO-TCD dataset as an example, the input feature map is flattened into a two-dimensional matrix with a shape of 16384×256, and Q, K, and V are generated through linear transformation. First, an attention matrix M is generated through the dot product operation of Q and K S ; then, elements with low attention values are masked out through an adaptive selection strategy, and only the first 50% of the elements are retained for Sigmoid activation, and the remaining elements are set to 0; finally, a dot product is performed with V to obtain a feature map with a shape of 16384×256, and the feature is reshaped into 128×128×256, which is the output F of the spatial domain branch S . In the frequency domain part, a self-attention structure is introduced to capture the relationships and interactions between different frequency bands. First, the input feature map is subjected to a fast Fourier transform to convert it to the frequency domain. The size of the output frequency domain feature map is still 128×128×256, but each element is a complex number representing frequency domain information. A linear transformation is performed on the frequency domain feature map to obtain Q f and Kf , V f . Then reshape Q f and K f into 16384×256, generate the frequency-domain attention map M through dot product operation F , and extract the real part and the imaginary part Then perform Softmax activation operations on the real part and the imaginary part respectively, and merge them to obtain the activated attention map Subsequently, use to optimize the weights on V f , convert it to the spatial domain using the inverse fast Fourier transform IFFT(·); finally, obtain the output F of the frequency-domain branch through modulo operation F , whose size is 128×128×256. The output F S of the spatial-domain branch and the output F F of the frequency-domain branch are initially feature fused through element-wise addition operation, that is and obtain the feature map F′ through two 2D convolution operations with a convolution kernel of 3×3 and a stride of 2. For the BIT-Vehicle dataset, obtain the feature map F′ with a size of 64×64×256; for the MIO-TCD dataset, obtain the feature map F′ with a size of 32×32×256. The basic capsule layer is obtained by 3×3×256 filters with a stride of 2, and each position contains 32 capsules with a size of 16×16×8.

[0081] Attachment Figure 4Intuitively demonstrates the capsule grouping process on the MIO-TCD dataset. In the original CapsNet, there is only one convolutional layer and one primary capsule layer before dynamic routing. Due to the limitations of expensive storage and computer resources, before dynamic routing, as many features as possible need to be extracted for subsequent capsule analysis. According to existing research, it is found that different types of capsules contribute differently to the classification accuracy. The present invention proposes an idea of capsule type grouping. The core idea of the capsule type grouping layer is to further divide the 32 capsules with a size of 16×16×8 obtained from the primary capsule layer into 8 groups in a snake-like sampling manner from top to bottom according to the serial number, so that each group has 4 different types of capsules. Among them, the 1st, 16th, 17th, and 32nd capsules are the first group, the 2nd, 15th, 18th, and 31st capsules are the second group, the 3rd, 14th, 19th, and 30th capsules are the third group, the 4th, 13th, 20th, and 29th capsules are the fourth group, the 5th, 12th, 21st, and 28th capsules are the fifth group, the 6th, 11th, 22nd, and 27th capsules are the sixth group, the 7th, 10th, 23rd, and 26th capsules are the seventh group, and the 8th, 9th, 24th, and 25th capsules are the eighth group. The starting point of such grouping is that the classification capsule layer is a combination of many different types of capsules, and capsule type i+1 contributes more to the classification of the classification capsule than capsule type i, where i is 1, 2,..., 32. In order to obtain a more balanced result, the concept of neutralization is used to group different types of capsules. Each group of capsules is merged along the channel dimension to obtain 8 capsules with a size of 32. The obtained low-level capsules of 16×16×32×8 are subjected to a 3D convolution operation with a convolution kernel of 3×3×3 and a stride of 2×2×1 to obtain a feature map of 8×8×32×8, and then through the dynamic routing process of 3D convolution, classification capsules with a size of 11×32 are obtained, where 11 represents 11 vehicle models in the MIO-TCD dataset.

[0082] Testing process:

[0083] To illustrate the technical effects of the present invention, two datasets, BIT-Vehicle and MIO-TCD, are used. The samples in the datasets are divided into two parts in a ratio of 4:1 for training and testing to verify the effectiveness of the present invention. Before training, first, the images containing multiple vehicles and their corresponding labels in the dataset are removed. To simplify the process, only rotation and flipping operations are considered for data augmentation. The present invention is trained and tested using the Pytorch framework and Python language on a device configured with an intel(R) i7-10700KF CPU and an NVIDIA RTX 2080Ti GPU, with a GPU memory of 11GB. During the training process, the Adam optimizer is used to optimize the model parameters, with the momentum set to 0.99, the initial learning rate set to 0.0001, the weight decay set to 0.0005, the batch size set to 2, and the number of iterations for routing is 3.

[0084] The mean average precision (mAP) and mean recall (mRe), which are widely used in tasks in the field of pattern recognition such as object detection and image classification, are used as performance indicators to evaluate the present invention. In addition, the kappa coefficient is also used to evaluate the consistency of the vehicle type classification task.

[0085] Table 1 shows the performance comparison of the present invention with methods named C-P CapsNet, S-L CapsNet, and sCapsNet on two vehicle type classification datasets. Among them, C-P CapsNet and S-L CapsNet are capsule-based methods for vehicle type classification; sCapsNet stacks convolutional layers in front of the original capsule layer to make its output size the same as that of the dual-domain feature extraction subnet.

[0086] Table 1

[0087]

[0088] As can be seen from Table 1, the results of the present invention are better than those of other methods in all evaluation indicators, which proves the effectiveness of the present invention and can be used as a benchmark for subsequent exploration of vehicle type classification based on capsule networks.

[0089] Figure 5 This is the vehicle classification confusion matrix of the present invention on the MIO-TCD dataset. The confusion matrix shows that the vehicle type classification method provided by the present invention can accurately classify rare categories such as trucks and articulated trucks. At the same time, for categories with an absolute dominant number in the dataset, such as cars and backgrounds, the present invention can classify them conveniently and correctly.

[0090] The present invention constructs a hybrid network with capsule type grouping, making it have higher accuracy and stronger robustness in the vehicle type classification task. In particular, the hybrid spatial domain and frequency domain feature processing technology is used to effectively characterize the global and local information of the target, thereby promoting more accurate vehicle type classification learning.

Claims

1. A vehicle type classification method with a capsule grouping hybrid network, characterized in that: Including the following steps: S1: Using the vehicle image I captured in the real monitoring scenario in the city as the input; S2: Adjusting the size of the input image I, and performing Conv+BN+ReLU operations on the adjusted image, where Conv represents the convolutional layer, BN is batch normalization, and ReLU is the activation function, to obtain the feature map F, which is used as the input of the dual-domain feature extraction subnet; S3: Processing the feature map F through the dual-domain feature extraction subnet to obtain the feature map F'; S4: The feature map F' passes through the basic capsule layer, mapping F' into the low-level capsule F D_Cap ; S5: Lower-level capsule F D_Cap It is further divided into multiple groups of intermediate capsules by the capsule type grouping layer, so that each group contains multiple types of capsules; S6: Perform 3D convolutional operations on each group of intermediate capsules to obtain a three-dimensional low-level feature map F 3D , and obtain the spatial relationship similar to that of an object in three-dimensional space through the 3D convolutional capsule layer; S7: 3D low-level feature map F 3D The low-level feature map is inferred to the high-level feature map through the dynamic routing process, and then the final vehicle type classification is achieved through the classification capsule layer; S8: After constructing the network model, using the objective function composed of the margin loss and the classification loss to train the network.

2. The vehicle type classification method with a capsule grouping and mixing network according to claim 1, wherein: In the step S1, the input is a large number of vehicle images captured in a real monitoring scenario in the city. Where W and H respectively represent the width and height of the image, and each vehicle image has a vehicle type annotation.

3. A vehicle type classification method with a capsule grouping hybrid network according to claim 1, characterized in that: In the step S3, the feature map F extracts global features and multi-scale local features from two dimensions of the frequency domain and the spatial domain respectively through a two-branch structure. The output F S of the spatial domain branch and the output F F of the frequency domain branch preliminarily integrate the frequency-spatial domain features through an element-wise addition operation, and generate a feature map F' through two layers of 2D convolution operations.

4. A vehicle type classification method with a capsule grouping and mixing network according to claim 3, characterized in that: The specific operation of the double-branch structure for simultaneously processing features in the frequency and spatial domains is as follows: One branch captures spatial-domain features using Top-K Token Selection Attention. Its structure is a self-attention mechanism composed of query Q, key K, and value V. First, the dimensions of Q and K are adjusted through a reshape operation, and a dot product operation is performed to generate the attention matrix M S ; then, through an adaptive selection strategy, elements with low attention values are masked, only the top 50% of the elements are retained, and the remaining elements are set to 0 for Sigmoid activation; finally, a dot product is performed with V to obtain the output F of the spatial-domain branch S ; The other branch first obtains the frequency-domain Q through a fast Fourier transform f = FFT(F), K f = FFT(F), V f = FFT(F), where FFT(·) represents the fast Fourier transform; then Q f and K f are subjected to a reshape operation, and a dot product operation is performed to generate the frequency-domain attention map M F , and the real part and the imaginary part are extracted. Then, Softmax activation operations are performed on the real and imaginary parts respectively, and they are combined to obtain the activated attention map Subsequently, use to optimize the weights on V f , and use the inverse fast Fourier transform IFFT(·) to convert it to the spatial domain; finally, a modulus operation is used to obtain the output F of the frequency-domain branch F .

5. A vehicle type classification method with a capsule grouping and mixing network according to claim 1, characterized in that: In the step S4, the basic capsule layer converts the feature map F' extracted by the dual-domain feature extraction sub-network into different low-level capsules containing vectorized feature information. Among them, N represents the number of capsules, k represents the number of neurons in each capsule, and each capsule is a vector composed of a group of scalars, reflecting the features of specific types of entities, including position, direction, size, and texture. The basic capsule layer is implemented by multiple filters with a fixed stride.

6. The vehicle type classification method with a capsule grouping hybrid network according to claim 1, characterized in that: In the step S5, the capsule type grouping layer further divides the N capsules obtained by the basic capsule layer into m groups from top to bottom according to the snake-shaped sampling strategy. Each group contains N / m capsules, and the m groups of different capsules containing N / m capsules are respectively concatenated along the channel dimension, so as to obtain m groups of intermediate capsules.

7. A vehicle type classification method with a capsule grouping and mixing network according to claim 1, characterized in that: In the step S6, the 3D convolutional capsule layer performs 3D convolutional operations on the m groups of intermediate capsules respectively. In this way, the dynamic routing object changes from the vector of the original capsule network to a three-dimensional feature map. The process of one 3D convolutional operation is expressed as: Among them, the dimension of the output feature map of layer l is w l represents the width and height of layer l, and c l represents the depth of the feature map obtained after a 3D convolution operation is performed by a convolution kernel, and n l represents the stride of the convolution kernel in the depth direction, that is, the depth of the convolution kernel, and F 3D is a series of feature maps obtained after 3D convolution.

8. A vehicle type classification method with a capsule grouping hybrid network according to claim 1, characterized in that: In the step S7, the dynamic routing process updates the coupling coefficient through the similarity between capsule vectors. The specific operation is as follows: Reshape the feature map obtained after 3D convolution into a new feature map, denoted as where n l+1 represents the number of neurons contained in each capsule. Then, use the capsules obtained after reshaping to route the classification capsules. The routing process is completed through the following formula: C [pqrs] = softmax3D(b [pqrs] ) V [pqrs] = squash3d(S [pqrs] ) Among them, C [pqrs] is the coupling coefficient, which is a four-dimensional matrix representing the contribution weight of the lower-level capsules to the higher-level capsules; its value can be determined after several iterations, and b [pqrs] represents a four-dimensional matrix, which is initialized to all 0s, and squash3D means compressing the lengths of all capsules in S [pqrs] to be between 0 and 1.

9. The vehicle type classification method with a capsule grouping and mixing network according to claim 1, wherein: In the step S8, the objective function is expressed as L total , which consists of a margin loss L k and a classification loss L c , and is specifically expressed as: L total = L k + αL c Among them, if the input image belongs to the k-th class, then T k is equal to 1, otherwise it is equal to 0, m + and m - are the upper bound for penalizing false positives and the lower bound for penalizing false negatives respectively. λ is used to control the effect of gradient backpropagation in the initial stage of training. In the classification loss, γ is set to 2. In the objective function, the margin loss dominates, and α is set to 0.

3.

10. A vehicle type classification method with a capsule grouping and mixing network according to claim 9, characterized in that: Use m + = 0.9 and m - = 0.1, λ = 0.5.