Small sample traffic abnormality image acquisition method and system based on multi-scale attention coupling mechanism

Through the method of collecting small sample traffic anomaly image with multi-scale attention coupling mechanism, the dependence problem of deep learning models on a large number of labeled data in the existing technology is solved, efficient detection and adaptability in new scenarios is achieved, and data labeling and training costs are reduced.

CN114898158BActive Publication Date: 2025-05-23HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210569646.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-24
Publication Date
2025-05-23
Estimated Expiration
2042-05-24

AI Technical Summary

Technical Problem

When existing deep learning models detect traffic anomalies in traffic monitoring, they need a large amount of labeled data for training, resulting in high training costs, insufficient generalization capabilities, and difficulty in quickly adapting to the detection of new scenarios.

Method used

A small sample traffic anomaly image acquisition method based on a multi-scale attention coupling mechanism is adopted to extract multi-scale input features through a backbone network, combine multi-domain attention and self-attention mechanism to form a feature multi-domain hierarchy, and image classification is realized through weighted aggregation of multi-scale measurement modules.

Benefits of technology

The correct classification and acquisition of traffic anomaly images can be achieved without a large amount of labeling data, which improves the detection ability and adaptability of the model in new scenarios, and reduces the cost of data labeling and training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114898158B_ABST
    Figure CN114898158B_ABST
Patent Text Reader

Abstract

The present invention discloses a small sample traffic anomaly image acquisition method and system based on a multi-scale attention coupling mechanism. The method of the present invention comprises the following steps: S1. labeling traffic anomaly situation images collected from cameras arranged earlier as labeled data sets, and further dividing them into training sets and test sets; S2. performing data processing on sample data and constructing scenario tasks, randomly sampling a small number of samples from the training set as support set samples and a certain number of similar samples as query sample images; S3. using a backbone network to extract features from images in each scenario task, and obtaining multi-scale input features; S4. combining two different levels of attention to form a feature multi-domain hierarchical structure; S5. weighting and aggregating measurement results of different scales, and realizing image classification according to the measurement scores between the final support set and query set samples; S6. performing end-to-end training using a loss function; S7. performing a test to retain the optimal training weights; and S8. model deployment and image acquisition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image analysis technology, and relates to a combination technology of deep learning and traffic control, and in particular to a small sample traffic abnormality image acquisition method and system based on a multi-scale attention coupling mechanism. Background Art

[0002] Traditional traffic monitoring cameras only save all monitoring images and send them back to the data center. Supervisors need to watch all monitoring images to determine whether there are any traffic anomalies on the road. Due to the huge number of cameras deployed, the review of traffic anomalies in the monitoring images is very time-consuming and inefficient, and it is impossible to dispatch relevant departments in time to handle various traffic anomalies.

[0003] With the continuous development of traditional deep learning, deep learning models have shown excellent performance in many computer vision fields such as image classification, object detection, and image segmentation. Deep learning models are also increasingly being deployed in smart cameras and used in the field of traffic for road supervision. After the pre-trained deep learning model is deployed in the smart camera, it can detect the camera's images in real time, annotate the abnormal traffic images, and then send them back to the data center. Detection of various types of traffic anomalies includes but is not limited to traffic violation detection and traffic accident detection. This means that supervisors no longer need to spend a lot of time manually searching for traffic anomalies in all monitoring images. They only need to further screen the images that are judged to be abnormal, thereby improving the efficiency of supervisors dispatching relevant departments to deal with various traffic anomalies.

[0004] However, the training of deep learning models requires a large amount of annotated data, which consumes a lot of manpower and time. The limited available data limits the availability and scalability of deep models. Although fine-tuning on a deep model pre-trained with a large amount of annotated data can correctly classify some images, in the absence of available annotated data, the model is prone to overfitting during the training process, and its practicality is limited. Due to the different deployment environments of cameras, the angles of the captured images, the brightness of the light, etc. are also different, which will cause the detection accuracy of deep learning models trained by specific data sets on cameras in some scenes to decrease, that is, the generalization ability is insufficient. At the same time, it is challenging to use a few abnormal traffic images taken on a specific road section in a short period of time to enable the deep learning model to quickly adapt to the capture work of the deployment scene.

[0005] When humans take on a new task, they can quickly master relevant skills based on experience. Inspired by this, a few-shot learning method was proposed in the field of image classification. Few-shot learning focuses on the problem of learning how to learn through various specific methods such as data augmentation, metric learning, and meta-learning. The meta-learning method is a popular and very effective few-shot learning method. It constructs a series of scenario tasks in the meta-training phase and uses a small number of support set samples (labeled samples) in each scenario task to build meta-knowledge and optimize the query sample (unlabeled sample) classification model. In the test phase, the same scenario setting is used to generalize the meta-knowledge to the new test task to complete the sample classification task. Summary of the invention

[0006] In view of the above status of the prior art, the present invention proposes a small sample traffic abnormality image acquisition method and system based on a multi-scale attention coupling mechanism.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] S1. Label the traffic abnormality images collected by the pre-arranged cameras as a labeled dataset, and further divide them into a training set and a test set;

[0009] S2. Process the sample data and construct scenario tasks, randomly sample a small number of samples from the training set as support set samples and a certain number of similar samples as query sample images;

[0010] S3. Use the backbone network to extract features from images in each scenario task and obtain multi-scale input features;

[0011] S4. Combine the two different levels of attention to form a feature multi-domain hierarchy;

[0012] S5. Weighted aggregation of different scale measurement results, and image classification based on the final measurement scores between the support set and query set samples;

[0013] S6. Perform end-to-end training using loss function;

[0014] S7. Perform test to retain the optimal training weights;

[0015] S8. Model deployment and image acquisition: Deploy the model with the optimal weights trained to the camera, place the camera in the new scene, collect traffic anomaly images for data annotation as support set samples, and use the subsequently collected images to be detected as query set samples to realize abnormal image classification and acquisition.

[0016] Furthermore, in step S1, various types of traffic abnormal images are labeled as a labeled dataset D label. Then there will be a labeled dataset D label The various types of traffic anomaly images are divided into training set D according to a certain ratio. base and the test set D test Preferably, the categories of traffic anomaly images in the two data subsets are different.

[0017] Furthermore, in step S2, the data processing includes image cropping and data enhancement. The scenario task T consists of N categories, each category has K samples (i.e., N-way K-shot setting), and the support set and queryset Composition, that is, T = (S, Q), where x i and x j are image samples in the support set and query set, respectively, i and j As a further optimization, in the training phase, it supports random sampling of samples from the training set D base , query samples are sampled from the same type of samples with the same traffic anomaly as the support set samples, and support samples in the test phase are randomly sampled from the test set D test ,The query samples are sampled from the same type of samples with the same traffic abnormality as the support set samples.

[0018] Furthermore, the idea of ​​demultiplexing is that a common input can be switched to multiple independent outputs, so a multi-scale feature demultiplexer is designed in step S3 of the present invention to obtain multi-scale input features. To extract features, in order to supplement the unique information of each scale feature, The phased output features that match their scales are extracted from the image, and feature fusion is performed through a convolution operation with a convolution kernel of 1×1. In addition, in order to reduce the number of parameters while retaining useful information, each scale feature will be subjected to a maximum pooling operation corresponding to the size of the feature map to obtain Z scale features. Where F = {f z}={f z,s ,f z,q}, z=1,...,Z.

[0019] Finally, it is necessary to calculate the class prototype of the sample category contained in each scenario task at each scale. Specifically, in the 1-shot or many-shot setting, the class prototype of category n can be represented by the mean embedding feature of the same sample:

[0020] Furthermore, the two different levels of attention in step S4 are: multi-domain attention weights focusing on inter-class diversity and self-attention weights that focus on intra-class correlations

[0021] Furthermore, the attention weights of multiple domains The adaptive spatial importance generator G with learnable parameters s And the sigmoid function is obtained, the formula is as follows:

[0022]

[0023]

[0024] where θ s is a learnable parameter, The mask for each domain is obtained by , and σ is the sigmoid function.

[0025] Furthermore, the prototype features By weighting The linear mapping layer is transformed into a query matrix, a key matrix, and a value matrix So the self-attention weight It can be expressed as:

[0026]

[0027] Furthermore, the attention coupling process in two different levels can be expressed as:

[0028]

[0029] in Represents element-wise product.

[0030] Furthermore, the output features of the multi-domain attention coupling module can be expressed as:

[0031]

[0032]

[0033] Among them, || represents the cascade operation, and FFN (Feed Forward Network) represents the feedforward network.

[0034] Furthermore, the multi-scale metric module in step S5 is composed of an adaptive weight generator G w Composition, prototype features at all scales and query feature f z,q are spliced ​​together and passed into G w The importance weights of the measurement results of each scale are generated, and the formula is as follows:

[0035]

[0036] where || represents the cascade operation, θ w is a learnable parameter. Under the constraint of the true label, through end-to-end training, G w It can be learned that the scale measurement results that are beneficial to the final classification result are assigned higher weights. The final measurement result can be expressed as:

[0037]

[0038] where d(·,·) represents the metric function.

[0039] Furthermore, the nearest neighbor algorithm can be used to obtain the query sample x according to the metric score. j The label prediction results

[0040]

[0041] Furthermore, in step S6, the method of the present invention can be optimized through loss function learning in an end-to-end setting during the training phase. The target loss function consists of three parts: multi-scale classification loss L cls , domain diversity loss And multi-scale balance loss

[0042] First, in order to accurately predict the query label, we use the conventional cross entropy loss for the classification loss of the multi-scale metric results:

[0043]

[0044] Where L CE represents the cross entropy loss.

[0045] Secondly, in order to prevent the domain perception at each scale from being concentrated on the more discriminative domain, the sparsity of the domain attention needs to be constrained to achieve the purpose of perceiving different domains at different scales. The specific formula is as follows:

[0046]

[0047] This loss uses cosine similarity to calculate the similarity between domains at each scale. When the domain similarity between the i-th and j-th scales is large, L div will be large. By minimizing L div Domain masks at different scales are encouraged to be discriminative.

[0048] Finally, in order to ensure that the prediction results of each scale can be optimized in the correct direction, a balanced loss function is used to constrain the prediction results of each scale:

[0049]

[0050] in Represents the query label prediction results at each scale.

[0051] The overall objective function combined with the above mentioned losses can be expressed as:

[0052]

[0053] where λ 1 and λ 2 Diversity loss and balance loss The balance parameters.

[0054] Furthermore, in step S7, multiple iterations of training will be performed during the training phase. After each iteration of the training data set, a test scenario task will be randomly sampled from the test data set, and accuracy testing will be performed using steps S3-S5 to obtain the image classification accuracy of the current training weight on the test set, and save the training weight with the highest accuracy.

[0055] The present invention also discloses a small sample traffic anomaly image acquisition system based on a multi-scale attention coupling mechanism, which includes the following modules:

[0056] Dataset creation module: Label the traffic abnormality images collected by the camera as a labeled dataset, and further divide them into training set and test set;

[0057] Construct scenario task module: process sample data and construct scenario tasks, randomly sample samples from the training set as support set samples and similar samples as query sample images;

[0058] Feature extraction module: Use the backbone network to extract features from images in each scenario task and obtain multi-scale input features;

[0059] Multi-domain attention coupling module: combines two different levels of attention to form a feature multi-domain hierarchical structure;

[0060] Multi-scale metric module: weights and aggregates the metric results of different scales, and implements image classification based on the metric scores between the final support set and query set samples;

[0061] Training module: end-to-end training using loss function;

[0062] Optimal training weight retention module: performs tests to retain the optimal training weights;

[0063] Model deployment and image acquisition module: deploy the model with the optimal weights trained to the camera, place the camera in the new scene, collect traffic anomaly images for data annotation as support set samples, and use the subsequently collected images to be detected as query set samples to realize abnormal image classification and acquisition.

[0064] Compared with the prior art, the small sample traffic anomaly image acquisition method and system based on the multi-scale attention coupling mechanism of the present invention can use a small number of labeled traffic anomaly images to correctly classify and collect unclassified traffic anomaly images, without the need for tedious large-scale data collection and annotation work, and can quickly adapt to the detection work of new scenes. At the same time, by constructing multi-scale and multi-domain feature relationships, the weak correlation of intra-class features in scenario tasks is enhanced, and the diversity of inter-class features is improved. More specifically, the self-attention mechanism perceives the intra-class correlation of the supported samples and realizes adaptive embedded feature enhancement. The multi-scale structure and multi-domain attention coupling module are responsible for generating domain importance weights through the domain perception module, and constructing a feature multi-domain hierarchical structure by combining the attention coupling structure with the self-attention weight, thereby achieving diversity guarantee between feature classes. The above two points are used to improve the accuracy of small sample image classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 This is a flow chart of a small sample traffic anomaly image acquisition method based on a multi-scale attention coupling mechanism provided in Example 1 of the present invention.

[0066] Figure 2 This is a flowchart of the multi-scale feature demultiplexer generating a multi-scale feature extraction in step S13 provided by an embodiment of the present invention.

[0067] Figure 3 It is a flowchart of the multi-domain attention coupling module in step S14 provided in an embodiment of the present invention.

[0068] Figure 4 This is a block diagram of a small sample traffic anomaly image acquisition system based on a multi-scale attention coupling mechanism according to a second embodiment of the present invention. DETAILED DESCRIPTION

[0069] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0070] The purpose of the present invention is to address the defects of the prior art and provide a small sample traffic abnormality image acquisition method and system based on a multi-scale attention coupling mechanism.

[0071] Embodiment 1

[0072] This embodiment provides a small sample traffic abnormality image acquisition method based on a multi-scale attention coupling mechanism, and its specific implementation process is as follows: Figure 1 As shown, the steps include:

[0073] S11. Label the traffic abnormality images collected by the cameras deployed earlier as a labeled data set, and further divide them into a training set and a test set;

[0074] S12. Process the sample data and construct scenario tasks, randomly sample a small number of samples from the training set as support set samples and a certain number of similar samples as query sample images;

[0075] S13. Extract features from images in each scenario task using the backbone network, and obtain multi-scale input features using a multi-scale feature demultiplexer;

[0076] S14. Use the multi-domain attention coupling module to combine the two different levels of attention to form a feature multi-domain hierarchical structure;

[0077] S15. Use the multi-scale metric module to weight and aggregate the metric results of different scales, and implement image classification based on the metric scores between the final support set and query set samples;

[0078] S16. Perform end-to-end training using loss function;

[0079] S17. Perform test to retain the optimal training weights;

[0080] S18. Model deployment and image acquisition.

[0081] The specific ideas of this embodiment are as follows: 1. Collect data and make a data set, label various types of traffic abnormality images as labeled data sets, and then divide various types of images in the labeled data sets into training sets and test sets at a ratio of 4:1; 2. In the training stage, center crop the impact and perform data enhancement, then randomly sample from the training data set to construct the scenario task, and in the test stage, randomly sample from the test data set to construct the scenario task; 3. Extract features from samples in the scenario task through a backbone network, and then use a multi-scale feature demultiplexer to obtain multi-scale input features; 4. Combine domain attention and self-attention in the multi-domain attention coupling module to enhance the simultaneous representation of feature representation. 5. Then use the multi-scale measurement module to weightedly aggregate the measurement results of different scales to obtain a final measurement result, and then use the nearest neighbor algorithm to obtain the image classification accuracy 6. Use the loss function for end-to-end training; 7. After each iteration of the training set, the current training weights will be used to test in the test set images, and the network weights with the highest test accuracy will be saved; 8. After the trained model weights and model are deployed to the smart camera, you only need to place the smart camera in the new scene, collect a small number of traffic anomaly images that you want to detect for data annotation as support set samples, and use the subsequent collected images to be detected as query set samples to realize abnormal image classification collection.

[0082] The steps of this embodiment are described in detail as follows:

[0083] In step S11, various traffic anomaly images are labeled as a labeled dataset D label , then there will be a labeled dataset D label The various types of traffic anomaly images are divided into training set D according to the ratio of 4:1 base and the test set D test The categories of traffic anomaly images in the two data subsets are different.

[0084] In step S12, the training set D base and the test set D test Center cropping was used to crop the image size to 512×512. base Data enhancement (such as random cropping, color jittering, horizontal flipping, etc.) is used, and the data enhancement method can adjust or change parameters according to the specific traffic anomaly images detected.

[0085] The method of constructing scenario task T is as follows: training set where x i represents the i-th picture, y i Represents x i The category label, C base Yes D base The set of label categories contained. Similarly, the test set Dtest The image sample labels contained in test , In the training phase, from the training dataset D base N classes are randomly selected from the dataset, each class contains K samples, that is, N-way K-shot setting, to form the support set Q query samples of the same category as the support samples constitute the query set The scenario task can then be expressed as T = (S, Q). The goal is to train a classifier that can be used in the test phase with a small number of labeled S∈D test Traffic anomaly images, accurately convert unlabeled samples Q∈D test Map to the correct label.

[0086] In step S13, the idea of ​​demultiplexing is used to construct multi-scale input features.

[0087] All samples x in scenario task T = {x i ,x j} i=1,...,N×K;j=1,...,Q Sent to the backbone network To obtain feature mapping, the ResNet34 feature extraction network is used in this embodiment. In addition, other lightweight feature extraction networks can also be used, and they can all be deployed in the terminal device. Figure 2 As shown, The final output features of are upsampled (nearest neighbor upsampling) to obtain larger scale features. At the same time, in order to supplement the unique information of each scale feature, The phased output features that match their scales are extracted from the image, and feature fusion is performed through a convolution operation with a convolution kernel of 1×1. In addition, in order to reduce the number of parameters while retaining useful information, each scale feature will be subjected to a maximum pooling operation corresponding to the size of the feature map to obtain Z scale features. Where F = {f z}={f z,s ,f z,q},

[0088] Finally, the class prototypes of the image sample categories contained in each scenario task are calculated at each scale. Specifically, in the 1-shot or many-shot setting, the class prototype of category n can be represented by the mean embedding feature of similar samples: Where n=1,...,N.

[0089] In step S14, Figure 3 As shown, the domain importance weights are generated by the domain-aware module in the multi-domain attention coupling module, and the feature multi-domain hierarchy is constructed by combining the attention coupling structure with the self-attention weights.

[0090] First, the prototype features at each scale are divided into multiple subspaces to form a multi-domain representation. Assuming the number of domain divisions is H, the features of each domain can be expressed as where h=1,...,H,L z =C z / H.

[0091] Then in the domain-aware module, this embodiment designs an adaptive spatial importance generator G with learnable parameters s , it is possible to derive a mask for each domain while suppressing irrelevant noise:

[0092]

[0093] where θ s is a learnable parameter, G s It consists of two fully connected layers. The attention weight formula of the domain is as follows:

[0094]

[0095] Where σ is the sigmoid function, This attention map can reflect the importance of different prototype features in different domains. The trained domain-aware module has good generalization and adaptability, and can effectively generate an attention map that best matches the importance distribution in the current domain.

[0096] Then the prototype features By weighting The linear mapping layer is converted into a query matrix, a key matrix, and a value matrix The self-attention weight formula is as follows:

[0097]

[0098] The domain weight is the result of weighing all the features. In order to form a multi-domain trade-off, the element product is used to couple it with to obtain a more effective feature pair relationship representation. The formula is as follows:

[0099]

[0100] in Represents the element-wise product. Domain weights can assign higher weights to domains with good sample feature representation and suppress domains that may have irrelevant noise.

[0101] Finally, the output of the multi-domain attention coupling model uses the conventional transformer output structure to strengthen the feature representation. The formula is as follows:

[0102]

[0103]

[0104] Where || represents the cascade operation, FFN (Feed Forward Network) represents the feedforward network,

[0105] In step S15, a multi-scale measurement module is used to integrate similarity measurement results between support samples and query samples at different scales through weighted aggregation to obtain a final measurement result.

[0106] It is not the most appropriate choice to weight the measurement results of features at different scales by simply equal weighting. In each independent scenario task, the contribution of measurement results at different scales is different. A suitable adaptive weighting method can maximize the use of measurement information at each scale. To achieve this goal, this embodiment designs an adaptive weight generator G w Specifically, the prototype features of each scale and query feature f z,q are spliced ​​together and passed into G w The importance weights of the measurement results of each scale are generated, and the formula is as follows:

[0107]

[0108] where || represents the cascade operation, θ w is a learnable parameter, G w It consists of two fully connected layers, where the first fully connected layer is followed by a LeakyReLU activation function. Under the constraint of the true label, G w It can be learned that the scale measurement results that are beneficial to the final classification result are assigned higher weights. The final measurement result can be expressed as:

[0109]

[0110] Wherein d(·,·) represents a metric function, and this embodiment uses the Euclidean distance function.

[0111] Using the nearest neighbor algorithm, we can get the query sample x j The label prediction results

[0112]

[0113] In step S16, the loss function consists of three parts: multi-scale classification loss L cls , domain diversity loss And multi-scale balance loss

[0114] First, in order to accurately predict the query label, the classification loss of the multi-scale metric results uses the cross entropy loss:

[0115]

[0116] Where L CE represents the cross entropy loss.

[0117] Secondly, in order to prevent the domain perception at each scale from being concentrated on the more discriminative domain, the sparsity of the domain attention needs to be constrained to achieve the purpose of perceiving different domains at different scales. The specific formula is as follows:

[0118]

[0119] This loss uses cosine similarity to calculate the similarity between domains at each scale. When the domain similarity between the i-th and j-th scales is large, L div will be large. By minimizing L div Domain masks at different scales are encouraged to be discriminative.

[0120] Finally, in order to ensure that the prediction results of each scale can be optimized in the correct direction, a balanced loss function is used to constrain the prediction results of each scale:

[0121]

[0122] in Represents the query label prediction results at each scale.

[0123] The overall objective function combined with the above mentioned losses can be expressed as:

[0124]

[0125] where λ 1 and λ 2 Diversity loss and balance loss The balancing parameters of are set to 0.1 and 1 respectively during training.

[0126] In step S17, after the complete network structure is built, the initial learning rate is set to 1×10 -4 The stochastic gradient descent optimizer is used for training. During the training process, after each iteration of the training data set, a test scenario task is randomly sampled from the test data set, and the accuracy test is performed using steps S13-S15. The model weight with the highest classification accuracy tested in the test set will be saved.

[0127] In step S18, after the trained model weights and model are deployed to the smart camera, it is only necessary to place the smart camera in a new scene, collect a small number of traffic anomaly images that you want to detect for data annotation as support set samples, and use the subsequently collected images to be detected as query set samples to achieve abnormal image classification and collection.

[0128] This embodiment proposes a small sample traffic anomaly image acquisition method based on a multi-scale attention coupling mechanism. First, data is collected and a data set is made. Then, a scenario task is constructed. The backbone network is used to extract sample features in the scenario task. Then, a multi-scale feature demultiplexer is used to obtain multi-scale input features. Then, a multi-domain attention coupling module and a multi-scale metric module are used to enhance weak correlation and improve feature diversity representation to improve image classification accuracy. Finally, the test set is used for testing and the optimal training weight is retained. In subsequent specific applications, only a small number of traffic anomaly images that are expected to be detected need to be collected from the newly arranged cameras for data annotation as support set samples, and the images to be detected collected later are used as query set samples to realize abnormal image classification acquisition.

[0129] Embodiment 2

[0130] like Figure 4 As shown, this embodiment provides a small sample traffic abnormality image acquisition system based on a multi-scale attention coupling mechanism, which includes the following modules:

[0131] Dataset creation module: Label the traffic abnormality images collected by the camera as a labeled dataset, and further divide them into training set and test set;

[0132] Construct scenario task module: process sample data and construct scenario tasks, randomly sample samples from the training set as support set samples and similar samples as query sample images;

[0133] Feature extraction module: Use the backbone network to extract features from images in each scenario task and obtain multi-scale input features;

[0134] Multi-domain attention coupling module: combines two different levels of attention to form a feature multi-domain hierarchical structure;

[0135] Multi-scale metric module: weights and aggregates the metric results of different scales, and implements image classification based on the metric scores between the final support set and query set samples;

[0136] Training module: end-to-end training using loss function;

[0137] Optimal training weight retention module: performs tests to retain the optimal training weights;

[0138] Model deployment and image acquisition module: deploy the model with the optimal weights trained to the camera, place the camera in the new scene, collect traffic anomaly images for data annotation as support set samples, and use the subsequently collected images to be detected as query set samples to realize abnormal image classification and acquisition.

[0139] In the dataset production module, various traffic anomaly images are labeled as labeled dataset D label , then there will be a labeled dataset D label The various types of traffic anomaly images are divided into training set D according to the ratio of 4:1 base and the test set D test The categories of traffic anomaly images in the two data subsets are different.

[0140] In the scenario task module, the training set D base and the test set D test Center cropping was used to crop the image size to 512×512. base Data enhancement (such as random cropping, color jittering, horizontal flipping, etc.) is used, and the data enhancement method can adjust or change parameters according to the specific traffic anomaly images detected.

[0141] The method of constructing scenario task T is as follows: training set where x i represents the i-th picture, y i Represents x i The category label, C base Yes D base The set of label categories contained. Similarly, the test set D test The image sample labels contained in test , In the training phase, from the training dataset D base N classes are randomly selected from the dataset, each class contains K samples, that is, N-way K-shot setting, to form the support set Q query samples of the same category as the support samples constitute the query set The scenario task can then be expressed as T = (S, Q). The goal is to train a classifier that can be used in the test phase with a small number of labeled S∈D test Traffic anomaly images, accurately convert unlabeled samples Q∈D test Map to the correct label.

[0142] In the feature extraction module, the idea of ​​demultiplexing is used to construct multi-scale input features.

[0143] All samples x in scenario task T = {x i ,x j} i=1,...,N×K;j=1,...,QSent to the backbone network To obtain feature mapping, the ResNet34 feature extraction network is used in this embodiment. In addition, other lightweight feature extraction networks can also be used, and they can all be deployed in the terminal device. Figure 2 As shown, The final output features of are upsampled (nearest neighbor upsampling) to obtain larger scale features. At the same time, in order to supplement the unique information of each scale feature, The phased output features that match their scales are extracted from the image, and feature fusion is performed through a convolution operation with a convolution kernel of 1×1. In addition, in order to reduce the number of parameters while retaining useful information, each scale feature will be subjected to a maximum pooling operation corresponding to the size of the feature map to obtain Z scale features. Where F = {f z}={f z,s ,f z,q},

[0144] Finally, the class prototypes of the image sample categories contained in each scenario task are calculated at each scale. Specifically, in the 1-shot or many-shot setting, the class prototype of category n can be represented by the mean embedding feature of similar samples: Where n=1,...,N.

[0145] In the multi-domain attention coupling module, Figure 3 As shown, the domain importance weights are generated by the domain-aware module in the multi-domain attention coupling module, and the feature multi-domain hierarchy is constructed by combining the attention coupling structure with the self-attention weights.

[0146] First, the prototype features at each scale are divided into multiple subspaces to form a multi-domain representation. Assuming the number of domain divisions is H, the features of each domain can be expressed as where h=1,...,H,L z =C z / H.

[0147] Then in the domain-aware module, this embodiment designs an adaptive spatial importance generator G with learnable parameters s , it is possible to derive a mask for each domain while suppressing irrelevant noise:

[0148]

[0149] where θ s is a learnable parameter, G s It consists of two fully connected layers. The attention weight formula of the domain is as follows:

[0150]

[0151] Where σ is the sigmoid function, This attention map can reflect the importance of different prototype features in different domains. The trained domain-aware module has good generalization and adaptability, and can effectively generate an attention map that best matches the importance distribution in the current domain.

[0152] Then the prototype features By weighting The linear mapping layer is transformed into a query matrix, a key matrix, and a value matrix The self-attention weight formula is as follows:

[0153]

[0154] The domain weight is the result of weighing all the features. In order to form a multi-domain trade-off, the element product is used to couple it with to obtain a more effective feature pair relationship representation. The formula is as follows:

[0155]

[0156] in Represents the element-wise product. Domain weights can assign higher weights to domains with good sample feature representation and suppress domains that may have irrelevant noise.

[0157] Finally, the output of the multi-domain attention coupling model uses the conventional transformer output structure to strengthen the feature representation. The formula is as follows:

[0158]

[0159]

[0160] Where || represents the cascade operation, FFN (Feed Forward Network) represents the feedforward network,

[0161] In the multi-scale measurement module, the similarity measurement results between support samples and query samples at different scales are integrated together through weighted aggregation to obtain the final measurement result.

[0162] It is not the most appropriate choice to weight the measurement results of features at different scales by simply equal weighting. In each independent scenario task, the contribution of measurement results at different scales is different. A suitable adaptive weighting method can maximize the use of measurement information at each scale. To achieve this goal, this embodiment designs an adaptive weight generator G wSpecifically, the prototype features of each scale and query feature f z,q are spliced ​​together and passed into G w The importance weights of the measurement results of each scale are generated, and the formula is as follows:

[0163]

[0164] where || represents the cascade operation, θ w is a learnable parameter, G w It consists of two fully connected layers, where the first fully connected layer is followed by a LeakyReLU activation function. Under the constraint of the true label, G w It can be learned that the scale measurement results that are beneficial to the final classification result are assigned higher weights. The final measurement result can be expressed as:

[0165]

[0166] Wherein d(·,·) represents a metric function, and this embodiment uses the Euclidean distance function.

[0167] Using the nearest neighbor algorithm, we can get the query sample x j The label prediction results

[0168]

[0169] In the training module, the loss function consists of three parts: multi-scale classification loss L cls , domain diversity loss And multi-scale balance loss

[0170] First, in order to accurately predict the query label, the classification loss of the multi-scale metric results uses the cross entropy loss:

[0171]

[0172] Where L CE represents the cross entropy loss.

[0173] Secondly, in order to prevent the domain perception at each scale from being concentrated on the more discriminative domain, the sparsity of the domain attention needs to be constrained to achieve the purpose of perceiving different domains at different scales. The specific formula is as follows:

[0174]

[0175] This loss uses cosine similarity to calculate the similarity between domains at each scale. When the domain similarity between the i-th and j-th scales is large, L divwill be large. By minimizing L div Domain masks at different scales are encouraged to be discriminative.

[0176] Finally, in order to ensure that the prediction results of each scale can be optimized in the correct direction, a balanced loss function is used to constrain the prediction results of each scale:

[0177]

[0178] in Represents the query label prediction results at each scale.

[0179] The overall objective function combined with the above mentioned losses can be expressed as:

[0180]

[0181] where λ 1 and λ 2 Diversity loss and balance loss The balancing parameters of are set to 0.1 and 1 respectively during training.

[0182] In the optimal training weight retention module, after the complete network structure is built, the initial learning rate is set to 1×10 -4 The model is trained using a stochastic gradient descent optimizer. During the training process, after each iteration of the training data set, the classification accuracy of the current training weights will be tested using the test set, and the model weight with the highest classification accuracy tested in the test set will be saved.

[0183] In the model deployment and image acquisition module, after the trained model weights and model are deployed to the smart camera, you only need to place the smart camera in the new scene, collect a small number of traffic anomaly images that you want to detect for data annotation as support set samples, and then use the subsequently collected images to be detected as query set samples to realize abnormal image classification and acquisition.

[0184] This embodiment ensures the ease of use and flexibility of the model to the greatest extent through modular design.

[0185] Compared with the prior art, the small sample traffic anomaly image acquisition method and system based on the multi-scale attention coupling mechanism of the present invention can use a small number of labeled traffic anomaly images to correctly classify and collect unclassified images, without the need for tedious large-scale data collection and annotation work, and can quickly adapt to the detection work of new scenes. At the same time, by constructing multi-scale and multi-domain feature relationships, the weak correlation of intra-class features in scenario tasks is enhanced, and the diversity of inter-class features is improved. Specifically, the self-attention mechanism perceives the intra-class correlation of the supported samples and realizes adaptive embedded feature enhancement. The multi-scale structure and multi-domain attention coupling module are responsible for generating domain importance weights through the domain perception module, and constructing a feature multi-domain hierarchical structure by combining the attention coupling structure with the self-attention weight, thereby achieving diversity guarantee between feature classes. The above two points are used to improve the accuracy of small sample image classification. The present invention also maximizes the ease of use and flexibility of the model through modular design.

[0186] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A small sample traffic anomaly image acquisition method based on multi-scale attention coupling mechanism, It is characterized in that Includes steps: S1. Label the traffic abnormality images collected by the camera as a labeled dataset, and further divide them into a training set and a test set; S2. Process the sample data and construct scenario tasks, randomly sample samples from the training set as support set samples and similar samples as query sample images; S3. Use the backbone network to extract features from images in each scenario task and obtain multi-scale input features; S4. Combine the two different levels of attention to form a feature multi-domain hierarchy; S5. Weighted aggregation of different scale measurement results, and image classification based on the final measurement scores between the support set and query set samples; S6. Perform end-to-end training using loss function; S7. Perform test to retain the optimal training weights; S8. Deploy the model with the best weight trained to the camera, place the camera in the new scene, collect traffic anomaly images for data annotation as support set samples, and use the subsequently collected images to be detected as query set samples to realize abnormal image classification and collection; In step S4, using the feature In the multi-domain attention coupling module, two different levels of attention are combined: the multi-domain attention weight W that focuses on inter-class diversity ss and the self-attention weight W that focuses on intra-class correlation sa ; Attention weights for multiple domains The adaptive spatial importance generator G is composed of a learning parameter s And the sigmoid function is obtained, the formula is as follows: Among them, θ s is a learnable parameter, is the mask for each domain, σ is the sigmoid function; The prototype feature By weighting The linear mapping layer is converted into a query matrix, a key matrix, and a value matrix So the self-attention weight It is expressed as: The attention coupling process in two different levels is expressed as: in, represents element-wise product; The output feature of the multi-domain attention coupling module is expressed as: Among them, || represents the cascade operation, and FFN represents the feedforward network; In step S5, the adaptive weight generator G in the multi-scale metric module is used w Get the importance weights of the measurement results at each scale: Among them, || represents the cascade operation, θ w is a learnable parameter; the final measurement result is expressed as: Where d(·,·) represents the metric function; Use the nearest neighbor algorithm to get the query sample x according to the metric score j The label prediction results 2. According to the method for acquiring small sample traffic abnormality images based on multi-scale attention coupling mechanism according to claim 1, It is characterized in that In step S1, multiple types of traffic anomaly images are labeled as a labeled dataset D label , then there will be a labeled dataset D label The multi-class traffic anomaly images are divided into training set D according to the proportion base and the test set D test .

3. According to the method for acquiring small sample traffic abnormality images based on multi-scale attention coupling mechanism according to claim 2, It is characterized in that In step S2, the image is cropped and data augmentation is performed at the same time; the scenario task T consists of N categories, each category has K samples, and the support set and queryset Composition, that is, T = (S, Q), where x i and x j are image samples in the support set and query set, respectively, i and j They are the corresponding labels respectively.

4. According to the method for acquiring small sample traffic abnormality images based on multi-scale attention coupling mechanism according to claim 3, It is characterized in that In the training phase, samples are randomly sampled from the training set D. base , query samples are sampled from the same type of samples with the same traffic anomaly as the support set samples, and support samples in the test phase are randomly sampled from the test set D test ,The query samples are sampled from the same type of samples with the same traffic abnormality as the support set samples.

5. According to the method for acquiring small sample traffic abnormality images based on multi-scale attention coupling mechanism according to claim 4, It is characterized in that In step S3, the backbone network is used for feature extraction, and a multi-scale feature demultiplexer is used to construct multi-scale input features. Where F = {f z }={f z,s ,f z,q }, z = 1, ..., Z; The class prototypes of the sample categories contained in each scenario task are calculated at each scale: In the 1-shot or many-shot setting, the class prototype of category n is represented by the mean of the embedding features of similar samples: Where n=1,...,N.

6. According to the method for acquiring small sample traffic abnormality images based on multi-scale attention coupling mechanism according to claim 5, It is characterized in that In step S6, the training phase is optimized and learned in an end-to-end setting; The target loss function consists of three parts: multi-scale classification loss L cls , domain diversity loss And multi-scale balance loss For the classification loss of multi-scale metric results, cross entropy loss is used: Among them, L CE represents the cross entropy loss; The sparsity of domain attention is constrained. The specific formula is as follows: A balanced loss function is used to constrain the prediction results of each scale: in, Indicates the query label prediction results at each scale; The overall objective function combined with the loss is expressed as: Among them, λ 1 and λ 2 Diversity loss and balance loss The balance parameters.

7. According to claim 6, a small sample traffic abnormality image acquisition method based on a multi-scale attention coupling mechanism, It is characterized in that In step S7, multiple iterations of training will be performed during the training phase. After each iteration of the training data set, a test scenario task will be randomly sampled from the test data set, and accuracy testing will be performed using steps S3-S5 to obtain the image classification accuracy of the current training weight on the test set, and save the training weight with the highest accuracy.

8. A small sample traffic anomaly image acquisition system based on multi-scale attention coupling mechanism, Its characteristics are Includes the following modules: Dataset creation module: Label the traffic abnormality images collected by the camera as a labeled dataset, and further divide them into training set and test set; Construct scenario task module: process sample data and construct scenario tasks, randomly sample samples from the training set as support set samples and similar samples as query sample images; Feature extraction module: Use the backbone network to extract features from images in each scenario task and obtain multi-scale input features; Multi-domain attention coupling module: combines two different levels of attention to form a feature multi-domain hierarchical structure; Multi-scale metric module: weights and aggregates the metric results of different scales, and implements image classification based on the metric scores between the final support set and query set samples; Training module: end-to-end training using loss function; Optimal training weight retention module: performs tests to retain the optimal training weights; Model deployment and image acquisition module: deploy the model with the best weights trained to the camera, place the camera in the new scene, collect traffic anomaly images for data annotation as support set samples, and use the subsequently collected images to be detected as query set samples to realize abnormal image classification and acquisition; In the multi-domain attention coupling module, the feature In the multi-domain attention coupling module, two different levels of attention are combined: the multi-domain attention weight W that focuses on inter-class diversity ss and the self-attention weight W that focuses on intra-class correlation sa ; Attention weights for multiple domains The adaptive spatial importance generator G is composed of a learning parameter s And the sigmoid function is obtained, the formula is as follows: Among them, θ s is a learnable parameter, is the mask for each domain, σ is the sigmoid function; The prototype feature By weighting The linear mapping layer is transformed into a query matrix, a key matrix, and a value matrix So the self-attention weight It is expressed as: The attention coupling process in two different levels is expressed as: in, represents element-wise product; The output feature of the multi-domain attention coupling module is expressed as: Among them, || represents the cascade operation, and FFN represents the feedforward network; In the multi-scale measurement module, the adaptive weight generator G in the multi-scale measurement module is used w Get the importance weights of the measurement results at each scale: Among them, || represents the cascade operation, θ w is a learnable parameter; the final measurement result is expressed as: Where d(·,·) represents the metric function; Use the nearest neighbor algorithm to get the query sample x according to the metric score j The label prediction results

Citation Information

Patent Citations

  • Image classification method based on matching network few-sample learning

    CN113537305A

  • Small sample remote sensing image scene classification method based on embedded smooth graph neural network

    CN114067160A