Small-sample Behavior Recognition Method and System Based on Subspace Classification

By constructing a subspace classifier, using depth estimation and feature fusion networks, the distance between query samples and subspace is directly calculated, which solves the problems of high computational complexity and low prediction accuracy in small sample behavior recognition, and achieves more efficient behavior recognition.

CN116189280BActive Publication Date: 2025-07-18BEIJING JIAOTONG UNIV +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211594956.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2025-07-18
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

The existing small sample behavior recognition methods have insufficient generalization capabilities, high computational complexity, and unreliable prototype representations, resulting in low classification prediction accuracy.

Method used

A small sample behavior recognition method based on subspace classification is adopted, and a subspace classifier is constructed through deep estimation network, feature extraction network, feature fusion network and recognition network, and all sample features of each category are supported to directly calculate the distance between query sample features and subspace, reducing the calculation amount and improving prediction accuracy.

Benefits of technology

Effectively utilizing all sample features of each category of the support set reduces the amount of calculation, improves the prediction accuracy of small sample behavior recognition, and solves the problems of high computational complexity and inaccurate prediction in the existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189280B_ABST
    Figure CN116189280B_ABST
Patent Text Reader

Abstract

The present invention provides a small-sample behavior recognition method and system based on subspace classification, belonging to the field of computer recognition technology, including: obtaining an image to be recognized; using a pre-trained small-sample behavior recognition model to process the obtained image to be recognized to obtain a behavior recognition result in the image; the small-sample behavior recognition model includes a depth estimation network, a feature extraction network, a feature fusion network, and a recognition network. The present invention makes full use of all sample features in each class of the support set, constructs a subspace for each class of behaviors for classification, rather than directly using the feature mean; condenses the sample features of each class into a subspace, and directly calculates the distance from the query sample feature to the subspace, rather than calculating the distance from the query sample feature to each sample in each class in turn, reducing the computational amount.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer recognition, and particularly relates to a small-sample behavior recognition method and system based on subspace classification. Background Art

[0002] Mobile devices ubiquitous in life facilitate the recording, storage, and transmission of video information, such as smartphones, surveillance videos, etc. With the construction of smart cities, video surveillance has been deployed in various public places, playing an important role in maintaining public safety. However, due to the huge amount of data, it is time-consuming and laborious to identify by manual means alone. Therefore, the research on behavior recognition has become increasingly important, and the recognition and capture of abnormal behaviors such as falls and fights are the key points of the intelligent construction. Facing the current situation of scarce abnormal behavior samples, small-sample behavior recognition has important research significance.

[0003] Currently, existing small-sample behavior recognition methods mostly adopt a network model of "embedding network" + "classifier". In the training stage, the data set is decomposed into different tasks to learn the generalization ability of the model under the condition of class change, so that the model can learn the common parts in different tasks, such as how to extract important features and compare sample similarities, etc. In the testing stage, in the face of a new class different from the training set, the existing model does not need to be changed to complete the classification. The classic small-sample classification method, Siamese network, extracts features from two images respectively using the same network structure, and predicts the sample class by calculating the L1 distance between the features. The Matching Network uses an LSTM network integrated with an attention mechanism to extract features, and classifies samples by measuring the gap between features through the cosine distance.

[0004] In the AMeFu-Net model method proposed by Fu et al., first, the depth information of the video data set is extracted, then a feature extractor is used to extract RGB features and depth features respectively, and the two features are fused to achieve the effect of feature enhancement. Finally, the support set features and query set features are sent into a small-sample classifier for classification. The small-sample classifier adopted by AMeFu-Net is a prototype network classifier, which needs to learn a symmetric function from high-dimensional data to implement the classifier. The symmetric function is realized by average pooling. The feature vectors of each class in the support set are averaged to create a prototype representation, and the Euclidean distance is used as the distance metric to calculate the similarity between the feature vector of the video to be queried and the prototype representations of all classes, so as to obtain the predicted class label of the video to be queried. However, for the case of a small number of samples, the prototype representation simply obtained by mean calculation is not reliable. When the background of the support set is cluttered and the difference from the query video is large, it is difficult to obtain the correct classification prediction label.

[0005] In summary, the Siamese network has the problem of complex process. When the number of shots in the task is greater than 1, it is necessary to compare the similarity between the sample features of the query set and the features of each sample in the support set, so as to analyze the category of the target. The Matching Network is restricted by non-parametric algorithms, and the computational cost of each iteration will increase rapidly with the increase in the number of support set samples, resulting in slow computational speed. Due to the scarcity of samples, the prototype representation directly calculated by the prototype network through the feature mean is not very representative, and there is a problem that support set samples similar to the query set samples may be misclassified due to the interference of other samples within the same class. Summary of the Invention

[0006] The purpose of the present invention is to provide a few-shot behavior recognition method and system based on subspace classification that can make full use of the features of all samples in each category of the support set, and does not need to repeatedly calculate the distance from the features of each sample within the same class, improving the prediction accuracy, so as to solve at least one of the technical problems existing in the above background technology.

[0007] In order to achieve the above purpose, the present invention adopts the following technical solutions:

[0008] On the one hand, the present invention provides a few-shot behavior recognition method based on subspace classification, including:

[0009] Obtain the image to be recognized;

[0010] Use a pre-trained few-shot behavior recognition model to process the obtained image to be recognized, and obtain the behavior recognition result in the image; wherein, the pre-trained few-shot behavior recognition model is obtained by training with a training set, and the training set includes multiple images and labels annotating the behavior distribution characteristics in the images; the few-shot behavior recognition model includes a depth estimation network, a feature extraction network, a feature fusion network, and a recognition network. The depth estimation network is used to perform depth estimation on the RGB image to obtain a depth image; the feature extraction network is used to extract the features of the RGB image and the features of the depth image; the feature fusion network is used to fuse the extracted RGB image features and depth image features, and the recognition network is used to perform few-shot behavior recognition calculations based on a subspace classifier in combination with the fused features.

[0011] Preferably, when training the few-shot behavior recognition model, the dataset used is composed of multiple videos containing multiple behavior categories. Each video sample in each category in the dataset is divided into a group of image frames RGB frames, and the number of frames n_frames in each group of image frames is counted; for each image in the framed dataset, first adjust its size, and then randomly crop it; in the depth estimation network, use the monodepth2 module as the depth estimator to perform depth estimation on the processed dataset images to obtain depth images.

[0012] Preferably, in the feature extraction network, the feature extractor ImageNet pretrained ResNet-50 is used to extract RGB image features and depth image features; among them, first, the RGB sub-model for extracting RGB image features and the depth sub-model for extracting depth image features are trained. The feature extraction networks of the RGB sub-model and the depth sub-model use ResNet-50 as the backbone network, and the last fully connected layer in ResNet-50 is replaced with their respective fully connected layers as classifiers. The feature information extraction layer is a feature encoder generated by a convolutional neural network, which extracts the image features required by the information processing layer from the input layer image to obtain the RGB feature map and the depth feature map.

[0013] Preferably, in the feature fusion network, the obtained RGB feature vector and depth feature vector are fused through the DGAdaINFusion Module to obtain the fused feature vector.

[0014] Preferably, the data obtained by the DGAdaIN Fusion Module are the extracted RGB feature vector and depth feature vector, and the fused feature vector is obtained after processing; the batch input of the module is represented as x ∈ R B×D×L , where B is the batch size, D is the number of picture frames into which a single video sample is divided, and L is the feature dimension of each frame;

[0015]

[0016] The input of the DGAdaIN module f(I rgb , I d ) is an RGB input batch I rgb and a depth input batch I d , g s (·) and g b (·) are learnable fully connected layers, and the output of f(I rgb , I d ) is processed as the scale factor γ and the shift factor β to deform the RGB features for adaptively learning the depth feature map.

[0017] Preferably, in the recognition network, the obtained support set fusion features are sent into the subspace classifier, and the feature means of all samples in each class in the support set are calculated; in the subspace classifier, the support set sample features are subtracted by the feature means of their respective classes to obtain a new sample representation set for each class of the support set; in the subspace classifier, the singular value decomposition is performed on the new sample representation set to obtain the subspace projection matrix; the query set sample features are sent into the subspace classifier, and the distances from the query sample features to each class subspace are calculated; the softmax function is used to calculate the probabilities of the query samples belonging to each behavior class; the Grassmann manifold is adopted to maximize the distances between different subspaces.

[0018] In a second aspect, the present invention provides a few-shot behavior recognition system based on subspace classification, including:

[0019] An acquisition module, configured to acquire an image to be recognized;

[0020] A recognition module, configured to process the acquired image to be recognized by using a pre-trained few-shot behavior recognition model to obtain a behavior recognition result in the image; wherein, the pre-trained few-shot behavior recognition model is obtained by training with a training set, the training set includes multiple images and labels for annotating the behavior distribution features in the images; the few-shot behavior recognition model includes a depth estimation network, a feature extraction network, a feature fusion network, and a recognition network, the depth estimation network is configured to perform depth estimation on an RGB image to obtain a depth image; the feature extraction network is configured to extract the features of the RGB image and the depth image; the feature fusion network is configured to fuse the extracted RGB image features and depth image features, and the recognition network is configured to perform few-shot behavior recognition calculation based on a subspace classifier in combination with the fused features.

[0021] In a third aspect, the present invention provides a non-transitory computer-readable storage medium, which is used to store computer instructions, and when the computer instructions are executed by a processor, the few-shot behavior recognition method based on subspace classification as described above is implemented.

[0022] In a fourth aspect, the present invention provides a computer program product, including a computer program, which when running on one or more processors, is used to implement the few-shot behavior recognition method based on subspace classification as described above.

[0023] Fifth aspect, the present invention provides an electronic device, comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes instructions for implementing the small-sample behavior recognition method based on subspace classification as described above.

[0024] Advantages of the present invention: Make full use of all sample features in each class of the support set, and classify by constructing a subspace for each class of behaviors instead of directly using the feature mean; condense the sample features of each class into a subspace, and directly calculate the distance from the query sample feature to the subspace instead of calculating the distance from the query sample feature to each sample in each class in turn, reducing the amount of calculation.

[0025] Advantages of the additional aspects of the present invention will be more clearly given in the following description part, or understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0027] Figure 1 It is a network architecture diagram of a small-sample behavior recognition model based on subspace classification according to an embodiment of the present invention.

[0028] Figure 2 It is an architecture diagram of a subspace classifier in a small-sample behavior recognition model based on subspace classification according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.

[0030] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the art in the field to which the present invention belongs.

[0031] It should also be understood that terms such as those defined in a general dictionary should be understood as having a meaning consistent with their meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as here.

[0032] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the description of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or groups thereof.

[0033] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0034] For ease of understanding the present invention, the following further explains the present invention with specific embodiments in conjunction with the accompanying drawings, and the specific embodiments do not limit the embodiments of the present invention.

[0035] Those skilled in the art should understand that the drawings are only schematic diagrams of the embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.

[0036] Embodiment 1

[0037] In this Embodiment 1, first, a small-sample behavior recognition system based on subspace classification is provided, and the system includes:

[0038] An acquisition module, configured to acquire an image to be recognized;

[0039] The recognition module is used to process the acquired image to be recognized by using a pre-trained few-shot behavior recognition model, and obtain the behavior recognition result in the image. Among them, the pre-trained few-shot behavior recognition model is obtained by training with a training set, and the training set includes multiple images and labels annotating the behavior distribution characteristics in the images. The few-shot behavior recognition model includes a depth estimation network, a feature extraction network, a feature fusion network, and a recognition network. The depth estimation network is used to perform depth estimation on the RGB image to obtain a depth image. The feature extraction network is used to extract the features of the RGB image and the depth image. The feature fusion network is used to fuse the extracted RGB image features and depth image features. The recognition network is used to perform few-shot behavior recognition calculation based on a sub-control classifier and in combination with the fused features.

[0040] In Embodiment 1 of the present invention, the few-shot behavior recognition method based on subspace classification is realized by using the above system, including:

[0041] Use the acquisition module to acquire the image to be recognized;

[0042] Use the recognition module to process the acquired image to be recognized based on a pre-trained few-shot behavior recognition model, and obtain the behavior recognition result in the image. Among them, the pre-trained few-shot behavior recognition model is obtained by training with a training set, and the training set includes multiple images and labels annotating the behavior distribution characteristics in the images. The few-shot behavior recognition model includes a depth estimation network, a feature extraction network, a feature fusion network, and a recognition network. The depth estimation network is used to perform depth estimation on the RGB image to obtain a depth image. The feature extraction network is used to extract the features of the RGB image and the depth image. The feature fusion network is used to fuse the extracted RGB image features and depth image features. The recognition network is used to perform few-shot behavior recognition calculation based on a sub-control classifier and in combination with the fused features.

[0043] Among them, when training the few-shot behavior recognition model, the dataset used is composed of multiple videos containing multiple behavior categories. Each video sample of each category in the dataset is divided into a group of image frames (RGB frames), and the number of frames n_frames in each group of image frames is counted. For each image in the dataset after frame division, first adjust its size and then randomly crop it. In the depth estimation network, use the monodepth2 module as the depth estimator to perform depth estimation on the processed dataset images to obtain depth images.

[0044] Among them, in the feature extraction network, the feature extractor ImageNetpretrained ResNet-50 is used to extract RGB image features and depth image features. First, the RGB sub-model for extracting RGB image features and the depth sub-model for extracting depth image features are trained. The feature extraction networks of the RGB sub-model and the depth sub-model use ResNet-50 as the backbone network, and the last fully connected layer in ResNet-50 is replaced with their respective fully connected layers as the classifier. The feature information extraction layer is a feature encoder generated by a convolutional neural network, which extracts the image features required by the information processing layer from the input layer images to obtain RGB feature maps and depth feature maps.

[0045] In the feature fusion network, the obtained RGB feature vector and depth feature vector are fused through the DGAdaIN FusionModule to obtain a fused feature vector. The data obtained by the DGAdaIN Fusion Module are the extracted RGB feature vector and depth feature vector, and the fused feature vector is obtained after processing. The batch processing input to the module is expressed as x ∈ R B ×D×L , where B is the batch size, D is the number of picture frames into which a single video sample is divided, and L is the feature dimension of each frame.

[0046]

[0047] The input of the DGAdaIN module f(I rgb , I d ) is an RGB input batch I rgb and a depth input batch I d , and g s (·) and g b (·) are learnable fully connected layers. The output of f(I rgb , I d ) is processed as the scale factor γ and the shift factor β to deform the RGB features for adaptively learning the depth feature maps.

[0048] In the recognition network, the obtained support set fused features are sent into the subspace classifier to calculate the feature mean of all samples in each class in the support set. In the subspace classifier, the support set sample features are subtracted by the feature mean of their respective classes to obtain a new sample representation set for each class of samples in the support set. In the subspace classifier, the new sample representation set is subjected to singular value decomposition to obtain the subspace projection matrix. The query set sample features are sent into the subspace classifier to calculate the distance from the query sample features to each class subspace. The softmax function is used to calculate the probability that the query sample belongs to each behavior class. The Grassmann manifold is adopted to maximize the distance between different subspaces.

[0049] Example 2

[0050] In this Example 2, a small-sample behavior recognition method based on subspace classification is provided. A subspace of a feature space is calculated for each class, the feature vector of the query sample is projected into the subspace, distance measurement is performed in the subspace to predict the sample class, and the classifier is optimized by maximizing the distance between different subspaces on the Grassmann manifold. The small-sample behavior recognition method based on subspace classification can make full use of the features of all samples in each class of the support set and does not need to repeatedly calculate the distance with the features of each sample within the same class, achieving good prediction accuracy.

[0051] In this Example 2, for the implementation of the small-sample behavior recognition method based on subspace classification, the configuration work of relevant links is first carried out, including installing the development environment of python 3.6 (and above versions), the deep framework of PyTorch 0.4.1 (and above versions), and Ubuntu 18.04 (and above versions). Since the algorithm used in this method is a deep learning-based model algorithm, and the training process of the model is required to be carried out in a GPU environment, it is necessary to install the GPU version and configure the CUDA (version 9.1 and above) parallel computing framework corresponding to the version. It is recommended to create a virtual environment of python 3.6.6 version.

[0052] See Figure 1 , a small-sample behavior recognition method based on subspace classification provided in this example includes the following stages, namely: data preprocessing stage, feature extraction and fusion stage, subspace classification and optimization stage.

[0053] Data preprocessing stage (S1~S3)

[0054] S1: Divide the video dataset into a training set, a validation set, and a test set, make each video sample of each class in the dataset into picture frames, and count the number of frames.

[0055] In this embodiment, the dataset used is a behavior recognition dataset UCF101 composed of real-world action videos, collected from YouTube, containing 13,320 videos from 101 behavior categories. The main behaviors are divided into five categories: human-object interaction, simple limb movements, human-human interaction, playing musical instruments, and sports. The dataset is segmented in the way that the training set includes 70 categories, the validation set includes 10 categories, and the test set includes 21 categories. Each video sample of each category in the dataset is divided into a group of image frames RGB frames, and the number of frames n_frames of each group of image frames is counted. For each image in the framed dataset, its size is first adjusted to 256x256, and then a region of size 224x224 is randomly cropped. In the test phase, since the behavior information in the image is usually in the middle of the image rather than at the edge, the edge information is preferentially discarded, and a central cropping method is used to obtain a region of size 224x224.

[0056] S2: Use the monodepth2 module as a depth estimator to perform depth estimation on the processed dataset images to obtain a Depth (depth) image.

[0057] In this embodiment, depth estimation is performed on all the RGB frames obtained in S1 using the monodepth2 module. The overall input of the monodepth2 module is all the RGB clips in the dataset. For a group of RGB frames containing n_frames frames, three consecutive frames It-1, It, It+1 are input into the Depth Network, where the t-th frame is the frame for which the depth is to be predicted, and the (t-1)-th and (t+1)-th frames are the subsequent frame and the previous frame of the t-th frame respectively. The Depth Network is implemented using the Unet structure and consists of an encoder module and a decoder module. The Depth Network outputs the depth map Dt of the t-th frame. Each frame in the RGB frames is processed in turn to recover the corresponding depth to obtain Depth frames.

[0058] S3: Sample and generate tasks from the dataset for episode training, and perform data augmentation through temporal asynchronous sampling.

[0059] Similar to a batch in traditional training methods, each episode of few-shot learning is a single training. An episode consists of a support set and a query set, with the aim of enabling the model to learn on the support set and then classify and predict the query set samples. One episode corresponds to one task.

[0060] In this example, the training process uses a 5-way, 1-shot task for model training, and the testing process can be carried out in cases such as 1-shot and 5-shot. To construct a C-way, K-shot task, first randomly select C classes from all classes in the dataset, and then randomly select K samples from the selected C classes. Each sample contains an RGB frame and a Depth frame, which constitutes the support set, with a total of C * K samples. Randomly select one sample from the remaining samples of the selected C classes as the query set.

[0061] As Figure 1 shown, for each modality, first divide the RGB frames and Depth frames of the randomly sampled samples into numseg equally long segments, and then randomly extract numf consecutive image frames from each segment to form a new RGB clip and a Depth clip. The Depth clip is divided into a clip that matches the RGB clip and a clip that does not match. During the training process, the model is trained using the perfectly matching RGB and depth clips, as well as the imperfectly matching clips to achieve the effect of feature enhancement.

[0062] Feature extraction and fusion stage (S4 - S5)

[0063] S4: Use the feature extractor ImageNet pretrained ResNet-50 to extract the features of RGB image frames and Depth image frames.

[0064] In this embodiment, due to the large difference between the RGB modality and the Depth modality, a two-stage training method is adopted to obtain features. First, the RGB sub-model and the Depth sub-model are trained. The feature extraction networks of the RGB clip and the Depth clip use ResNet-50 as the backbone network, and the last fully connected layer in ResNet-50 is replaced with its own fully connected layer as the classifier. A task is generated in step S3, and the samples of the support set and the query set are used as the input of step S4. The RGB image and the Depth image are sent into the feature extraction network, and the feature information extraction layer is a feature encoder generated by a convolutional neural network, which extracts rich image features required by the information processing layer from the images of the input layer to obtain the RGB feature map and the Depth feature map. For the RGB sub-model, it is fine-tuned for 6 epochs in the training stage. The learning rate of ResNet-50 is set to lr1 = 0.00001, and the learning rate of the fully connected layer is set to lr2 = 0.001. For the depth sub-model, since it is difficult to extract the features of the depth frame, it is fine-tuned for 60 epochs, and the learning rates lr1 and lr2 are both set to 0.00001, and are reduced by 10% after 30 epochs. The pre-trained sub-models are used as feature extractors to extract the feature vectors of the RGB clip and the Depth clip respectively.

[0065] S5: The RGB feature vector and the Depth feature vector obtained in S4 are subjected to feature fusion through the DGAdaIN Fusion Module to obtain the fused feature vector.

[0066] Traditional CNN convolutions can only extract low-level features, rather than high-level abstract features. Image style transfer uses a correlation matrix to represent the style of an image, extracts common abstract features, and transfers them to other images to achieve style transformation of the image while keeping the content unchanged. In the single-model multi-style framework of Image style transfer, the algorithm controls the conversion of multiple styles by learning the affine transformation coefficients of instance normalization. These affine parameters are replaced by the variance and mean of the style map image itself to generate images of any style. Inspired by the successful application of instance normalization in Image style transfer, the DGAdaIN Fusion Module considers that the RGB modality contains most of the visual information and the Depth modality contains rich scene information. The RGB feature map is regarded as the content image in Image style transfer, and the Depth feature map is regarded as the style image to be transferred. The Depth modality is used to supplement the RGB modality with scene information, and multi-modal information is fused to achieve feature enhancement.

[0067] In this example, since some behaviors such as mountain biking and playing basketball are closely related to the scene, a method of adaptively fusing depth modal features such as RGB modal features is adopted to strengthen the feature background information. The data obtained by the DGAdaINFusion Module are the RGB feature vectors and depth feature vectors extracted from S2, and the fused feature vectors are obtained after processing. The batch input to the module is represented as x ∈ R B×D×L , where B is the batch size, D is the number of picture frames into which a single video sample is divided, L is the feature dimension of each frame, and γ and β are used to adaptively learn the depth feature map.

[0068]

[0069] In formula (1), the input of the DGAdaIN module f(I rgb , I d ) is an RGB input batch I rgb and a depth input batch I d , g s (·) and g b (·) are learnable fully connected layers, and the output of f(I rgb , I d ) is processed as the scale factor γ and the shift factor β to deform the RGB features.

[0070]

[0071]

[0072] Formulas 2 and 3 are the calculation formulas for the mean and variance of I rgb , and both μ b,d and σ b,d are calculated along the L dimension. DGAdaIN is fine-tuned in 6 epochs, each epoch contains 2000 scenario trainings, the task is set to 5-way, 1-shot, and the learning rate lr3 = 0.00002. When training the DGAdaIN module, the parameters of the two feature extraction sub-modules in S3 are no longer updated.

[0073] Subspace Classification and Optimization Phase (S6~S11)

[0074] S6: Send the fused features of the support set obtained in S5 into the subspace classifier to find the feature mean of all samples in each class in the support set.

[0075] The idea of subspace classification is to construct N subspaces according to the sample data points Each subspace Z i has a set of bases B i= [b1, b2,..., b n ∈ R D×n , where n < D, and B i T ×B i = I n . The basis of a class can be obtained by various methods such as PCV (Principal Component Analysis), SVD (Singular Value Decomposition), etc. In this embodiment, SVD is used to reduce the dimension of the data to obtain the basis of the subspace. The process is as follows Figure 2 .

[0076] In this embodiment, first, the feature mean of all samples in each category of the support set is calculated. The sample features are from the fused features output in S5.

[0077]

[0078] In Formula 4, K is the number of samples included in the c-th class behavior. By default, K is the shot value, x i is the i-th sample in the c-th class, F θ (x i ) is the fused feature vector obtained after steps S2, S3, S4, and S5 for x i , and μ c is the sample feature mean calculated in this step.

[0079] S7: In the subspace classifier, subtract the feature mean of the class to which the support set sample belongs from the support set sample features to obtain a new sample representation set for each class of the support set

[0080] For the support set sample set, subtract the feature mean of its belonging class obtained in S7 from the fused features obtained in S6 to obtain a new support set sample representation set.

[0081]

[0082] In Formula 5, F θ (x c , k) represents the k-th sample feature in the c-th class, and μ c is the sample feature mean of the c-th class.

[0083] S8: In the subspace classifier, perform singular value (SVD) decomposition on to obtain the subspace projection matrix P c .

[0084] SVD directly performs singular value decomposition on the original data to find the largest possible eigenvalues in the original data, and uses the eigenvectors corresponding to these eigenvalues as new features.

[0085] In this embodiment, taking as the original data, performing singular value decomposition on it to obtain matrix U. Selecting the first n dimensions in U to obtain the truncated matrix P c , that is, the subspace projection matrix.

[0086] Define the SVD of as Formula 6.

[0087]

[0088] In Formula 6, U is an m×m matrix, Σ is an m×n matrix with all elements being 0 except for the elements on the main diagonal, and each element on the main diagonal is called a singular value. V is an n×m matrix. Both U and V are unitary matrices, that is, satisfying U T U = I, V T V = I.

[0089] Next, solve the three matrices U, Σ, and V.

[0090] Taking the transpose of and performing matrix multiplication to obtain an n×n square matrix, and performing eigenvalue decomposition on the square matrix to obtain the eigenvalue and eigenvector formula 7.

[0091]

[0092] The eigenvalue decomposition yields the n eigenvalues of matrix and the corresponding n eigenvectors v. Spanning all the eigenvectors of into an n×n matrix V, that is, obtaining the matrix V in the SVD formula.

[0093] Similarly, taking and the transpose of performing matrix multiplication to obtain an m×m square matrix, and performing eigenvalue decomposition on it to obtain the eigenvalue and eigenvector satisfying Formula 8.

[0094]

[0095] The eigenvalue decomposition yields the m eigenvalues of matrix and the corresponding m eigenvectors u. Spanning all the eigenvectors of into an m×m matrix V, that is, obtaining the matrix U in the SVD formula.

[0096] The diagonal of the singular value matrix Σ is the singular value, and other positions are all 0. The process is shown in Formula 9.

[0097]

[0098] Each singular value is obtained through the above formula, and then the singular value matrix Σ is obtained.

[0099] In this embodiment, the first n dimensions in the matrix U are selected to obtain the truncated matrix P c , from The process of obtaining the subspace P c is the truncated singular value decomposition (TSVD). Different from the general SVD, TSVD can generate a decomposition matrix of a specified dimension to achieve dimensionality reduction.

[0100] S9: Send the query set sample features into the subspace classifier to calculate the distance from the query sample features to each class subspace.

[0101] In the subspace, a feasible classification metric basis is the projection distance from the sample point to the subspace. After obtaining P c , the projection matrix can be used to calculate the distance from any query sample feature to the projection point.

[0102] In this embodiment, the distance between the fused features of the query samples obtained in S5 and the class subspaces obtained in S8 is calculated through Formula 10.

[0103] d c (q) = -||(I - M c )(F θ (q) - u c )|| 2 (10)

[0104] In Formula 10, M c = P c P c T , μ c is the class feature mean, which can be understood as the offset between the point and the subspace.

[0105] S10: Use the softmax function to calculate the probability p c,q .

[0106] In this embodiment, the softmax function is used to define the probability assigned to the queries of class c.

[0107]

[0108] S11: Adopt the Grassmann manifold to maximize the distance between different subspaces obtained in S8.

[0109] To make the subspaces more discriminative and thus better complete the classification task, the Grassmann manifold is adopted to maximize the distance between different subspaces. Given the bases of two linear subspaces Pi and Pj, the projection metric is defined as:

[0110]

[0111] Maximize the projection metric distance, that is, minimize The loss function is obtained as:

[0112] where NM is the total number of samples in the query set, which is 1 here. p(c,q) is the class probability obtained in S10.

[0113] Example 3

[0114] Example 3 of the present invention provides an electronic device, including a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute a small-sample behavior recognition method based on subspace classification. The method includes the following process steps:

[0115] Obtain an image to be recognized;

[0116] Process the obtained image to be recognized by a pre-trained small-sample behavior recognition model to obtain a behavior recognition result in the image; wherein, the pre-trained small-sample behavior recognition model is trained by a training set, the training set includes multiple images and labels for annotating the behavior distribution characteristics in the images; the small-sample behavior recognition model includes a depth estimation network, a feature extraction network, a feature fusion network and a recognition network, the depth estimation network is used to perform depth estimation on the RGB image to obtain a depth image; the feature extraction network is used to extract the features of the RGB image and the depth image; the feature fusion network is used to fuse the extracted RGB image features and depth image features, and the recognition network is used to perform small-sample behavior recognition calculation based on a sub-control classifier in combination with the fused features.

[0117] Example 4

[0118] Example 4 of the present invention provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements a small-sample behavior recognition method based on subspace classification. The method includes the following process steps:

[0119] Obtain an image to be recognized;

[0120] The obtained image to be recognized is processed by using a pre-trained few-shot behavior recognition model to obtain the behavior recognition result in the image. Among them, the pre-trained few-shot behavior recognition model is trained by a training set, and the training set includes multiple images and labels annotating the behavior distribution characteristics in the images. The few-shot behavior recognition model includes a depth estimation network, a feature extraction network, a feature fusion network, and a recognition network. The depth estimation network is used to perform depth estimation on the RGB image to obtain a depth image. The feature extraction network is used to extract the features of the RGB image and the depth image. The feature fusion network is used to fuse the extracted RGB image features and depth image features. The recognition network is used to perform few-shot behavior recognition calculation based on a sub-control classifier in combination with the fused features.

[0121] Embodiment 5

[0122] Embodiment 5 of the present invention provides a computer device, including a memory and a processor. The processor and the memory communicate with each other. The memory stores program instructions executable by the processor. The processor calls the program instructions to execute a few-shot behavior recognition method based on subspace classification. The method includes the following steps:

[0123] Obtain an image to be recognized;

[0124] The obtained image to be recognized is processed by using a pre-trained few-shot behavior recognition model to obtain the behavior recognition result in the image. Among them, the pre-trained few-shot behavior recognition model is trained by a training set, and the training set includes multiple images and labels annotating the behavior distribution characteristics in the images. The few-shot behavior recognition model includes a depth estimation network, a feature extraction network, a feature fusion network, and a recognition network. The depth estimation network is used to perform depth estimation on the RGB image to obtain a depth image. The feature extraction network is used to extract the features of the RGB image and the depth image. The feature fusion network is used to fuse the extracted RGB image features and depth image features. The recognition network is used to perform few-shot behavior recognition calculation based on a sub-control classifier in combination with the fused features.

[0125] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0126] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more flows and / or one or more blocks in the flow. Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks.

[0127] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one or more flows and / or one or more blocks in the flow. Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks.

[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing device, and a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or one or more blocks in the flow. Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks.

[0129] Although the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions disclosed in the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts should be covered within the protection scope of the present invention.

Claims

1. A small-sample behavior recognition method based on subspace classification, characterized in that, Including: Obtain the image to be recognized; Process the obtained image to be recognized by using a pre-trained few-shot behavior recognition model to obtain the behavior recognition result in the image. Among them, the pre-trained few-shot behavior recognition model is obtained by training with a training set, and the training set includes multiple images and labels annotating the behavior distribution characteristics in the images. The few-shot behavior recognition model includes a depth estimation network, a feature extraction network, a feature fusion network, and a recognition network. The depth estimation network is used to perform depth estimation on the RGB image to obtain a depth image. The feature extraction network is used to extract the features of the RGB image and the depth image. The feature fusion network is used to fuse the extracted RGB image features and depth image features. The recognition network is used to perform few-shot behavior recognition calculations based on a subspace classifier and in combination with the fused features; Among them, in the feature fusion network, the obtained RGB feature vector and depth feature vector are fused through the DGAdaINFusionModule to obtain a fused feature vector; The data obtained by the DGAdaIN Fusion Module are the extracted RGB feature vectors and depth feature vectors, and the fused feature vectors are obtained after processing; the batch input to the module is represented as x ∈ R B×D×L , where B is the batch size, D is the number of frames of a set of picture frames into which a single video sample is divided, and L is the feature dimension of each frame; The input of the DGAdaIN module f(I rgb ,I d ) is an RGB input batch I rgb and a depth input batch I d , g s (·) and g b (·) are learnable fully connected layers. The output of f(I rgb ,I d ) is processed as the scale factor γ and the shift factor β to deform the RGB features for adaptively learning the depth feature map; In the recognition network, the obtained support set fused features are sent into the subspace classifier to calculate the feature mean of all samples in each category in the support set. In the subspace classifier, subtract the feature mean of the category to which the support set sample belongs from the support set sample feature to obtain a new sample representation set for each category of support set samples. In the subspace classifier, perform singular value decomposition on the new sample representation set to obtain a subspace projection matrix. Send the query set sample features into the subspace classifier to calculate the distance from the query sample features to each category subspace. Use the softmax function to calculate the probability that the query sample belongs to each behavior category. Adopt the Grassmann manifold to maximize the distance between different subspaces.

2. The small-sample behavior recognition method based on subspace classification according to claim 1, wherein, When training the few-shot behavior recognition model, the dataset used is composed of multiple videos containing multiple behavior categories. Each video sample in each category in the dataset is divided into a group of image frames RGB frames, and the number of frames n_frames in each group of image frames is counted. For each image in the dataset after frame division, first adjust its size and then randomly crop it. In the depth estimation network, use the monodepth2 module as the depth estimator to perform depth estimation on the processed dataset images to obtain depth images.

3. The small-sample behavior recognition method based on subspace classification according to claim 2, wherein In the feature extraction network, use the feature extractor ImageNetpretrained ResNet-50 to extract the features of the RGB image and the depth image. Among them, first train the RGB sub-model for extracting the features of the RGB image and the depth sub-model for extracting the features of the depth image. The feature extraction networks of the RGB sub-model and the depth sub-model use ResNet-50 as the backbone network, and replace the last fully connected layer in ResNet-50 with their respective fully connected layers as classifiers. The feature information extraction layer is a feature encoder generated by a convolutional neural network, which extracts the image features required by the information processing layer from the input layer image to obtain the RGB feature map and the depth feature map.

4. A small-sample behavior recognition system based on subspace classification, characterized in that, Comprising: An acquisition module, configured to acquire an image to be recognized; A recognition module, configured to process the acquired image to be recognized by using a pre-trained few-shot behavior recognition model, and obtain a behavior recognition result in the image; wherein, the pre-trained few-shot behavior recognition model is obtained by training with a training set, the training set includes multiple images and labels annotating the behavior distribution characteristics in the images; the few-shot behavior recognition model includes a depth estimation network, a feature extraction network, a feature fusion network, and a recognition network, the depth estimation network is configured to perform depth estimation on an RGB image to obtain a depth image; the feature extraction network is configured to extract features of the RGB image and the depth image; the feature fusion network is configured to fuse the extracted RGB image features and depth image features, and the recognition network is configured to perform few-shot behavior recognition calculation based on a subspace classifier and in combination with the fused features; Wherein, in the feature fusion network, the obtained RGB feature vector and depth feature vector are subjected to feature fusion through a DGAdaIN FusionModule to obtain a fused feature vector; The data obtained by the DGAdaIN Fusion Module are the extracted RGB feature vectors and depth feature vectors, and the fused feature vectors are obtained after processing; the batch input to the module is represented as x ∈ R B×D×L , where B is the batch size, D is the number of picture frames into which a single video sample is divided, and L is the feature dimension of each frame; The input of the DGAdaIN module f(I rgb ,I d ) is an RGB input batch I rgb and a depth input batch I d , g s (·) and g b (·) are learnable fully connected layers. The output of f(I rgb ,I d ) is processed as the scale factor γ and the shift factor β to deform the RGB features for adaptively learning the depth feature map; In the recognition network, the obtained support set fused features are sent into a subspace classifier, and the feature mean of all samples in each category in the support set is calculated; in the subspace classifier, the support set sample features are subtracted from the feature mean of the class to which they belong to obtain a new sample representation set for each category of support set samples; in the subspace classifier, singular value decomposition is performed on the new sample representation set to obtain a subspace projection matrix; the query set sample features are sent into the subspace classifier, and the distances from the query sample features to the subspaces of each category are calculated; the probability that the query sample belongs to each behavior category is calculated by using a softmax function; the Grassmann manifold is adopted to maximize the distance between different subspaces.

5. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is configured to store computer instructions, and when the computer instructions are executed by a processor, the method for few-shot behavior recognition based on subspace classification according to any one of claims 1-3 is implemented.

6. A computer program product, characterized in that, Including a computer program, which when running on one or more processors, is configured to implement the method for few-shot behavior recognition based on subspace classification according to any one of claims 1-3.

7. An electronic device, characterized in that, Comprising: A processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes instructions for implementing the method for few-shot behavior recognition based on subspace classification according to any one of claims 1-3.

Citation Information

Patent Citations

  • Image recognition method for pre-judging existence of illegal behaviors based on character micro-expressions and actions

    CN111062243A

  • Human body behavior classification method based on multi-task learning model

    CN111488840A