An intelligent recognition and classification method for small-sample video behaviors

Through the combination of backbone network and motion activation module, motion information between video frames is extracted, and the attention mechanism is used to identify and classify video behaviors in small samples, solving the problem of video behavior recognition and classification under small samples conditions, and achieving efficient video behavior recognition and classification.

CN115424174BActive Publication Date: 2025-07-18BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211039693.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-07-18
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

Under small sample conditions, existing deep learning models are difficult to effectively identify and classify video behaviors, especially when samples are scarce or the cost of labeling samples is too high, traditional methods cannot effectively identify and classify video behaviors.

Method used

The backbone network is used to combine the motion activation module and the interval motion activation module. By extracting the motion information between the video frame slices, a feature representation is constructed, and the attention mechanism is used to identify and classify small-sample video behaviors, including the combination of the ResNet network, the motion activation module and the interval motion activation module, and the global average pooling and full connection layer are used for feature alignment and classification.

Benefits of technology

In complex scenarios, the recognition and classification accuracy of character actions is improved, the receptive field of motion information is enhanced, the efficiency and accuracy of classification processing is improved, and key technical support is provided for the review of video content and personalized recommendation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115424174B_ABST
    Figure CN115424174B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of pattern recognition, and particularly relates to an intelligent recognition and classification method for small-sample video behaviors. The method includes: extracting the input sequence of each small-sample video, dividing the small-sample data set into a training set, a validation set, and a test set and generating the corresponding task set; constructing a backbone network, extracting the feature representations of all videos of a single task in the task set, and performing global average pooling one by one to obtain two-dimensional tensor features; aligning the two-dimensional tensor features of the support set in the single task with the two-dimensional tensor features of the query set, and obtaining a set of two-dimensional tensor alignment features corresponding to each query set video; using the two-dimensional tensor alignment features as input to obtain the average similarity score and the predicted class label of all query set videos. The method can capture the motion information between video frame slices and can accurately complete the behavior recognition and classification of small-sample videos in complex human action scenarios with extremely scarce samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of pattern recognition and relates to an intelligent recognition and classification method for small-sample video behaviors. Technical Background

[0002] With the rapid development of computer technology and mobile Internet technology, watching and sharing videos has become a part of people's daily lives, and video data has also become an important information carrier. It not only brings great convenience and pleasure to people, but also enables people to have a new lifestyle. Today, with the continuous increase in the amount of video data, how to better understand and classify events and actions in video data, so as to provide key technical support for downstream tasks such as video content review and personalized recommendation, has become a challenging and problem to be solved in the fields of pattern recognition and computer vision.

[0003] In the current stage of artificial intelligence, when using deep learning methods for video behavior recognition, a large number of samples with classification labels often need to be provided in advance to enable the neural network to complete parameter training well. When it learns a completely new class, the existing model usually needs to be changed and a large number of sample trainings need to be carried out again. In contrast, when thinking about the learning process of humans for a thing, only a small number of samples are needed to complete the learning process of the video. In most scenarios in life, it is very likely that due to reasons such as the scarcity of sample videos or the high cost of labeling samples, a large number of training samples with classification labels cannot be provided to the neural network. In this case, traditional deep learning models will become helpless. Therefore, in view of the problem of sample scarcity in many scenarios, it is of great significance to draw on the human cognitive process and study how to perform effective video behavior recognition under the condition of only a small number of classification labels. Summary of the Invention

[0004] The purpose of the present invention is to propose an intelligent recognition and classification method for small-sample video behaviors in view of the technical defects of difficult action classification and low recognition rate caused by sample scarcity under small-sample conditions.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] The intelligent recognition and classification method takes a single task in the task set as input relying on a backbone network, and extracts the feature representations of all small-sample videos in the single task;

[0007] The backbone network includes a ResNet network, a motion activation module, and an interval motion activation module;

[0008] The ResNet network is respectively connected to the motion activation module and the interval motion activation module. Specifically: the outputs of the first four layers of the ResNet network are connected to the motion activation module, the output of the motion activation module is input to the fifth layer of the ResNet network, and the output of the fifth layer is given to the interval motion activation module;

[0009] Among them, ResNet18, ResNet34, ResNet50, ResNet101, and ResNet152 can all be used as the ResNet network in the backbone network;

[0010] The ResNet network includes five layers. Among them, the first layer is composed of a convolutional layer and a pooling layer in series; the second layer, the third layer, the fourth layer, and the fifth layer are respectively composed of a, b, c, and d residual blocks in series;

[0011] Among them, when ResNet18 is used as the ResNet network, a, b, c, and d are respectively equal to 2, 2, 2, and 2; when ResNet34 / ResNet50 is used as the ResNet network, a, b, c, and d are respectively equal to 3, 4, 6, and 3; when ResNet101 is used as the ResNet network, a, b, c, and d are respectively equal to 3, 4, 23, and 3; when ResNet152 is used as the ResNet network, a, b, c, and d are respectively equal to 3, 8, 36, and 3;

[0012] Specifically, when the ResNet network is ResNet18 / ResNet34, the residual block includes two 3×3 convolutional layers in series, and the output feature of the residual block is obtained by using the shortcut connection method; when the ResNet network is ResNet50 / ResNet101 / ResNet152, the residual block includes three 1×1 convolutional layers, 3×3 convolutional layers, and 1×1 convolutional layers in series, and the output feature of the residual block is obtained by using the shortcut connection method;

[0013] The shortcut connection method means that the input feature of the first convolutional layer of the residual block is added element by element to the output feature of the last convolutional layer of the residual block, and the obtained added feature is used as the output feature of the residual block;

[0014] The motion activation module and the interval motion activation module work together to capture the motion information between video frame slices, including the meta-training stage and the meta-testing stage; in the meta-training stage, the training set is used for model training and the validation set is used for model verification; in the meta-testing stage, the test set is used to complete the test and the test results are output;

[0015] The intelligent recognition and classification method includes the following steps:

[0016] S1: Divide the action datasets of all small-sample videos into a training set, a validation set, and a test set;

[0017] Both the training set, the validation set, and the test set contain small-sample videos, and it is necessary to ensure that the categories to which the small-sample videos belong have no intersection;

[0018] The number of frames in each small-sample video is P, and P is greater than or equal to 5;

[0019] S2: Extract the input sequences corresponding to each small-sample video from all frames of each small-sample video included in the training set, the validation set, and the test set;

[0020] For the extraction of the input sequence corresponding to each small-sample video, the specific frame extraction method for a single small-sample video is as follows: evenly divide all P frames of a single small-sample video into T segments, and randomly extract one frame from each segment to form the input sequence of the single small-sample video, so that the input sequence corresponding to each small-sample video contains T frames;

[0021] The T is less than or equal to P;

[0022] S3: Use the input sequences corresponding to the small-sample videos to construct the task sets belonging to the training set, the validation set, and the test set respectively;

[0023] Among them, the task sets belonging to the training set, the validation set, and the test set all include multiple single tasks;

[0024] The task set is the model input of the small-sample network;

[0025] The single task is created according to the N-way K-shot small-sample problem setting and includes a support set and a query set;

[0026] Both the support set and the query set contain N identical sample video categories;

[0027] Among them, each category in the support set contains K small-sample videos, and each category in the query set contains M small-sample videos. It is necessary to ensure that the sample videos contained in the support set and the query set have no intersection;

[0028] The K small-sample videos are specifically the input sequences of K small-sample videos, and the M small-sample videos are specifically the input sequences of M small-sample videos;

[0029] S4: Construct a backbone network, input the single tasks in the task set into the backbone network, and extract the feature representations of all small-sample videos in the single tasks;

[0030] The backbone network includes a ResNet network, a motion activation module, and an interval motion activation module;

[0031] Taking a small-sample video in a single task as an example, extract the feature representation of this small-sample video in the single task, which specifically includes the following sub-steps:

[0032] S41: Input the input sequence of a small-sample video in the single task into the first four layers of the ResNet network to obtain the first output feature;

[0033] The first output feature has four dimensions: time, number of channels, length, and width;

[0034] S42: Input the first output feature into the motion activation module to obtain the motion activation output feature, specifically:

[0035] S421: Divide the first output feature along the time dimension to obtain the motion pre-activation features at each moment;

[0036] Each moment corresponds to each frame in the T frames of this small-sample video;

[0037] S422: Compress the motion pre-activation features at each moment using the first convolutional layer of the motion activation module to obtain the motion activation compressed features at each moment;

[0038] The compression using the first convolutional layer of the motion activation module has a compression coefficient of r1;

[0039] The number of channels of the motion activation compressed features at each moment is reduced to 1 / r1 of the number of channels of the first output feature;

[0040] S423: Traverse the motion activation compressed features at each moment in the order of the small-sample video input sequence, and use the second convolutional layer of the motion activation module to align the motion activation compressed features of adjacent frames with the motion activation compressed features of the current frame to obtain the motion activation alignment features of adjacent frames relative to the current frame;

[0041] Among them, when traversing the motion activation compressed feature of the last moment, there is no adjacent frame, so the alignment of the motion activation compressed feature of the adjacent frame with the motion activation compressed feature of the current frame is not performed at the last moment;

[0042] The current frame refers to the video frame corresponding to the current traversal moment;

[0043] The adjacent frame refers to the video frame corresponding to the next moment of the current frame;

[0044] The obtaining of the motion activation alignment features of adjacent frames relative to the current frame specifically means: input the motion activation compressed feature of the adjacent frame into the second convolutional layer of the motion activation module, and output to obtain the motion activation alignment features of adjacent frames relative to the current frame;

[0045] S424: Traverse the motion activation compressed features at each moment in the order of the small sample video input sequence. For the current traversed moment, calculate the frame difference between the motion activation alignment feature of the adjacent frame relative to the current frame and the motion activation compressed feature of the current frame to obtain the motion activation difference frame feature of the current frame;

[0046] Among them, when traversing the motion activation compressed feature at the last moment, there is no adjacent frame, so the frame difference calculation result at the last moment is replaced by a zero tensor;

[0047] The frame difference calculation is specifically the element-wise subtraction between two feature tensors;

[0048] S425: Input the motion activation difference frame features at each moment into the global average pooling layer of the motion activation module, the third convolutional layer of the motion activation module, and the activation layer of the motion activation module in sequence to obtain the motion activation channel attention weights;

[0049] The motion activation channel attention weights are a one-dimensional tensor, and the number of elements contained is equal to the number of channels of the first output feature;

[0050] S426: Perform a residual operation on the motion activation channel attention weights and the first output feature along the channel dimension to obtain the motion activation output feature;

[0051] The specific residual operation is: along the channel dimension, multiply the first output feature by the motion activation channel attention weights, and then add the obtained result to the first output feature element-wise to obtain the motion activation output feature;

[0052] The feature dimension of the motion activation output feature is consistent with the feature dimension of the first output feature;

[0053] S43: Input the motion activation output feature into the fifth layer of the ResNet network and obtain the second output feature;

[0054] The second output feature has four dimensions: time, number of channels, length, and width;

[0055] S44: Input the second output feature into the interval motion activation module to obtain the interval motion activation output feature, specifically:

[0056] S441: Divide the second output feature along the time dimension to obtain the interval motion pre-activation features at each moment;

[0057] Each moment corresponds to each frame in the T frames of the small sample video;

[0058] S442: Compress the interval motion pre-activation features at each moment using the first convolutional layer of the interval motion activation module to obtain the interval motion activation compressed features at each moment;

[0059] The first convolutional layer using the interval motion activation module is used for compression, and the compression coefficient is r2;

[0060] The number of channels of the interval motion activation compressed features at each moment is reduced to 1 / r2 of the number of channels of the second output features;

[0061] S443: Traverse the interval motion activation compressed features at each moment in the order of the small sample video input sequence, and use the second convolutional layer of the interval motion activation module to align the interval motion activation compressed features of the interval frames with the interval motion activation compressed features of the current frame, so as to obtain the interval motion activation alignment features of the interval frames relative to the current frame;

[0062] Among them, when traversing the interval motion activation compressed features of the last two moments, there are no interval frames, so the alignment of the interval motion activation compressed features of the interval frames with the interval motion activation compressed features of the current frame is not performed at the last two moments;

[0063] The current frame refers to the video frame corresponding to the current traversed moment;

[0064] The interval frame refers to the video frame corresponding to the next next moment of the current frame;

[0065] The obtained interval motion activation alignment features of the interval frames relative to the current frame specifically refer to: inputting the interval motion activation compressed features of the interval frames into the second convolutional layer of the interval motion activation module, and outputting the interval motion activation alignment features of the interval frames relative to the current frame;

[0066] S444: Traverse the interval motion activation compressed features at each moment in the order of the small sample video input sequence, and perform frame difference calculation on the interval motion activation alignment features of the interval frames relative to the current frame and the interval motion activation compressed features of the current frame in the currently traversed moment, so as to obtain the interval motion activation difference frame features of the current frame;

[0067] Among them, when traversing the interval motion activation compressed features of the last two moments, there are no interval frames, so the frame difference calculation results at the last two moments are replaced by zero tensors;

[0068] The frame difference calculation is specifically to perform element-wise subtraction on two feature tensors;

[0069] S445: Input the interval motion activation difference frame features at each moment into the global average pooling layer of the interval motion activation module, the third convolutional layer of the interval motion activation module, and the activation layer of the interval motion activation module in sequence, so as to obtain the interval motion activation channel attention weights;

[0070] The interval motion activation channel attention weights are a one-dimensional tensor, and the number of elements contained is equal to the number of channels of the second output features;

[0071] S446: Perform a residual operation on the spatio-temporal motion activation channel attention weights and the second output feature along the channel dimension to obtain the spatio-temporal motion activation output feature;

[0072] Specifically, the residual operation is as follows: along the channel dimension, multiply the second output feature by the spatio-temporal motion activation channel attention weights, and then add the result of the multiplication to the second output feature element-wise to obtain the spatio-temporal motion activation output feature;

[0073] The feature dimension of the spatio-temporal motion activation output feature is the same as that of the second output feature, which is called the feature representation of the few-shot video;

[0074] S45: For all few-shot videos in a single task, repeat S41 to S44 to obtain the feature representations of all few-shot videos in the single task;

[0075] S5: Along the channel dimension, perform global average pooling on the feature representation of each few-shot video in the single task, and each few-shot video corresponds to a two-dimensional tensor feature;

[0076] Among them, a single task obtains N×K + N×M two-dimensional tensor features in total;

[0077] For the two-dimensional tensor feature, the first dimension corresponds to the number of input sequence frames T of the few-shot video, and the second dimension corresponds to the number of channels of the second output feature;

[0078] S6: Align the two-dimensional tensor features corresponding to all few-shot videos in the support set of the single task with the two-dimensional tensor features corresponding to each few-shot video in the query set of the single task to obtain N×M groups of two-dimensional tensor alignment features;

[0079] Specifically, aligning the two-dimensional tensor features corresponding to all few-shot videos in the support set of the single task with the two-dimensional tensor features corresponding to each few-shot video in the query set of the single task means that if there are N×M few-shot videos in the query set of the single task, the two-dimensional tensor features corresponding to all few-shot videos in the support set will be aligned with the two-dimensional tensor features corresponding to each few-shot video in the query set. A single task performs N×M alignments in total. Each query set few-shot video will generate a group of two-dimensional tensor alignment features, and a single task obtains N×M groups of two-dimensional tensor alignment features in total;

[0080] Taking a few-shot video in the query set and combining all few-shot videos in the support set as an example, perform one alignment to generate a group of two-dimensional tensor alignment features. The alignments of the remaining few-shot videos in the query set of the single task are the same as this. Specifically:

[0081] S61: Use all the small-sample videos of the support set and the N×K + 1 two-dimensional tensor features corresponding to the small-sample videos of the query set as inputs. Each two-dimensional tensor feature is successively subjected to multi-group feature recombination and multi-group feature dimensionality reduction to obtain N×K + 1 groups of multi-group recombined two-dimensional tensor features;

[0082] The N×K + 1 two-dimensional tensor features include: N×K two-dimensional tensor features corresponding to all the small-sample videos of the support set, and one two-dimensional tensor feature corresponding to the small-sample video of the query set;

[0083] The N×K + 1 groups of multi-group recombined two-dimensional tensor features include: N×K groups of multi-group recombined two-dimensional tensor features corresponding to all the small-sample videos of the support set, and one group of multi-group recombined two-dimensional tensor features corresponding to the small-sample video of the query set;

[0084] Taking the multi-group feature recombination and multi-group feature dimensionality reduction of one two-dimensional tensor feature among the N×K + 1 two-dimensional tensor features as an example, the multi-group feature recombination and multi-group feature dimensionality reduction of the remaining two-dimensional tensor features among the N×K + 1 two-dimensional tensor features are the same, and include the following sub-steps:

[0085] S611: Use this two-dimensional tensor feature as an input, and split out multiple one-dimensional tensors row by row according to the first time dimension;

[0086] There are a total of T such one-dimensional tensors;

[0087] S612: Perform multi-group feature recombination on all the one-dimensional tensors corresponding to this two-dimensional tensor feature to obtain a group of multi-group three-dimensional tensor features;

[0088] The multi-group specifically refers to all ω i -tuples, When k > 0, ω0 ≠ ω1 ≠.... ≠ ω k ;

[0089] Among them, T represents the number of sampled frames, represents a natural number;

[0090] The ω i -tuple is called a unit tuple, and a group of multi-groups contains a total of k unit tuples;

[0091] The group of multi-group three-dimensional tensor features contains a total of k unit group three-dimensional tensor features;

[0092] The multi-group feature recombination includes a total of k unit group feature recombinations. Specifically: perform k unit group feature recombinations on all the one-dimensional tensors corresponding to this two-dimensional tensor feature, and each unit group feature recombination will respectively obtain a unit group three-dimensional tensor feature;

[0093] The recombination of the unit group features specifically involves: rearranging and sorting all the one-dimensional tensors corresponding to the two-dimensional tensor features according to the set of combination sequences corresponding to the unit group to obtain a three-dimensional tensor feature of the unit group;

[0094] The set of combination sequences corresponding to the unit group specifically refers to: Let the unit group be ω i tuple. For ω i the set of corresponding combination sequences is

[0095] wherein, the set of combination sequences corresponding to ω i tuple contains a total of elements;

[0096] The three-dimensional tensor feature of the unit group specifically refers to: Let the unit group be ω i tuple. The three-dimensional tensor feature of the unit group is the three-dimensional tensor feature of ω i tuple. The first dimension corresponds to the i possible combinations of ω tuple, the second dimension corresponds to the value of ω i tuple, and the third dimension corresponds to the number of channels of the second output feature;

[0097] S613: Perform dimensionality reduction on the multi-tuple three-dimensional tensor feature to obtain a set of recombined two-dimensional tensor features of the multi-tuple;

[0098] The set of recombined two-dimensional tensor features of the multi-tuple contains a total of k recombined two-dimensional tensor features of the unit group;

[0099] The dimensionality reduction of the multi-tuple feature includes a total of k times of dimensionality reduction of the unit group feature. Specifically, it involves performing dimensionality reduction on the k three-dimensional tensor features of the unit group contained in the multi-tuple three-dimensional tensor feature respectively. Each time of dimensionality reduction of the unit group feature will correspondingly obtain a recombined two-dimensional tensor feature of the unit group;

[0100] The dimensionality reduction of the unit group feature specifically involves: Let the unit group be ω i tuple. The three-dimensional tensor feature of the unit group is the three-dimensional tensor feature of ω i tuple. Merge the second and third dimensions of the three-dimensional tensor feature of the unit group to obtain a recombined two-dimensional tensor feature of the unit group;

[0101] The recombined two-dimensional tensor feature of the unit group specifically refers to: Let the unit group be ω i tuple. The recombined two-dimensional tensor feature of the unit group is the recombined two-dimensional tensor feature of ω i tuple. The first dimension corresponds to the i possible combinations of ω tuple, and the second dimension corresponds to ω iThe product of the value and the number of second output feature channels;

[0102] S614: For all the few-shot videos in the support set and the N×K + 1 two-dimensional tensor features corresponding to the few-shot videos in the query set, repeat S611 to S613 to obtain N×K + 1 groups of multi-tuple reconstructed two-dimensional tensor features;

[0103] Among them, a group of multi-tuple reconstructed two-dimensional tensor features contains k unit-group reconstructed two-dimensional tensor features in total; the N×K + 1 groups of multi-tuple reconstructed two-dimensional tensor features contain k×(N×K + 1) unit-group reconstructed two-dimensional tensor features in total;

[0104] S62: Use all the few-shot videos in the support set and the k×(N×K + 1) unit-group reconstructed two-dimensional tensor features corresponding to the few-shot videos in the query set as inputs, and divide them by taking the unit-group to which each unit-group reconstructed two-dimensional tensor feature belongs, to obtain k groups of same-tuple reconstructed two-dimensional tensor features;

[0105] Each group of the k groups of same-tuple reconstructed two-dimensional tensor features corresponds to the same unit-group respectively;

[0106] Each group of the same-tuple reconstructed two-dimensional tensor features contains N×K + 1 unit-group reconstructed two-dimensional tensor features belonging to the same ω i tuple;

[0107] The N×K + 1 unit-group reconstructed two-dimensional tensor features belonging to the same ω i tuple contain in total: N×K unit-group reconstructed two-dimensional tensor features belonging to the ω i tuple corresponding to all the few-shot videos in the support set, and one unit-group reconstructed two-dimensional tensor feature belonging to the ω i tuple corresponding to the few-shot video in the query set;

[0108] S63: Perform k times of same-tuple feature alignment on the k groups of same-tuple reconstructed two-dimensional tensor features to obtain a group of two-dimensional tensor alignment features;

[0109] Among them, a group of two-dimensional tensor alignment features contains k groups of same-tuple two-dimensional tensor alignment features in total, that is, one same-tuple reconstructed two-dimensional tensor feature will perform one same-tuple feature alignment to obtain a group of same-tuple two-dimensional tensor alignment features;

[0110] The group of same-tuple two-dimensional tensor alignment features contains N×K + 1 unit-group two-dimensional tensor alignment features in total;

[0111] The two-dimensional tensor alignment features of the N×K+1 unit groups include: the two-dimensional tensor alignment features of the N×K unit groups corresponding to the small-sample videos of all support sets, and the two-dimensional tensor alignment feature of one unit group corresponding to the small-sample video of the query set;

[0112] The first-order same-tuple feature alignment specifically means: Let the unit group to which the same-tuple belongs be ω i tuple, input the i reorganized two-dimensional tensor feature of the same-tuple corresponding to the ω i tuple, perform the first-order same-tuple feature alignment, and obtain a group of two-dimensional tensor alignment features belonging to the ω

[0113] S631: Among the group of reorganized two-dimensional tensor features of the same-tuple corresponding to the ω i tuple, input the N×K unit-group reorganized two-dimensional tensor features corresponding to the small-sample videos of all support sets and belonging to the ω i tuple into the fully connected layer, and each small-sample video in the support set will correspondingly obtain a key value

[0114] Among them, represents the number of neurons in the fully connected layer, and the set of combined sequences corresponding to the ω i tuple contains a total of elements;

[0115] S632: Among the reorganized two-dimensional tensor features of the same-tuple corresponding to the ω i tuple, input the one unit-group reorganized two-dimensional tensor feature corresponding to the small-sample video of the query set and belonging to the ω i tuple into the fully connected layer, and a query value will be correspondingly obtained

[0116] S633: Among the group of reorganized two-dimensional tensor features of the same-tuple corresponding to the ω i tuple, input the N×K unit-group reorganized two-dimensional tensor features corresponding to the small-sample videos of all support sets and belonging to the ω i tuple into the fully connected layer, and each small-sample video in the support set will correspondingly obtain an s_value value

[0117] Among them, represents the number of neurons in the fully connected layer;

[0118] S634: For ω iAmong the two-dimensional tensor features of the same tuple recombination corresponding to the tuple, the small-sample videos in the query set belong to ω i One unit tuple recombination two-dimensional tensor feature of the tuple is input into the fully connected layer, and a q_value will be obtained correspondingly

[0119] The q_value is the unit tuple two-dimensional tensor alignment feature of the tuple corresponding to the small-sample video in the query set that belongs to ω i ;

[0120] S635: Layer normalization is performed on the key values corresponding to the small-sample videos of each support set and the query values corresponding to the small-sample videos of the query set respectively to obtain the normalized key values corresponding to the small-sample videos of each support set and the normalized query values corresponding to the small-sample videos of the query set;

[0121] S636: The normalized key values corresponding to the small-sample videos of each support set and the normalized query values corresponding to the small-sample videos of the query set are respectively subjected to a dot product operation of two-dimensional tensors, and the results of each dot product operation are respectively mapped by the Softmax function to obtain the attention weights corresponding to the small-sample videos of each support set;

[0122] Among them, the dot product operation of two-dimensional vectors and the Softmax function mapping are both performed N×K times respectively;

[0123] There are N×K attention weights in total, and the dimension of each attention weight is two-dimensional, and both the first dimension and the second dimension are

[0124] The N×K corresponds to the number of small-sample videos of all support sets;

[0125] S637: The attention weights corresponding to the small-sample videos of each support set are subjected to a dot product operation of two-dimensional tensors with the corresponding s_value , and each small-sample video of each support set will correspondingly obtain the unit tuple two-dimensional tensor alignment feature belonging to ω i ;

[0126] The unit tuple two-dimensional tensor alignment feature There are N×K in total;

[0127] S638: Integrate the unit tuple two-dimensional tensor alignment features belonging to ω i to obtain a group of same-tuple two-dimensional tensor alignment features belonging to ω i ;

[0128] The belonging to ω iA set of two-dimensional tensor alignment features of the same tuples of tuples contains N×K + 1 belonging to ω i Two-dimensional tensor alignment features of the unit tuples of the tuples;

[0129] The N×K + 1 belonging to ω i Two-dimensional tensor alignment features of the unit tuples of the tuples, specifically referring to: N×K and one

[0130] S639: Use the two-dimensional tensor features of k groups of recombined same tuples as input, repeat S631 to S638, to obtain a set of two-dimensional tensor alignment features;

[0131] Among them, a set of two-dimensional tensor alignment features contains k groups of two-dimensional tensor alignment features of the same tuples;

[0132] S64: Align the two-dimensional tensor features corresponding to all small-sample videos in the support set of the single task with the two-dimensional tensor features corresponding to each small-sample video in the query set of the single task, repeat S61 to S63, to obtain N×M groups of two-dimensional tensor alignment features;

[0133] S7: Perform N×M small-sample classification calculations on N×M groups of two-dimensional tensor alignment features to obtain N×M average similarity scores and N×M predicted class labels for all small-sample videos in the query set of the single task;

[0134] The average similarity score is specifically a one-dimensional tensor, which includes a total of N elements, corresponding to the N sample classes in the single task;

[0135] The predicted class labels are N×M in total, corresponding to the N×M small-sample videos in the query set of the single task;

[0136] The N×M small-sample classification calculations on N×M groups of two-dimensional tensor alignment features are specifically: taking a small-sample video in a query set combined with all small-sample videos in the support set to perform a small-sample classification calculation, generating an average similarity score and a predicted class label as an example, including the following sub-steps:

[0137] S71: Use the two-dimensional tensor alignment features corresponding to this query set as input, divide them by unit tuples to obtain k groups of two-dimensional tensor alignment features of the same tuples;

[0138] S72: Calculate k similarity scores using the k groups of two-dimensional tensor alignment features of the same tuples corresponding to this query set;

[0139] The calculation of the k similarity scores, taking a set of two-dimensional tensor alignment features of the same tuples belonging to ω i to calculate a similarity score as an example, includes the following sub-steps:

[0140] Among them, a group belongs to ω i The two-dimensional tensor alignment features of the same tuple of the tuple belonging to ω in the small-sample videos of all support sets include: the N×K unit-group two-dimensional tensor alignment features of the tuples corresponding to the small-sample videos of all support sets belonging to ω i And the one unit-group two-dimensional tensor alignment feature of the tuple corresponding to the small-sample video of this query set belonging to ω i ;

[0141] S721: Use the N×K unit-group two-dimensional tensor alignment features of the tuples corresponding to the small-sample videos of all support sets belonging to ω as input, divide them according to whether they belong to the same sample class, and obtain N groups of same-class unit-group two-dimensional tensor alignment features; i The N groups of same-class unit-group two-dimensional tensor alignment features contain a total of K unit-group two-dimensional tensor alignment features;

[0142] ;

[0143] S722: For the N groups of same-class unit-group two-dimensional tensor alignment features, perform element-wise average calculation between features within each group respectively, and correspondingly obtain N class prototypes;

[0144] The N class prototypes respectively represent the average features of N sample classes;

[0145] S723: Calculate the Euclidean distance between the unit-group two-dimensional tensor alignment feature of the tuple corresponding to the small-sample video of this query set belonging to ω and the N class prototypes respectively, and correspondingly obtain the similarity scores of the small-sample videos of this query set belonging to ω i for the tuple; i ;

[0146] The similarity scores are specifically a one-dimensional tensor, including a total of N elements, corresponding to the N sample classes in the single task;

[0147] S724: Use the k groups of two-dimensional alignment features of the same tuple corresponding to this query set, repeat S721 to S723 k times, and obtain k similarity scores;

[0148] S73: Perform element-wise average calculation on the k similarity scores to obtain the average similarity score corresponding to the small-sample video of this query set;

[0149] S74: Select the sample class corresponding to the maximum average score from the average similarity scores corresponding to the small-sample videos of this query set as the predicted class label of the small-sample video of this query set;

[0150] S75: Use the N×M groups of two-dimensional tensor alignment features as input, repeat S71 to S74 N×M times, and obtain the N×M average similarity scores and M predicted class labels of all small-sample videos in the single task;

[0151] So far, from S1 to S7, an intelligent recognition and classification method for small-sample video behaviors has been completed.

[0152] Beneficial Effects

[0153] An intelligent recognition and classification method for small-sample video behaviors proposed by the present invention has the following beneficial effects compared with the prior art:

[0154] 1. The intelligent recognition and classification method is the first to propose a method for enhancing the capture of short-term spatio-temporal information features from the perspective of extracting feature-level motion information;

[0155] 2. The intelligent recognition and classification method constructs a motion-activated backbone network, expands the receptive field for capturing dynamic information in a hierarchical manner, and amplifies the original cumulative time window from two frames to four frames, greatly increasing the sensitivity of the backbone network to motion information;

[0156] 3. The method uses an attention mechanism and a classifier for small-sample video behavior recognition tasks. The embedded features output by the backbone network are aligned and measured in the classifier to obtain the average similarity score and the predicted class label to complete small-sample classification;

[0157] 4. The method realizes high-precision recognition and classification of human actions in complex scenarios, effectively improves the efficiency and accuracy of classification processing, and provides key technical support for downstream tasks such as video content review and personalized recommendation. Description of the Drawings

[0158] Figure 1 is a schematic flowchart of an intelligent recognition and classification method for small-sample video behaviors and an embodiment according to the present invention;

[0159] Figure 2 is an internal architecture diagram of an intelligent recognition and classification method for small-sample video behaviors relying on a motion activation module and an inter-frame motion activation module in the backbone network. Detailed Embodiments

[0160] The following further describes the content of an intelligent recognition and classification method for small-sample video behaviors of the present invention in conjunction with the drawings and examples.

[0161] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanations for the present application. The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0162] Example 1

[0163] Taking the application scenario with a video resolution of 224×224 and a 5-way 5-shot few-shot learning setting as an example, a method for intelligent recognition and classification of few-shot video behaviors of the present invention is described in detail. In the embodiment, the number of support sets for a single task is 25 in total. One query set for a single task is randomly selected from the videos of each class outside the support set videos, with a total of 5 query set videos. The input sequence of each few-shot video contains T = 8 frames. The backbone network used consists of a ResNet50 network, a motion activation module, and an interval-frame motion activation module. The compression factors r1 = r2 = 16, {ω0 = 2, ω1 = 3, i = 0, 1}.

[0164] The recognition method for classifying few-shot video behaviors in this embodiment executes the steps as Figure 1 shown, specifically including the following sub-steps:

[0165] S1: Divide the action data set of all few-shot videos into a training set, a validation set, and a test set;

[0166] Both the training set, the validation set, and the test set contain few-shot videos and must ensure that the classes to which the few-shot videos belong have no intersections;

[0167] The number of frames in each few-shot video is P, and P is greater than or equal to 5;

[0168] Specifically in implementation, P is greater than or equal to 8;

[0169] S2: Extract the input sequence corresponding to each few-shot video from all the frames of each few-shot video included in the training set, the validation set, and the test set;

[0170] For the extraction of the input sequence corresponding to each few-shot video, the frame extraction method for a single few-shot video is specifically as follows: Divide all P frames of a single few-shot video equally into T segments, and randomly extract one frame from each segment to form the input sequence of the single few-shot video, so that the input sequence corresponding to each few-shot video contains T frames;

[0171] T is less than or equal to P;

[0172] Specifically in implementation, T is equal to 8;

[0173] S3: Use the input sequences corresponding to the few-shot videos to construct the task sets belonging to the training set, the validation set, and the test set respectively;

[0174] Among them, the task sets belonging to the training set, the validation set, and the test set all include multiple single tasks;

[0175] The task set is the model input of the few-shot network;

[0176] The single task is created according to the N-way K-shot few-shot problem setting, and includes a support set and a query set;

[0177] Both the support set and the query set contain N identical sample video categories;

[0178] Among them, each category in the support set contains K few-shot videos, and each category in the query set contains M few-shot videos. It must be ensured that the sample videos contained in the support set and the query set have no intersection;

[0179] The K few-shot videos are specifically the input sequences of K few-shot videos, and the M few-shot videos are specifically the input sequences of M few-shot videos;

[0180] Specifically in implementation, N is equal to 5, K is equal to 5, and M is equal to 1;

[0181] S4: Construct a backbone network, and input the single task in the task set into the backbone network to extract the feature representations of all few-shot videos in the single task;

[0182] The backbone network includes a ResNet50 network, a motion activation module, and an interval motion activation module;

[0183] The structure of the motion activation module is as shown in Figure 2 Part 1, and the structure of the interval frame motion activation module is as shown in Figure 2 Part 2;

[0184] In specific implementation, the first layer in the ResNet50 network is composed of a conv2D convolutional layer with 64 7×7 convolutional kernels in series with a 3×3 max pooling layer; the second layer is composed of 3 Bottleneck_1 residual blocks in series, where Bottleneck_1 is composed of a conv2D convolutional layer with 64 1×1 convolutional kernels, a conv2D convolutional layer with 64 3×3 convolutional kernels, and a conv2D convolutional layer with 256 1×1 convolutional kernels in series; the third layer is composed of 4 Bottleneck_2 residual blocks in series, where Bottleneck_2 is composed of a conv2D convolutional layer with 128 1×1 convolutional kernels, a conv2D convolutional layer with 128 3×3 convolutional kernels, and a conv2D convolutional layer with 512 1×1 convolutional kernels in series; the fourth layer is composed of 6 Bottleneck_3 residual blocks in series, where Bottleneck_3 is composed of a conv2D convolutional layer with 256 1×1 convolutional kernels, a conv2D convolutional layer with 256 3×3 convolutional kernels, and a conv2D convolutional layer with 1024 1×1 convolutional kernels in series; the fifth layer includes 3 Bottleneck_4 residual blocks in series, where Bottleneck_4 is composed of a conv2D convolutional layer with 512 1×1 convolutional kernels, a conv2D convolutional layer with 512 3×3 convolutional kernels, and a conv2D convolutional layer with 2048 1×1 convolutional kernels in series;

[0185] Taking a small-sample video in a single task as an example, the feature representation of the small-sample video in the single task is extracted, which specifically includes the following sub-steps:

[0186] S41: Input the input sequence of a small-sample video in the single task into the first four layers of the ResNet50 network to obtain the first output feature;

[0187] The first output feature has four dimensions: time, number of channels, length, and width;

[0188] In specific implementation, the time dimension is 8, the number of channels is 1024, the length is 14, and the width is 14;

[0189] S42: Input the first output feature into the motion activation module to obtain the motion activation output feature, specifically:

[0190] S421: Divide the first output feature according to the time dimension to obtain the motion pre-activation features at each moment;

[0191] Each moment corresponds to each frame in the T frames of the small-sample video;

[0192] In specific implementation, T is equal to 8;

[0193] S422: Compress the motion pre-activation features at each moment using the first convolutional layer of the motion activation module to obtain the motion activation compressed features at each moment;

[0194] The compression using the first convolutional layer of the motion activation module has a compression coefficient of r1;

[0195] The number of channels of the motion activation compressed features at each moment is reduced to 1 / r1 of the number of channels of the first output feature;

[0196] In specific implementation, r1 is equal to 16 and 1 / r1 is equal to 1 / 16;

[0197] S423: Traverse the motion activation compressed features at each moment in the order of the small-sample video input sequence, and use the second convolutional layer of the motion activation module to align the motion activation compressed features of adjacent frames with the motion activation compressed features of the current frame to obtain the motion activation alignment features of adjacent frames relative to the current frame;

[0198] Among them, when traversing to the motion activation compressed feature at the last moment, there is no adjacent frame, so the alignment of the motion activation compressed feature of the adjacent frame with the motion activation compressed feature of the current frame is not performed at the last moment;

[0199] The current frame refers to the video frame corresponding to the current traversed moment;

[0200] The adjacent frame refers to the video frame corresponding to the next moment of the current frame;

[0201] The obtaining of the motion activation alignment features of adjacent frames relative to the current frame specifically means: inputting the motion activation compressed features of the adjacent frame into the second convolutional layer of the motion activation module, and outputting to obtain the motion activation alignment features of adjacent frames relative to the current frame;

[0202] S424: Traverse the motion activation compressed features at each moment in the order of the small-sample video input sequence, and perform frame difference calculation on the motion activation alignment features of adjacent frames relative to the current frame and the motion activation compressed features of the current frame in the currently traversed moment to obtain the motion activation difference frame features of the current frame;

[0203] Among them, when traversing to the motion activation compressed feature at the last moment, there is no adjacent frame, so the frame difference calculation result at the last moment is replaced by a zero tensor;

[0204] The frame difference calculation is specifically the element-wise subtraction between two feature tensors;

[0205] S425: Input the motion activation difference frame features at each moment into the global average pooling layer of the motion activation module, the third convolutional layer of the motion activation module, and the activation layer of the motion activation module in sequence to obtain the motion activation channel attention weights;

[0206] The motion activation channel attention weight is a one-dimensional tensor, and the number of elements it contains is equal to the number of channels of the first output feature;

[0207] S426: Perform a residual operation on the motion activation channel attention weight and the first output feature along the channel dimension to obtain the motion activation output feature;

[0208] The specific residual operation is as follows: along the channel dimension, multiply the first output feature by the motion activation channel attention weight, and then add the resulting product to the first output feature element by element to obtain the motion activation output feature;

[0209] The feature dimension of the motion activation output feature is consistent with that of the first output feature;

[0210] S43: Input the motion activation output feature into the fifth layer of the ResNet50 network and obtain the second output feature;

[0211] The second output feature has four dimensions: time, number of channels, length, and width;

[0212] In specific implementation, the time dimension is 8, the number of channels is 2048, the length is 7, and the width is 7;

[0213] S44: Input the second output feature into the interval motion activation module to obtain the interval motion activation output feature, specifically:

[0214] S441: Divide the second output feature along the time dimension to obtain the interval motion pre-activation features at each moment;

[0215] Each moment corresponds to each frame in the T frames of the small sample video;

[0216] In specific implementation, T is equal to 8;

[0217] S442: Compress the interval motion pre-activation features at each moment using the first convolutional layer of the interval motion activation module to obtain the interval motion activation compressed features at each moment;

[0218] The compression using the first convolutional layer of the interval motion activation module has a compression coefficient of r2;

[0219] The number of channels of the interval motion activation compressed features at each moment is reduced to 1 / r2 of the number of channels of the second output feature;

[0220] In specific implementation, r2 is equal to 16, and 1 / r2 is equal to 1 / 16;

[0221] S443: Traverse the interval motion activation compressed features at each time interval in the order of the small sample video input sequence. Use the second convolutional layer of the interval motion activation module to align the interval motion activation compressed features of the interval frames with the interval motion activation compressed features of the current frame, and obtain the interval motion activation alignment features of the interval frames relative to the current frame;

[0222] Among them, when traversing the interval motion activation compressed features at the last two time intervals, there are no interval frames, so the alignment of the interval motion activation compressed features of the interval frames with the interval motion activation compressed features of the current frame is not performed at the last two time intervals;

[0223] The current frame mentioned above refers to the video frame corresponding to the currently traversed time;

[0224] The interval frame mentioned above refers to the video frame corresponding to the time two moments after the current frame;

[0225] The obtaining of the interval motion activation alignment features of the interval frames relative to the current frame specifically means: input the interval motion activation compressed features of the interval frames into the second convolutional layer of the interval motion activation module, and output the interval motion activation alignment features of the interval frames relative to the current frame;

[0226] S444: Traverse the interval motion activation compressed features at each time interval in the order of the small sample video input sequence. Calculate the frame difference between the interval motion activation alignment features of the interval frames relative to the current frame and the interval motion activation compressed features of the current frame at the currently traversed time, and obtain the interval motion activation difference frame features of the current frame;

[0227] Among them, when traversing the interval motion activation compressed features at the last two time intervals, there are no interval frames, so the frame difference calculation results at the last two time intervals are replaced with zero tensors;

[0228] The frame difference calculation is specifically the element-wise subtraction of two feature tensors;

[0229] S445: Input the interval motion activation difference frame features at each time into the global average pooling layer of the interval motion activation module, the third convolutional layer of the interval motion activation module, and the activation layer of the interval motion activation module, and obtain the interval motion activation channel attention weights;

[0230] The interval motion activation channel attention weights are a one-dimensional tensor, and the number of elements contained is equal to the number of channels of the second output feature;

[0231] S446: Perform a residual operation on the interval motion activation channel attention weights and the second output feature along the channel dimension to obtain the interval motion activation output feature;

[0232] The specific residual operation is as follows: in the channel dimension, multiply the second output feature by the spatio-temporal motion activation channel attention weight, and then add the result of the multiplication to the second output feature element by element to obtain the spatio-temporal motion activation output feature;

[0233] The feature dimension of the spatio-temporal motion activation output feature is the same as that of the second output feature, which is called the feature representation of this small-sample video;

[0234] S45: For all small-sample videos in a single task, repeat S41 to S44 to obtain the feature representations of all small-sample videos in the single task;

[0235] S5: In the channel dimension, perform global average pooling on the feature representation of each small-sample video in the single task, and each small-sample video corresponds to a two-dimensional tensor feature;

[0236] Among them, a single task obtains N×K + N×M two-dimensional tensor features in total;

[0237] For the two-dimensional tensor feature, the first dimension corresponds to the number of input sequence frames T of the small-sample video, and the second dimension corresponds to the number of channels of the second output feature;

[0238] In specific implementation, N is equal to 5, K is equal to 5, M is equal to 1, and T is equal to 8;

[0239] S6: Align the two-dimensional tensor features corresponding to all small-sample videos in the support set of the single task with the two-dimensional tensor features corresponding to each small-sample video in the query set of the single task to obtain N×M groups of two-dimensional tensor alignment features;

[0240] The alignment of the two-dimensional tensor features corresponding to all small-sample videos in the support set of the single task with the two-dimensional tensor features corresponding to each small-sample video in the query set of the single task is specifically as follows: if there are N×M small-sample videos in the query set of this single task, the two-dimensional tensor features corresponding to all small-sample videos in the support set will be aligned with the two-dimensional tensor features corresponding to each small-sample video in the query set. A single task performs N×M alignments in total. Each query set small-sample video will generate a group of two-dimensional tensor alignment features, and a single task obtains N×M groups of two-dimensional tensor alignment features in total;

[0241] In specific implementation, N is equal to 5 and M is equal to 1;

[0242] Taking a small-sample video in the query set and using all small-sample videos in the support set as an example, perform one alignment to generate a group of two-dimensional tensor alignment features. The alignments of the remaining small-sample videos in the query set of the single task are the same as this, specifically:

[0243] S61: Take the small-sample videos of all support sets and the N×K + 1 two-dimensional tensor features corresponding to the small-sample videos of the query set as inputs. For each two-dimensional tensor feature, perform multi-tuple feature recombination and multi-tuple feature dimensionality reduction in sequence to obtain N×K + 1 groups of multi-tuple recombined two-dimensional tensor features;

[0244] The N×K + 1 two-dimensional tensor features include: N×K two-dimensional tensor features corresponding to the small-sample videos of all support sets, and one two-dimensional tensor feature corresponding to the small-sample video of the query set;

[0245] The N×K + 1 groups of multi-tuple recombined two-dimensional tensor features include: N×K groups of multi-tuple recombined two-dimensional tensor features corresponding to the small-sample videos of all support sets, and one group of multi-tuple recombined two-dimensional tensor features corresponding to the small-sample video of the query set;

[0246] Regarding the multi-tuple feature recombination and multi-tuple feature dimensionality reduction, taking one two-dimensional tensor feature among the N×K + 1 two-dimensional tensor features as an example for multi-tuple feature recombination and multi-tuple feature dimensionality reduction, the multi-tuple feature recombination and multi-tuple feature dimensionality reduction of the remaining two-dimensional tensor features among the N×K + 1 two-dimensional tensor features are the same, and include the following sub-steps:

[0247] S611: Take this two-dimensional tensor feature as an input, and split out multiple one-dimensional tensors row by row according to the first time dimension;

[0248] There are a total of T such one-dimensional tensors;

[0249] Specifically, when implemented, T is equal to 8;

[0250] S612: Perform multi-tuple feature recombination on all the one-dimensional tensors corresponding to this two-dimensional tensor feature to obtain a group of multi-tuple three-dimensional tensor features;

[0251] The multi-tuple specifically refers to all ω i -tuples, When k > 0, ω0 ≠ ω1 ≠.... ≠ ω k ;

[0252] Among them, T represents the number of sampled frames, represents a natural number;

[0253] Specifically, when implemented, {ω0 = 2, ω1 = 3, i = 0, 1};

[0254] The ω i -tuple is called a unit tuple, and a group of multi-tuples contains a total of k unit tuples;

[0255] Specifically, when implemented, k is equal to 2;

[0256] The set of multi - tuple three - dimensional tensor features contains k unit - group three - dimensional tensor features in total;

[0257] The multi - tuple feature recombination contains k unit - group feature recombinations in total. Specifically: all the one - dimensional tensors corresponding to the two - dimensional tensor feature are subjected to k unit - group feature recombinations, and each unit - group feature recombination will respectively obtain a unit - group three - dimensional tensor feature;

[0258] The unit - group feature recombination is specifically: all the one - dimensional tensors corresponding to the two - dimensional tensor feature are recombined and sorted according to the combination sequence set corresponding to the unit - group, and a unit - group three - dimensional tensor feature is obtained;

[0259] The combination sequence set corresponding to the unit - group specifically refers to: Let the unit - group be ω i tuple. For ω i tuple, the set of combination sequences corresponding to it is

[0260] wherein, the set of combination sequences corresponding to ω i tuple contains elements in total;

[0261] Specifically in implementation, the set of combination sequences corresponding to ω0 tuple contains 28 elements, and the set of combination sequences corresponding to ω1 tuple contains 56 elements;

[0262] The unit - group three - dimensional tensor feature specifically refers to: Let the unit - group be ω i tuple. The unit - group three - dimensional tensor feature is the ω i tuple three - dimensional tensor feature. The first dimension corresponds to the i possible combinations of ω tuple, the second dimension corresponds to the value of ω i tuple, and the third dimension corresponds to the number of channels of the second output feature;

[0263] Specifically in implementation, the output dimension of the combination sequence corresponding to ω0 tuple is 28×2×2048, and the output dimension of the combination sequence corresponding to ω1 tuple is 56×3×2048;

[0264] S613: Perform multi - tuple feature dimensionality reduction on the multi - tuple three - dimensional tensor features to obtain a set of multi - tuple recombined two - dimensional tensor features;

[0265] The set of multi - tuple recombined two - dimensional tensor features contains k unit - group recombined two - dimensional tensor features in total;

[0266] The multi - tuple feature dimensionality reduction contains k unit - group feature dimensionality reductions in total. Specifically: the k unit - group three - dimensional tensor features contained in the multi - tuple three - dimensional tensor features are respectively subjected to unit - group feature dimensionality reduction, and each unit - group feature dimensionality reduction will correspondingly obtain a unit - group recombined two - dimensional tensor feature;

[0267] Dimensionality reduction of the unit group features is specifically as follows: Let the unit group be ω i tuple, and the three-dimensional tensor feature of the unit group is ω i tuple three-dimensional tensor feature. The second and third dimensions of the three-dimensional tensor feature of the unit group are merged to obtain a reorganized two-dimensional tensor feature of the unit group;

[0268] The reorganized two-dimensional tensor feature of the unit group specifically refers to: Let the unit group be ω i tuple, and the reorganized two-dimensional tensor feature of the unit group is ω i tuple reorganized two-dimensional tensor feature, where the first dimension corresponds to ω i tuple's number of possible combinations, and the second dimension corresponds to the product of the ω i value and the number of channels of the second output feature;

[0269] In specific implementation, the output dimension of the combination sequence corresponding to the ω0 tuple is 28×4096, and the output dimension of the combination sequence corresponding to the ω1 tuple is 56×6144;

[0270] S614: For all the small-sample videos in the support set and the N×K + 1 two-dimensional tensor features corresponding to the small-sample videos in the query set, repeat S611 to S613 to obtain N×K + 1 groups of reorganized two-dimensional tensor features of multi-tuples;

[0271] Among them, one group of reorganized two-dimensional tensor features of multi-tuples contains k reorganized two-dimensional tensor features of unit groups; N×K + 1 groups of reorganized two-dimensional tensor features of multi-tuples contain k×(N×K + 1) reorganized two-dimensional tensor features of unit groups;

[0272] In specific implementation, N is equal to 5, K is equal to 5, and k is equal to 2;

[0273] S62: Use the k×(N×K + 1) reorganized two-dimensional tensor features of unit groups corresponding to all the small-sample videos in the support set and the small-sample videos in the query set as input, and divide them by the unit group to which each reorganized two-dimensional tensor feature of the unit group belongs to obtain k groups of reorganized two-dimensional tensor features of the same tuple;

[0274] Each group of reorganized two-dimensional tensor features of the same tuple in the k groups of reorganized two-dimensional tensor features of the same tuple corresponds to the same unit group;

[0275] Each group of reorganized two-dimensional tensor features of the same tuple contains N×K + 1 reorganized two-dimensional tensor features of unit groups belonging to the same ω i tuple;

[0276] The N×K + 1 reorganized two-dimensional tensor features of unit groups belonging to the same ω iThe unit group recombination two-dimensional tensor features of the tuple, which altogether include: the small-sample videos of all support sets that belong to ω i The N×K unit group recombination two-dimensional tensor features of the tuple, and the small-sample videos of the query set that belong to ω i One unit group recombination two-dimensional tensor feature of the tuple;

[0277] In specific implementation, N is equal to 5, K is equal to 5, and k is equal to 2;

[0278] S63: Perform k times of same-tuple feature alignment on k groups of same-tuple recombination two-dimensional tensor features to obtain a group of two-dimensional tensor alignment features;

[0279] Among them, a group of two-dimensional tensor alignment features altogether includes k groups of same-tuple two-dimensional tensor alignment features, that is, one same-tuple recombination two-dimensional tensor feature will perform one same-tuple feature alignment to obtain a group of same-tuple two-dimensional tensor alignment features;

[0280] The said group of same-tuple two-dimensional tensor alignment features altogether includes N×K+1 unit group two-dimensional tensor alignment features;

[0281] The N×K+1 unit group two-dimensional tensor alignment features altogether include: the N×K unit group two-dimensional tensor alignment features of the small-sample videos of all support sets, and one unit group two-dimensional tensor alignment feature of the small-sample videos of the query set;

[0282] In specific implementation, N is equal to 5, K is equal to 5, and k is equal to 2;

[0283] The said one same-tuple feature alignment specifically refers to: assuming that the unit group to which this same-tuple belongs is ω i Tuple, input ω i The same-tuple recombination two-dimensional tensor feature corresponding to the tuple, perform one same-tuple feature alignment to obtain a group of same-tuple two-dimensional tensor alignment features that belong to ω i Tuple, including the following sub-steps:

[0284] S631: Among the group of same-tuple recombination two-dimensional tensor features corresponding to the ω i Tuple, input the N×K unit group recombination two-dimensional tensor features of the small-sample videos of all support sets that belong to ω i Tuple into the Fully connected layer, and each small-sample video of the support set will correspondingly obtain a key value

[0285] Among them, Represents The number of neurons in the fully connected layer, and the set of combined sequences corresponding to the ω i Tuple altogether includes Elements;

[0286] In a specific embodiment, the equals 1152, and Γ belongs to equals 28; the equals 1152, and Γ belongs to equals 56;

[0287] S632: In the two-dimensional tensor feature obtained by reorganizing the same tuples corresponding to the ω i tuple, input the two-dimensional tensor feature of a single-unit tuple corresponding to the small-sample video of the query set that belongs to the ω i tuple into the fully connected layer, and a query value will be obtained correspondingly

[0288] In a specific embodiment, Υ corresponding to the ω0 tuple belongs to equals 28; Υ corresponding to the ω1 tuple belongs to equals 56;

[0289] S633: In the two-dimensional tensor feature obtained by reorganizing the same tuples corresponding to the ω i tuple, input the two-dimensional tensor features of N×K single-unit tuples corresponding to the small-sample videos of all support sets that belong to the ω i tuple into the fully connected layer, and each small-sample video of the support set will obtain an s_value value correspondingly

[0290] wherein, represents the number of neurons in the fully connected layer;

[0291] In a specific embodiment, the equals 1152, and Λ s belongs to equals 28; the equals 1152, and Λ s belongs to equals 56;

[0292] S634: In the two-dimensional tensor feature obtained by reorganizing the same tuples corresponding to the ω i tuple, input the two-dimensional tensor feature of a single-unit tuple corresponding to the small-sample video of the query set that belongs to the ω i tuple into the fully connected layer, and a q_value value will be obtained correspondingly

[0293] The q_value is the two-dimensional tensor alignment feature of the unit tuple belonging to ω corresponding to the small-sample video of the query set; i tuple;

[0294] In a specific embodiment, Λ corresponding to the ω0 tuple q belongs to is equal to 28; Λ corresponding to the ω1 tuple q belongs to is equal to 56;

[0295] S635: Perform layer normalization on the key value corresponding to the small-sample video of each support set and the query value corresponding to the small-sample video of the query set respectively, to obtain the normalized key value corresponding to the small-sample video of each support set and the normalized query value corresponding to the small-sample video of the query set;

[0296] S636: Perform a dot product operation of two-dimensional tensors on the normalized key value corresponding to the small-sample video of each support set and the normalized query value corresponding to the small-sample video of the query set respectively, and map the result of each dot product operation with the Softmax function respectively, to obtain the attention weight corresponding to the small-sample video of each support set;

[0297] Among them, the dot product operation of two-dimensional vectors and the mapping with the Softmax function are both performed N×K times respectively;

[0298] There are N×K attention weights in total, and the dimension of each attention weight is two-dimensional, and both the first dimension and the second dimension are

[0299] The N×K corresponds to the number of small-sample videos of all support sets;

[0300] In a specific implementation, N is equal to 5 and K is equal to 5; corresponding to the ω0 tuple is equal to 28; corresponding to the ω1 tuple is equal to 56;

[0301] S637: Perform a dot product operation of two-dimensional tensors on the attention weight corresponding to the small-sample video of each support set and the corresponding s_value value, and each small-sample video of each support set will correspondingly obtain the two-dimensional tensor alignment feature of the unit tuple belonging to ω i tuple

[0302] The two-dimensional tensor alignment feature of the unit tuple There are N×K in total;

[0303] In a specific embodiment, the equals 1152, belongs to equals 28; the equals 1152, belongs to equals 56;

[0304] S638: Integrate the two-dimensional tensor alignment features of the unit groups belonging to the ω i tuples to obtain a set of two-dimensional tensor alignment features of the same-tuple tuples belonging to the ω i tuples;

[0305] The set of two-dimensional tensor alignment features of the same-tuple tuples belonging to the ω i tuples altogether contains N×K + 1 two-dimensional tensor alignment features of the unit groups belonging to the ω i tuples;

[0306] The N×K + 1 two-dimensional tensor alignment features of the unit groups belonging to the ω i tuples specifically refer to: N×K and one

[0307] In a specific embodiment, the belonging to Λ q belongs to equals 28; the belonging to Λ q belongs to equals 56;

[0308] S639: Use the k sets of two-dimensional tensor features of the reorganized same-tuple tuples as the input, repeat S631 to S638, to obtain a set of two-dimensional tensor alignment features;

[0309] Among them, a set of two-dimensional tensor alignment features altogether contains k sets of two-dimensional tensor alignment features of the same-tuple tuples;

[0310] S64: Align the two-dimensional tensor features corresponding to all the small-sample videos in the support set of the single task with the two-dimensional tensor features corresponding to each small-sample video in the query set of the single task, repeat S61 to S63, to obtain N×M sets of two-dimensional tensor alignment features;

[0311] Specifically in implementation, N equals 5 and M equals 1;

[0312] S7: Perform N×M times of few-shot classification calculations on the aligned features of N×M groups of two-dimensional tensors to obtain N×M average similarity scores and N×M predicted class labels for all query-set few-shot videos in a single task;

[0313] The average similarity score is specifically a one-dimensional tensor, which includes N elements in total, corresponding to the N sample classes in a single task;

[0314] There are N×M predicted class labels in total, corresponding to the N×M query-set few-shot videos in a single task;

[0315] Specifically in implementation, N equals 5 and M equals 1;

[0316] The performing N×M times of few-shot classification calculations on the aligned features of N×M groups of two-dimensional tensors is specifically as follows: Taking the few-shot video of one query set combined with the few-shot videos of all support sets to perform one few-shot classification calculation to generate one average similarity score and one predicted class label as an example, it includes the following sub-steps:

[0317] S71: Take the aligned features of the two-dimensional tensor corresponding to this query set as input, and divide it by unit groups to obtain k groups of aligned features of the same-element tensors;

[0318] Specifically in implementation, k equals 2;

[0319] S72: Calculate k similarity scores by using the k groups of aligned features of the same-element tensors corresponding to this query set;

[0320] The calculating k similarity scores, taking calculating one similarity score by using a group of aligned features of the same-element tensors belonging to ω i tuple as an example, includes the following sub-steps:

[0321] Among them, a group of aligned features of the same-element tensors belonging to ω i tuple contains in total: N×K groups of aligned features of the unit tensors corresponding to the few-shot videos of all support sets belonging to ω i tuple, and one group of aligned features of the unit tensor corresponding to the few-shot video of this query set belonging to ω i tuple;

[0322] S721: Take the N×K groups of aligned features of the unit tensors corresponding to the few-shot videos of all support sets belonging to ω i tuple as input, and divide them according to whether they belong to the same sample class to obtain N groups of aligned features of the unit tensors of the same class;

[0323] The N groups of aligned features of the unit tensors of the same class contain K groups of aligned features of the unit tensors in total;

[0324] In a specific embodiment, N is equal to 5 and K is equal to 5;

[0325] S722: For the two-dimensional tensor alignment features of N groups of same-category unit groups, perform element-wise average calculation between features within each group respectively, and correspondingly obtain N class prototypes;

[0326] The N class prototypes respectively represent the average features of N sample classes;

[0327] In a specific embodiment, 5 class prototypes are correspondingly obtained;

[0328] S723: For the two-dimensional tensor alignment features of the unit groups corresponding to the small-sample videos in the query set that belong to the ω i tuple, perform Euclidean distance calculation with the N class prototypes respectively, and correspondingly obtain the similarity scores of the small-sample videos in the query set that belong to the ω i tuple;

[0329] The similarity scores are specifically a one-dimensional tensor, which altogether includes N elements, corresponding to the N sample classes in the single task;

[0330] S724: Utilize the k groups of two-dimensional alignment features corresponding to the query set, repeat S721 to S723 k times, and obtain k similarity scores;

[0331] In specific implementation, k is equal to 2;

[0332] S73: Perform element-wise average calculation on the k similarity scores to obtain the average similarity score corresponding to the small-sample videos in the query set;

[0333] S74: From the average similarity scores corresponding to the small-sample videos in the query set, select the sample class corresponding to the maximum average score as the predicted class label of the small-sample videos in the query set;

[0334] S75: Take the N×M groups of two-dimensional tensor alignment features as the input, repeat S71 to S74 N×M times, and obtain the N×M average similarity scores and M predicted class labels of all the small-sample videos in the query set in the single task;

[0335] In a specific embodiment, take 5 groups of two-dimensional tensor alignment features as the input, repeat S71 to S74 5 times, and obtain 5 average similarity scores and 5 predicted class labels of all the small-sample videos in the query set in the single task;

[0336] So far, from S1 to S7, an intelligent recognition and classification method for small-sample video behaviors is completed.

[0337] In summary, by means of the above technical solution of the present invention, the feature-level motion information between small-sample videos can be effectively captured. This method is simple and effective, and improves the accuracy of small-sample video recognition.

[0338] The above description is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intelligent recognition and classification method for small-sample video behaviors, characterized in that, Its working process includes the following steps: S1: Divide the action datasets of all small-sample videos into a training set, a validation set, and a test set; Among them, the training set, the validation set, and the test set all contain small-sample videos, and it must be ensured that the categories to which the small-sample videos belong have no intersections; S2: Extract the input sequences corresponding to each small-sample video from all frames of each small-sample video included in the training set, the validation set, and the test set; Among them, when extracting the input sequences corresponding to each small-sample video, the specific frame extraction method for a single small-sample video is as follows: evenly divide all P frames of a single small-sample video into T segments, and randomly extract one frame from each segment to form the input sequence of the single small-sample video, so that the input sequence corresponding to each small-sample video contains T frames; The P is greater than or equal to 5, and the T is less than or equal to P; S3: Use the input sequences corresponding to the small-sample videos to construct the task sets belonging to the training set, the validation set, and the test set respectively; Among them, the task sets belonging to the training set, the validation set, and the test set all include multiple single tasks; The task set is the model input of the small-sample network; The single task is created according to the N-way K-shot small-sample problem setting, and includes a support set and a query set; Both the support set and the query set contain N identical sample video categories; Among them, each category in the support set contains K small-sample videos, and each category in the query set contains M small-sample videos. It must be ensured that the sample videos contained in the support set and the query set have no intersections; The K small-sample videos are specifically the input sequences of K small-sample videos, and the M small-sample videos are specifically the input sequences of M small-sample videos; S4: Construct a backbone network, input the single tasks in the task set into the backbone network, and extract the feature representations of all small-sample videos in the single tasks; S5: Perform global average pooling on the feature representations of each small-sample video in the single task from the channel dimension, and each small-sample video correspondingly obtains a two-dimensional tensor feature; S6: Align the two-dimensional tensor features corresponding to all small-sample videos in the support set of the single task with the two-dimensional tensor features corresponding to each small-sample video in the query set of the single task to obtain N×M groups of two-dimensional tensor alignment features; S7: Perform N×M times of small-sample classification calculations on the N×M groups of two-dimensional tensor alignment features to obtain N×M average similarity scores and N×M predicted class labels for all small-sample videos in the query set of the single task.

2. The intelligent recognition and classification method according to claim 1, characterized in that: In S4, the backbone network includes a ResNet network, a motion activation module, and an inter-frame motion activation module; the ResNet network is respectively connected to the motion activation module and the inter-frame motion activation module. Specifically: the outputs of the first four layers of the ResNet network are connected to the motion activation module, the output of the motion activation module is used as the input of the fifth layer of the ResNet network, and the output of the fifth layer is given to the inter-frame motion activation module.

3. The intelligent recognition and classification method according to claim 1, wherein: S4 specifically includes the following sub-steps. Taking a small-sample video in a single task as an example, extract the feature representation of this small-sample video in the single task: S41: Input the input sequence of a small-sample video in a single task into the first four layers of the ResNet network to obtain the first output feature; S42: Input the first output feature into the motion activation module to obtain a motion activation output feature; The feature dimension of the motion activation output feature is the same as that of the first output feature; S43: Input the motion activation output feature into the fifth layer of the ResNet network and obtain a second output feature; The second output feature has four dimensions: time, number of channels, length, and width; S44: Input the second output feature into the interval motion activation module to obtain an interval motion activation output feature; The feature dimension of the interval motion activation output feature is the same as that of the second output feature, which is the feature representation of the small sample video; S45: For all small sample videos in a single task, repeat S41 to S44 to obtain the feature representations of all small sample videos in the single task; For the two-dimensional tensor feature, the first dimension corresponds to the number of input sequence frames T of the small sample video, and the second dimension corresponds to the number of channels of the second output feature.

4. The first output feature is input into the motion activation module according to S42 in claim 3 to obtain a motion activation output feature, characterized in that: Specifically, it includes the following sub-steps: S421: Divide the first output feature along the time dimension to obtain motion pre-activation features at each moment; Each moment corresponds to each frame in the T frames of the small sample video; S422: Compress the motion pre-activation features at each moment using the first convolutional layer of the motion activation module to obtain motion activation compressed features at each moment; When using the first convolutional layer of the motion activation module for compression, the compression coefficient is r1, which will reduce the number of channels of the compressed feature to 1 / r1 of the number of channels of the first output feature; S423: Traverse the motion activation compressed features at each moment in the order of the input sequence of the small sample video, and use the second convolutional layer of the motion activation module to align the motion activation compressed features of adjacent frames with the motion activation compressed feature of the current frame to obtain the motion activation alignment feature of adjacent frames relative to the current frame; S424: Traverse the motion activation compressed features at each moment in the order of the input sequence of the small sample video, and calculate the frame difference between the motion activation alignment feature of adjacent frames relative to the current frame and the motion activation compressed feature of the current frame at the currently traversed moment to obtain the motion activation difference frame feature of the current frame; S425: Input the motion activation difference frame features at each moment into the global average pooling layer, the third convolutional layer, and the activation layer of the motion activation module in sequence to obtain the motion activation channel attention weight; S426: Perform a residual operation on the motion activation channel attention weight and the first output feature along the channel dimension to obtain the motion activation output feature.

5. According to the sub-step S42 described in claim 4, it is characterized in that: In S423, when traversing to the motion activation compressed feature at the last moment, there is no adjacent frame, so the alignment of the motion activation compressed feature of the adjacent frame with the motion activation compressed feature of the current frame is not performed at the last moment; the current frame refers to the video frame corresponding to the currently traversed moment; the adjacent frame refers to the video frame corresponding to the next moment of the current frame; obtaining the motion activation alignment feature of adjacent frames relative to the current frame specifically means: inputting the motion activation compressed feature of the adjacent frame into the second convolutional layer of the motion activation module and outputting the motion activation alignment feature of adjacent frames relative to the current frame; In S424, when traversing to the motion activation compression feature at the last moment, there is no adjacent frame, so the frame difference calculation result at the last moment is replaced by a zero tensor; the frame difference calculation is specifically the element-wise subtraction between two feature tensors; in S425, the motion activation channel attention weight is a one-dimensional tensor, and the number of elements it contains is equal to the number of channels of the first output feature.

6. The second output feature is input into the intermittent motion activation module according to S44 in claim 3 to obtain an intermittent motion activation output feature, characterized in that: Specifically, it includes the following sub-steps: S441: Divide the second output feature along the time dimension to obtain the motion pre-activation features at each time interval; Each moment corresponds to each frame in the T frames of this small sample video; S442: Compress the motion pre-activation features at each time interval using the first convolutional layer of the interval motion activation module to obtain the motion activation compression features at each time interval; When compressing using the first convolutional layer of the interval motion activation module, the compression coefficient is r2, which will reduce the number of channels of the compression feature to 1 / r2 of the number of channels of the second output feature; S443: Traverse the motion activation compression features at each time interval in the order of the input sequence of the small sample video, and use the second convolutional layer of the interval motion activation module to align the motion activation compression features of the interval frames with the motion activation compression features of the current frame to obtain the motion activation alignment features of the interval frames relative to the current frame; S444: Traverse the motion activation compression features at each time interval in the order of the input sequence of the small sample video, and calculate the frame difference between the motion activation alignment features of the interval frames relative to the current frame and the motion activation compression features of the current frame at the currently traversed moment to obtain the motion activation difference frame features of the current frame; S445: Input the motion activation difference frame features at each time into the global average pooling layer, the third convolutional layer, and the activation layer of the interval motion activation module in sequence to obtain the motion activation channel attention weight; S446: Perform a residual operation on the motion activation channel attention weight and the second output feature along the channel dimension to obtain the motion activation output feature.

7. According to the sub-step S44 described in claim 6, it is characterized in that: In S443, when traversing to the motion activation compression features at the last two moments, there are no interval frames, so the alignment between the motion activation compression features of the interval frames and the motion activation compression features of the current frame is not performed at the last two moments; the current frame refers to the video frame corresponding to the currently traversed moment; the interval frame refers to the video frame corresponding to the next next moment of the current frame; obtaining the motion activation alignment features of the interval frames relative to the current frame specifically means: inputting the motion activation compression features of the interval frames into the second convolutional layer of the interval motion activation module, and outputting the motion activation alignment features of the interval frames relative to the current frame; In S444, when traversing to the motion activation compression features at the last two moments, there are no interval frames, so the frame difference calculation results at the last two moments are replaced by zero tensors; the frame difference calculation is specifically the element-wise subtraction between two feature tensors; in S445, the motion activation channel attention weight is a one-dimensional tensor, and the number of elements it contains is equal to the number of channels of the second output feature.

Citation Information

Patent Citations

  • Multi-modal data-oriented small sample machine learning method and system, and medium

    CN110363239A

  • Human body behavior recognition method and system based on extracted video spatio-temporal information

    CN113887419A