A Behavior Transfer Learning Recognition Method for Small-Sample Scenarios
By using the behavioral transfer learning recognition method based on 3D convolutional neural network in small sample scenarios, a positive and negative sample data set was constructed and the transfer learning model improved by SlowFast was adopted, transfer learning from the source domain to the target domain was realized, and the problem of low behavior recognition rate in small sample scenarios was solved, and the recognition accuracy was improved.
Patent Information
- Application Number
- CN202310508297.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-08
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2043-05-08
AI Technical Summary
The existing behavior recognition methods are not highly recognized in small sample scenarios, and it is difficult to complete the behavior recognition task through a small number of samples.
The behavioral transfer learning recognition method based on 3D convolutional neural network is adopted. By constructing positive and negative sample data sets, the transfer learning model improved by SlowFast is used to realize transfer learning from the source domain to the target domain, and improve the accuracy of behavior detection and recognition in small sample scenarios.
The accuracy of behavior detection and recognition in small sample scenarios is improved, and the problem of low recognition rate of 3D convolutional neural networks in small sample scenarios is solved.
Smart Images

Figure CN116778570B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine vision and relates to a method for behavior transfer learning recognition for small-sample scenarios. Background Art
[0002] In recent years, with the development of machine learning and artificial intelligence, computer vision has made rapid progress and has begun to be applied in different fields, bringing great changes to human life. Due to the urgent application requirements in fields such as human-computer interaction and intelligent security, behavior recognition has become one of the research hotspots in the field of computer vision in recent years. Whether it is intelligent security monitoring, customer shopping behavior analysis, smart home systems, motion sensing games, or action recognition of pedestrians on the road during unmanned driving, it all depends on a highly efficient and accurate behavior recognition system. And the purpose of behavior recognition is to classify and recognize the behaviors or actions of one or more people in a video. Its research object is often a series of video sequences, rather than being limited to the analysis of single-frame images.
[0003] However, existing behavior recognition methods rely on a large number of data samples to train the behavior recognition model. However, in many real-world scenarios, it is often difficult to obtain effective training samples, resulting in many behavior recognition methods being unable to complete the behavior recognition task with a small number of samples. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a behavior transfer learning recognition method based on a 3D convolutional neural network for detecting video behavior recognition in small-sample scenarios, to solve the problem of low recognition rate of the 3D convolutional neural network in the case of insufficient training samples, and to improve the accuracy of behavior detection and recognition in small-sample scenarios.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] A method for behavior transfer learning recognition for small-sample scenarios, comprising the following steps:
[0007] S1: Construct a positive sample data set, including a source domain positive sample data set with rich data and a target domain positive sample data set with small samples;
[0008] S2: Perform spatial information transformation and temporal information shuffling on the positive sample data set to construct a negative sample data set;
[0009] S3: Construct a behavior recognition transfer learning model improved based on SlowFast to enhance the temporal information of the video and achieve the function of migrating from the source domain to the target domain;
[0010] S4: Use the positive and negative sample datasets of the source domain and the target domain to train the improved behavior recognition transfer learning model based on SlowFast to obtain a behavior recognition model that generalizes on the target domain;
[0011] S5: Use the trained behavior recognition model to perform behavior recognition in the actual scenario.
[0012] Further, the step S1 specifically includes the following steps:
[0013] S11: Use the acquisition device to collect video of the target behavior; form all the target and behavior videos into the target domain behavior video set TargetV{vt 1 ,..., vt i ,..., vt N}, where represents the target domain behavior video, i ∈ {1,..., N}, and N represents the total number of target domain behavior videos collected; frame each behavior video vt i in the TargetV at 30 frames per second to obtain the frame sequence seqt i corresponding to vt i {p 1 ,..., p 2 ,..., p 30}; after frame extraction of all behavior videos in TargetV, obtain the target domain frame sequence set Target_Seq{seqt 1 ,..., seqt i ,..., seqt N};
[0014] Collect videos of the same behavior category from the public dataset to obtain the source domain behavior video vs j , and construct the source domain behavior video set SourceV{vs 1 ,..., vs j ,..., vs M}, where vs j represents the source domain behavior video, j ∈ {1,..., M}, and M represents the total number of source domain behavior videos; perform frame extraction on all behavior videos in sourceV to obtain the source domain frame sequence set Source_Seq{seqs 1 ,..., seqs j ,..., seqs M};
[0015] S12: Establish domain labels, positive and negative sample labels, and action category labels for the frame sequences in Target_Seq and Source_Seq, respectively, to obtain the positive sample datasets for the source domain and the target domain. Among them, the domain label is denoted as d, where d = 0 or 1. When d = 0, it indicates that the sample belongs to the source domain; when d = 1, it indicates that the sample belongs to the target domain. The positive and negative sample label is denoted as b, where b = 0 or 1. When b = 0, it indicates that the sample is a negative sample; when b = 1, it indicates that the sample is a positive sample. The action category label is denoted as y, where y = 1, 2, 3,..., k, and each number represents an action category. Attach labels to each seqt in the target domain frame sequence set Target_Seq, and the target domain positive sample is denoted as target_pos i = {seqt i , d, b, y}, where d = 1, b = 1, and y is the number corresponding to the action category to which the video belongs; obtain the target domain positive sample dataset
[0016] Target_Pos = {target_pos 1 ,..., target_pos i ,..., target_pos N}; Attach labels to each seqs in Source_Seq j , and the source domain positive sample is denoted as source_pos j = {seqs j , d, b, y}, where d = 0, b = 1, and y is the number corresponding to the action category, to obtain the source domain positive sample dataset
[0017] Source_Pos = {source_pos 1 ,..., source_pos j ,..., source_pos M}.
[0018]
[0019] Furthermore, the step S2 specifically includes the following steps:
[0020] S21: Shuffle each frame sequence seqt in the target domain frame sequence set Target_Seq both spatially and temporally, where seqt i = {p i , p 1 ,..., p 2}; Reverse and randomly crop the first frame image p 30 in seqt i with the 17th frame image p 1 to obtain p′1 With p' 17 , then except for p' 1 With p' 17 , shuffle all the frames except p' to obtain the target domain negative sample frame sequence seqt_neg i ; The set of target domain negative sample frame sequences is denoted as Target_Seq_Neg = {seqt_neg 1 ,..., seqt_neg i ,..., seqt_neg N};
[0020] Similarly, shuffle each frame sequence seqs in Source_Seq j to obtain the source domain negative sample frame sequence set, denoted as Source_Seq_Neg = {seqs_neg 1 ,..., seqs_neg j ,..., seqs_neg M};
[0021] S22: Establish domain labels, positive and negative sample labels, and action category labels for the frame sequences in Source_Seq_Neg and Target_Seq_Neg, respectively, to obtain the negative sample data sets for the source domain and the target domain; Label each seqt_neg in the target domain negative sample frame sequence set Target_Seq_Neg i , and the target domain negative sample is denoted as target_neg i = {seqt_neg i , d, b, y}, where d = 1, b = 0, and y is the number corresponding to the action category to which the video belongs; Obtain the target domain negative sample data set Target_Neg = {target_neg 1 ,..., target_neg i ,..., target_neg N}; Label each seqs_neg in Source_Seq_Neg j , and the source domain negative sample is denoted as source_neg j = {seqs_neg j , d, b, y}, where d = 0, b = 0, and y is the number corresponding to the action category, to obtain the source domain negative sample data set Source_Neg = {source_neg 1 ,..., source_neg j ,..., source_neg M}.
[0022] Further, in step S3, a transfer learning model based on SlowFast is constructed, which specifically includes the following steps:
[0023] S31: Based on the network architecture of the Siamese neural network, the model is divided into two sub-networks network1 and network2 with shared weights. Both sub-networks are of the SlowFast model structure. The source domain positive sample dataset Source_Pos and the source domain negative sample dataset Source_Neg are used as the inputs of network1, and the target domain positive sample dataset Target_Pos and the target domain negative sample dataset Target_Neg are used as the inputs of network2. After passing through the feature extraction part of SlowFast, the output feature f1 = {feature1, d, b, y} is obtained, where feature1 is the feature tensor obtained after the input sample passes through the SlowFast feature extraction.
[0024] S32: The feature f1 in step S31 is fed into the fully connected layer fc6 to obtain the output feature f2 = {feature2, d, b, y}, where feature2 is the feature tensor obtained after feature1 passes through fc6. Calculate the positive and negative sample losses within the domain to enhance the temporal information of the video and the action information of the person. The loss function formula is as follows:
[0025]
[0026] where, x i , x j , x p = {feature2, d, b, y} are different input samples, which contain a feature tensor feature2 and three labels d, y, b; S is the batch size, which controls the number of comparison samples; is the indicator function. When the input is True, returns 1, otherwise returns 0; means that when the sample x i and x j have the same class label y and the positive and negative sample label b = 1, it returns True, otherwise returns False; Θ{A, B} is the cosine similarity, which is used to measure the distance between features. The formula is as follows:
[0027]
[0028] where, τ is the hyperparameter;
[0029] S33: Add a fully-connected layer fc7 after the fc6 layer. After passing through fc7, f2 obtains the output feature f3 = {feature3, d, b, y}; calculate the inter-domain loss between the source domain and the target domain. To achieve the purpose of migrating the source domain to the few-shot target domain, the loss function formula is as follows:
[0030]
[0031] Where x h , x l , x t = {feature3, d, b, y} are different input samples, which contain a feature tensor feature3 and three labels d, y, b; ρ h,t Indicates that when the domain labels d of samples x h and x t are different and the positive and negative sample label b = 1, return True, otherwise return False;
[0032] S34: Add a fully-connected layer fc8 after the fc7 layer. After passing through fc8, f3 obtains the output feature f4, and use the cross-entropy function to obtain the classification loss. The total loss of the model is
[0033] Furthermore, in step S4, after executing the specified number of training epochs, the loss reaches convergence, completing the model training, and obtaining a behavior recognition model for the few-shot scenario.
[0034] Furthermore, in step S5, use the object detection model to calibrate the position box of the person, and then use the behavior recognition model trained in S4 to detect the behavior of the person in the position box.
[0035] The beneficial effects of the present invention are as follows:
[0036] (1) The present invention constructs a negative sample dataset by performing spatial transformation and temporal shuffling on the dataset, and obtains the contrast loss by inputting the positive and negative sample datasets into the SlowFast network, enhancing the temporal information of the data and the action information of the person, thereby improving the accuracy of the model.
[0037] (2) The present invention designs a transfer learning network model based on the twin neural network architecture, enabling the knowledge learned by the model in the scenario with sufficient data volume to be applied to the few-shot scenario with less data volume, thereby solving the problems of low recognition accuracy and easy overfitting of the behavior recognition model in the few-shot scenario.
[0038] Other advantages, objects, and features of the present invention will be set forth in part in the following description, and in part will be obvious to those skilled in the art upon examination of the following, or may be learned from the practice of the present invention. The objects and other advantages of the present invention may be realized and obtained by the following description of the specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to make the objectives, technical solutions, and advantages of the present invention more clear, the present invention will be described in detail with reference to the accompanying drawings. Among them:
[0040] Figure 1 It is a flowchart for identifying the transfer learning behavior of personnel based on the SlowFast network in the small-sample scenario of the present invention;
[0041] Figure 2 It is a transfer learning structure diagram based on the SlowFast model proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0043] Among them, the accompanying drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as limiting the present invention; in order to better illustrate the embodiments of the present invention, some components in the accompanying drawings will be omitted, enlarged, or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.
[0044] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the accompanying drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0045] Please refer to Figures 1 to 2 , in order to improve the accuracy of action behavior recognition for different categories in small-sample scenarios, as Figure 1 shown, the present invention provides a transfer learning method based on a 3D convolutional neural network, including the following steps:
[0046] S1: Construct a positive sample data set, including a source domain positive sample data set with rich data and a target domain positive sample data set with small samples, specifically as follows:
[0047] S11: Construct a set of video frame sequences for the target domain and the source domain. Use a collection device with a frame rate of not less than 30fps, specify the installation location of the collection device, and require the collection device to be able to clearly capture the action behavior of personnel. Conduct video collection of the target behavior. Combine all target and behavior videos into a set to obtain the target domain behavior video set TargetV{vt 1 ,..., vt i ,..., vt N}, where vt i represents the target domain behavior video, i∈{1,..., N}, and N represents the total number of target domain behavior videos collected. Collect videos of the same behavior category from publicly available action recognition data sets such as Kinetics400 and UCF101 to obtain the source domain behavior video vs j , and construct the source domain behavior video set SourceV{vs 1 ,..., vs j ,..., vs M}, where vs j represents the source domain behavior video, j∈{1,..., M}, and M represents the total number of source domain behavior videos, and it is required that M is not less than 100. Extract frames from each behavior video vt i in TargetV at a rate of 30 frames per second to obtain the frame sequence seqt i {p i , p 1 ,..., p 2 ,..., p 30} corresponding to it. After extracting frames from all behavior videos in TargetV, obtain the target domain frame sequence set Target_Seq{seqt 1 ,..., seqt i ,..., seqt N}. Similarly, extract frames from all behavior videos in SourceV to obtain the source domain frame sequence set Source_Seq{seqs 1 ,..., seqs j ,..., seqs M, uniformly process the video frames into the size of 224*224;
[0048] S12: Establish domain labels, positive and negative sample labels, and action category labels for the frame sequences in Target_Seq and Source_Seq, respectively obtaining the positive sample datasets for the source domain and the target domain. Among them, the domain label is denoted as d, d = 0 or 1. When d = 0, it indicates that the sample belongs to the source domain; when d = 1, it indicates that the sample belongs to the target domain. The positive and negative sample label is denoted as b, b = 0 or 1. When b = 0, it indicates that the sample is a negative sample; when b = 1, it indicates that the sample is a positive sample. The action category label is denoted as y, y = 1, 2, 3,..., k, and each number represents an action category. Label each seqt in the target domain frame sequence set Target_Seq i with labels, and the positive samples in the target domain can be expressed as target_pos i ={seqt i , d, b, y}, where d = 1, b = 1, and y is the number corresponding to the action category to which the video belongs. Obtain the target domain positive sample dataset Target_Pos{target_pos 1 ,..., target_pos i ,..., target_pos N}. Similarly, label each seqs in Source_Seq j with labels, and the positive samples in the source domain can be expressed as source_pos j ={seqs j , d, b, y}, where d = 0, b = 1, and y is the number corresponding to the action category, obtaining the source domain positive sample dataset
[0049] Source_Pos={source_pos 1 ,..., source_pos j ,..., source_pos M};
[0050] S2: Construct a negative sample action recognition dataset, specifically as follows:
[0051] S21: Construct the negative sample frame sequence sets for the source domain and the target domain. Shuffle each frame sequence seqt in the target domain frame sequence set Target_Seq i spatially and temporally, where seqt i ={p 1 , p 2 ,..., p 30}. Take the first frame picture p in seqt i 1 Perform operations such as inversion and random cropping on the 17th frame picture p 17 to obtain p'. 1 For p' 17 , except for p' 1 and p' 17 , shuffle all other frames to obtain the target domain negative sample frame sequence seqt_neg i . The set of target domain negative sample frame sequences can be represented as Target_Seq_Neg = {seqt_neg 1 ,..., seqt_neg i ,..., seqt_neg N}. Similarly, shuffle each frame sequence seqs j in Source_Seq. Obtain the source domain negative sample frame sequence set, which can be represented as Source_Seq_Neg = {seqs_neg 1 ,..., seqs_neg j ,..., seqs_neg M}.;
[0052] S22: Establish domain labels, positive and negative sample labels, and action category labels for the frame sequences in Source_Seq_Neg and Target_Seq_Neg, respectively, to obtain the negative sample data sets for the source domain and the target domain. Label each seqt_neg i in the target domain negative sample frame sequence set Target_Seq_Neg. The target domain negative sample can be represented as target_neg i = {seqt_neg i , d, b, y}, where d = 1, b = 0, and y is the number corresponding to the action category to which the video belongs. Obtain the target domain negative sample data set Target_Neg = {target_neg 1 ,..., target_neg i ,..., target_neg N}. Similarly, label each target_neg i in Source_Seq_Neg. The source domain negative sample can be represented as source_neg j = {seqs_neg j , d, b, y}, where d = 0, b = 0, and y is the number corresponding to the action category, to obtain the source domain negative sample data set Source_Neg = {source_neg 1 ,..., source_neg j ,..., source_negM};
[0053] S3: Construct an improved SlowFast transfer learning model, specifically as follows:
[0054] S31: Construct a transfer learning network framework based on the SlowFast model, specifically as follows: Based on the network architecture of the Siamese neural network, the model is divided into two sub-networks network1 and network2 with shared weights, and both sub-networks are of the SlowFast model structure. The source domain datasets Source_Pos and Source_Neg are used as the inputs of network1, and the target domain datasets Target_Pos and Target_Neg are used as the inputs of network2. After the input data passes through the feature extraction part of SlowFast, the feature map is flattened to obtain the output feature f1 = {feature1, d, b, y}, where feature 1 is a one-dimensional tensor obtained by flattening the feature map, and its size is {1×1×1764}.
[0055] S32: Add a fully connected layer fc6 after the SlowFast feature extraction part. The fc6 layer has 1024 neurons. Use the feature f1 in S31 as the input of the fully connected layer fc6 to obtain the output feature f2 = {feature2, d, b, y}. feature2 is the feature tensor obtained by passing feature1 through fc6, and its size is {1×1×1024}. Calculate the positive and negative sample losses of f2 in network1 and network2 respectively Average the two losses to obtain the intra-domain contrast loss
[0056] S33: Add a fully connected layer fc7 after the fc6 layer. The fc7 layer has 512 neurons. After f2 passes through fc7, the output feature f3 = {feature3, d, b, y} is obtained, and the size of feature3 is {1×1×512}. Calculate the inter-domain loss of f3
[0057] S34: Add a fully connected layer fc8 after the fc7 layer. The fc8 layer has k neurons, where k is the number of behavior categories included in the dataset. After f3 passes through fc8, the output feature f4 = {feature4, d, b, y} is obtained, and the size of feature4 is {1×1×k}. Use the cross-entropy function to calculate the classification loss The intra-domain loss The inter-domain loss The classification loss Are weighted and added together to obtain the total loss of the model Set α = 0.15, β = 0.15, γ = 0.7.
[0058] S4: Use the positive sample datasets of the source domain and the target domain and the negative sample dataset constructed in S2 to perform transfer learning training on the improved SlowFast model;
[0059] In the training phase, the number of training epochs = 100. Use learning rate warm start. The initial learning rate (learningrate) is set to 0.001. The optimization strategy optimizing_method: sgd (stochastic gradient descent). The number of epochs for learning rate warm start = 5, the decay rate weight_decay = 1e - 7, and the batch size is set to 64. In the first 5 training epochs, learning rate warm start is performed. After 5 epochs, the learning rate reaches a steady state. Then, in the next 15 epochs, the model is trained relatively stably. After training is completed, a behavior recognition model for the small sample scenario is obtained, achieving the purpose of accurately recognizing the target behavior in the case of a small amount of sample data or lack of data.
[0060] S5: Use the behavior recognition model obtained after training to perform behavior detection in the actual scenario;
[0061] In the detection process of the actual scenario, the object detection model in the MMAction library can be used to detect the position of the person in the video and establish a position bounding box. The object detection model can be selected from Faster - RCNN, Yolov4, and Yolov5; then use the model trained in step S4 to perform behavior recognition on the person in the input bounding box, and the behavior of the target person can be displayed in real time and the behavior can be recorded in the server log for easy monitoring;
[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A behavior transfer learning recognition method for small sample scenarios, characterized in that: It includes the following steps: S1: Construct a positive sample dataset, which includes a source domain positive sample dataset with rich data and a target domain positive sample dataset with small samples; S2: Perform spatial information transformation and temporal information shuffling on the positive sample dataset to construct a negative sample dataset; S3: Construct a behavior recognition transfer learning model improved based on SlowFast to enhance the temporal information of the video and achieve the function of migrating from the source domain to the target domain; Constructing a transfer learning model based on SlowFast specifically includes the following steps: S31: Based on the network architecture of the Siamese neural network, the model is divided into two sub-networks network1 and network2 with shared weights, and both sub-networks are SlowFast model structures; The source domain positive sample dataset Source_Pos and the source domain negative sample dataset Source_Neg are used as the inputs of network1, and the target domain positive sample dataset Target_Pos and the target domain negative sample dataset Target_Neg are used as the inputs of network2; After passing through the feature extraction part of SlowFast, the output feature f1 = {feature1, d, b, y} is obtained, where feature1 is the feature tensor obtained after the input sample passes through SlowFast feature extraction; S32: Feed the feature f1 in step S31 into the fully connected layer fc6 to obtain the output feature f2 = {feature2, d, b, y}, where feature2 is the feature tensor obtained by passing feature1 through fc6; calculate the positive and negative sample losses within the domain To enhance the temporal information of the video and the action information of the person, the loss function formula is as follows: where x i , x j , x p = {feature2, d, b, y} are different input samples, which contain a feature tensor feature2 and three labels d, y, b; S is the batch size, controlling the number of contrast samples; is an indicator function that returns 1 when the input is True and 0 otherwise; means that when the class label y of the sample x i is the same as that of x j and the positive and negative sample label b = 1, it returns True, otherwise it returns False; Θ{a, B} is the cosine similarity, used to measure the distance between features, and the formula is as follows: where τ is a hyperparameter; S33: Add a fully connected layer fc7 after the fc6 layer. After passing through fc7, f2 obtains the output feature f3 = {feature3, d, b, y}; calculate the inter-domain loss between the source domain and the target domain To achieve the purpose of migrating the source domain to the few-shot target domain, the loss function formula is as follows: where x h , x l , x t = {feature3, d, b, y} are different input samples, which contain a feature tensor feature3 and three labels d, y, b; ρ h,t represents that when the domain labels d of samples x h and x t are different and the positive and negative sample label b = 1, return True, otherwise return False; S34: Add a fully connected layer fc8 after the fc7 layer. After passing through fc8, f3 obtains the output feature f4, and the classification loss is obtained by using the cross-entropy function for it. The total loss of the model is S4: Use the positive sample dataset and negative sample dataset of the source domain and target domain to train the behavior recognition transfer learning model improved based on SlowFast to obtain a behavior recognition model that generalizes on the target domain; S5: Use the trained behavior recognition model to perform behavior recognition in the actual scenario.
2. The behavior transfer learning recognition method for small sample scenarios according to claim 1, characterized in that: The step S1 specifically includes the following steps: S11: Use a collection device to collect videos of target behaviors; combine all target and behavior videos to form a target domain behavior video set TargetV{vt 1 ,..., vt i ,..., vt N}, where vt i represents a target domain behavior video, i ∈ {1,..., N}, and N represents the total number of collected target domain behavior videos; for each behavior video vt i in targetV, extract frames at a rate of 30 frames per second to obtain the frame sequence seqt i corresponding to vt i {p 1 , p 2 ,..., p 30}; after extracting frames from all behavior videos in TargetV, obtain the target domain frame sequence set Target_Seq{seqt 1 ,..., seqt i ,..., seqt N}; Collect videos of the same behavior category from the public dataset to obtain the source domain behavior video $v_s$. j , construct the source domain behavior video set $SourceV = \{v_s$ 1 ,..., $v_s$ j ,..., $v_s$ M}, where $v_s$ j represents the source domain behavior video, $j\in\{1,...,M\}$, and $M$ represents the total number of source domain behavior videos; extract frames from all behavior videos in $sourceV$ to obtain the source domain frame sequence set $Source\_Seq = \{seqs$ 1 ,..., $seqs$ j ,..., $seqs$ M}; S12: Establish domain labels, positive and negative sample labels, and behavior category labels for the frame sequences in Target_Seq and Source_Seq, respectively, to obtain the positive sample datasets for the source domain and the target domain. Among them, the domain label is denoted as d, d = 0 or 1. When d = 0, it indicates that the sample belongs to the source domain; when d = 1, it indicates that the sample belongs to the target domain. The positive and negative sample label is denoted as b, b = 0 or 1. When b = 0, it indicates that the sample is a negative sample; when b = 1, it indicates that the sample is a positive sample. The behavior category label is denoted as y, y = 1, 2, 3,..., k, and each number represents a behavior category. Label each seqt in the target domain frame sequence set Target-Seq. The positive sample in the target domain is denoted as target_pos i ={seqt i , d, b, y}, where d = 1, b = 1, and y is the number corresponding to the behavior category to which the video belongs. Obtain the target domain positive sample dataset Target_Pos = {target_pos i , …, target_pos 1 ,..., target_pos i ,... target_pos N}; Label each seqs in Source_Seq. The positive sample in the source domain is denoted as source_pos j ={seqs j , d, b, y}, where d = 0, b = 1, and y is the number corresponding to the behavior category. Obtain the source domain positive sample dataset Source_Pos = {source_pos j ,..., source_pos 1 ,..., source_pos j ,... source_pos M}.
3. The behavior transfer learning recognition method for small sample scenarios according to claim 1, characterized in that: The step S2 specifically includes the following steps: S21: For each frame sequence seqt in the target domain frame sequence set Target_Seq i perform shuffling in both space and time, where seqt i = {p 1 , p 2 , …, p 30}; Reverse and randomly crop the first frame image p i and the 17th frame image p 1 in seqt 17 to obtain p’ 1 and p’ 17 , and then randomly shuffle all frames except p’ 1 and p’ 17 to obtain the target domain negative sample frame sequence seqt-neg i ; The target domain negative sample frame sequence set is represented as Target_Seq_Neg = {seqt_neg 1 ,..., seqt_neg i ,..., seqt_neg N}; Similarly, for each frame sequence seqs in Source_Seq j shuffle it to obtain a set of source domain negative sample frame sequences, denoted as Source_Seq_Neg = {seqs_neg 1 , …, seqs_neg j , …, seqs_neg M}; S22: Establish domain labels, positive and negative sample labels, and action category labels for the frame sequences in Source_Seq_Neg and Target_Seq_Neg, respectively, to obtain the negative sample datasets for the source domain and the target domain; for each seqt_neg in the target domain negative sample frame sequence set Target_Seq_Neg i attach a label, and the target domain negative sample is represented as target_neg i ={seqt_neg i , d, b, y}, where d = 1, b = 0, and y is the number corresponding to the action category to which the video belongs; obtain the target domain negative sample dataset Target_Neg = {target_neg 1 ,..., target_neg i ,..., target_neg N}; for each seqs_neg in Source_Seq_Neg j attach a label, and the source domain negative sample is represented as source_neg j ={seqs_neg j , d, b, y}, where d = 0, b = 0, and y is the number corresponding to the action category, and obtain the source domain negative sample dataset Source_Neg = {source_neg 1 , …, source_neg j ,...,, source_neg M}.
4. The behavior transfer learning recognition method for small sample scenarios according to claim 1, characterized in that: In step S4, after executing the specified number of training epochs, the loss reaches convergence, the model training is completed, and a behavior recognition model for small sample scenarios is obtained.
5. The behavior transfer learning recognition method for small sample scenarios according to claim 1, characterized in that: In step S5, use the object detection model to calibrate the position box of the person, and then use the trained behavior recognition model to detect the behavior of the person in the position box.
Citation Information
Patent Citations
Motion normalization detection method and device based on time consistency contrast learning
CN114648723A
Cross-domain video action recognition method, device and equipment and computer readable storage medium
CN115439791A