A Semi-Supervised Human Behavior Recognition Method Based on Contrastive Learning

By dividing the video sample data into labeled and labelless data, using the dual-path time comparison learning framework for semi-supervised training, and assigning pseudo-labels to labelless data, the problem of insufficient information on limited labeled data in the prior art human behavior recognition method is solved, improving the recognition performance and reducing costs.

CN115100738BActive Publication Date: 2025-07-04GUILIN UNIV OF ELECTRONIC TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210638655.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-07
Publication Date
2025-07-04
Estimated Expiration
2042-06-07

AI Technical Summary

Technical Problem

Existing human behavior recognition methods obtain less human behavior information on limited labeled data, resulting in excessive labor and time costs.

Method used

The video sample data is divided into labeled data and labelless data, and the dual-path time comparison learning framework is used for training. The recognition model is semi-supervised through comparison learning and supervised learning, and pseudo-labels are assigned to the labelless data to obtain the final model.

Benefits of technology

By mining more information in the video sample data, the performance of human behavior recognition is improved, the dependence on labeled data is reduced, and labor and time costs are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115100738B_ABST
    Figure CN115100738B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of video processing technology, and in particular to a semi-supervised human behavior recognition method based on contrastive learning, comprising dividing video sample data into labeled data and unlabeled data; based on a dual-path time contrastive learning framework, training a recognition model with labeled data and unlabeled data in turn to obtain a secondary training model; after assigning pseudo labels to the unlabeled data, performing supervised training on the secondary training model to obtain a final model; and evaluating the model with test data to obtain the recognition performance of the model. The present invention mines more information in video sample data by constructing a semi-supervised model with contrastive learning, thereby solving the problem that the existing human behavior recognition method obtains less human behavior information on limited labeled data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and particularly to a semi-supervised human behavior recognition method based on contrast learning. Background Art

[0002] Human behavior information contained in video data.

[0003] Currently, generally, human labor is used to process and analyze video data to obtain human behavior information in the video.

[0004] However, this method requires a large amount of labor cost and time cost, and cannot meet the demand for processing human behavior information in video data. Summary of the Invention

[0005] The purpose of the present invention is to provide a semi-supervised human behavior recognition method based on contrast learning, aiming to solve the problem that the existing human behavior recognition methods obtain less human behavior information on limited labeled data.

[0006] To achieve the above purpose, the present invention provides a semi-supervised human behavior recognition method based on contrast learning, including the following steps:

[0007] Divide video sample data into labeled data and unlabeled data;

[0008] Based on a dual-path temporal contrast learning framework, and successively use the labeled data and the unlabeled data to train an identification model to obtain a secondarily trained model;

[0009] Assign pseudo-labels to the unlabeled data and then perform supervised training on the secondarily trained model to obtain a final model;

[0010] Use test data to evaluate the final model to obtain the recognition performance of the final model.

[0011] Among them, the dual-path temporal contrast learning framework includes a basic path and an auxiliary path.

[0012] Among them, the specific method of using the labeled data and the unlabeled data to train the identification model based on the dual-path temporal contrast learning framework to obtain a secondarily trained model is as follows:

[0013] On the basic path, use the labeled data to perform supervised training on the identification model to obtain a preliminarily trained model;

[0014] Perform semi-supervised training on the preliminarily trained model based on the basic path and the auxiliary path through contrast learning and supervised learning to obtain a secondarily trained model.

[0015] The specific method of performing semi-supervised training on the preliminary training model based on the basic path and the auxiliary path through contrastive learning and supervised learning to obtain the secondary training model is:

[0016] The model outputs of the preliminary training models on the basic path and the auxiliary path are constrained by instance contrast learning using the batch number of the unlabeled data, and the similarity between the video samples of the unlabeled data is modeled by similarity contrast learning. At the same time, the preliminary training model is semi-supervisedly trained using the labeled data to obtain a secondary training model.

[0017] The specific method of performing supervised training on the secondary training model after assigning pseudo labels to the unlabeled data to obtain the final model is:

[0018] Assigning a pseudo label to the unlabeled data to obtain the unlabeled data with the pseudo label;

[0019] The secondary training model is supervisedly trained on the basic path using the labeled data and the unlabeled data with pseudo labels to obtain a final model.

[0020] The present invention discloses a semi-supervised human behavior recognition method based on contrastive learning, which divides video sample data into labeled data and unlabeled data; based on a dual-path time contrastive learning framework, the recognition model is trained using the labeled data and the unlabeled data in turn to obtain a secondary training model; after assigning pseudo labels to the unlabeled data, the secondary training model is supervised and trained to obtain a final model; the final model is evaluated using test data to obtain the recognition performance of the final model; the present invention constructs the semi-supervised model using contrastive learning to mine more task-related and valuable information in the video sample data, thereby solving the problem that the existing human behavior recognition method obtains less human behavior information on limited labeled data. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0022] Figure 1 It is a flow chart of a semi-supervised human behavior recognition method based on contrastive learning provided by the present invention.

[0023] Figure 2It is a flowchart for training an identification model based on a dual-path temporal contrast learning framework, and successively using the labeled data and the unlabeled data to obtain a secondarily trained model.

[0024] Figure 3 It is a flowchart for performing supervised training on the secondarily trained model after assigning pseudo-labels to the unlabeled data to obtain a final model. Detailed implementation manners

[0025] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation of the present invention.

[0026] Please refer to Figures 1 to 3 , the present invention provides a semi-supervised human behavior recognition method based on contrast learning, including the following steps:

[0027] S1 Divide the video sample data into labeled data and unlabeled data;

[0028] Specifically, divide the video sample data of the training set into labeled data and unlabeled data where N1 and N2 respectively represent the number of samples of the labeled data and the number of samples of the unlabeled data.

[0029] S2 Construct an identification model, based on a dual-path temporal contrast learning framework, and successively use the labeled data and the unlabeled data to train the identification model to obtain a secondarily trained model;

[0030] Specifically, the dual-path temporal contrast learning framework includes a basic path and an auxiliary path.

[0031] The specific manner of training the identification model based on the dual-path temporal contrast learning framework and successively using the labeled data and the unlabeled data to obtain a secondarily trained model is as follows:

[0032] S21 Perform supervised training on the identification model using the labeled data on the basic path to obtain a preliminarily trained model;

[0033] Specifically, the following cross-entropy loss needs to be minimized during the training process:

[0034]

[0035] where V i and y irespectively represent the i-th video sample in the labeled data and its corresponding class label, and g(V i ) represents the model output obtained when V i is used as the input of the model. C represents the total number of classes.

[0036] S22 performs semi-supervised training on the training model based on the basic path and the auxiliary path through contrastive learning and supervised learning to obtain a secondarily trained model.

[0037] Specifically, the batch number of the unlabeled data is used to constrain the model outputs of the preliminary training model on the basic path and the auxiliary path through instance contrastive learning, and the similarity between the video samples of the unlabeled data is modeled through similarity contrastive learning. At the same time, the labeled data is used to perform semi-supervised training on the preliminary training model to obtain a secondarily trained model;

[0038] The batch number of the unlabeled data is used to constrain the model outputs of the preliminary training model on the basic path and the auxiliary path through instance contrastive learning:

[0039] Specifically, in the process of contrastive learning, mainly consider a small batch of B unlabeled data, and sample different numbers of video frames for the video U in the unlabeled data i Two different forms of videos can be obtained by sampling different numbers of video frames (which can be called fast videos and slow videos respectively): and where both M and N represent the number of video frames, and M > N;

[0040] Take and as the inputs of the basic path and the auxiliary path respectively, and their corresponding outputs are and The different outputs corresponding to the same video are called positive pairs, and the outputs corresponding to different videos are called negative pairs. The model outputs on the basic path and the auxiliary path are constrained through instance contrastive learning. In order to maximize the consistency between the two model outputs and corresponding to the same video sample, and minimize the consistency between the model outputs corresponding to different video samples, it is necessary to minimize the following instance contrastive loss during the model training process:

[0041]

[0042]

[0043] where, and respectively represent two different forms of video data corresponding to the i-th video sample in the unlabeled data, τ represents a hyperparameter, and I {*} represents the indicator function, and when the condition in the curly brackets * holds, I {*} = 1, otherwise I {*} = 0. The final instance contrast loss can be obtained by calculating the instance contrast loss between all positive pairs (including and );

[0044] Model the similarity between video samples of the unlabeled data through similarity contrast learning:

[0045] Specifically, model the similarity between video samples of the unlabeled data through similarity contrast learning. To ensure that the similarity between videos on the base path and the auxiliary path in the dual-path temporal contrast learning framework is consistent, the following similarity contrast loss needs to be minimized during model training:

[0046] L sc = ||A f - A s || F ,

[0047] where ||*|| F represents the Frobenius norm, A f ∈R B×B represents the similarity matrix obtained by calculating the similarity between videos on the base branch, A s ∈R B×B represents the similarity matrix obtained by calculating the similarity between videos on the auxiliary branch, A f and A s The element in the i-th row and j-th column of can be obtained through the following calculation method:

[0048]

[0049] where F represents the dimension of, and respectively represent and the k-th element of, represents the average value of all elements of, represents the average value of all elements of.

[0050] Use the unlabeled data and the labeled data to perform semi-supervised training on the model to obtain a secondarily trained model:

[0051] Specifically, when performing semi-supervised training on the recognition model, the following loss needs to be minimized:

[0052] L = L sup + γ * L ic + η * L sc ,

[0053] where γ and β represent the weights of the instance contrast loss and the similarity contrast loss respectively.

[0054] S3 performs supervised training on the secondary training model after assigning pseudo-labels to the unlabeled data to obtain the final model;

[0055] The specific method is as follows:

[0056] S31 assigns pseudo-labels to the unlabeled data to obtain unlabeled data with pseudo-labels;

[0057] Specifically, the secondary training model is used to assign pseudo-labels to the unlabeled data. During the process of assigning pseudo-labels, pseudo-labels will only be assigned to the unlabeled data when the confidence of the pseudo-labels is greater than a certain threshold.

[0058] S32 uses the labeled data and the unlabeled data with pseudo-labels to perform supervised training on the secondary training model on the basic path to obtain the final model.

[0059] S4 uses the test data to evaluate the final model to obtain the recognition performance of the final model.

[0060] Specifically, in order to test the recognition performance of the model, the trained model is evaluated using the test data on the basic path of the dual-path temporal contrast learning framework, and the human behavior recognition of the video data to be recognized is performed through the final model.

[0061] Supervised deep learning methods have been widely applied to various visual tasks due to their excellent performance. Such methods usually require a large amount of labeled data to improve the generalization performance. In the case of only a small amount of labeled data, the generalization performance of supervised deep learning methods is easily restricted. Compared with supervised deep learning methods, semi-supervised deep learning methods can use unlabeled data to mine some useful information during the model training process, thus effectively avoiding the problem of using a large amount of labeled data.

[0062] Although semi-supervised deep learning methods have achieved remarkable results in the field of images, their development in the field of videos has been relatively slow. On the one hand, videos are more difficult to process than images. Generally speaking, when processing images, only the information in the spatial dimension needs to be concerned, while when processing videos, in order to analyze the content in the video more precisely, not only the information in the spatial dimension needs to be concerned, but also the information in the temporal dimension needs to be concerned. On the other hand, applying semi-supervised deep learning methods to the video field requires a large amount of computing resources.

[0063] Contrastive learning methods have attracted research interest in recent years. They do not rely on labeled data and mainly use the data itself as supervision information to learn more valuable feature representations.

[0064] A semi-supervised human behavior recognition method based on contrastive learning of the present invention divides video sample data into labeled data and unlabeled data; based on a dual-path temporal contrastive learning framework, and successively uses the labeled data and the unlabeled data to train an identification model to obtain a secondarily trained model; assigns pseudo-labels to the unlabeled data and then conducts supervised training on the secondarily trained model to obtain a final model; uses test data to evaluate the final model to obtain the recognition performance of the final model; the present invention solves the problem that the existing human behavior recognition methods obtain less human behavior information on limited labeled data by using contrastive learning to construct a semi-supervised model to mine more task-related and valuable information in video sample data.

[0065] The above-disclosed is only a preferred embodiment of a semi-supervised human behavior recognition method based on contrastive learning of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. A semi-supervised human behavior recognition method based on contrastive learning, characterized in that, Including the following steps: Dividing the video sample data into labeled data and unlabeled data; Based on the dual-path temporal contrastive learning framework, and successively using the labeled data and the unlabeled data to train the recognition model to obtain a secondarily trained model; Assigning pseudo-labels to the unlabeled data and then performing supervised training on the secondarily trained model to obtain a final model; Evaluating the final model using test data to obtain the recognition performance of the final model; The dual-path temporal contrastive learning framework includes a basic path and an auxiliary path; The specific method of successively using the labeled data and the unlabeled data to train the recognition model based on the dual-path temporal contrastive learning framework to obtain a secondarily trained model is as follows: Performing supervised training on the recognition model using the labeled data on the basic path to obtain a preliminarily trained model; Performing semi-supervised training on the preliminarily trained model based on the basic path and the auxiliary path through contrastive learning and supervised learning to obtain a secondarily trained model; The specific method of performing semi-supervised training on the preliminarily trained model based on the basic path and the auxiliary path through contrastive learning and supervised learning to obtain a secondarily trained model is as follows: Constraining the model outputs of the preliminarily trained model on the basic path and the auxiliary path through instance contrastive learning using the batch number of the unlabeled data, and modeling the similarity between the video samples of the unlabeled data through similarity contrastive learning, and at the same time performing semi-supervised training on the preliminarily trained model using the labeled data to obtain a secondarily trained model; Modeling the similarity between the video samples of the unlabeled data through similarity contrastive learning. To ensure that the similarity between the videos on the basic path and the auxiliary path in the dual-path temporal contrastive learning framework is consistent, the following similarity contrast loss needs to be minimized during model training: , Among them, represents the Frobenius norm, represents the similarity matrix obtained by calculating the similarity between videos on the basic branch, represents the similarity matrix obtained by calculating the similarity between videos on the auxiliary branch, and The element in the i th row and j th column of is obtained through the following calculation method: , Among them, represents the dimension of and respectively represent and the k th element of represents the average value of all elements, represents the average value of all elements.

2. The semi-supervised human behavior recognition method based on contrastive learning according to claim 1, wherein The specific method of assigning pseudo-labels to the unlabeled data and then performing supervised training on the secondarily trained model to obtain a final model is as follows: Assigning pseudo-labels to the unlabeled data to obtain unlabeled data with pseudo-labels; Performing supervised training on the secondarily trained model on the basic path using the labeled data and the unlabeled data with pseudo-labels to obtain a final model.