A behavior recognition method, device and equipment based on a behavior detection model applicable to fuzzy labels
By building a behavior detection model containing the main path and auxiliary branch, using obfuscated data samples for training and collaborative distillation, the problem of insufficient detection accuracy in the prior art is solved, and high-precision behavior recognition is achieved.
Patent Information
- Application Number
- CN202411512109.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-10-28
AI Technical Summary
The existing behavior recognition methods cannot train high-precision recognition models, resulting in insufficient detection accuracy when processing fuzzy label data sets, which cannot meet the needs of real-time monitoring.
By building an initial behavior detection model containing the main path model and auxiliary branch, using obfuscated data samples for training, obtaining auxiliary branch loss function, branch features and potential distribution soft labels, combining multi-dimensional loss function for collaborative distillation training, adjusting weight information to improve model accuracy.
It effectively improves the detection accuracy of the behavior detection model for fuzzy label data sets, meets the needs of real-time monitoring, and solves the problem of insufficient detection accuracy in the existing technology.
Smart Images

Figure CN119580347B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of behavior recognition technology, and in particular to a behavior recognition method, device and equipment based on a behavior detection model suitable for fuzzy labels. Background Art
[0002] At present, as a key research direction in the field of computer vision, behavioral target detection has made significant progress in recent years driven by deep learning technology. Researchers have provided a basis for model training and evaluation by constructing and annotating a large number of behavioral datasets. In terms of model architecture, innovations continue to emerge, including the integration of convolutional neural networks, recurrent neural networks, and attention mechanisms to improve the accuracy and robustness of behavior recognition.
[0003] In practical application scenarios, such as behavior recognition in examination rooms, the system is usually required to be able to monitor the behavior of examinees in real time and quickly identify violations, such as cheating or whispering, to ensure the fairness of the examination. This requires the behavior detection model to be able to process the video stream in a very short time and accurately identify the relevant behavior. The detection accuracy of the behavior detection model in the existing scheme usually depends on the annotation quality of the data set (i.e., sample data). The model training is carried out using a high-quality data set so that the trained model can recognize various behaviors. However, the atomic action set of human behavior in the sample data usually has behaviors that are not obvious for a certain feature presented in the image due to the lack of information about the previous and next frame sequences. In addition, different annotators have different annotation standards, and the same sample data often gives different labels, and there are fuzzy labels (uncertain labels). As a result, no matter how strong the detection model is, it is impossible to train a high-precision model. Due to insufficient model accuracy, video streams with fuzzy behaviors cannot be detected with high precision. It can be seen that the existing behavior recognition method cannot accurately detect the input data with high precision, and the detection efficiency is low. Summary of the invention
[0004] The present application provides a behavior recognition method, device and equipment based on a behavior detection model suitable for fuzzy labels. In the model training stage, auxiliary branch training and main path collaborative distillation training are completed by controlling the weight of the loss function, thereby reducing the impact of the original fuzzy labels on the model performance. A high-precision behavior detection model can be trained for annotated data sets with poor quality, meeting actual use needs, effectively improving the detection accuracy and on-site use effect of the behavior detection model for confused label data sets, thereby improving the accuracy of the behavior detection algorithm and the behavior detection model, and solving the problem of insufficient behavior recognition detection accuracy caused by the inability to train a high-precision recognition model in existing behavior recognition methods.
[0005] In a first aspect, the present application provides a behavior recognition method based on a behavior detection model applicable to fuzzy labels, including:
[0006] Construct an initial behavior detection model and obtain confused data samples. The initial behavior detection model includes a main path model and at least one auxiliary branch, and the auxiliary branches correspond one-to-one to the behavior categories in the confused data samples;
[0007] Train the auxiliary branches according to the confused data samples to obtain an auxiliary branch loss function, branch features, and potential distribution soft labels corresponding to the behavior categories;
[0008] Perform feature extraction and feature fusion training on the main path model based on the confused data samples to obtain main path fusion features;
[0009] Perform multi-dimensional loss constraints on the main path model according to the branch features, the main path fusion features, and the potential distribution soft labels to obtain a divergence loss function and a distillation loss function;
[0010] Based on the auxiliary branch loss function, the divergence loss function, and the distillation loss function, perform model training on the initial behavior detection model through preset weight information to obtain a target behavior detection model, and the weight information is adjusted correspondingly based on the training cycle of the initial behavior detection model;
[0011] Perform feature recognition and behavior detection processing on an experimental video stream containing fuzzy behaviors through the target behavior detection model to obtain a behavior recognition result.
[0012] Optionally, the training of the auxiliary branches according to the confused data samples to obtain an auxiliary branch loss function, branch features, and potential distribution soft labels corresponding to the behavior categories includes:
[0013] Extract sample images and sample labels corresponding to the sample images from the confused data samples, and input the sample images and the sample labels into the corresponding auxiliary branches;
[0014] Perform classification prediction and feature extraction on the sample images and the sample labels through the auxiliary branches to obtain prediction anchor boxes, prediction categories corresponding to the prediction anchor boxes, potential distribution soft labels of the confused data samples on the behavior categories, and branch features output by each auxiliary branch;
[0015] For the prediction anchor boxes and the prediction categories, perform classification loss constraints on the sample images and the sample labels through the auxiliary branches to obtain an auxiliary branch category loss function and an auxiliary branch anchor box regression loss function;
[0016] Determine the auxiliary branch loss function based on the auxiliary branch class loss function and the auxiliary branch anchor box regression loss function.
[0017] Optionally, the class loss constraint of the sample image and the sample label by the auxiliary branch to obtain the auxiliary branch class loss function and the auxiliary branch anchor box regression loss function includes:
[0018] According to the formula Determine the auxiliary branch class loss function;
[0019] According to the formula Determine the auxiliary branch anchor box regression loss function;
[0020] Among them, is the auxiliary branch class loss function, p t is the predicted box class confidence, α is a parameter for controlling the positive and negative sample balance, γ is a modulation factor parameter for reducing the proportion of easy-to-separate samples in model training, is the auxiliary branch anchor box regression loss function, Pred bbox is the predicted anchor box, GT box is the labeled anchor box, and the CIoU loss is used to minimize the deviation between the predicted anchor box and the labeled anchor box.
[0021] Optionally, the multi-dimensional loss constraint on the main path model based on the branch feature, the main path fusion feature, and the latent distribution soft label to obtain the divergence loss function and the distillation loss function includes:
[0022] Based on the latent distribution soft label, perform spatial classification loss constraint on the main path model through a preset main path detection algorithm to obtain the divergence loss function, and the divergence loss function is used to minimize the deviation between the confidence of the predicted target box of the main path model and the latent distribution soft label;
[0023] Perform loss constraint on the main path model in the feature space dimension based on the branch feature and the main path fusion feature to obtain the distillation loss function.
[0024] Optionally, the performing spatial classification loss constraint on the main path model through a preset main path detection algorithm based on the latent distribution soft label to obtain the divergence loss function includes:
[0025] Obtain the branch predicted target box predicted by the auxiliary branch for the confused data sample, and obtain the main path predicted target box predicted by the main path model for the latent distribution soft label and the class confidence corresponding to the main path predicted target box;
[0026] Perform intersection over union (IoU) matching based on the branch prediction target box and the main path prediction target box, and calculate the divergence loss function according to the class confidence.
[0027] Optionally, the performing intersection over union (IoU) matching based on the branch prediction target box and the main path prediction target box, and calculating the divergence loss function includes:[[]]
[0028] Calculate the divergence loss function according to the formula and perform intersection over union (IoU) matching according to ;
[0029] where D KL is the KL loss function, p(x) is the latent distribution soft label, and q(x) is the class confidence distribution of the main path prediction target box.
[0030] Optionally, the performing loss constraint on the main path model in the feature space dimension based on the branch feature and the main path fusion feature to obtain the distillation loss function includes:[[]]
[0031] Perform loss constraint in the feature space dimension according to the formula to obtain the distillation loss function;
[0032] where L SP is the distillation loss, G main is the main path fusion feature, G bi is each branch feature, and the distillation loss function is used to make the feature extraction ability of the target branch approach that of multiple auxiliary branches.
[0033] Optionally, the weight information includes the auxiliary branch training weight and the main path training weight. The training the initial behavior detection model based on the auxiliary branch loss function, the divergence loss function, and the distillation loss function through the preset weight information includes:[[]]
[0034] Determine the auxiliary branch loss function according to ;
[0035] Determine the main branch loss function according to ;
[0036] Determine the calculated loss function according to and train the initial behavior detection model;
[0037] where is the main branch class classification loss, is the main branch regression loss, $\mathcal{L}_{aux}$ is the loss of the auxiliary branch feature map, $\varepsilon$ is the training weight of the auxiliary branch, $(1 - \varepsilon)$ is the training weight of the main path. The training weight of the auxiliary branch and the training weight of the main path are in an inverse proportional relationship and are adjusted based on the training period of the initial behavior detection model.
[0038] In a second aspect, the present application provides a behavior recognition device based on a behavior detection model applicable to fuzzy labels, including:
[0039] A construction module, configured to construct an initial behavior detection model and obtain confused data samples. The initial behavior detection model includes a main path model and at least one auxiliary branch, and the auxiliary branches correspond one-to-one with the behavior categories in the confused data samples;
[0040] A first training module, configured to train the auxiliary branches according to the confused data samples to obtain an auxiliary branch loss function, branch features, and potential distribution soft labels corresponding to the behavior categories;
[0041] An extraction module, configured to perform feature extraction and feature fusion training on the main path model based on the confused data samples to obtain main path fusion features;
[0042] A loss constraint module, configured to perform multi-dimensional loss constraints on the main path model according to the branch features, the main path fusion features, and the potential distribution soft labels to obtain a divergence loss function and a distillation loss function;
[0043] A second training module, configured to perform model training on the initial behavior detection model based on the auxiliary branch loss function, the divergence loss function, and the distillation loss function through preset weight information to obtain a target behavior detection model, and the weight information is correspondingly adjusted based on the training period of the initial behavior detection model;
[0044] A behavior recognition module, configured to perform feature recognition and behavior detection processing on an experimental video stream including fuzzy behaviors through the target behavior detection model to obtain a behavior recognition result.
[0045] In a third aspect, the present application provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0046] The memory is used to store a computer program;
[0047] The processor is configured to, when executing the program stored in the memory, implement the steps of the behavior recognition method for a behavior detection model applicable to fuzzy labels according to any one of the embodiments in the first aspect.
[0048] In summary, in the model training stage of the behavior detection model in the embodiments of the present application, by obtaining confused data samples and constructing an initial behavior detection model including a main path model and an auxiliary branch, and making the auxiliary branch correspond to the behavior categories in the confused data samples one by one, then training the auxiliary branch in the initial behavior detection model according to the confused data samples to obtain an auxiliary branch loss function, branch features, and potential distribution soft labels corresponding to the behavior categories, so as to perform feature extraction and feature fusion training on the main path model based on the confused data samples to obtain main path fusion features. Subsequently, the main path model is subjected to multi-dimensional loss constraints using the branch features, main path fusion features, and potential distribution soft labels to obtain a divergence loss function and a distillation loss function. Finally, based on the auxiliary branch loss function, divergence loss function, and distillation loss function, the initial behavior detection model is trained through preset weight information to obtain a target behavior detection model. Thus, by controlling the weights of multiple loss functions, collaborative distillation training of the auxiliary branch and the main path in the initial behavior detection model is realized, so that the trained target behavior detection model can be used to perform feature recognition and high-precision behavior detection on an experimental video stream containing fuzzy behaviors to obtain behavior recognition results. The present application reduces the impact of the original fuzzy labels on the model performance through collaborative distillation learning, effectively improves the detection accuracy and practical application effect of the behavior detection model on the confused label dataset, so that the behavior detection model can be used to perform feature recognition and behavior detection on the input video stream with high precision, improving the accuracy and detection efficiency of behavior detection, and solving the problem of insufficient behavior recognition and detection accuracy caused by the inability to train a high-precision recognition model in the existing behavior recognition methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0051] Figure 1 It is a schematic flowchart of a behavior recognition method based on a behavior detection model applicable to fuzzy labels provided by an embodiment of the present application;
[0052] Figure 2 It is a schematic step flowchart of a behavior recognition method based on a behavior detection model applicable to fuzzy labels provided by an optional embodiment of the present application;
[0053] Figure 3It is the model structure diagram of a behavior detection model provided by an optional example of the present application;
[0054] Figure 4 It is the structural block diagram of a behavior recognition device based on a behavior detection model applicable to fuzzy labels provided by an embodiment of the present application;
[0055] Figure 5 It is the schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0057] In the related art, in addition to being applied to scenarios such as examination room behavior recognition, behavior target detection also has extensive applications in the field of security detection. This requires that the behavior detection model needs to quickly respond to potential security threats, such as illegal intrusion or abnormal activities, so as to take actions in a timely manner. These may all involve techniques such as lightweight network design, using more efficient convolution operations, introducing attention mechanisms to focus on key areas, and adopting multi-scale processing. Through these methods, the model can achieve fast processing while maintaining high accuracy, meeting the requirements of real-time monitoring.
[0058] To achieve high detection accuracy, existing behavior detection models mainly rely on the annotation quality of the dataset during the training stage. However, the dataset used for training the behavior detection model has various defects, including but not limited to: ① For the atomic action set of human behaviors, due to the lack of information on the front and back frame sequences, for a behavior with an unclear feature presented in the image, different annotators often give different labels, resulting in the situation of fuzzy labels (uncertainty labels); ② Each annotator has different criteria for judging behaviors. For example, for "tilting the head left and right", the criteria (tilting amplitude) for the annotator to determine whether this action is "tilting the head left and right" are also different; ③ When the number of targets in the same image is too large, the situations of missing annotation and misannotation will inevitably occur; ④ The long-tail data problem. The above factors are the main reasons for the poor quality of the annotated dataset, which also leads to the inability to train a high-precision model no matter how strong the detection model is. Existing solutions can only continuously repeat the inspection of data annotation to reduce fuzzy labels in an attempt to improve the quality of the dataset, wasting human resources and unable to meet the actual needs at the same time.
[0059] To solve the problem of insufficient behavior recognition accuracy of the behavior detection model, this application proposes a behavior recognition method based on the behavior detection model suitable for fuzzy labels. In the model training stage, by innovatively improving the framework of the behavior detection model, based on existing object detection and anomaly detection algorithms, N auxiliary branches (assuming the number of categories is N) are constructed to extract the potential feature distribution and attributes (labels) of the confused data samples, and the training of the auxiliary branches and the collaborative distillation training of the main path are completed by controlling the loss function weights, so as to effectively improve the detection accuracy and practical application effect of the basic algorithm for the confused label data set, and obtain a target behavior detection model that can accurately recognize fuzzy behaviors. The target behavior detection model is used to perform high-precision behavior detection and recognition on the input video stream, solving the problem of insufficient behavior recognition and detection accuracy caused by the existing behavior recognition method being unable to train a high-precision recognition model.
[0060] To facilitate the understanding of the embodiments of this application, the following will further explain with reference to the drawings and specific embodiments. The embodiments do not constitute a limitation to the embodiments of this application.
[0061] Figure 1 The flow chart of a method provided by an embodiment of this application is as follows Figure 1 As shown, the method provided by the embodiment of this application can specifically include the following steps:
[0062] Step 110, construct an initial behavior detection model and obtain confused data samples.
[0063] Among them, the initial behavior detection model includes a main path model and at least one auxiliary branch, and the auxiliary branch corresponds to the behavior category in the confused data sample one by one;
[0064] Specifically, the confused data sample can be composed of sample images / videos and sample annotations corresponding to multiple behavior categories. To improve the detection accuracy of the behavior detection model, the sample labels in the confused data sample can include a part of fuzzy labels.
[0065] When constructing the initial behavior detection model in the embodiment of this application, the main path model (or main branch) and multiple auxiliary branches can be constructed in the initial behavior detection model at the same time. The auxiliary branches can correspond to the behavior categories one by one. For example, when there are N behavior categories in the confused data sample, N auxiliary branches can be constructed in the initial behavior model, and the N auxiliary branches correspond to the N behavior categories one by one.
[0066] Step 120, train the auxiliary branch according to the confused data sample to obtain the auxiliary branch loss function, branch features, and the potential distribution soft label corresponding to the behavior category.
[0067] Specifically, the auxiliary branch loss function can be composed of a classification loss and an anchor box regression loss. Since the output of the auxiliary branch is the anchor box coordinates and the predicted category, in this embodiment, the loss function of the auxiliary branch is defined as the classification loss and the anchor box coordinate regression loss. When classifying and predicting the confused data samples through the auxiliary branch, the predicted category and the anchor box coordinates of the output can be obtained. The classification loss can correspond to the predicted category, and the anchor box coordinate regression loss can correspond to the anchor box coordinates. The classification loss is determined through the predicted category, and the anchor box coordinate regression loss is determined through the anchor box coordinates. Subsequently, the auxiliary branch loss function is determined by combining the classification loss and the anchor box regression loss.
[0068] In a specific implementation, the auxiliary branch can be pre-trained for feature extraction using data samples, so that the auxiliary branch can possess the feature extraction ability. Thus, each auxiliary branch can perform feature recognition and feature extraction based on the input data samples of the corresponding behavior category to obtain branch features.
[0069] In actual processing, the sample annotation corresponding to the sample image can be understood as the target box category label. In this embodiment, the auxiliary branch identifies the potential distribution soft label of the corresponding behavior category based on the input data samples. For a sample with the target box category label of class i, the potential distribution of the corresponding target box in each behavior category, that is, the potential distribution soft label, is obtained by inputting it into N auxiliary branches. Thus, the embodiment of the present application realizes the classification loss constraint in the sample label space, and uses N single-category detectors to predict the potential distribution of the sample in each category. Among them, for a sample with the annotated category of i, the potential distribution is predicted by N - 1 single-category detectors other than detector i.
[0070] Step 130, perform feature extraction and feature fusion training on the main path model based on the confused data samples to obtain main path fusion features.
[0071] In a specific implementation, in addition to pre-training the auxiliary branch for feature extraction, this embodiment can also pre-train the main path model for feature extraction and feature fusion using data samples, so that the main path model can possess both feature extraction ability and feature fusion ability. Thus, when performing collaborative distillation training on the main path model and the auxiliary branch, the confused data samples can be input into the main path model and the auxiliary branch at the same time. Through feature recognition and extraction, the main path fusion features output by the main path model and the branch features output by the auxiliary branch are obtained, and the features output by the two different paths are used for subsequent collaborative distillation training.
[0072] Step 140, perform multi-dimensional loss constraint on the main path model according to the branch features, the main path fusion features, and the potential distribution soft label to obtain a divergence loss function and a distillation loss function.
[0073] In this embodiment, the divergence loss function is also called the KL (Kullback-Leibler Divergence) loss function. In this embodiment, the latent distribution soft label is used as the co-distillation learning training soft label for the main path detection algorithm, and the KL loss function is added to minimize the deviation between the confidence of the main path prediction target box and the latent distribution soft label. The latent distribution soft label is used to perform classification loss constraint on the sample label space to construct the KL loss function.
[0074] In actual implementation, for the branch features output by the auxiliary branch and the main path fusion features output by the main path model, this embodiment uses the distillation loss function to constrain both types of features at the same time, and makes the feature extraction ability of the target branch approach the N auxiliary branches by using the distillation loss function.
[0075] Step 150, based on the auxiliary branch loss function, the divergence loss function, and the distillation loss function, train the initial behavior detection model through preset weight information to obtain a target behavior detection model.
[0076] Wherein, the weight information is adjusted correspondingly based on the training cycle of the initial behavior detection model.
[0077] Specifically, the weight information may include the weight corresponding to the auxiliary branch and the weight corresponding to the main path model, and this embodiment does not limit this. In the training stage of the model, this embodiment adjusts the above two weights correspondingly according to the training cycle of the initial behavior detection model. Preferably, the two weights may be inversely proportional, so that during model training, as the training cycle changes, the training focus can gradually shift from the auxiliary branch to the main path model, or from the main path model to the auxiliary branch, and this embodiment does not limit this.
[0078] In specific implementation, after constructing the auxiliary branch loss function, the divergence loss function, and the distillation loss function in this embodiment, these loss functions can be combined together to construct a new constraint function for model training. As can be seen from the foregoing, different loss functions have different constraint meanings. In this embodiment, multiple loss functions are combined together to perform co-distillation learning on the auxiliary branch and the main path model, and the training focus of the auxiliary branch and the main path model is controlled by weights, so as to reduce the impact of the original fuzzy label on the model performance and effectively improve the accuracy of the behavior detection algorithm.
[0079] Step 160, perform feature recognition and behavior detection processing on the experimental video stream containing fuzzy behaviors through the target behavior detection model to obtain a behavior recognition result.
[0080] In this embodiment, after the model training of the behavior detection model is completed, the target behavior detection model with the highest recognition accuracy can be selected. The experimental video stream containing fuzzy behaviors or without model behaviors is input into the target behavior detection model, and the target behavior detection model performs processing such as feature recognition extraction and behavior detection on the input experimental video stream to obtain the behavior recognition result.
[0081] It can be seen that the embodiments of the present application aim to improve the behavior recognition accuracy of the behavior detection model for the input video stream. By obtaining the confused data samples and constructing an initial behavior detection model including a main path model and an auxiliary branch, and making the auxiliary branch correspond one-to-one with the behavior categories in the confused data samples, then training the auxiliary branch in the initial behavior detection model according to the confused data samples to obtain the auxiliary branch loss function, branch features, and the potential distribution soft label corresponding to the behavior category, so as to perform feature extraction and feature fusion training on the main path model based on the confused data samples to obtain the main path fusion features. Subsequently, the main path model is subjected to multi-dimensional loss constraints using the branch features, main path fusion features, and potential distribution soft labels to obtain the divergence loss function and the distillation loss function. Finally, based on the auxiliary branch loss function, divergence loss function, and distillation loss function, the initial behavior detection model is trained through the preset weight information to obtain the target behavior detection model. Thus, by controlling the weights of multiple loss functions, the collaborative distillation training of the auxiliary branch and the main path in the initial behavior detection model is realized. After the model training in this embodiment is completed to obtain the target behavior detection model, the target behavior detection model can be used to perform feature recognition and behavior detection processing on the experimental video stream containing fuzzy behaviors to obtain the behavior recognition result. Therefore, in this embodiment, the influence of the original fuzzy label on the model performance is reduced through collaborative distillation learning, effectively improving the detection accuracy and practical application effect of the behavior detection model for the confused label data set. In the application stage of the model, the model can be used to perform high-precision detection on the fuzzy behaviors in the input video stream, realizing fast processing and meeting the requirements of real-time monitoring, and solving the problem that the training method of the existing behavior detection model is limited by the quality of the labeled data set and cannot train a high-precision detection model.
[0082] Refer to Figure 2 , which shows a schematic flowchart of the steps of a behavior recognition method based on a behavior detection model applicable to fuzzy labels provided by an optional embodiment of the present application. The behavior recognition method based on the behavior detection model applicable to fuzzy labels may specifically include the following steps:
[0083] Step 210, construct an initial behavior detection model and obtain confused data samples.
[0084] Among them, the initial behavior detection model includes a main path model and at least one auxiliary branch, and the auxiliary branch corresponds one-to-one with the behavior categories in the confused data samples.
[0085] In actual implementation, refer to Figure 3 As shown, the behavior detection model in this embodiment is mainly composed of the main path model and the auxiliary branch. The RetinaNet algorithm is used as the main path for target detection. The auxiliary branch does not participate in the calculation during actual deployment. The basic detection algorithms used in this embodiment include but are not limited to common models such as RetinaNet, YOLO series, SSD, and RCNN series.
[0086] In this embodiment, the main path model can be composed of Backbone (feature extraction network), FPN (feature pyramid network) and Head (regression, classification head). The function is to input the image to the main path model to obtain the behavior target prediction box and category prediction score. In addition, combined with Figure 3 ,The auxiliary branch can also include Backbone and FPN for feature extraction and feature fusion.
[0087] In the behavior detection task, various behaviors are not mutually exclusive, and the detection algorithm is improved to multi-category detection. Therefore, assuming that the behavior categories are divided into N categories, there is a label 1 containing: [behavior 1, behavior 2, behavior 3...behavior n], the corresponding position of the existing behavior is set to 1, and the position without behavior is set to 0. For example, the multi-hot encoding corresponding to label 1 can be [1,0,1,...,0].
[0088] For auxiliary branches. Assuming that the behaviors are divided into N categories, this embodiment can use N single-category detectors (i.e., N auxiliary branches) to predict the potential distribution of samples in each category. Preferably, samples labeled with category i use N-1 single-category detectors other than detector i to predict the potential distribution. Taking into account computational efficiency, sharing low-level features (such as backbone and FPN in this embodiment), all branches have the same structure, and the Head module adopts the SSH (Single Stage Headless) model structure.
[0089] In this embodiment, N behavior categories can be understood as N types of detection tasks. In this embodiment, when training the model of the auxiliary branch, each auxiliary branch can be trained using a data sample of a separate category, that is, for the obfuscated data samples of the same behavior category, the corresponding auxiliary branch is input for training. Exemplarily, for the i-th behavior category, the obfuscated data samples of the i-th behavior category are input into the corresponding i-th auxiliary branch. At this time, the i-th auxiliary branch only trains the obfuscated data samples of the i-th behavior category, thereby effectively reducing the bias effect of long-tail data and avoiding the problem of insufficient model generalization ability caused by category imbalance.
[0090] Step 220: Extract the sample image and the sample label corresponding to the sample image from the obfuscated data sample, and input the sample image and the sample label into the corresponding auxiliary branch.
[0091] Step 230: Perform classification prediction and feature extraction on the sample image and the sample label through the auxiliary branch to obtain predicted anchor boxes, the predicted categories corresponding to the predicted anchor boxes, the soft labels of the potential distribution of the obfuscated data sample in the behavior category, and the branch features output by each auxiliary branch.
[0092] Step 240: For the predicted anchor boxes and the predicted categories, perform classification loss constraint on the sample image and the sample label through the auxiliary branch to obtain the auxiliary branch category loss function and the auxiliary branch anchor box regression loss function.
[0093] A unified description of Steps 220 - 240:
[0094] In actual implementation, referring to Figure 3 , the sample image and the sample label can be input into the auxiliary branch and the main path model simultaneously. The sample image can first pass through the auxiliary branch for feature extraction (exemplarily, ResNet50 can be used for feature extraction in the experiment), and then the extracted features are input into the FPN pyramid for feature fusion to obtain the feature feature map_i as the branch feature, which is input into N auxiliary branches. The auxiliary branches can share data with the main path model to construct the loss function.
[0095] In specific implementation, the auxiliary branch performs classification prediction on the input data to obtain multiple items of data, including but not limited to: predicted anchor boxes, the predicted categories corresponding to the predicted anchor boxes, soft labels of the potential distribution, the confidence corresponding to the predicted categories, and the branch features used to define the loss function subsequently. Specifically, in this embodiment, the data sample is input into the auxiliary branch, and each of the N auxiliary branches trains and detects one category. When the j-th auxiliary branch is training and calculating, only the images of the j-th category and the j-th category label are used for training. The loss of the auxiliary branch uses Focal Loss as the classification loss and CIoU loss as the anchor box loss. The predicted anchor boxes, predicted categories, confidence, etc. obtained by performing classification prediction on the input data by the auxiliary branch are used to define the auxiliary branch category loss function and the auxiliary branch anchor box regression loss function.
[0096] In this embodiment, the obfuscated data samples are input into the corresponding auxiliary branches according to the behavior categories, and the auxiliary branches predict and train the sample images and sample labels according to the input data samples. Since the output of the auxiliary branch is the anchor box coordinates and the predicted categories, to ensure that the auxiliary branch can accurately identify and extract features from the input data, this embodiment defines a loss function for the auxiliary branch and further defines it as a classification loss and an anchor box coordinate regression loss. Among them, the classification loss can be determined using the auxiliary branch category loss function, corresponding to the output predicted category; the anchor box coordinate regression loss can use the auxiliary branch anchor box regression loss function, corresponding to the output anchor box coordinates.
[0097] Based on the above analysis, at this time, the samples with the target box category label of class i are input into N branches to obtain the potential distribution of the corresponding target box in each behavior category.
[0098] In an alternative embodiment, the embodiment of the present application constrains the classification loss of the sample image and the sample label through the auxiliary branch to obtain an auxiliary branch category loss function and an auxiliary branch anchor box regression loss function, which may specifically include: according to the formula Determine the auxiliary branch category loss function; according to the formula Determine the auxiliary branch anchor box regression loss function; where is the auxiliary branch category loss function, p t is the confidence of the predicted box category, α is a parameter used to control the balance between positive and negative samples, γ is a modulation factor parameter, a parameter used to reduce the proportion of easily separable samples in model training, is the auxiliary branch anchor box regression loss function, Pred bbox is the predicted anchor box, GT box is the labeled anchor box, and the CIoU loss is used to minimize the deviation between the predicted anchor box and the labeled anchor box. In this embodiment, Pred bbox is obtained by the auxiliary branch detecting and classifying according to the input data, and the category loss function and the anchor box regression loss function of the auxiliary branch are used to improve the ability of the auxiliary branch to accurately classify the input data.
[0099] Preferably, α = 0.25 and γ = 2.
[0100] Step 250, based on the auxiliary branch category loss function and the auxiliary branch anchor box regression loss function, determine the auxiliary branch loss function.
[0101] In this embodiment, by combining the auxiliary branch class loss function and the auxiliary branch anchor box regression loss function, the loss function for training the auxiliary branch can be obtained. When each auxiliary branch can stably detect the target box coordinates and class confidence of the corresponding class, the training of the auxiliary branch is basically completed.
[0102] Step 260, according to the potential distribution soft label, perform spatial classification loss constraint on the main path model through a preset main path detection algorithm to obtain a divergence loss function.
[0103] Among them, the divergence loss function is used to minimize the deviation between the confidence of the predicted target box of the main path model and the potential distribution soft label.
[0104] Specifically, the divergence loss function can be understood as performing classification loss constraint from the sample label space. In this embodiment, the potential distribution is used as the soft label for collaborative distillation learning training of the main path detection algorithm, and the KL loss is used to minimize the deviation between the confidence of the main path predicted target box and the soft label.
[0105] In actual implementation, the main path model can predict the input data to obtain the confidence distribution of the predicted target box category. Using this confidence distribution, combined with the potential distribution soft label, the KL loss function is defined to minimize the deviation between the confidence of the main path predicted target box and the soft label.
[0106] Optionally, the above-mentioned performing spatial classification loss constraint on the main path model through a preset main path detection algorithm according to the potential distribution soft label to obtain a divergence loss function may specifically include the following sub-steps:
[0107] Sub-step 2601, obtain the branch predicted target box predicted by the auxiliary branch for the confused data sample, and obtain the main path predicted target box predicted by the main path model for the potential distribution soft label and the class confidence corresponding to the main path predicted target box.
[0108] Sub-step 2602, perform intersection over union (IoU) matching based on the branch predicted target box and the main path predicted target box, and calculate the divergence loss function according to the class confidence.
[0109] A unified description of sub-step 2601 - sub-step 2602:
[0110] In actual implementation, considering that the predicted boxes output by object detection are disordered, in this embodiment, the target boxes predicted by the auxiliary branch can be matched with the target boxes predicted by the main path through IoU (Intersection over Union), and then the KL loss is calculated for the classes of the boxes. Through the defined KL loss, the feature extraction capabilities of the main branch and the auxiliary branch are respectively improved.
[0111] In an alternative embodiment, the embodiment of the present application performs intersection over union (IoU) matching based on the branch prediction target box and the main path prediction target box, and calculates a divergence loss function according to the class confidence, which may specifically include: according to the formula calculate the divergence loss function, and according to perform IoU matching; where D KL is the KL loss function, p(x) is the latent distribution soft label, and q(x) is the class confidence distribution of the main path prediction target box.
[0112] Step 270, perform loss constraint on the main path model in terms of the feature space dimension based on the branch feature and the main path fusion feature, to obtain a distillation loss function.
[0113] In a specific implementation, the above KL loss function has performed classification loss constraint in the sample label space. For the object detection algorithm, this embodiment additionally considers the feature space dimension, and uses the similarity-preserving knowledge distillation loss to constrain the feature maps of the main path and the auxiliary branches, so that similar semantic features can be extracted from the image on different branches.
[0114] In an alternative embodiment, the embodiment of the present application performs loss constraint on the main path model in terms of the feature space dimension based on the branch feature and the main path fusion feature, to obtain a distillation loss function, which may specifically include: according to the formula perform loss constraint in the feature space dimension to obtain a distillation loss function; where L SP is the distillation loss, G main is the main path fusion feature, G bi is the feature of each branch, and the distillation loss function is used to make the ability of the target branch to extract features approach multiple auxiliary branches.
[0115] Step 280, based on the auxiliary branch loss function, the divergence loss function, and the distillation loss function, perform model training on the initial behavior detection model through preset weight information, to obtain a target behavior detection model.
[0116] Wherein, the weight information is correspondingly adjusted based on the training period of the initial behavior detection model.
[0117] In an alternative embodiment, when the weight information in the embodiments of the present application includes the training weight of the auxiliary branch and the training weight of the main path, the above-mentioned initial behavior detection model is trained based on the auxiliary branch loss function, the divergence loss function, and the distillation loss function through preset weight information, which may specifically include: According to Determine the auxiliary branch loss function; According to Determine the main branch loss function; According to Determine the calculated loss function and train the initial behavior detection model; where, is the main branch class classification loss, ε is the training weight of the auxiliary branch, (1 - ε) is the training weight of the main path, the training weight of the auxiliary branch and the training weight of the main path are in an inverse relationship, and are adjusted based on the training cycle of the initial behavior detection model.
[0118] In a specific implementation, the training weight of the auxiliary branch and the training weight of the main path can be adjusted as the training cycle (epoch) increases. In the initial stage of model training, since the auxiliary branch is still exploring the potential distribution of the data, a higher weight can be assigned to promote its learning and improve the feature extraction ability of the auxiliary branch. At this time, the training weight of the auxiliary branch can be greater than the training weight of the main path. As the training progresses, once the performance of the auxiliary branch is significantly improved, the focus of training can gradually shift to the main path to further improve the overall performance of the model. Thus, the embodiments of the present application achieve reducing the impact of the original fuzzy labels on the model performance through collaborative distillation learning, and effectively improve the accuracy of the behavior detection algorithm.
[0119] Specifically, in the overall algorithm training, in the initial stage, the value of ε can be increased, and the loss function of the auxiliary branch is mainly trained. When each auxiliary branch can stably detect the target box coordinates and class confidence of the corresponding class, the training of the auxiliary branch is basically completed; at this time, the algorithm starts to adjust the value of ε to make the loss function gradually tend to train the main branch. Among them, the losses on the main branch include the regression loss between the target box and the labeled box class loss the loss between the intermediate layer feature map and the auxiliary branch feature map and the class loss of the target boxes predicted by the main branch and the auxiliary branch
[0120] It should be noted that when calculating the KL loss, if on a single image, the auxiliary branch predicts two bounding boxes with classes A1 and B1 respectively, and the main path also predicts two bounding boxes, A2 and B2 respectively. When calculating the class loss, it can be calculated for A1 and A2, and B1 and B2. Through the IoU matching method, the bounding boxes of the corresponding objects predicted by the auxiliary and main paths are matched, and then the loss is calculated.
[0121] Step 290, perform feature recognition and behavior detection processing on the experimental video stream containing fuzzy behaviors through the target behavior detection model to obtain a behavior recognition result.
[0122] In summary, in the training stage of the behavior detection model of the present application, by constructing an initial behavior detection model and obtaining confused data samples, the confused data samples are input into the corresponding auxiliary branches. Through the auxiliary branches, classification prediction and feature extraction are performed on the sample images and sample labels in the confused data samples to obtain predicted anchor boxes, predicted classes corresponding to the predicted anchor boxes, soft labels of the potential distribution of the confused data samples in terms of behavior classes, and branch features output by each auxiliary branch. Then, for the predicted anchor boxes and predicted classes, classification loss constraints are imposed on the sample images and sample labels through the auxiliary branches to obtain an auxiliary branch class loss function and an auxiliary branch anchor box regression loss function. Subsequently, based on the soft labels of the potential distribution, spatial classification loss constraints are imposed on the main path model through a preset main path detection algorithm to obtain a divergence loss function, and loss constraints are imposed on the main path model in terms of the feature space dimension based on the branch features and the main path fusion features to obtain a distillation loss function. Finally, based on the auxiliary branch loss function, the divergence loss function, and the distillation loss function, the initial behavior detection model is trained through preset weight information to obtain a target behavior detection model that can accurately perform feature recognition and behavior detection on video streams containing fuzzy behaviors, thereby realizing the use of fuzzy labels to train the behavior detection model. By collaborative distillation learning, the impact of the original fuzzy labels on the model performance is reduced, the accuracy of the behavior detection algorithm is effectively improved, and the impact of the original fuzzy labels on the model performance is reduced. For poorly labeled data sets, a behavior detection model with high accuracy can also be trained to meet the actual use requirements, effectively improving the detection accuracy and practical use effect of the behavior detection model for confused label data sets, and further improving the accuracy of the behavior detection algorithm and the behavior detection model, solving the problem of insufficient behavior recognition and detection accuracy caused by the inability to train a high-precision recognition model in existing behavior recognition methods. In addition, through an innovative model training method, the present application solves: ① the problem that the training method of the existing behavior detection model is limited by the quality of the labeled data set and cannot train a high-precision detection model; ② the problem that the existing technology can only continuously repeat the inspection of data annotation to reduce fuzzy labels in order to improve the quality of the data set, resulting in a waste of human resources and unable to meet the actual needs at the same time.
[0123] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of the present application are not limited by the described action sequences, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously.
[0124] As Figure 4 shown, the embodiments of the present application also provide a behavior recognition device 400 based on a behavior detection model applicable to fuzzy labels, including:
[0125] A construction module 410, configured to construct an initial behavior detection model and obtain confused data samples, where the initial behavior detection model includes a main path model and at least one auxiliary branch, and the auxiliary branches correspond one-to-one to the behavior categories in the confused data samples;
[0126] A first training module 420, configured to train the auxiliary branches according to the confused data samples to obtain an auxiliary branch loss function, branch features, and potential distribution soft labels corresponding to the behavior categories;
[0127] An extraction module 430, configured to perform feature extraction and feature fusion training on the main path model based on the confused data samples to obtain main path fusion features;
[0128] A loss constraint module 440, configured to perform multi-dimensional loss constraints on the main path model according to the branch features, the main path fusion features, and the potential distribution soft labels to obtain a divergence loss function and a distillation loss function;
[0129] A second training module 450, configured to perform model training on the initial behavior detection model based on the auxiliary branch loss function, the divergence loss function, and the distillation loss function through preset weight information to obtain a target behavior detection model, and the weight information is adjusted correspondingly based on the training cycle of the initial behavior detection model;
[0130] A behavior recognition module 460, configured to perform feature recognition and behavior detection processing on an experimental video stream containing fuzzy behaviors through the target behavior detection model to obtain a behavior recognition result.
[0131] Optionally, the first training module 420 includes:
[0132] An extraction sub-module, configured to extract sample images and sample labels corresponding to the sample images from the confused data samples, and input the sample images and the sample labels into the corresponding auxiliary branches;
[0133] The classification prediction and feature extraction sub-module is used to perform classification prediction and feature extraction on the sample image and the sample label through the auxiliary branch, and obtain the predicted anchor box, the predicted category corresponding to the predicted anchor box, the potential distribution soft label of the confusion data sample on the behavior category, and the branch feature output by each auxiliary branch;
[0134] The classification loss constraint sub-module is used to perform classification loss constraint on the sample image and the sample label through the auxiliary branch for the predicted anchor box and the predicted category, and obtain the auxiliary branch category loss function and the auxiliary branch anchor box regression loss function;
[0135] The auxiliary branch loss function determination sub-module is used to determine the auxiliary branch loss function based on the auxiliary branch category loss function and the auxiliary branch anchor box regression loss function.
[0136] Optionally, the classification loss constraint sub-module is specifically used to: according to the formula Determine the auxiliary branch category loss function; according to the formula Determine the auxiliary branch anchor box regression loss function; where, is the auxiliary branch category loss function, p t is the confidence of the predicted box category, α is a parameter used to control the positive and negative sample balance, γ is a modulation factor parameter, a parameter used to reduce the proportion of easy-to-separate samples in model training, is the auxiliary branch anchor box regression loss function, Pred bbox is the predicted anchor box, GT box is the labeled anchor box, and the CIoU loss is used to minimize the deviation between the predicted anchor box and the labeled anchor box.
[0137] Optionally, the loss constraint module 440 includes:
[0138] The first loss constraint sub-module is used to perform spatial classification loss constraint on the main path model through a preset main path detection algorithm according to the potential distribution soft label, and obtain the divergence loss function, which is used to minimize the deviation between the confidence of the predicted target box of the main path model and the potential distribution soft label;
[0139] The second loss constraint sub-module is used to perform loss constraint on the main path model in the feature space dimension based on the branch feature and the main path fusion feature, and obtain the distillation loss function.
[0140] Optionally, the first loss constraint sub-module includes:
[0141] A prediction unit, configured to obtain the branch prediction target box predicted by the auxiliary branch for the confused data sample, and obtain the main path prediction target box predicted by the main path model for the potential distribution soft label and the class confidence corresponding to the main path prediction target box;
[0142] An intersection over union (IoU) matching unit, configured to perform IoU matching based on the branch prediction target box and the main path prediction target box, and calculate a divergence loss function according to the class confidence.
[0143] Optionally, the IoU matching unit is specifically configured to: calculate the divergence loss function according to the formula and perform IoU matching according to ; where D KL is the KL loss function, p(x) is the potential distribution soft label, and q(x) is the class confidence distribution of the main path prediction target box.
[0144] Optionally, the second loss constraint sub-module is specifically configured to: perform loss constraint in the feature space dimension according to the formula to obtain a distillation loss function; where L SP is the distillation loss, G main is the main path fusion feature, G bi is the feature of each branch, and the distillation loss function is used to make the ability of the target branch to extract features approach multiple auxiliary branches.
[0145] Optionally, the weight information includes the training weight of the auxiliary branch and the training weight of the main path. The second training module 450 is specifically configured to: determine the loss function of the auxiliary branch according to ; determine the loss function of the main branch according to ; determine the calculation loss function according to and perform model training on the initial behavior detection model; where, is the main branch class classification loss, is the main branch regression loss, is the auxiliary branch feature map loss, ε is the training weight of the auxiliary branch, (1 - ε) is the training weight of the main path, the training weight of the auxiliary branch and the training weight of the main path are in an inverse relationship, and are adjusted based on the training cycle of the initial behavior detection model.
[0146] It should be noted that the behavior recognition device based on the behavior detection model applicable to fuzzy labels provided in the embodiments of the present application can execute the behavior recognition method based on the behavior detection model applicable to fuzzy labels provided in any embodiment of the present application, and has the corresponding functions and beneficial effects for executing the method.
[0147] In a specific implementation, the above-described behavior recognition device based on a behavior detection model applicable to fuzzy labels can be integrated into a device, enabling the device to be based on existing object detection and anomaly detection algorithms, combined with N auxiliary branches and a main path, and complete the training of the auxiliary branches and the collaborative distillation training of the main path by controlling the loss function weights. As an electronic device, it can improve the detection accuracy and practical application effect of the basic algorithm for the confused label dataset. The electronic device can be composed of two or more physical entities or a single physical entity. For example, the electronic device can be a personal computer (PC), a computer, a server, etc., and the embodiments of the present application do not make specific limitations on this.
[0148] As Figure 5 shown, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114; the memory 113 is used to store a computer program; when the processor 111 executes the program stored on the memory 113, it implements the steps of the behavior recognition method based on a behavior detection model applicable to fuzzy labels provided by any one of the foregoing method embodiments. Exemplarily, the steps of the behavior recognition method based on a behavior detection model applicable to fuzzy labels may include the following steps: constructing an initial behavior detection model and obtaining confused data samples, where the initial behavior detection model includes a main path model and at least one auxiliary branch, and the auxiliary branch corresponds one-to-one with the behavior categories in the confused data samples; training the auxiliary branch according to the confused data samples to obtain an auxiliary branch loss function, branch features, and potential distribution soft labels corresponding to the behavior categories; performing feature extraction and feature fusion training on the main path model based on the confused data samples to obtain main path fusion features; performing multi-dimensional loss constraints on the main path model according to the branch features, the main path fusion features, and the potential distribution soft labels to obtain a divergence loss function and a distillation loss function; training the initial behavior detection model based on the auxiliary branch loss function, the divergence loss function, and the distillation loss function through preset weight information to obtain a target behavior detection model, and the weight information is adjusted correspondingly based on the training cycle of the initial behavior detection model; performing feature recognition and behavior detection processing on an experimental video stream containing fuzzy behaviors through the target behavior detection model to obtain a behavior recognition result.
[0149] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the behavior recognition method based on a behavior detection model applicable to fuzzy labels provided in any of the foregoing method embodiments are implemented.
[0150] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0151] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A behavior recognition method based on a behavior detection model applicable to fuzzy labels, characterized in that, Including: Constructing an initial behavior detection model and obtaining obfuscated data samples, where the initial behavior detection model includes a main path model and at least one auxiliary branch, and the auxiliary branches correspond one-to-one with the behavior categories in the obfuscated data samples; Training the auxiliary branches according to the obfuscated data samples to obtain auxiliary branch loss functions, branch features, and potential distribution soft labels corresponding to the behavior categories; Performing feature extraction and feature fusion training on the main path model based on the obfuscated data samples to obtain main path fusion features; Constraining the main path model with multi-dimensional losses according to the branch features, the main path fusion features, and the potential distribution soft labels to obtain a divergence loss function and a distillation loss function; Training the initial behavior detection model based on the auxiliary branch loss functions, the divergence loss function, and the distillation loss function through preset weight information to obtain a target behavior detection model, where the weight information is adjusted correspondingly based on the training cycle of the initial behavior detection model; Performing feature recognition and behavior detection processing on an experimental video stream containing ambiguous behaviors through the target behavior detection model to obtain behavior recognition results.
2. The method according to claim 1, wherein The training the auxiliary branches according to the obfuscated data samples to obtain auxiliary branch loss functions, branch features, and potential distribution soft labels corresponding to the behavior categories includes: Extracting sample images and sample labels corresponding to the sample images from the obfuscated data samples, and inputting the sample images and the sample labels into the corresponding auxiliary branches; Classifying and predicting and extracting features from the sample images and the sample labels through the auxiliary branches to obtain predicted anchor boxes, predicted categories corresponding to the predicted anchor boxes, potential distribution soft labels of the obfuscated data samples in the behavior categories, and branch features output by each auxiliary branch; Performing classification loss constraint on the sample images and the sample labels through the auxiliary branches for the predicted anchor boxes and the predicted categories to obtain an auxiliary branch category loss function and an auxiliary branch anchor box regression loss function; Determining an auxiliary branch loss function based on the auxiliary branch category loss function and the auxiliary branch anchor box regression loss function.
3. The method according to claim 2, wherein The performing classification loss constraint on the sample images and the sample labels through the auxiliary branches to obtain an auxiliary branch category loss function and an auxiliary branch anchor box regression loss function includes: Determine the loss function of the auxiliary branch category according to the formula According to the formula to determine the auxiliary branch anchor box regression loss function; Among them, is the loss function of the auxiliary branch category, p t is the confidence of the predicted box category, α is the parameter used to control the balance of positive and negative samples, γ is the modulation factor parameter, which is used to reduce the proportion of easy-to-separate samples in model training. is the regression loss function of the auxiliary branch anchor box, Pred bbox is the predicted anchor box, GT box is the labeled anchor box, and the CIoU loss is used to minimize the deviation between the predicted anchor box and the labeled anchor box.
4. The method according to claim 1, wherein The constraining the main path model with multi-dimensional losses according to the branch features, the main path fusion features, and the potential distribution soft labels to obtain a divergence loss function and a distillation loss function includes: Constraining the main path model with a spatial classification loss according to the potential distribution soft labels through a preset main path detection algorithm to obtain a divergence loss function, where the divergence loss function is used to minimize the deviation between the confidence of the predicted target boxes of the main path model and the potential distribution soft labels; Constraining the main path model with a loss in the feature space dimension based on the branch features and the main path fusion features to obtain a distillation loss function.
5. The method according to claim 4, characterized in that Constraining the spatial classification loss of the main path model according to the potential distribution soft label by a preset main path detection algorithm to obtain a divergence loss function, including: Obtaining the branch prediction target boxes predicted by the auxiliary branch for the confused data samples, and obtaining the main path prediction target boxes predicted by the main path model for the potential distribution soft label and the class confidence corresponding to the main path prediction target boxes; Performing intersection over union (IoU) matching based on the branch prediction target boxes and the main path prediction target boxes, and calculating the divergence loss function according to the class confidence.
6. The method according to claim 5, characterized in that, The performing intersection over union (IoU) matching based on the branch prediction target boxes and the main path prediction target boxes, and calculating the divergence loss function according to the class confidence, including: Calculate the divergence loss function according to the formula and perform intersection over union matching according to ; Among them, D KL is the KL loss function, p(x) is the soft label of the latent distribution, and q(x) is the confidence distribution of the predicted target box category on the main path.
7. The method according to claim 6, characterized in that Constraining the loss of the main path model in the feature space dimension based on the branch features and the main path fusion features to obtain a distillation loss function, including: According to the formula perform loss constraint on the feature space dimension to obtain the distillation loss function; Among them, L SP is the distillation loss, G main is the main path fusion feature, G bi is the feature of each branch, and the distillation loss function is used to make the feature extraction ability of the target branch approach that of multiple auxiliary branches.
8. The method according to claim 7, wherein The weight information includes the auxiliary branch training weight and the main path training weight. Based on the auxiliary branch loss function, the divergence loss function, and the distillation loss function, training the initial behavior detection model by preset weight information, including: According to Determine the loss function of the auxiliary branch According to Determine the main branch loss function; According to Determine the calculation of the loss function and perform model training on the initial behavior detection model; Among them, is the main branch category classification loss, is the main branch regression loss, is the auxiliary branch feature map loss, ε is the auxiliary branch training weight, (1 - ε) is the main path training weight. The auxiliary branch training weight and the main path training weight are in an inverse relationship and are adjusted based on the training cycle of the initial behavior detection model.
9. A behavior recognition device based on a behavior detection model applicable to fuzzy labels, characterized in that, Including: A construction module, configured to construct an initial behavior detection model and obtain confused data samples. The initial behavior detection model includes a main path model and at least one auxiliary branch, and the auxiliary branches correspond to the behavior categories in the confused data samples one by one; A first training module, configured to train the auxiliary branches according to the confused data samples to obtain an auxiliary branch loss function, branch features, and the potential distribution soft label corresponding to the behavior categories; An extraction module, configured to perform feature extraction and feature fusion training on the main path model based on the confused data samples to obtain main path fusion features; A loss constraint module, configured to perform multi-dimensional loss constraints on the main path model according to the branch features, the main path fusion features, and the potential distribution soft label to obtain a divergence loss function and a distillation loss function; A second training module, configured to train the initial behavior detection model by preset weight information based on the auxiliary branch loss function, the divergence loss function, and the distillation loss function to obtain a target behavior detection model, where the weight information is adjusted correspondingly based on the training period of the initial behavior detection model; A behavior recognition module, configured to perform feature recognition and behavior detection processing on an experimental video stream including fuzzy behaviors through the target behavior detection model to obtain a behavior recognition result.
10. An electronic device, characterized in that, Including a processor, a communication interface, a memory, and a communication bus. The processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used for storing a computer program; The processor, when executing the program stored on the memory, implements the steps of the behavior recognition method based on the behavior detection model applicable to fuzzy labels according to any one of claims 1-8.
Citation Information
Patent Citations
Knowledge distillation-based recognition model training method and device
CN116206275A
Sewer line defect detection method and system based on confusion matrix
CN117893818A