Multi-person Head Pose Estimation Method and Device Based on Multi-task Loss Balance
By using the YOLO network based on the Mamba model and the dynamic weight loss function in the multi-person head pose estimation, the problems of low detection efficiency and insufficient feature extraction capabilities are solved, and efficient, real-time and accurate multi-person head pose estimation is achieved.
Patent Information
- Application Number
- CN202510408793.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The prior art has low detection efficiency in multi-person head pose estimation, insufficient feature extraction capability of the CNN model, and imbalance in multi-task loss value affects model accuracy and reliability.
The YOLO network based on the Mamba model is used to build a multi-person head pose estimation model, combining the PAFPN module and detection head, and using a loss function of dynamic weights to balance multi-task losses, improving feature extraction capabilities and model accuracy.
It realizes efficient and real-time multi-person head posture estimation, improves detection accuracy and reliability, and facilitates engineering deployment and application.
Smart Images

Figure CN119919500B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of head pose estimation, and particularly to a multi-person head pose estimation method and device based on multi-task loss balance. Background Art
[0002] Head pose is a very important factor for judging human behavior and is also of great help for understanding human behavior. Head pose estimation refers to a computer analyzing and predicting an input image or video sequence to determine the position and pose parameters of a person's head in three-dimensional space. Head pose estimation can be applied in various application fields such as human-computer interaction, biometrics, virtual reality, and fatigue detection.
[0003] Currently, head pose estimation can be mainly divided into two directions: single-person head pose estimation and multi-person head pose estimation. Most of the research mainly focuses on single-person head pose estimation. Although the single-person head pose estimation algorithm has high accuracy, in the multi-person head pose estimation scenario, its detection efficiency is low, which is not conducive to engineering deployment and application. For multi-person head pose estimation, in the existing technology, a single-stage multi-person head pose estimation algorithm based on deep learning models such as CNN (Convolutional Neural Network) is usually adopted, which can improve the efficiency of the multi-person head pose estimation algorithm. However, the backbone feature extraction ability of the CNN model is weak, and it cannot fully extract the detailed features of the head region. Moreover, since multiple learning tasks such as class prediction, localization, and pose estimation need to be performed during the detection process, and fixed weights are usually adopted for each task algorithm during the learning process, this will lead to an imbalance between the loss values of each learning task, resulting in an impact on the accuracy and reliability of the model's head pose estimation. Summary of the Invention
[0004] The technical problem to be solved by the present invention lies in: aiming at the technical problems existing in the prior art, the present invention provides a multi-person head pose estimation method and device based on multi-task loss balance with a simple implementation method, low cost, high detection accuracy and reliability, and strong real-time performance, which can realize end-to-end multi-person head pose estimation based on multi-task loss balance, and improve the estimation accuracy while maintaining the detection real-time performance.
[0005] To solve the above technical problem, the technical solution proposed by the present invention is:
[0006] A multi-person head pose estimation method based on multi-task loss balance, the steps include:
[0007] Build a multi-person head pose estimation model using the YOLO network based on the Mamba model. The multi-person head pose estimation model includes a backbone network, a PAFPN (Pyramid Attention Feature Pyramid Network) module, and a detection head. The backbone network is an ODMamba network structure used to process the input image and extract preliminary features. The PAFPN module is used to extract and fuse features of different scales. The detection head is used to perform tasks such as class prediction, head position detection, and head pose estimation;
[0008] Train the multi-person head pose estimation model using an image training set containing multi-person head objects. During the training process, use a loss function based on dynamic weights to control the training process. The loss function based on dynamic weights is obtained by weighting the loss functions of the head object prediction classification task, the head position detection task, and the head pose estimation task using dynamic weights. The dynamic weights of each task are determined according to the loss values of the corresponding tasks to balance the losses of each task;
[0009] Obtain a to-be-detected picture containing multiple people's heads, and input it into the trained multi-person head pose estimation model to obtain the detection results of multiple head objects.
[0010] Further, the ODMamba network structure includes a picture division module, an ODSSBlock module, a visual clue merging module, and an SPPF module. The picture division module is used to preliminarily process the input image through multiple convolutional layers. The ODSSBlock module is used to further extract features. The visual clue merging module is used to further fuse and optimize the features output by the ODSSBlock module. The SPPF module is used to aggregate features at multiple scales. The core module in the PAFPN module is the ODSSBlock module. Convolution operations are performed using ACConv convolution at the input end of the ODSSBlock module, and depthwise separable DW-ACConv convolution is used for depthwise separable convolution operations in the LS Block module, the SS2D module, and the RG Block module in the ODSSBlock module. The SS2D module is used to scan and merge image features to extract multi-directional global information. The LS Block module is used to extract local spatial information. The RG Block module is used to capture local dependencies by combining a gating mechanism and a residual connection.
[0011] Further, the calculation expression of the loss function based on dynamic weights is:
[0012]
[0013] Among them, represents the loss function based on dynamic weights, represents the batch size of the multi-person head pose estimation model during training, and and respectively represent and and dynamic weights. The values of the dynamic weights of each dynamic weight change with the change rate of the loss value of the training epoch. represents the loss function of the head position detection task, represents the loss function of the class prediction task, represents the loss function of the head pose estimation task.
[0014] Furthermore, the dynamic weights of each branch are calculated according to the following expression:
[0015]
[0016]
[0017] where represents the dynamic weight of the p-th task, p = 1, 2, 3, corresponding to the class prediction task, the head position detection task, and the head pose estimation task respectively. t represents the number of training times, and K represents the total number of tasks. represents the loss value of the p-th task, T represents a preset temperature used to control the softness of task weighting.
[0018] Furthermore, the loss function of the head position detection task is constructed based on the focal regression box loss of gamma transformation, and the calculation expression is:
[0019]
[0020]
[0021]
[0022]
[0023] where N represents the number of head detections in the image to be detected, and are the i-th head prediction detection box and the ground truth label respectively, represents the regression loss based on Siou, represents a preset threshold with a value range of (0, 1), represents the loss value of the center point coordinates and the box length in the head detection box. Represents the focal regression box loss based on gamma transformation, Represents the intersection over union (IoU) between the head region detection box and the ground truth head box, Represents the IoU between the reconstructed head region detection box and the ground truth head box, Represents the gamma transformation adjustment factor, Represents the gamma transformation adjustment factor.
[0024] Furthermore, the loss function of the head object prediction classification task is calculated according to the following formula:
[0025]
[0026] Where, Represents the predicted head probability of the i-th head detection box, Represents the true head label of the i-th head detection box, Represents the overlap degree between the i-th head detection box and the ground truth head box.
[0027] Furthermore, the loss function of the head pose estimation task is constructed based on the tangent function, and the calculation expression is:
[0028]
[0029]
[0030]
[0031] Where, i represents the i-th dimension of the head pose matrix, Represents the loss weight on the i-th dimension of the head pose matrix, Represents the head pose estimation matrix, Represents the true head pose matrix, 、 Represents 、 Elements in.
[0032] Furthermore, it also includes the step of obtaining the head pose estimation matrix , including:
[0033] Using the 6D rotation representation parameters of the head pose 6D decoupled detection head set in the multi-person head pose estimation model to detect the head pose ;
[0034] Converting the two three-dimensional vectors corresponding to the 6D rotation representation parameters of the head pose of each head object 、 to obtain the elements in the head pose estimation matrix :
[0035] ; ; ;
[0036] Among them, are respectively u at x , y , z the coordinate values on the axis, are respectively v at x , y , z the coordinate values on the axis, , , respectively represent the three three-dimensional unit vectors in the 3×3 rotation matrix formed by transforming the 6D rotation representation parameter matrix , , , , ~ respectively represent , , the elements in, represents a vector perpendicular to in three-dimensional space.
[0037] An electronic device, comprising a processor and a memory, the memory is used for storing a computer program, and the processor is used for executing the computer program to execute the method as described above.
[0038] A computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the method as described above.
[0039] Compared with the prior art, the advantages of the present invention are as follows:
[0040] 1. By adopting the YOLO architecture based on the Mamba model as the basic model for head pose estimation, the present invention constructs a multi-person head pose estimation model, which can expand the context semantic connection of each pixel, improve the feature extraction ability of the head pose estimation model for the head region, and at the same time adopts a dynamic weighting-based strategy to determine the loss value to balance the loss values between various tasks, reduce the impact of the imbalance between the loss values of various tasks in the head pose estimation model on the model accuracy, and realize end-to-end multi-person head pose estimation based on multi-task loss balance, which can improve the accuracy of the end-to-end multi-person head pose estimation algorithm while maintaining its detection real-time performance, and is convenient for actual engineering deployment and application.
[0041] 2. For the loss function of the further head position detection task of the present invention, by adopting the focal regression box loss based on gamma transformation, it is possible to focus on regression samples of different difficulties, enhance the model's attention to the regression boxes of difficult samples, improve the regression accuracy of the head detection boxes, and thus effectively improve the performance of the detector in different detection tasks.
[0042] 3. For the loss function of the further head pose estimation task of the present invention, it is constructed based on the tangent function, which can effectively update and optimize the head pose estimation parameters, enable the model to pay more attention to the learning of difficult dimension vectors, and further improve the head pose estimation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 FIG. is a schematic flowchart of the implementation of the multi-person head pose estimation method based on multi-task loss balance in this embodiment.
[0044] Figure 2 FIG. is a schematic diagram of the structural principle of the multi-person head pose estimation model constructed using the YOLO network based on the Mamba model in this embodiment.
[0045] Figure 3 FIG. is a schematic diagram of the structural principle of the ODSSBlock module in this embodiment.
[0046] Figure 4 FIG. is a schematic diagram of the structural principle of the LS Block module in this embodiment.
[0047] Figure 5 FIG. is a schematic diagram of the structural principle of the SS2D module in this embodiment.
[0048] Figure 6 FIG. is a schematic diagram of the structural principle of the RG Block module in this embodiment.
[0049] Figure 7 FIG. is a schematic diagram of the structural principle during the training of ACConv in this embodiment.
[0050] Figure 8 FIG. is a schematic diagram of the structural principle during the inference of ACConv in this embodiment.
[0051] Figure 9 FIG. is a schematic diagram of the principle of convolutional reparameterization during the inference of ACConv in this embodiment.
[0052] Figure 10 FIG. is a schematic diagram of the structural principle of the Scan module in this embodiment.
[0053] Figure 11 FIG. is a schematic diagram of the tangent function curve.
[0054] Figure 12It is a schematic diagram of the implementation process of the non-maximum suppression algorithm that does not depend on confidence and IoU and is used in this embodiment for post-processing. Detailed implementation manners
[0055] The present invention will be further described below in conjunction with the accompanying drawings of the specification and specific preferred embodiments, but the protection scope of the present invention is not limited thereby.
[0056] For ease of understanding, first, the relevant technical background related to the present invention will be introduced exemplarily.
[0057] Taking the multi-person head pose estimation algorithm based on Yolov5 in the prior art as an example, this type of method adds an Euler angle detection head branch on the basis of YOLOv5 and uses MSE (mean square error) as the loss function to optimize the parameters of the head pose prediction branch, so as to realize the training and recognition of the end-to-end multi-person head pose recognition model, which can simplify the training and recognition process of the head pose recognition algorithm model, improve the real-time performance on various platforms, and also prove that embedding the head position detection and head pose estimation into a feature representation layer is beneficial for the network to learn the internal connection between the two. However, this algorithm model has the following problems:
[0058] 1. The feature extraction ability of the model backbone is weak
[0059] In the downstream tasks of object detection, mainly CNN and Transformer structures are adopted. CNN and its improvements can provide fast execution speed while ensuring accuracy, but the convolutional operation of CNN can only perceive the features of the local area in each layer and is difficult to capture long-range dependencies. For a multi-task model such as multi-person head pose estimation, this will result in it being difficult to establish context connections between the target pixels and a larger range of surrounding pixels, which is not conducive to the model to perform more effective feature extraction on the head area, thereby affecting the performance of the end-to-end multi-person head pose estimation model. Transformer performs well in global modeling, can effectively capture long-range dependencies, and significantly expands the receptive field of the model. For example, the DETR series relies on the powerful global modeling ability of the self-attention mechanism to solve the problem of the small receptive field of CNN and improve the model's feature extraction ability. However, the quadratic complexity of its self-attention mechanism will increase the computational burden of the model, which is not conducive to the real-time processing and engineering deployment of the multi-person head pose estimation model.
[0060] 2. The imbalance between the loss values of multiple learning tasks
[0061] End-to-end multi-person head pose estimation algorithms usually involve multiple learning tasks such as classification, localization, and pose estimation. In traditional end-to-end multi-person head pose estimation algorithms, the loss values of each learning task are usually directly summed with fixed weights as the final loss value, ignoring the imbalance problem between the loss values of each learning task, which in turn affects the model accuracy.
[0062] To address the above problems of the one-stage multi-person head pose estimation method in the prior art, the present invention constructs a multi-person head pose estimation model by using the YOLO architecture based on the Mamba model as the basic model for head pose estimation, which can expand the contextual semantic connections of each pixel and improve the feature extraction ability of the head pose estimation model for the head region, solving the limitations of models such as traditional CNN (with local receptive field limitations in capturing long-distance information) or Transformer (excellent in global modeling but high in computational complexity). At the same time, a strategy based on dynamic weighting is adopted to determine the loss value to balance the loss values between various tasks, reducing the impact of the imbalance between the loss values of various tasks in the head pose estimation model on the model accuracy, realizing end-to-end multi-person head pose estimation based on multi-task loss balance, improving the accuracy of the end-to-end multi-person head pose estimation algorithm while maintaining its detection real-time performance, and facilitating practical engineering deployment and application.
[0063] As Figure 1 shown, the steps of the multi-person head pose estimation method based on multi-task loss balance in this embodiment include:
[0064] Step S01. Use the YOLO network based on the Mamba model to construct a multi-person head pose estimation model. The multi-person head pose estimation model includes a backbone network, a PAFPN module, and a detection head. The backbone network is an ODMamba network structure for processing the input image and extracting preliminary features. The PAFPN module is used to extract and fuse features of different scales. The detection head is used to implement class prediction, human head position detection, and head pose parameter estimation.
[0065] The Mamba model is a method based on the State Space Model (Selective State Space Model, SSM). By dynamically adjusting the model behavior through context awareness of the input content, it can effectively capture long-range dependencies while effectively alleviating the problems brought by quadratic complexity, enabling the multi-person head pose estimation model to have both local feature extraction capabilities, global modeling capabilities, and maintain a certain real-time performance. In this embodiment, by adopting the YOLO network based on the Mamba model, a new backbone network Mamba-YOLO that combines the advantages of SSM and CNN is formed. The feature extraction backbone of the network uses the Mamba structure as the basic network of the end-to-end multi-person head pose estimation model, making it have a strong feature extraction ability for the head region while maintaining a certain inference real-time performance, thus facilitating engineering deployment.
[0066] As Figure 2 shown, the multi-person head pose estimation model constructed using the YOLO network based on the Mamba model in this embodiment includes a backbone network, a PAFPN module, and a detection head. The backbone network (feature extraction network part) of Mamba-YOLO is the ODMamba network structure (ODMamba Backbone), which is used to process the input image and extract preliminary features. It is the basic part of feature extraction and is used to generate sufficient information for subsequent modules to further process. The PAFPN module is a feature pyramid network used to fuse and refine features of different scales. It combines upsampling and downsampling with additional convolutional layers to enhance the model's detection ability for objects of various scales, especially improving the detection ability in multi-scale object detection tasks. The model sets three different detection layers according to the different sizes of the feature maps, which are respectively responsible for detecting small, medium, and large-sized head object targets (each grid includes a total of 11 channels for the head prediction probability Cls, the position detection box Box, and the head pose parameter Head Pos) and generating a dense head object output. Each detection head uses a decoupled detection head to perform three tasks: class prediction, human head position detection, and head pose parameter estimation in three paths using sliding convolution. The method of using a decoupled detection head can accelerate the model convergence speed and improve the model detection accuracy.
[0067] As Figure 2As shown, the ODMamba network structure consists of four modules: image partitioning, ODSSBlock, visual cue merging, and SPPF. The image partitioning module is used to preliminarily process the input image through multiple convolutional layers. Specifically, the image partitioning module contains multiple convolutional layers for preliminary processing of the input image, such as scaling and basic feature extraction, to reduce the dimension of the input data and thus lighten the burden for subsequent deeper feature extraction. ODSSBlock is the most core module in the ODMamba network structure, used to process the input image and extract preliminary features. The SPPF module is the Spatial Pyramid Pooling Fast layer for feature extraction. The visual cue merging module is used to further fuse and optimize the features output by the ODSSBlock module to ensure that the features extracted from different levels can be effectively combined, enhancing the detection ability for small and large targets. The SPPF module is used to aggregate features at multiple scales, and by aggregating features at multiple scales, it helps the model better understand and process objects of different scales. The core module in the PAFPN module is ODSSBlock, which is used to abstract image features, enabling the module to maintain a good balance between accurate feature expression and efficient extraction. The PAFPN module also includes replacing traditional convolutions with dilated convolutions to further abstract image features and establish semantic connections between feature pixels and a larger range of surrounding pixels, improving the context semantic connection of the output features and thus further enhancing the feature expression ability of the output features.
[0068] In this embodiment, the structure of the ODSSBlock module is as Figure 3As shown, it includes an LS (LocalSpatial) Block module, an SS2D (Scanning and Merging) module, and an RG (ResGated) Block module. The SS2D module is used to scan and merge image features to extract multi-directional global information. The LS Block module is used to extract local spatial information to enhance the model's ability to capture local features. The RG Block module is used to capture local dependencies by combining a gating mechanism and residual connections to enhance the model's robustness. At the same time, an ACConv (Asymmetric ConvolutionBlock, asymmetric convolution block) is used instead of a traditional ordinary two-dimensional convolution for convolution operations at the input end of the ODSSBlock structure, and depthwise separable DW-ACConv convolutions are used for convolution operations in each module (the LS Block module, the SS2D module, and the RG Block module). The depthwise separable DW-ACConv convolution is formed by replacing the ordinary convolution unit in the depthwise separable convolution DW-Conv with an ACConv convolution unit. The ACConv convolution decomposes the traditional symmetric convolution kernel into two non-symmetric convolution kernels in two directions, which can reduce the computational amount and the number of parameters while maintaining or improving the performance of the model. ODSSBlock captures and processes image features by using a state space model, enabling the LS Block module, the SS2D module, and the RG Block module to deeply extract complex features in the image, including but not limited to shape, texture, and dynamic changes, etc., which can help the model perform better classification and localization in subsequent steps.
[0069] Specifically, the data processing flow of ODSSBlock includes:
[0070] S1.11. Initially abstract the input features using a depthwise separable convolution DW-Conv, then perform batch normalization on the output features, and use an activation function to better maintain the distribution of the feature map information.
[0071] S1.12. Input the output features into the LS Block module to enhance the model's ability to capture local features, perform a normalization operation on the feature map, and input the normalized feature map into the SS2D module to achieve global feature extraction and fusion of the input features, enhance the context semantic information of the feature layer, improve the feature expression ability of the feature layer, and finally obtain the output Z through residual concatenation and summation with the input feature Z l-2 sum l-1 .
[0072] S1.13 Input Z l-1Normalization is performed, and the normalized post-processed output is input to the RG Block module to capture more global features and is cascaded with the input feature Z through the residual l-1 The final output Z is obtained by summing l .
[0073] The Mamba architecture can capture long-range ground dependencies, but is not suitable for extracting local features when processing tasks involving complex scale changes. This embodiment uses the LS Block module in the ODSSBlock module to enhance the capture of local features. Specifically, Figure 4 As shown in the figure, for a given input feature, the LS Block module first performs a depth-separable DW-ACConv convolution to replace the traditional depth-separable convolution DW-conv. By adopting this convolution, each input channel can be operated separately without mixing channel information. Compared with the traditional depth-separable convolution DW-conv, the local spatial information of the input feature map can be extracted more effectively by adopting DW-ACConv, and the amount of calculation will not be increased. Compared with the traditional convolution, the computational cost and the number of parameters can be reduced at the same time by adopting DW-ACConv. After the depth-separable DW-ACConv convolution, batch normalization is performed to provide a certain degree of regularization effect while reducing overfitting. The intermediate state mixes channel information through 1×1 convolution, and better maintains the distribution of information through activation functions, so that the model can learn more complex feature representations. These feature representations can extract rich multi-scale contextual information from the input feature map, and then further use 1×1 convolution to change the number of channels of the feature without changing the spatial dimension, thereby enhancing the feature representation. Finally, the original input is fused with the processed features through residual concatenation, which enables the model to understand and integrate features of different dimensions in the image, thereby improving the robustness to scale changes.
[0074] like Figure 5 As shown, in this embodiment, the SS2D module first performs a linear transformation on the input features, then uses the deeply separable DW-ACConv to abstract the features, and uses the activation function to better maintain the distribution of information, so that the model can learn more complex feature representations; and then uses the Scan module to realize global feature extraction and fusion of the input features to enhance the contextual semantic information of the feature layer and improve the feature expression ability of the feature layer.
[0075] The RG Block module can capture more global features while only slightly increasing the computational cost. Figure 6As shown, in this embodiment, the RG Block module creates two branches from the input and implements a fully connected layer in the form of 1×1 convolution on each branch; on the right branch, a depthwise separable DW-ACConv convolution is used as the position encoding module, and during training, the gradient is more effectively reflected through residual concatenation, resulting in a lower computational cost, and the performance is significantly improved by preserving and using the spatial structure information of the image, which can effectively reduce the computational cost while enhancing the performance of the model. Specifically, the RG block can use the non-linear GeLU as the activation function to control the information flow at each level, then merge with the left branch through element-wise multiplication, then refine the global features through 1×1 convolution to mix the channel information, and finally sum the features in the original input through residual concatenation. The Scan module realizes global feature extraction and fusion of the input features to enhance the context semantic information of the feature layer and improve the feature expression ability of the feature layer. Finally, the output of the Scan module is batch-normalized, providing a certain degree of regularization effect while reducing overfitting.
[0076] In a specific application embodiment, during the ACConv training phase, as Figure 7 shown, the model is trained by adopting a multi-branch structure of 3×3-BN, 1×3-BN, and 3×1-BN to replace the traditional 3×3-BN structure. -BN means that a batch normalization layer is set after the convolutional layer, and the output feature layers (feature layers 1 to 3) of each branch are summed to obtain the final feature layer 4, and then a non-linear transformation is performed through the GELU activation function; after training, as Figure 8 shown, during ACConv inference, each 1×3-BN and 3×1-BN asymmetric kernel and the corresponding BN layer parameters are added to the backbone and the number of BN layers of the 3×3-BN branch, that is, the overlapping part of the square kernel, to achieve ACConv inference convolution reparameterization, as Figure 9 shown; then the three parallel convolutional models during training are converted back to the same structure as the original ordinary convolution. Through the above reparameterized convolution process, the ordinary convolution can enhance the feature recognition ability of the convolutional backbone part without increasing the calculation and changing the original structure, thereby further improving the overall feature representation ability of the model for scene images.
[0077] In a specific application embodiment, the structural principle of the Scan module is as Figure 10 shown, and the specific execution process of the Scan module includes:
[0078] S1.21. Scanning expansion: The input image is expanded into a series of sub-images through the scanning expansion operation. Each sub-image represents a specific direction, and when observed from a diagonal viewpoint, the scanning expansion operation is processed along four symmetric directions.
[0079] Specifically, the above four directions can be upper left, lower right, lower left, and upper right respectively. Through the above scanning and expanding layout method, not only can all regions of the input image be comprehensively covered, but also through the system's direction transformation, a rich multi-dimensional information library can be provided for subsequent feature extraction, thereby improving the efficiency and comprehensiveness of multi-dimensional capture of image features.
[0080] S1.22. Send the 4 one-dimensional vectors obtained in step S1.21 into the S6 block module for S6 operation.
[0081] Specifically, the S6 operation is an efficient selective information retention and parallel scanning algorithm that can independently process the flattened one-dimensional vectors, ensure that the information in each direction is thoroughly scanned, thereby capturing features in different directions, forming a global receptive field without increasing the linear computational complexity, and effectively extracting the global features of the image.
[0082] S1.23. Scanning and merging: Merge the 4 one-dimensional vectors obtained in step S1.22 into a two-dimensional feature output through the scanning and merging module.
[0083] Step S02. Use an image training set containing multiple human head objects to train a multi-person head pose estimation model. During the training process, use a loss function based on dynamic weights to control the training process. The loss function based on dynamic weights is obtained by weighting the loss functions of the head object prediction classification task, the head position detection task, and the head pose estimation task using dynamic weights. The dynamic weights of each task are determined according to the loss values of the corresponding tasks to balance the losses of each task.
[0084] The end-to-end head pose estimation algorithm involves three learning tasks: classification, localization, and pose estimation. The traditional detection method directly uses the sum of fixed weights as the final loss value for the loss values of each learning task, which will cause imbalance between the loss values of each learning task and thus affect the model accuracy. In this embodiment, the detection head includes three tasks: class prediction (detecting the head prediction probability), head position detection (detecting the head detection box), and head pose estimation. By adopting a strategy of dynamic weighting combined with the cube root, the loss function in the model training process is determined, that is, by performing cube root processing on the loss values of the three tasks and using dynamic weights for weighting to form the total loss function, so as to reduce the impact of the imbalance of the loss value magnitudes and loss rates of different tasks on the model, solve the problem of imbalance of the losses of the three learning tasks in the head pose estimation model, and thus further improve the accuracy of the head pose estimation model.
[0085] As an optional implementation manner, the loss function can specifically adopt the following calculation expression:
[0086] (1)
[0087] Among them, represents the loss function based on dynamic weights, represents the batch size of the multi-person head pose estimation model during training, , , respectively represent the , and dynamic weights of the three at the t-th training. The values of the dynamic weights of each dynamic weight will change with the change of the change rate of the loss value of the training rounds, represents the loss function of the head detection branch, represents the loss function of the head prediction probability branch, represents the loss function of the head pose estimation branch.
[0088] The dynamic weights of each branch are calculated according to the following expression:
[0089] (2)
[0090] (3)
[0091] Among them, represents the relative decrement rate of the loss value of the p-th task, represents the dynamic weight of the p-th task, where p represents the order of the task, and p = 1, 2, 3, corresponding to the class prediction task, the head position detection task, and the head pose estimation task respectively, represents the number of training times, represents the total number of learning tasks, represents the loss value of the p-th task; T represents the preset temperature, which is a fixed value used to control the softness of task weighting. A large T will make the distribution between different tasks more uniform. If T is large enough, then there is ≈ 1, that is, the weights of the tasks are equal.
[0092] According to the above formulas (2) and (3), the dynamic weights corresponding to the three tasks can be determined respectively , , , and through these three real-time dynamic weights , , , it can effectively reduce the impact of the imbalance in learning rates between different learning tasks on the model. Specifically, by calculating the relative decrement rate of the task loss value through Equation (3), the overall loss value of the model can show a downward trend during training. If the loss value of a certain task decreases significantly during training, its loss value will become much smaller, and its relative decrement rate will be relatively large. Combining with the dynamic weight of this task obtained through Equation (2), the dynamic weight will become larger. Therefore, by adjusting the dynamic weights of different tasks at this time, the loss values between different tasks can be dynamically balanced, and the impact of the imbalance in loss rates on the model can be reduced.
[0093] The loss function of the head detection branch is constructed based on the focal regression box loss with gamma transformation, and its calculation expression is:
[0094] (4)
[0095] (5)
[0096] (6)
[0097] (7)
[0098] Among them, N represents the number of head detections in the image to be detected, , respectively represent the predicted detection box and the ground truth label of the i-th head, represents the regression loss based on Siou, represents a preset threshold with a value range of (0, 1), represents the loss value of the center point coordinates and the box length in the head detection box, represents the focal regression box loss with gamma transformation, represents the intersection over union (IoU) between the head region detection box and the ground truth head box, represents the IoU between the reconstructed head region detection box and the ground truth head box, represents the gamma transformation adjustment factor, represents the gamma transformation adjustment factor.
[0099] Bounding box regression plays a crucial role in the field of object detection. The positioning accuracy of object detection largely depends on the loss function of bounding box regression. In traditional end-to-end head pose estimation algorithms, the head detection box branch usually adopts a loss function based on IOU (Intersection over Union) and its improvements. For example, the geometric relationship between bounding boxes is used to improve the regression performance. This type of loss function ignores the influence of the distribution of easy and hard samples on bounding box regression, thereby further affecting the accuracy of head pose estimation. Hard samples refer to samples containing rich information, which can be manifested as a relatively large loss value for this sample during model training, making it difficult for the model to learn. Such samples are more beneficial for model learning due to their rich information content. Easy samples refer to samples containing less information, which can be manifested as a relatively small loss value for this sample during model training, making it easy for the model to learn. Therefore, an excessive number of easy samples will affect the training accuracy of the model. The loss function for the head position detection task in this embodiment adopts a focal regression box loss based on gamma transformation (as shown in Equation (6)), enabling attention to regression samples of different difficulties, enhancing the model's attention to the regression boxes of hard samples, improving the regression accuracy of the head detection box, and thus effectively improving the performance of the detector in different detection tasks.
[0100] As shown in Equation (7), when the value is relatively large, it indicates that the model has learned the sample detection regression box well at this time. Then, the sample at this time tends to be an easy sample. If the at this time, then after transformation, the result > Substituting the result of Equation (7) into Equation (6), it can be seen that as increases, its L Gama-Siou is a decreasing function and the loss decrease rate will increase. At this time, for easy samples, their loss proportion will decrease; by analogy, when the value is relatively small, it indicates that the model has learned the sample detection regression box poorly at this time. Then, the sample at this time tends to be a hard sample. If the < d, then after transformation by Equation (7), the result < Substituting it into Equation (6), it can be seen that as becomes smaller, its L Gama-Siou will be an increasing function and its loss increase rate will increase. At this time, for hard samples, their loss proportion will increase. Through the above mechanism, it can effectively make different difficult regression samples be concerned during object detection.
[0101] The head object prediction classification task is to determine whether the current object is a head region. As an alternative implementation, the loss function for the head object prediction classification task can be calculated according to the following formula:
[0102] (8)
[0103] wherein, represents the predicted head probability of the i-th head detection box, represents the true head label of the i-th head detection box, represents the overlap degree between the i-th head detection box and the true head box. According to the above formula (8), the classification prediction loss of the head object can be calculated, that is, to judge whether the current object is the head area.
[0104] Considering that the rotation representation method with less than four dimensions is discontinuous and not suitable for neural network learning, this embodiment adopts the 6D rotation representation method to estimate the head pose parameters based on the matrix of 6D rotation representation, which can realize the continuous rotation matrix representation, thus facilitating neural network learning and further improving the accuracy of end-to-end one-stage multi-person head pose estimation. Based on the Mamba-YOLO framework, this embodiment adds a 6D decoupled detection head for head pose estimation to predict the head pose rotation matrix, and the 6D branch at each anchor point will predict 6 values , and these 6 values form two three-dimensional vectors:
[0105]
[0106]
[0107] Then, the above two three-dimensional vectors are transformed according to the following formula;
[0108] ; ; (9)
[0109] wherein, are respectively u at x , y , z axis coordinate values, are respectively v at x , y , z axis coordinate values, , , respectively represent the three three-dimensional unit vectors in the 3×3 rotation matrix formed by transforming the 6D rotation representation parameter matrix , , , , ~ respectively represent , , each element in, represents A vector that is perpendicular in three-dimensional space.
[0110] The two three-dimensional vectors corresponding to the 6D rotation parameters of each head object's head posture are , According to the above conversion, we can get the head posture estimation matrix By deriving the elements in , we can get an orthogonal matrix for head posture estimation prediction:
[0111]
[0112] The label rotation matrix obtained by Euler angle is also an orthogonal matrix:
[0113]
[0114] in, , , are three mutually orthogonal unit vectors of the true head posture.
[0115] The multi-person head posture model needs a loss function to measure the distance between the two matrices, so that the model can continuously optimize parameters and predict the head posture rotation matrix, and then obtain the head posture Euler angle. In order not to destroy the SO(3) popular geometry structure of the head posture 3×3 rotation matrix, this embodiment uses a method based on cosine similarity to characterize the head posture estimation matrix. With the real head pose matrix The similarity between .
[0116] Specifically, this embodiment uses a dense detection network to directly predict a group of head objects ,in Represents the head prediction score and position parameters, Represents the head pose 6D rotation representation parameters); The specific parameters are , Represents the probability of the current prediction being the head, Represents the center position of the current predicted head detection box, Represents the width and height of the current predicted head detection box; The specific parameters are , represents the 6D rotation representation parameter, where u and v represent two non-parallel vectors. The 6D rotation representation parameter is converted into a 3×3 rotation matrix according to the above formula (9). As an attribute of the head object, compare it with Combine to form the head prediction object , it is possible to achieve end-to-end one-stage multi-person head pose high-precision estimation based on 6D rotation representation.
[0117] To calculate the similarity between two 3D rotation matrices and form the loss function of the head pose estimation branch, considering that the loss value of easy samples should account for a smaller proportion and the loss value of difficult samples should account for a larger proportion, and when x is closer to , the value of y is , and its characteristics are consistent with those of the tangent function. The loss function of the head pose estimation task in this embodiment is constructed based on the tangent function to update and optimize the head pose estimation parameters, enabling the model to pay more attention to the learning of difficult dimension vectors and further improving the head pose estimation accuracy.
[0118] Optionally, the calculation expression of the loss function of the head pose estimation task is:
[0119] (10)
[0120] where i represents the i-th dimension of the head pose matrix, represents the loss weight on the i-th dimension of the head pose matrix, represents the head pose estimation matrix, represents the true head pose matrix, , represents , elements in
[0121] As Figure 11 shown, when the abscissa x of the tangent function is in the interval , the range of its value y is , and as x increases, the growth rate of y becomes larger. This function change trend conforms to the requirement for difficult sample mining (the loss value of easy samples should account for a smaller proportion, the loss value of difficult samples should account for a larger proportion, and when x is closer to , the value of y is ), so using the tangent function to construct the loss function of the head pose estimation task can enable the model to pay more attention to the learning of difficult dimension vectors. As shown in Equation (10), considering the boundary problem of the value of y, the maximum value of x is set to ( value range [0,1]), so that the value of y is within a certain range. At the same time, by setting the ratio among the three can be further dynamically adjusted. If the similarity of a certain dimension component in the two matrices is small, the loss contribution of it will further increase, so that the head pose model can further pay attention to difficult samples and achieve the purpose of further improving the head pose estimation accuracy.
[0122] Step S03. Obtain the to-be-detected image containing multiple people's heads, input it into the trained multi-person head pose estimation model, and obtain the detection results of multiple head objects.
[0123] After the multi-person head pose algorithm model performs inference, a series of head detection objects will be generated. It is necessary to perform relevant post-processing on this series of detection objects to obtain the final real detection objects. Non-maximum suppression (NMS) is a commonly used post-processing method adopted by many object detection models. However, since the multi-person head pose estimation scenario belongs to dense object detection and there are different degrees of overlap between objects, the traditional non-maximum suppression algorithm (NMS) only depends on the highest confidence of object classification and IoU, resulting in certain inaccuracies in positioning and missed detections in overlapping detection scenarios. In this embodiment, a non-maximum suppression algorithm that does not depend on confidence and IoU is adopted to improve the detection accuracy for overlapping scenarios, thereby improving the head pose estimation accuracy. As Figure 12 shown, the specific algorithm steps are as follows:
[0124] 1) Filter out the detection boxes with confidence less than the preset threshold;
[0125] 2) Calculate the proximity P between every two detection boxes used: First, normalize the relevant coordinates, calculate the center coordinates, length, and width of the detection boxes, and then calculate the proximity using the Manhattan distance based on logarithmic transformation. Its calculation formula is shown in Equation 11;
[0126] (11)
[0127] In the above formula, x1 represents the x coordinate of the center point of the first detection box, y1 represents the y coordinate of the center point of the first detection box,
[0128] w1 represents the width of the first detection box, h1 represents the height of the first detection box, and the same applies to x2, y2, w2, and h2.
[0129] 3) Group the detection boxes with proximity P <= d (d is the proximity threshold, which can be determined according to the specific scenario) into one cluster;
[0130] 4) Calculate its confidence-weighted proximity;
[0131] 5) Find the detection box with the smallest confidence-weighted proximity in one cluster, save it, and filter it out from the total boxes;
[0132] 6) Filter out other detection boxes with proximity less than the preset threshold k (d is the proximity threshold, which must be determined according to the specific scenario);
[0133] 7) For the remaining boxes, repeat steps 2 - 6.
[0134] This embodiment further provides an electronic device, including a processor and a memory. The memory is used to store a computer program, and the processor is used to execute the computer program to execute the method as described above.
[0135] It can be understood that the above method of this embodiment can be executed by a single device, such as a computer or a server, etc., or can also be applied to a distributed scenario where multiple devices cooperate with each other to complete. In the case of a distributed scenario, one of the multiple devices can only execute one or more steps of the above method of this embodiment, and the multiple devices interact with each other to complete the above method. The processor can be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., and is used to execute relevant programs to implement the above method of this embodiment. The memory can be implemented in the form of a read-only memory ROM, a random access memory RAM, a static storage device, and a dynamic storage device, etc. The memory can store an operating system and other application programs. When implementing the above method of this embodiment through software or firmware, the relevant program codes are stored in the memory and are called and executed by the processor.
[0136] This embodiment further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the method as described above.
[0137] Those skilled in the art should understand that the above embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions in one Figure 1 one flow or multiple flows and / or blocks Figure 1The functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes and / or boxes Figure 1 One process or more processes and / or boxes Figure 1 Steps for implementing the functions specified in one box or more boxes.
[0138] The above is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed above in a preferred embodiment, it is not intended to limit the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for multi-person head pose estimation based on multi-task loss balance, characterized in that the steps include: A multi-person head posture estimation model is constructed using a YOLO network based on the Mamba model. The multi-person head posture estimation model includes a backbone network, a PAFPN module, and a detection head. The backbone network is an ODMamba network structure for processing input images and extracting preliminary features. The PAFPN module is used to extract and fuse features of different scales. The detection head is used to achieve category prediction, head position detection, and head posture estimation tasks. Using an image training set containing multiple head objects to train the multiple head posture estimation model, during the training process, a loss function based on dynamic weights is used to control the training process, wherein the loss function based on dynamic weights is obtained by weighting the loss function of the head object prediction classification task, the loss function of the head position detection task, and the loss function of the head posture estimation task using dynamic weights, and the dynamic weight of each task is determined according to the loss value of the corresponding task to balance the loss of each task; Obtain a picture to be detected containing multiple heads, input it into the trained multi-person head posture estimation model, and obtain the detection results of multiple head objects; the calculation expression of the loss function based on dynamic weight is: in, represents the loss function based on dynamic weights, Indicates the batchsize of the multi-person head pose estimation model during training. , , Respectively represent the tth training time , and The dynamic weights of each dynamic weight change with the change rate of the loss value in the number of training rounds. represents the loss function of the head position detection task, represents the loss function of the category prediction task, Represents the loss function for the head pose estimation task.
2. The method for multi-person head pose estimation based on multi-task loss balance according to claim 1, characterized in that: The ODMamba network structure includes an image segmentation module, an ODSSBlock module, a visual cue merging module and an SPPF module. The image segmentation module is used to perform preliminary processing on the input image through multiple convolutional layers, and the ODSSBlock module is used to perform feature extraction; the visual cue merging module is used to fuse and optimize the features output by the ODSSBlock module, and the SPPF module is used to aggregate features at multiple scales; the core module in the PAFPN module is the ODSSBlock module; ACConv convolution is used to perform convolution operations at the input end of the ODSSBlock module, and deep separable DW-ACConv convolution is used to perform deep separation convolution operations in the LS Block module, SS2D module and RG Block module in the ODSSBlock module, the SS2D module is used to scan and merge image features to extract multi-directional global information, the LS Block module is used to extract local spatial information, and the RG Block module is used to capture local dependencies by combining a gating mechanism and a residual connection.
3. The method for multi-person head pose estimation based on multi-task loss balance according to claim 1, characterized in that: The dynamic weight of each branch is calculated according to the following expression: in, represents the dynamic weight of the pth task, p=1, 2, 3, corresponding to the category prediction task, head position detection task, and head posture estimation task, respectively. represents the number of training times, Indicates the total number of tasks, represents the loss value of the pth task, T Indicates the preset temperature used to control the softness of task weighting.
4. The method for multi-person head pose estimation based on multi-task loss balance according to claim 1, characterized in that: The loss function of the head position detection task is constructed based on the focus regression box loss of gamma transform, and the calculation expression is: Where N represents the number of head detections in the image to be detected. , The i-th head prediction detection box and the true label are respectively, represents the Siou-based regression loss, Indicates the preset threshold value with a value range of (0,1). Indicates the loss value of the center point coordinates and frame length of the head detection frame, represents the focus regression box loss based on gamma transform, It represents the intersection-over-union ratio between the head region detection frame and the real head frame. It represents the intersection-over-union ratio between the reconstructed head region detection frame and the real head frame. represents the gamma shift adjustment factor, represents the gamma shift adjustment factor.
5. The method for multi-person head pose estimation based on multi-task loss balance according to claim 1, 3 or 4, characterized in that: The loss function of the head object prediction classification task is calculated according to the following formula: in, represents the probability of the head predicted by the i-th head detection box, represents the real label of the head of the i-th head detection frame, Indicates the overlap between the i-th head detection frame and the true head frame.
6. The method for multi-person head pose estimation based on multi-task loss balance according to claim 1, 3 or 4, characterized in that: The loss function of the head posture estimation task is constructed based on the tangent function, and the calculation expression is: Among them, i represents the i-th dimension of the head posture matrix, represents the loss weight on the i-th dimension of the head posture matrix, represents the head pose estimation matrix, represents the true head pose matrix, , express , The elements in .
7. The method for multi-person head pose estimation based on multi-task loss balance according to claim 6, characterized in that: It also includes obtaining the head pose estimation matrix The steps include: Use the head pose estimation 6D decoupled detection head set in the multi-person head pose estimation model to detect the head pose 6D rotation representation parameters ; The two unrelated three-dimensional vectors corresponding to the 6D rotation parameters of the head pose of each head object are , Convert to get the head pose estimation matrix The elements in: ; ; ; in, They are u exist x , y , z The coordinate values of the axis, They are v exist x , y , z The coordinate values of the axis, , , Respectively represent the 6D rotation parameter matrix Transformed into three three-dimensional unit vectors in a 3×3 rotation matrix, , , , ~ Respectively , , Each element in Representation and A vector that is perpendicular in three-dimensional space.
8. An electronic device comprising a processor and a memory, wherein the memory is used to store a computer program, wherein: The processor is configured to execute the computer program to perform the method according to any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
End-to-end multi-person head posture estimation method and device based on 6D rotation representation
CN118762075A
Multi-person head posture estimation method and device based on geodesic line loss and medium
CN118865453A