Personnel posture intelligent estimation method for complex industrial scene

By using the data augmentation training set and transfer learning method generated by real occlusion in complex industrial scenarios, the problem of performance degradation in human pose estimation models in industrial scenarios is solved, and high-precision pose estimation is achieved.

CN120032394AInactive Publication Date: 2025-05-23HANGZHOU DIANZI UNIV

Patent Information

Application Number
CN202510505890.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In complex industrial scenarios, the performance of the human posture estimation model decreases when facing occlusion factors, and cannot accurately identify the key points that are blocked, resulting in a large deviation in the pose estimation results.

Method used

The intelligent estimation method of personnel poses for complex industrial scenarios is adopted. By obtaining industrial scenario videos, labeling human key points, learning posture characteristics, using data generated based on real occlusions to enhance the training set, and adjusting model parameters through transfer learning methods to improve the model's ability to identify occlusion areas.

Benefits of technology

It effectively improves the accuracy of the human posture estimation model under occlusion conditions in complex industrial scenarios, ensures the accuracy and reliability of the posture estimation results, and meets the practical application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032394A_ABST
    Figure CN120032394A_ABST
Patent Text Reader

Abstract

The invention discloses a personnel posture intelligent estimation method for a complex industrial scene, and the method comprises the steps: firstly obtaining an industrial scene personnel operation behavior video, and marking human body key points for the behavior video; secondly, using the human body key points to learn industrial scene personnel operation posture features, and obtaining an industrial scene human body posture estimation model; and then shielding human body key points in the behavior video according to a detection result of the human body posture estimation model. And finally, using a data enhancement training set generated by shielding, adjusting parameters of the human body posture estimation model through a transfer learning method, and repeating the operation until the precision of the human body posture estimation model is not improved any more or the improvement change rate is smaller than a set threshold value. According to the method, human body semantic information of different scales is fully utilized, loss of detail information is reduced, the feature alignment effect of the occlusion area is improved, continuous optimization of model performance is ensured, and accurate personnel attitude estimation is carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is applicable to the cross-technical field of computer vision and artificial intelligence, and in particular relates to a method for intelligently estimating a person's posture in complex industrial scenarios. Background Art

[0002] Human posture estimation is a basic task in the field of computer vision. It aims to estimate the position of body joints and predict the coordinates of key points of the human body from images. It can provide support for multiple downstream tasks in the field of computer vision, such as action recognition, human-computer interaction, etc. In industrial scenarios, human posture estimation ensures production safety by analyzing and understanding the actions and postures of operators. Accurate posture estimation is crucial to preventing safety accidents.

[0003] However, when faced with occlusion factors, the human pose estimation model suffers from severe performance degradation. There are a large number of occlusions in industrial environments, such as industrial equipment, stacked goods, billboards, etc. The model often cannot accurately identify the occluded key points, resulting in large deviations in the pose estimation results, which cannot meet the needs of real applications.

[0004] Therefore, it is urgent to develop an intelligent personnel posture estimation method for complex industrial scenes that can effectively deal with occlusion to meet the needs of practical applications. Summary of the invention

[0005] The purpose of the present invention is to solve the problem of reduced performance of human posture estimation models facing occlusion in complex industrial scenes, and to provide a method for intelligent estimation of human posture for complex industrial scenes.

[0006] The method for intelligently estimating a person's posture in complex industrial scenarios comprises the following steps:

[0007] S1, obtaining a video of a person's working behavior in an industrial scene, and marking key points of the human body in the video of the behavior.

[0008] S2 uses human body key points to learn the working posture characteristics of personnel in industrial scenes, obtains the human body posture estimation model of industrial scenes, and obtains the personnel posture estimation results.

[0009] S3: According to the detection result of the human posture estimation model, the key points of the human body in the behavior video are blocked.

[0010] S4, using a data enhancement training set generated based on real occluders, and adjusting parameters of the human posture estimation model through a transfer learning method.

[0011] S5, repeat S3 and S4 until the accuracy of the human posture estimation model is no longer improved or the improvement rate is less than a threshold .

[0012] S6, deploying the intelligent personnel posture estimation method for complex industrial scenarios to the system.

[0013] Further, step S2 includes the following steps:

[0014] S21, using human key points, trains a quantized autoencoder network consisting of an encoder, a codebook, and a decoder;

[0015] S22, using the improved Swin-Transformer and two residual convolution blocks to build a transformer module to extract the personnel work posture features in the personnel work behavior video frame, and convert them to the feature space where the trained codebook is located, and reconstruct the key point coordinates through the trained decoder to achieve human posture estimation;

[0016] The improved Swin-Transformer is specifically implemented as follows: , , , They are the feature outputs of each stage of the Swin-Transformer model. The feature outputs of each stage are downsampled to the same target size through bilinear interpolation and integrated, and then enter the convolutional neural network for channel number conversion to obtain the feature , to fuse information of different scales of the model:

[0017]

[0018] In the formula, , , and are the feature outputs of each stage of the Swin-Transformer model, It is a two-dimensional convolutional neural network. It is the integration and splicing of the matrix. Downsampling for bilinear interpolation.

[0019] Further, step S3 includes the following steps:

[0020] S31, extracts a certain number of common human key point occluders from industrial scene video data, including industrial equipment, stacked goods, billboards, paint buckets, carts, tires, and pipelines;

[0021] S32, obtaining the detection result of the human body posture estimation model on the training set, evaluating the recognition accuracy of the model for each human body key point on the training set, and using the MSE error to evaluate the accuracy of the human body posture estimation model for detecting each key point on the entire training set:

[0022]

[0023] In the formula, represents the MSE error, Indicates The first image of the training set Personal key points, is the real coordinate of the key point of the human body in the picture, are the key point coordinates predicted by the posture estimation model;

[0024] S33, using the common human key point occluders, sorting the recognition accuracy of the entire training set from high to low, % The key points of the human body are occluded to generate a data enhancement training set based on real occluders.

[0025] set up represents the original image, Represents the occluded object image, represent To convert Integration In the first Perform morphological erosion to obtain , and then use Get the mask The edge of the object , then The pixels contained are set to 191, and the grayscale mask is obtained (non-object: 0, edge: 191, object: 255), and finally use Divide by 255 to get a non-integer mask , and use Will Integration middle:

[0026]

[0027] In the formula, Represents a data augmented image generated by occlusion.

[0028] Further, step S4 includes the following steps:

[0029] S41, freezing the backbone network parameters of the trained converter module;

[0030] S42, using the data enhancement training set generated based on the real occluders, fine-tuning the parameters of the fully connected layer and the linear layer of the converter module.

[0031] Further, step S5 includes the following steps:

[0032] S51, iteratively generating the data enhancement training set based on real occlusion objects according to the detection results of each key point on the training set by the human posture estimation model in multiple rounds, and fine-tuning the model by using the transfer learning method;

[0033] The training set of the human posture estimation model in each iteration round is the one with the highest accuracy of key point detection on the entire training set based on the model recognition results of the previous round, that is, The smallest front %Data enhancement training set generated after occluding key points of the human body;

[0034] In the Nth round of iteration, the initial stage of the model is based on the backbone network after the N-1th round of training. The weights and biases of the source task (N-1th round) are transferred to the human posture recognition task in this stage (Nth round). During the transfer learning process, only the parameters of the fully connected layer and linear layer of the converter module are updated, and the other networks of the model are frozen;

[0035] If the accuracy of the human posture estimation model no longer improves or the rate of change is less than a threshold , then terminate the training and get the final model parameters;

[0036] S52, using an early stopping method for each round of fine-tuning to prevent overfitting of the human body posture estimation model;

[0037] In each round of fine-tuning, the average accuracy AP (Average Precision) on the validation set of each epoch is calculated. If the validation set performance improvement rate is less than , then stop this round of training.

[0038] Further, step S6 includes the following steps:

[0039] S61, using image acquisition devices such as cameras to capture human behavior data as input to a human key point detection model;

[0040] S62, using the edge computing device to run the human body key point detection model to identify the human body key points;

[0041] S63, using a display as an output device to visualize the model prediction results to achieve real-time display.

[0042] Compared with the prior art, the beneficial effects and advantages of the present invention are:

[0043] 1. The present invention discloses an improved quantized autoencoder human posture estimation model for industrial scenarios, which integrates the outputs of different stages of the feature extraction network, fully utilizes human semantic information of different scales, and reduces the loss of detail information.

[0044] 2. The present invention discloses a key point-guided adaptive data enhancement method, which uses occlusion objects based on real industrial scenes to perform targeted occlusion on human key points according to the recognition results of a human posture estimation model, simulates the occlusion conditions in actual industrial scenes, and realizes lightweight and high-quality data enhancement.

[0045] 3. The present invention discloses a multi-round iterative fine-tuning method based on transfer learning, which helps to improve the feature alignment effect of the occluded area, ensure the continuous optimization of the model performance, and perform accurate personnel posture estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The present invention will be described in detail below in conjunction with the accompanying drawings. The above or other aspects of the present invention will become clearer and easier to understand through the detailed description made in conjunction with the following accompanying drawings, in which:

[0047] Figure 1 It is a flowchart of the present invention;

[0048] Figure 2 It is a structural diagram of the first stage of the present invention;

[0049] Figure 3 It is a structural diagram of the second stage of the present invention;

[0050] Figure 4 It is a structural diagram of a human posture estimation model based on a quantized autoencoder network provided by the present invention;

[0051] Figure 5 is a structural diagram of an improved feature extraction network model provided by an embodiment of the present invention;

[0052] Figure 6 is a flow chart of model iteration based on transfer learning provided by an embodiment of the present invention;

[0053] Figure 7 This is a schematic diagram of the results of the intelligent personnel posture estimation method based on complex industrial scenarios. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0055] like Figure 1As shown, the embodiment of the present invention proposes a method for intelligent estimation of human posture for complex industrial scenes. The method is based on an improved quantized autoencoder for human posture estimation, and proposes a data enhancement method based on real scenes and a multi-round iterative fine-tuning method of the model. This method overcomes the problem of performance degradation of the human posture estimation model in the face of occlusion, improves the practicability of the model, and realizes accurate human posture prediction in industrial scenes. The specific steps are as follows:

[0056] S1, obtaining a video of a person's work behavior in an industrial scene, and marking key points of the human body in the video of the behavior;

[0057] S2, using the human body key points to learn the working posture characteristics of personnel in industrial scenes, obtain an industrial scene human body posture estimation model, and obtain a personnel posture estimation result;

[0058] S3, according to the detection result of the human posture estimation model, using industrial equipment, billboards, paint buckets, carts, tires, and pipelines to block the key points of the human body in the behavior video;

[0059] S4, using a data augmentation training set generated based on real occluders, and adjusting parameters of the human posture estimation model through a transfer learning method;

[0060] S5, repeat S3 and S4 until the accuracy of the human posture estimation model is no longer improved or the improvement rate is less than a threshold ;

[0061] S6, deploying the intelligent personnel posture estimation method for complex industrial scenarios to the system.

[0062] Specifically, step S1 includes the following steps:

[0063] S11, obtains video data of personnel and related environments in industrial scenes, including two work scenes and operation types: chemical industry and shipbuilding;

[0064] S12, select a certain number of representative sample videos, and annotate the key points of the human body from the image frames of the sample videos, including nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left hand, right hand, left hip, right hip, left knee, right knee, left foot and right foot, a total of 17 key points of the human body.

[0065] Step S2 constructs Figure 4 The human posture estimation model based on the quantized autoencoder network shown in FIG. 1 includes two stages, and the steps of each stage are as follows:

[0066] S21, such as Figure 2 As shown, using the key points of the human body, a quantized autoencoder network consisting of an encoder, a codebook, and a decoder is trained, and the original posture is defined as , Represents human joints. is the joint dimension (when the pose is 2D, ; When the posture is 3D, ), the quantized encoder uses the labeled human key points to learn and convert the posture into Token features:

[0067]

[0068] In the formula, It's the human body posture. is the quantized encoder, which converts the human body posture into Token features, It is token features, each token feature corresponds to a substructure of human posture;

[0069] Constructing a codebook , discretize the extracted features and record them, where is the number of discretized feature entries recorded in the codebook, is the dimension of the discretized feature;

[0070] S22, such as Figure 3 As shown, a converter module is constructed using Swin-Transformer and two residual convolution blocks to extract the working posture features of the personnel in the industrial scene from the input image data and convert them to the feature space where the codebook is located. The key point coordinates are reconstructed through the decoder to achieve human posture estimation;

[0071] like Figure 5 As shown, the original feature extraction network of the converter module is improved. , , , They are the feature outputs of each stage of the Swin-Transformer model. The feature outputs of each stage are downsampled to the same target size through bilinear interpolation and integrated, and then enter the convolutional neural network for channel number conversion to obtain the feature , to fuse information of different scales of the model:

[0072]

[0073] In the formula, , , , are the feature outputs of each stage of the Swin-Transformer model, It is a two-dimensional convolutional neural network. It is the integration and splicing of the matrix. Downsampling for bilinear interpolation;

[0074] The extracted features are transformed into The output is token features, and quantize each token feature using nearest neighbor search in the constructed codebook :

[0075]

[0076] In the formula, It is the original posture. It is The most recent feature index that a token matches in the codebook. All tokens share the same codebook. ;

[0077] Finally, the quantized token features are fed into the decoder to restore the original posture:

[0078]

[0079] In the formula, is the quantized token feature set, is the decoder, using the quantized Tokens reconstruct the coordinates of the key points of the human body .

[0080] Specifically, step S3 includes the following steps:

[0081] S31, extracts a certain number of common human key point occluders from industrial scene video data, including industrial equipment, billboards, paint buckets, carts, tires, and pipelines;

[0082] S32, obtaining the detection result of the human body posture estimation model on the training set, evaluating the recognition accuracy of the model for each human body key point on the training set, and using the MSE error to evaluate the accuracy of the human body posture estimation model for detecting each key point on the entire training set:

[0083]

[0084] In the formula, represents the MSE error, Indicates The first image of the training set Personal key points, is the real coordinate of the key point of the human body in the picture, are the key point coordinates predicted by the posture estimation model;

[0085] Using the common human key point occluders, the recognition accuracy of the entire training set is the highest, that is, The smallest front % The key points of the human body are occluded to generate a data enhancement training set based on real occluders.

[0086] S33, using the common human key point occluders, the front end with the highest recognition accuracy in the entire training set %Occlude key points of the human body to generate a data enhancement training set based on real occluders;

[0087] set up represents the original image, Represents the occluded object image, represent To convert Integration In the first Perform morphological erosion to obtain , and then use Get the mask The edge of the object , then The pixels contained are set to 191, and the grayscale mask is obtained (non-object: 0, edge: 191, object: 255), and finally use Divide by 255 to get a non-integer mask of 0, 0.75, 1 , and use Will Integration middle:

[0088]

[0089] In the formula, Represents a data augmented image generated by occlusion.

[0090] Specifically, step S4 includes the following steps:

[0091] S41, freezing the backbone network parameters of the trained converter module;

[0092] S42, using the data enhancement training set generated based on the real occluders, fine-tuning the parameters of the fully connected layer and the linear layer of the converter module.

[0093] Specifically, step S5 includes the following steps:

[0094] S51, iteratively generating the data enhancement training set based on real occlusion objects according to the detection results of each key point on the training set by the human posture estimation model in multiple rounds, and fine-tuning the model by using the transfer learning method;

[0095] S52, using an early stopping method for each round of fine-tuning to prevent overfitting of the human body posture estimation model;

[0096] In each round of fine-tuning, the average accuracy AP (Average Precision) on the validation set of each epoch is calculated. If the validation set performance improvement rate is less than , then stop this round of training.

[0097] like Figure 6 As shown, step S52 includes the following steps:

[0098] The training set of the human posture estimation model in each iteration round is the one with the highest accuracy of key point detection on the entire training set based on the model recognition results of the previous round, that is, The smallest front %Data enhancement training set generated after occluding key points of the human body;

[0099] In the Nth round of iteration, the initial stage of the model is based on the backbone network after the N-1th round of training. The weights and biases of the source task (N-1th round) are transferred to the human posture recognition task in this stage (Nth round). During the transfer learning process, only the parameters of the fully connected layer and linear layer of the converter module are updated, and the other networks of the model are frozen;

[0100] If the accuracy of the human posture estimation model no longer improves or the rate of change is less than a threshold , the training is terminated and the final model parameters are obtained.

[0101] Specifically, step S6 includes the following steps:

[0102] S61, using image acquisition devices such as cameras to capture human behavior data as input to a human key point detection model;

[0103] S62, using the edge computing device to run the human body key point detection model to identify the human body key points;

[0104] S63, using a display as an output device to visualize the model prediction results to achieve real-time display.

[0105] According to the records of the above method, the present invention gives the results on the self-built human posture estimation datasets of two working scenes and types of operations: chemical industry and shipbuilding. The evaluation index is defined by the target key point similarity (OKS):

[0106]

[0107] In the formula, For the human body Key points, is the Euclidean distance between the key points obtained by the test and the true value of the label, Is a sign of whether the true value of the label is visible, Represents the size of the object. is the control attenuation constant of each key point, For visibility. Using the standard average accuracy AP, the standard average accuracy AP with an OKS threshold of 0.5 0.5 and OKS threshold is 0.75 standard average accuracy AP 0.75 , as well as the standard average recall rate AR, the standard average recall rate AR with an OKS threshold of 0.5 0.5 and OKS threshold is 0.75 standard average recall AR 0.75 As the evaluation index of the human posture estimation model, the results are shown in Table 1 below.

[0108] Table 1 Experimental comparison results of the present invention and other advanced methods on self-built data sets

[0109]

[0110] Among them, IPR, ViTPose and PCT, which are regression methods for predicting key point coordinates, are all existing algorithms in the art.

[0111] This paper uses an improved quantized autoencoder to perform structured modeling of key points of the human body to enhance the ability of the posture estimation model to resist partial occlusion of the worker's body. Considering the difficulty in constructing a worker occlusion dataset, a dynamic data enhancement training method is proposed. During the model training process, the posture estimation results of the model are evaluated and worker occlusion images are dynamically generated for the next model training. Figure 7 As shown, the experimental results prove that the method proposed in the present invention can effectively deal with the problem of human body occlusion in complex industrial scenes.

[0112] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, so the present invention can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0113] The above implementation methods have been described in detail. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A method for intelligent estimation of personnel posture in complex industrial scenes, characterized in that: The following steps are involved: S1, obtain the working behavior video of personnel in industrial scenes and mark the key points of the human body in the behavior video; S2, using human key points, learns the characteristics of the working posture of personnel in industrial scenes, obtains the human posture estimation model of industrial scenes, and obtains the personnel posture estimation results; S3, based on the detection results of the human posture estimation model, the key points of the human body in the behavior video are blocked; S4, uses the data generated by occlusion to enhance the training set and adjusts the parameters of the human pose estimation model through transfer learning method; S5, repeat S3 and S4 until the accuracy of the human posture estimation model no longer improves or the improvement rate is less than the set threshold .

2. According to claim 1, a method for intelligent estimation of personnel posture for complex industrial scenes is characterized in that: The specific implementation process of step S2 is as follows: S21, using human key points, trains a quantized autoencoder network consisting of an encoder, a codebook, and a decoder; S22, a converter module is constructed using the improved Swin-Transformer and two residual convolution blocks to extract the personnel work posture features in the personnel work behavior video frames and convert them to the feature space where the trained codebook is located. The key point coordinates are reconstructed through the trained decoder to realize human posture estimation.

3. According to claim 2, a method for intelligent estimation of personnel posture for complex industrial scenes is characterized in that: The improved Swin-Transformer is specifically implemented as follows: , , and They are the feature outputs of each stage of the Swin-Transformer model. The feature outputs of each stage are downsampled to the same target size through bilinear interpolation and integrated, and then enter the convolutional neural network for channel number conversion to obtain the feature , integrating information at different scales.

4. The method for intelligently estimating personnel posture in complex industrial scenes according to claim 3 is characterized in that: The specific implementation process of step S3 is as follows: S31, extracts human key point occluders from industrial scene video data, including industrial equipment, stacked goods, billboards, paint buckets, carts, tires, and pipelines; S32, obtaining the detection result of the human body posture estimation model on the training set, evaluating the recognition accuracy of the model for each human body key point on the training set, and using the MSE error to evaluate the accuracy of the human body posture estimation model for detecting each key point on the entire training set; S33, using human key point occluders, sort the recognition accuracy of the entire training set from high to low, % The key points of the human body are occluded to generate a data enhancement training set based on real occluders.

5. The method for intelligently estimating a person's posture in a complex industrial scene according to claim 4 is characterized in that: The shielding in step S33 is specifically implemented as follows: represents the original image, Represents the occluded object image, represent The binary mask of Integration In the first Perform morphological erosion to obtain , and then use Get the mask The edge of the object , then The pixels contained are set to 191, and the grayscale mask is obtained ; Finally, use Divide by 255 to get a non-integer mask , and use Will Integration middle: ; In the formula, Represents a data augmented image generated by occlusion.

6. The method for intelligently estimating a person's posture in a complex industrial scene according to claim 5 is characterized in that: The specific implementation process of step S5 is as follows: S51, iteratively generate the data enhancement training set based on real occluders based on the detection results of each key point on the training set according to the human posture estimation model in multiple rounds, and use the transfer learning method to fine-tune the model: The training set of the human posture estimation model in each iteration round is the previous model with the minimum MSE error based on the model recognition result of the previous round. %Data enhancement training set generated after occluding key points of the human body; In the Nth round of iteration, the initial stage of the human posture estimation model is based on the network after the N-1th round of training. The source task, that is, the weights and biases of the N-1th round are transferred to the human posture recognition task of the Nth round. During the transfer learning process, only the parameters of the fully connected layer and the linear layer of the converter module are updated, and other networks are frozen; If the accuracy of the human posture estimation model no longer improves or the rate of change is less than a set threshold , then terminate the training and get the final model parameters; S52, an early stopping method is used for each round of fine-tuning to prevent the human posture estimation model from overfitting.

Citation Information

Patent Citations

  • Single-person posture estimation method based on novel high-resolution network model

    CN110175575A

  • Data processing method and device, computer equipment and storage medium

    CN110889503A

  • Human body shape and posture estimation method for object occlusion scene

    CN111339870A

  • Cherry grading detection method and system based on deep convolutional neural network

    CN113077450A

  • Engineering drawing review recognition processing method and device for manual instrument drawing

    CN113658095A

Cited By

  • 3D human body posture optimization estimation method based on generative data enhancement

    CN120612716A