A method for detecting compliance of personal protective equipment
By building a human perception and shape reconstruction network, combining video surveillance means, analyzing laboratory video stream data, the problem that traditional monitoring methods are difficult to accurately detect the wear compliance of personal protective equipment in complex scenarios, achieving efficient and accurate detection results.
Patent Information
- Application Number
- CN202510237910.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-03
AI Technical Summary
In modern laboratories, the compliance of monitoring personal protective equipment (PPE) faces problems of inefficiency and uncertainty in the use of inefficiency, especially in complex scenarios, it is difficult to accurately determine whether the experimenter wears all protective equipment in a standardized manner.
By constructing a two-dimensional human perception network, a three-dimensional human shape reconstruction network and a layered protective equipment judgment network, combining modern image recognition technology and video surveillance means, analyzing video stream data to determine whether the experimental personnel wear personal protective equipment in a standardized manner.
It realizes efficient and accurate personal protective equipment compliance inspection in complex scenarios, greatly improving the detection efficiency and reliability of results.
Smart Images

Figure CN119741639B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and in particular to a compliance detection method for personal protective equipment. Background Art
[0002] With the continuous development of science and technology and the increasing complexity of modern laboratory environments, laboratory safety issues have gradually attracted widespread attention. Personal protective equipment (PPE) is an important measure to protect the lives of laboratory personnel and prevent accidents. The compliance of its use directly affects the safety of the laboratory and the accuracy of experimental results.
[0003] However, in modern laboratories, monitoring the compliance of PPE use faces many technical challenges. Traditional manual inspection and manual recording are not only inefficient, but also easily affected by subjective factors, making it difficult to ensure real-time and accuracy. Therefore, how to improve the level of laboratory safety management through intelligent means and ensure the correct use of PPE has become an urgent problem to be solved in the current laboratory information management system (LIMS).
[0004] In order to solve the above technical problems, some intelligent monitoring technologies currently attempt to determine the wearing status of PPE by installing sensors in the laboratory. These sensors usually include hardware connected to protective equipment, such as position sensors, temperature sensors, pressure sensors, etc., which are designed to monitor the wearer's status in real time and ensure that the PPE equipment is worn correctly. However, in the actual laboratory environment, due to the frequent activities of personnel and the complex laboratory scenes, it is difficult to accurately determine whether the experimenters are wearing all protective equipment in a standardized manner by relying on a single detection method of sensors. For example, when experimenters wear multiple protective equipment (such as protective glasses, gloves, masks, protective clothing, etc.), sensors often cannot fully and effectively monitor the wearing status of all equipment, and are prone to misjudgment or missed judgments. At the same time, the installation and maintenance of these sensor devices requires a lot of capital investment, which increases the construction and operation costs of the system.
[0005] Therefore, how to combine modern image recognition technology and video surveillance methods to determine whether the experimenters are wearing personal protective equipment properly by analyzing video stream data has become a technical problem that needs to be solved urgently. Summary of the invention
[0006] The main purpose of the present invention is to provide a personal protective equipment compliance detection method, which aims to combine modern image recognition technology and video monitoring means to determine whether the experimenter is wearing personal protective equipment in a standardized manner by analyzing video stream data.
[0007] In order to achieve the above object, the present invention proposes a personal protective equipment compliance detection method, comprising the following steps:
[0008] Construct a two-dimensional human perception network, a three-dimensional human shape reconstruction network, and a hierarchical protective equipment judgment network;
[0009] Obtain video data, extract features from the video data through the two-dimensional human perception network, and generate accurate coordinates of human bone key points and body joints;
[0010] Based on the human bone key points, generate the working state result of the target person through the three-dimensional human shape reconstruction network;
[0011] Based on the human bone key points and video data, generate the protective equipment wearing feature through the hierarchical protective equipment judgment network, and judge whether the working state result of the target person meets the preset result. If so, continue to judge whether the protective equipment wearing feature meets the preset condition. If the protective equipment wearing feature meets the preset condition, output that the personal protective equipment is worn in compliance.
[0012] In an embodiment of the present application, when extracting features from the video data through the two-dimensional human perception network to generate accurate coordinates of human bone key points and body joints, the following steps are further included:
[0013] Split the video data into continuous video frames, perform human detection on each video frame, and segment the region of interest corresponding to the human target;
[0014] Extract features within each region of interest to obtain global pose information and local pose information;
[0015] Extract bone key points according to the global pose information and local pose information;
[0016] Perform feature fusion on the bone key points through a multi-layer perceptron to generate accurate coordinates of body joints.
[0017] In an embodiment of the present application, when generating the working state result of the target person through the three-dimensional human shape reconstruction network based on the human bone key points, the following steps are further included:
[0018] Project the bone key points into three-dimensional space to form a skeleton representation;
[0019] Based on the skeleton representation, generate the body shape parameters and joint angle parameters of the target person through the SMPL model, and generate an initial mesh based on the body shape parameters and joint angle parameters of the target person;
[0020] Optimize the initial mesh to generate a detailed human mesh;
[0021] Generate the working state result of the target person according to the detailed human mesh, global pose information, and the corresponding video frame.
[0022] In one embodiment of the present application, optimizing the initial human body network includes the following steps:
[0023] Select local meshes in the initial mesh, and align the geometric consistency and surface smoothness of the local mesh vertices with the initial image perception of the corresponding video frame;
[0024] Judge the visibility of each joint of the human body in the initial mesh. When a joint is invisible, deduce the position information of the corresponding invisible joint according to the prior knowledge encoded in the SMPL model and the context information of the visible joints to generate a detailed human body mesh.
[0025] In one embodiment of the present application, based on the human body bone key points and video data, generating protection equipment wearing features through a hierarchical protection equipment judgment network includes the following steps:
[0026] Determine the key parts of the human body through the bone key points;
[0027] Generate an overall mask of the target person according to the key parts through the SAM algorithm, so as to realize the segmentation of the target person from the background;
[0028] Obtain local masks according to the precise coordinates of the body joints and the overall mask;
[0029] Generate protection equipment wearing features according to the local masks and the corresponding video frames.
[0030] In one embodiment of the present application, the local mask includes at least one of a face mask, a hand mask, a trunk mask, and a foot mask.
[0031] In one embodiment of the present application, the two-dimensional human body perception network is constructed based on the AlphaPose architecture.
[0032] In one embodiment of the present application, the three-dimensional human body shape reconstruction network is constructed based on the Hybrik framework.
[0033] In one embodiment of the present application, after constructing the two-dimensional human body perception network, the three-dimensional human body shape reconstruction network, and the hierarchical protection equipment judgment network, it further includes:
[0034] Obtain training videos, preprocess the training data to generate enhanced data;
[0035] Input the enhanced data into the two-dimensional human body perception network and the three-dimensional human body shape reconstruction network, and set training hyperparameters;
[0036] Through a two-dimensional human perception network, a three-dimensional human shape reconstruction network, and a hierarchical protective equipment judgment network, obtain the working state and protective equipment wearing state of the target person, input the working state and protective equipment wearing state of the target person into the loss function, and train the two-dimensional human perception network and the three-dimensional human shape reconstruction network through backpropagation.
[0037] In an embodiment of the present application, the preprocessing includes at least one of disturbing, flipping, and stretching the training video.
[0038] Adopting the above technical solution, by constructing a two-dimensional human perception network, a three-dimensional human shape reconstruction network, and a hierarchical protective equipment judgment network, it is possible to accurately extract the accurate coordinates of human bone key points and body joints from video data, and generate the working state result of the target person through the three-dimensional human shape reconstruction network. The hierarchical protective equipment judgment network further combines the working state and the wearing characteristics of the protective equipment, and can accurately judge whether the wearing of the protective equipment meets the preset conditions and output a compliance result. This method can achieve efficient and accurate personal protective equipment compliance detection in complex scenarios, greatly improving the detection efficiency and the reliability of the results. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The present invention will be described in detail below with reference to specific embodiments and drawings, where:
[0040] Figure 1 It is a schematic structural diagram of the first embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be described in detail below with reference to the drawings and embodiments. It should be understood that the following specific embodiments are only used to explain the present invention and do not limit the present invention.
[0042] As Figure 1 shown, in order to achieve the above purpose, the present invention proposes a personal protective equipment compliance detection method, including the following steps:
[0043] Construct a two-dimensional human perception network, a three-dimensional human shape reconstruction network, and a hierarchical protective equipment judgment network;
[0044] Obtain video data, extract features from the video data through the two-dimensional human perception network, and generate the accurate coordinates of human bone key points and body joints;
[0045] Based on the human bone key points, generate the working state result of the target person through the three-dimensional human shape reconstruction network;
[0046] Based on the human body bone key points and video data, through a hierarchical protective equipment judgment network, generate protective equipment wearing characteristics, and judge whether the working state result of the target person meets the preset result. If so, continue to judge whether the protective equipment wearing characteristics meet the preset conditions. If the protective equipment wearing characteristics meet the preset conditions, output that the personal protective equipment is worn in compliance.
[0047] Specifically, construct a two-dimensional human body perception network, a three-dimensional human body shape reconstruction network, and a hierarchical protective equipment judgment network. The two-dimensional human body perception network is modified and extended based on the AlphaPose architecture and can generate the accurate coordinates of human body bone key points and body joints. The three-dimensional human body shape reconstruction network is extended based on the Hybrik (inverse dynamics algorithm) framework and is used to generate a three-dimensional human body mesh and the working state result of the target person. The hierarchical protective equipment judgment network is constructed through a multi-layer perceptron algorithm and can generate wearing characteristics for protective equipment in different parts and realize the judgment of the working state of the target person and the compliance of wearing protective equipment.
[0048] Obtain video data, split the video data into frames and input them into the two-dimensional human body perception network. The two-dimensional human body perception network generates the accurate coordinates of human body bone key points and body joints from each frame by extracting the multi-scale spatial features of the video frames.
[0049] Based on the above-generated human body bone key point data, input it into the three-dimensional human body shape reconstruction network. The three-dimensional human body shape reconstruction network first maps the two-dimensional bone key points to the canonical three-dimensional space to generate a three-dimensional human body skeleton structure. Then, combined with the human body mesh reconstruction algorithm, it generates a three-dimensional mesh representation of the human body and extracts three-dimensional working state features. By fusing the human body mesh information and the global pose features of the video frames, finally generate the working state result of the target person.
[0050] Based on the above human body bone key points and video data, input them into the hierarchical protective equipment judgment network. The hierarchical protective equipment judgment network generates protective equipment wearing characteristics according to the protective equipment-related features extracted from the video frames. Specifically, the hierarchical protective equipment judgment network combines the wearing areas of the protective equipment (such as the face, hands, torso, etc.) and the target working state result to judge whether the working state result of the target person meets the preset result. If the working state result meets the preset result, then further judge whether the protective equipment worn by the target person meets the preset conditions according to the protective equipment wearing characteristics. If the conditions are met, output the result that the personal protective equipment is worn in compliance.
[0051] By adopting the above technical solution, through constructing a two-dimensional human perception network, a three-dimensional human shape reconstruction network, and a hierarchical protective equipment judgment network, the accurate extraction of the precise coordinates of human bone key points and body joints from video data is realized, and the working state result of the target person is generated through the three-dimensional human shape reconstruction network. The hierarchical protective equipment judgment network further combines the working state and the wearing characteristics of the protective equipment, and can accurately judge whether the wearing of the protective equipment meets the preset conditions and output the compliance result. This method can achieve efficient and accurate compliance detection of personal protective equipment in complex scenarios, greatly improving the detection efficiency and the reliability of the result.
[0052] In an embodiment of the present application, the video data is subjected to feature extraction by the two-dimensional human perception network to generate the precise coordinates of human bone key points and body joints, and the following steps are further included:
[0053] The video data is split into consecutive video frames, and for each video frame, human detection is performed to segment the region of interest corresponding to the human target;
[0054] Feature extraction is performed within each region of interest to obtain global pose information and local pose information;
[0055] Bone key points are extracted according to the global pose information and local pose information;
[0056] The bone key points are subjected to feature fusion through a multi-layer perceptron to generate the precise coordinates of body joints.
[0057] Specifically, the obtained video data is decomposed into multiple consecutive video frames according to the frame rate, and each frame of data represents an independent time slice for subsequent frame-by-frame processing.
[0058] Human detection is performed on each video frame, and a target detection algorithm is used to locate the human region contained in the frame. According to the detection result, the region of interest corresponding to the human target is segmented to eliminate background interference and focus on the human target.
[0059] Multi-scale feature extraction is performed in each region of interest, and the two-dimensional human perception network is used to extract global pose information and local pose information. Among them, the global pose information captures the overall pose characteristics of the human body through large-scale features, and the local pose information captures the fine features of key parts of the human body through small-scale features, such as the detailed poses of the head, hands, and torso.
[0060] Based on the global pose information and local pose information, the two-dimensional human perception network is applied to generate human bone key points. The extraction of bone key points is based on the prior knowledge of the human joint structure, and a multi-stage regression method is used to gradually optimize the prediction result of the key points to ensure the high precision and stability of the bone key points.
[0061] The extracted skeletal key points are input into a multi-layer perceptron to fuse the features of the skeletal key points, combining the spatial correlation and context information between the key points. Through hierarchical feature learning, the multi-layer perceptron further optimizes and improves the data representation of the skeletal key points, generating body joint data containing precise coordinates.
[0062] By adopting the above technical solution, by splitting the video data into consecutive video frames and performing human detection and region of interest segmentation, the focus of feature extraction and the minimization of background interference are ensured. Global and local pose information is extracted within the region of interest, and the features of the skeletal key points are fused by combining with a multi-layer perceptron, realizing the high-precision extraction of human skeletal key points and the generation of precise coordinates of body joints. This method has high robustness, can adapt to background changes and occlusion situations in complex scenarios, and effectively improves the extraction efficiency and accuracy of human pose information.
[0063] In an embodiment of the present application, based on the human skeletal key points, through a three-dimensional human shape reconstruction network, a working state result of the target person is generated, and the following steps are further included:
[0064] Project the skeletal key points into a three-dimensional space to form a skeleton representation;
[0065] Based on the skeleton representation, through the SMPL model, generate the body shape parameters and joint angle parameters of the target person, and generate an initial mesh based on the body shape parameters and joint angle parameters of the target person;
[0066] Optimize the initial mesh to generate a detailed human mesh;
[0067] Generate a working state result of the target person according to the detailed human mesh, global pose information, and the corresponding video frame.
[0068] Specifically, the skeletal key point data extracted by the two-dimensional human perception network is input into the three-dimensional human shape reconstruction network, and the two-dimensional skeletal key points are projected into a canonical three-dimensional space through a three-dimensional space mapping algorithm to generate a three-dimensional skeleton representation. The three-dimensional skeleton representation accurately describes the joint positions of the target person and their geometric relationships in the three-dimensional space.
[0069] Based on the three-dimensional skeleton representation, combined with the SMPL (Skinned Multi-Person Linear) model, calculate the body shape parameters and joint angle parameters of the target person. The SMPL model uses prior knowledge of human anatomy and parametric representation to model the body type characteristics (such as height, weight, etc.) of the target person, generating complete body shape parameters, and at the same time generating joint angle parameters in combination with the skeletal joint positions. Based on the above parameters, an initial three-dimensional mesh of the target person is generated.
[0070] The initial mesh is adjusted using a geometric consistency optimization algorithm, including optimizing the geometric details of local regions to ensure the smoothness of the mesh surface and the perceptual consistency with the original video data. At the same time, for occluded or partially visible joint regions in the video frames, the prior knowledge of the SMPL model and the context information of the visible joints are used to infer the reasonable positions of the unseen joints, thereby generating a complete and detailed human mesh.
[0071] The optimized detailed human mesh is combined with the global pose information, and the pose and motion characteristics of the target person are analyzed through a 3D human shape reconstruction network. Combining the corresponding video frame data, the 3D features of the target person are extracted and input into a multi-layer perceptron, and the working state result of the target person is generated through a classification or regression model. For example, it is recognized whether the target person is in working states such as standing, walking, bending, carrying, etc., and an accurate working state classification result or motion parameters are generated.
[0072] Adopting the above technical solution, by projecting the bone key points into the 3D space and combining the SMPL model to generate the body shape parameters and joint angle parameters, the 3D human shape of the target person can be accurately restored. Through the optimization process of the initial mesh, it is ensured that the generated detailed human mesh has high precision and surface smoothness, and can adapt to occlusion and illumination changes in complex scenarios. Further combining the 3D mesh, global pose information and video frame data to analyze the motion characteristics of the target person, and generating an accurate working state result.
[0073] In an embodiment of the present application, the optimization of the human initial network includes the following steps:
[0074] Select local meshes in the initial mesh, and align the geometric consistency and surface smoothness of the local mesh vertices with the initial image perception of the corresponding video frame;
[0075] Judge the visibility of each joint of the person's body in the initial mesh. When a joint is invisible, according to the prior knowledge encoded in the SMPL model and the context information of the visible joints, deduce the position information of the corresponding invisible joint to generate a detailed human mesh.
[0076] Specifically, in the initial mesh generated by the 3D human shape reconstruction network, specific local meshes are selected for optimization. The local meshes are joint parts (such as shoulders, knees, elbows, etc.) or regions with large pose changes in the human mesh. For these local meshes, the vertex positions are adjusted through a geometric consistency optimization algorithm to ensure that the geometric structure of these local meshes is consistent with the human target shape in the video frame. The optimization process also includes the adjustment of surface smoothness to ensure that the surface of the mesh presents a natural and smooth effect, avoiding obvious geometric defects.
[0077] During the process of local mesh optimization, align the optimized local mesh vertices with the initial image perception data corresponding to the video frame. Utilize the image perception information of the human target area in the video frame, combine the texture features and color information of the human target, and further refine the vertex positions of the mesh through the image features extracted by a deep learning model (such as a convolutional neural network) to make it more conform to the shape of the person in the video frame. This process not only ensures geometric consistency but also improves the adaptability of the mesh to the pose changes of the person in the video frame.
[0078] Judge the visibility of each joint position in the initial mesh. By analyzing the image features and the positions of the skeletal key points in the video frame, determine whether each joint is visible. For invisible joints, deduce their positions using the prior knowledge based on the SMPL model and the context information of visible joints.
[0079] The process is as follows: First, infer the reasonable positions of the joints according to the human skeleton structure and the relative joint position relationship encoded in the SMPL model; then, combine the spatial position information and motion patterns of visible joints to deduce the reasonable positions of invisible joints and update the corresponding joint coordinates in the mesh. Through this deduction mechanism, the occluded areas in the image can be effectively filled, generating a more complete and realistic three-dimensional mesh representation.
[0080] During the optimization process, through local adjustment of the initial mesh and joint visibility deduction, a detailed three-dimensional human mesh is finally generated. This mesh not only has an accurate geometric shape but also can accurately reflect the motion pose and detailed features of the target person. The generated detailed mesh is consistent in space and time and can be continuously used in multi-frame video data to ensure the smoothness and stability of the mesh over time.
[0081] Adopting the above technical solution, by optimizing the local mesh vertices in the initial human mesh, geometric consistency and surface smoothness are ensured. At the same time, through alignment with the image perception in the video frame, the accuracy and naturalness of the mesh are further improved. By judging the visibility of joints and deducing the positions of invisible joints, the occlusion problem is solved, ensuring the integrity of the mesh.
[0082] In an embodiment of the present application, based on the human skeletal key points and video data, generate the wearing characteristics of protective equipment through a hierarchical protective equipment judgment network, including the following steps:
[0083] Determine the key parts of the human body through the skeletal key points;
[0084] Generate an overall mask of the target person according to the key parts through the SAM algorithm, thereby realizing the segmentation of the target person from the background;
[0085] Obtain a local mask based on the precise coordinates of the body joints and the overall mask;
[0086] Generate the wearing characteristics of protective equipment based on the local mask and the corresponding video frame.
[0087] Specifically, using the bone key point data extracted in the previous steps, first determine the key parts of the human body, including important parts such as the head, torso, arms, and legs. Through the spatial relationship of the bone key points, accurately locate the relative positions and geometric shapes of each key part, providing an accurate reference for subsequent mask generation. For example, the head can be determined by the bone key points of the top of the head and the neck, and the hand can be determined by the key points of the shoulder and the wrist.
[0088] Based on the key parts determined above, use the SAM (Segment Anything Model) algorithm to generate the overall mask of the target person. The SAM algorithm constructs an accurate contour of the human target by extracting the context information of the bone key points and the target parts. This mask not only effectively distinguishes the human body and the background regions but also can handle target detection tasks in complex backgrounds. Through the generation of the overall mask, the segmentation of the target person and the background is successfully achieved, providing a clean image area for subsequent local mask generation and protective equipment judgment.
[0089] According to the precise coordinates of the body joints and the generated overall mask, further extract the local mask of the human body. The local mask is extracted for specific parts of the human body (such as the head, hands, torso, etc.) to minimize background interference when analyzing the wearing characteristics of protective equipment in these areas. For example, for the local mask of the hand, the mask area of the hand can be extracted through the bone key point information of the wrist and the elbow, and this area is corresponded to the actual image in the video frame to detect the wearing situation of protective equipment such as gloves and wristbands.
[0090] Combine the extracted local mask with the corresponding video frame data to generate the wearing characteristics of protective equipment. This process extracts fine features related to protective equipment through a convolutional neural network (CNN) or other deep learning methods. For different protective equipment (such as face masks, gloves, goggles, etc.), the network will analyze features such as their coverage range, shape, and material in the local mask area to determine whether the protective equipment is worn correctly. For example, the face mask detection branch will evaluate whether the facial area is covered and judge the integrity of the coverage; the glove detection branch will check whether the hand area is completely covered by the glove and whether the wearing meets the safety standards. The finally generated wearing characteristics of protective equipment are used to determine whether the target person meets the preset protective equipment wearing standards.
[0091] By adopting the above technical solution, through the precise positioning of skeletal key points, the SAM algorithm is used to generate an overall mask to ensure the precise segmentation of the human body from the background and reduce background interference. By extracting the key parts of the human body through local masks and combining video frame data, the wearing characteristics of protective equipment are generated, achieving an accurate assessment of the wearing status of protective equipment.
[0092] In an embodiment of the present application, the local mask includes at least one of a face mask, a hand mask, a torso mask, and a foot mask.
[0093] Specifically, in the foregoing steps, the overall mask generated by the skeletal key points and the SAM algorithm can be further divided into multiple local masks, specifically including masks for multiple local regions such as a face mask, a hand mask, a torso mask, and a foot mask.
[0094] The mask for the face region is determined by facial skeletal key points (such as key positions like eyes, nose, and mouth), which is particularly suitable for the wearing detection of protective equipment such as face shields and masks.
[0095] The mask region for the hand is determined by the skeletal key points of the hand (such as the wrist and elbow), which is applicable to the judgment of glove wearing.
[0096] The mask for the torso region is extracted through the skeletal key points related to the torso (such as the shoulders, chest, and abdomen), which is applicable to the detection of equipment such as protective clothing and chest protectors.
[0097] The mask for the foot region is extracted through the skeletal key points of the foot (such as the ankle and knee), which is applicable to the detection of footwear or ankle protection equipment.
[0098] By adopting the above technical solution, by subdividing local masks such as the face, hands, torso, and feet, the wearing situation of protective equipment can be evaluated more precisely.
[0099] In an embodiment of the present application, the two-dimensional human perception network is constructed based on the AlphaPose architecture.
[0100] Specifically, AlphaPose is an efficient and accurate multi-person pose estimation algorithm widely used in human key point detection tasks. The two-dimensional human perception network improves the detection ability of human skeletal key points through the modification and extension of the AlphaPose architecture. AlphaPose extracts features from images through a deep convolutional neural network (CNN) and predicts the coordinates of each human part through key point regression.
[0101] With the above technical solution, through the two-dimensional human perception network constructed based on the AlphaPose architecture, the human skeleton key points and the coordinates of key parts in the video frame can be efficiently and accurately extracted. By modifying and expanding the AlphaPose architecture, the performance of the network in complex scenarios is optimized, and high accuracy can still be maintained in the case of dense crowds or large pose changes.
[0102] In an embodiment of the present application, the three-dimensional human shape reconstruction network is constructed based on the Hybrik framework.
[0103] Specifically, the Hybrik framework is a deep learning architecture for human three-dimensional reconstruction, which can accurately reconstruct the three-dimensional shape of the human body by combining image data and pose information. This framework integrates the processing ability of multi-modal information, including two-dimensional image information and three-dimensional spatial bone data, making it have efficient performance in human shape reconstruction.
[0104] The Hybrik framework adopts a multi-stage neural network design, which includes a deep convolutional neural network (CNN) for extracting features in the input video frame and mapping these features into three-dimensional space. This framework establishes a mapping relationship between image features and three-dimensional human parameters, enabling the network to generate an accurate three-dimensional human mesh according to the two-dimensional human skeleton key point data.
[0105] First, based on the bone key point data generated by the two-dimensional human perception network, these key points are projected into three-dimensional space through the Hybrik framework. The projected skeleton representation serves as the basis for three-dimensional shape reconstruction. This projection process not only calculates the skeleton structure according to the joint positions, but also combines a three-dimensional prior model (such as the SMPL model) to generate a more accurate initial shape.
[0106] Next, based on the three-dimensional skeleton representation, the framework estimates the body shape parameters and joint angle parameters of the target human body through a regression network. These parameters mainly include the body type information of the human body, the angles of each joint, and the relative positions of different body parts. By regressing these parameters, the framework generates an initial three-dimensional mesh model of the target person.
[0107] The generated initial mesh then enters the optimization stage. Through a local optimization algorithm, the mesh is adjusted for geometric consistency and surface smoothness, making the mesh more conform to the actual image perception in the input video frame. In this stage, the framework also adjusts the temporal consistency of the mesh according to the dynamic changes of the bone key points and the continuity of the human pose.
[0108] By introducing a Recurrent Neural Network (RNN) or Long Short-Term Memory Network (LSTM), the Hybrik framework can effectively model the temporal dependencies between video frames. When processing consecutive video frames, the network can maintain the consistency of the three-dimensional grid in the temporal dimension, avoiding grid deformation or unnatural reconstruction results caused by pose changes. During the optimization process, the Hybrik framework also uses a refinement algorithm to further adjust the detailed parts of the three-dimensional grid, ensuring the accuracy and realism of the grid. This process will involve fine-tuning of joints and key parts to accurately reflect the changes in human poses in dynamic videos.
[0109] After completing the reconstruction of the three-dimensional grid, based on the detailed three-dimensional grid, global pose information, and input data in the video frame, the three-dimensional human shape reconstruction network can generate the working state results of the target person. These working state results can determine whether the target person is in a specific working state (such as standing, walking, sitting, etc.) through the analysis of the three-dimensional grid.
[0110] Adopting the above technical solution, the three-dimensional human shape reconstruction network based on the Hybrik framework can efficiently map the bone key point data generated by the two-dimensional human perception network into the three-dimensional space and accurately reconstruct the three-dimensional shape of the target person. By combining multi-modal features, the Hybrik framework can not only generate high-precision three-dimensional grids but also ensure the temporal consistency of the grids.
[0111] In an embodiment of the present application, after constructing a two-dimensional human perception network, a three-dimensional human shape reconstruction network, and a hierarchical protective equipment judgment network, it further includes:
[0112] Obtain training videos, preprocess the training data to generate enhanced data;
[0113] Input the enhanced data into the two-dimensional human perception network and the three-dimensional human shape reconstruction network, and set training hyperparameters;
[0114] Through the two-dimensional human perception network, the three-dimensional human shape reconstruction network, and the hierarchical protective equipment judgment network, obtain the working state and protective equipment wearing state of the target person, input the working state and protective equipment wearing state of the target person into the loss function, and train the two-dimensional human perception network and the three-dimensional human shape reconstruction network through backpropagation.
[0115] Specifically, first, collect video data containing the activities of the target person. The video data can come from different scenarios and different working environments, including but not limited to indoor, outdoor, workplaces, construction sites, etc. These video data should cover the diverse performances of the target person in different poses and different working states.
[0116] For the obtained original training videos, necessary preprocessing operations are carried out, including video frame extraction, image size adjustment, color normalization, background removal, etc. Through these preprocessings, it is ensured that the video data can reach a high quality when entering the network model, avoiding the influence of noise and redundant information on the training effect of the model.
[0117] To enhance the robustness of the model, especially in the case of dealing with large pose changes and complex environments, a variety of data augmentation methods are used to expand the original training data. Common data augmentation techniques include:
[0118] Image perturbation, randomly changing parameters such as the brightness, contrast, and color of video frames to simulate scenes under different lighting conditions.
[0119] Image flipping and rotation, performing horizontal or vertical flipping, rotation, etc. on the input image to simulate different perspective changes.
[0120] Scale stretching and cropping, randomly cropping the image and adjusting the scale to simulate target persons of different sizes and poses.
[0121] Noise addition, adding Gaussian noise, etc. to the image to increase the model's tolerance to noise.
[0122] Through these enhancement operations, the generated enhanced data can effectively increase the diversity of training samples during the training process, enabling the trained model to have stronger generalization ability and adapt to various pose and environmental changes.
[0123] The generated enhanced data is input into the constructed two-dimensional human perception network and three-dimensional human shape reconstruction network for network training. The enhanced data will provide diverse input information, enabling the network to learn multi-level features such as human key points, bone positions, and body postures.
[0124] During the training process, appropriate hyperparameters are set, including the learning rate, batch size, optimization algorithm (such as Adam, SGD, etc.), number of training epochs, loss function, etc.
[0125] Through the two-dimensional human perception network and three-dimensional human shape reconstruction network, combined with the bone key points, body joint information, and three-dimensional mesh representation extracted from the video frames, the working state result of the target person is generated. Through the hierarchical protective equipment judgment network, combined with the pose information of the target person and the protective equipment features, the wearing state of the protective equipment is generated. This network can identify whether the target person is wearing compliant protective equipment.
[0126] During the training process, the working state and the wearing state of the protective equipment output by the network will be compared with the preset standard labels, and the loss function will be calculated. The loss function consists of multiple parts, including: classification loss, regression loss, smoothing loss, and cosine similarity loss. Through the backpropagation algorithm, according to the gradients calculated from the loss function, the weights of the two-dimensional human perception network, the three-dimensional human shape reconstruction network, and the hierarchical protective equipment judgment network are updated. Through multiple rounds of iterative training, the network gradually adjusts its parameters to improve the accuracy of working state recognition and protective equipment wearing detection.
[0127] Adopting the above technical solutions, through effective data preprocessing and data augmentation techniques, the diversity of training data can be significantly improved, so that the model can still maintain high accuracy when dealing with different working environments and complex postures. Through the joint training of the two-dimensional human perception network and the three-dimensional human shape reconstruction network, the working state of the target person can be accurately recognized, and combined with the hierarchical protective equipment judgment network, the wearing situation of the protective equipment can be effectively evaluated.
[0128] In an embodiment of the present application, the preprocessing includes at least one of perturbation, flipping, and stretching of the training video.
[0129] Specifically, the perturbation is to randomly perturb parameters such as brightness, contrast, and saturation for each frame image in the training video to simulate different lighting conditions.
[0130] Flipping is to perform horizontal or vertical flipping on the training video. By flipping the image horizontally, the reverse situation of observing the target person can be simulated.
[0131] Stretching refers to stretching each frame image in the training video, that is, changing the proportion of the target person by compressing or expanding the horizontal or vertical dimensions of the image.
[0132] Adopting the above technical solutions, through data augmentation processing such as perturbing, flipping, and stretching the training video, the diversity of training data can be greatly increased, and the generalization ability of the model can be improved. The enhanced data can cover various situations such as different lighting conditions, shooting angles, and scale changes, improving the model's ability to handle complex scenarios in the actual environment.
[0133] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structural transformation made under the inventive concept of the present invention, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present invention.
Claims
1. A method for detecting compliance of personal protective equipment, characterized in that: The following steps are involved: Construct a two-dimensional human perception network, a three-dimensional human shape reconstruction network, and a layered protective equipment judgment network; obtain video data, extract features from the video data through the two-dimensional human perception network, and generate accurate coordinates of human skeleton key points and body joints; based on the human skeleton key points, generate the working status results of the target person through the three-dimensional human shape reconstruction network; based on the human skeleton key points and video data, generate protective equipment wearing features through the layered protective equipment judgment network, and judge whether the working status results of the target person meet the preset results. If so, continue to judge whether the protective equipment wearing features meet the preset conditions. If the protective equipment wearing features meet the preset conditions, output personal protective equipment wearing compliance; extract features from the video data through the two-dimensional human perception network to generate accurate coordinates of human skeleton key points and body joints, and also include the following steps: The video data is split into continuous video frames, and human body detection is performed on each video frame to segment out the region of interest corresponding to the human body target; feature extraction is performed in each region of interest to obtain global posture information and local posture information; skeleton key points are extracted according to the global posture information and local posture information; feature fusion is performed on the skeleton key points through a multi-layer perceptron to generate precise coordinates of body joints; based on the human skeleton key points and video data, protective equipment wearing features are generated through a layered protective equipment judgment network, including the following steps: determining key parts of the human body through skeleton key points; generating an overall mask of the target person according to the key parts through a SAM algorithm, thereby achieving segmentation of the target person and the background; obtaining a local mask according to the precise coordinates of the body joints and the overall mask; generating protective equipment wearing features according to the local mask and the corresponding video frames.
2. The personal protective equipment compliance detection method according to claim 1, characterized in that: Based on the human skeleton key points, a working state result of the target person is generated through a three-dimensional human body shape reconstruction network, and the following steps are also included: Projecting the skeleton key points into three-dimensional space to form a skeleton representation; Based on the skeleton representation, generating body shape parameters and joint angle parameters of the target person through the SMPL model, and generating an initial mesh based on the body shape parameters and joint angle parameters of the target person; Optimizing the initial mesh to generate a detailed human body mesh; The working status results of the target person are generated based on the detailed human body mesh, global posture information, and the corresponding video frames.
3. The method for detecting compliance of personal protective equipment according to claim 2, characterized in that: Optimizing the human body initial network comprises the following steps: Selecting a local mesh from the initial mesh, and aligning the geometric consistency and surface smoothness of the local mesh vertices with the initial image perception of the corresponding video frame; The visibility of each joint of the character's body in the initial mesh is judged. When the joint is invisible, the position information of the corresponding invisible joint is derived according to the prior knowledge encoded in the SMPL model and the contextual information of the visible joints to generate a detailed human body mesh.
4. The method for detecting compliance of personal protective equipment according to claim 1, characterized in that: The local mask includes at least one of a face mask, a hand mask, a torso mask, and a foot mask.
5. The method for detecting compliance of personal protective equipment according to claim 1, characterized in that: The two-dimensional human body perception network is constructed based on the Alpha gesture architecture.
6. The method for detecting compliance of personal protective equipment according to claim 1, characterized in that: The three-dimensional human body shape reconstruction network is constructed based on the Hybrik framework.
7. The method for detecting compliance of personal protective equipment according to claim 1, characterized in that: After building a 2D human perception network, a 3D human shape reconstruction network, and a layered protective equipment judgment network, it also includes: Obtain training videos, preprocess the training data, and generate enhanced data; Input the enhanced data into the 2D human perception network and 3D human shape reconstruction network, and set the training hyperparameters; Through the two-dimensional human perception network, three-dimensional human shape reconstruction network, and layered protective equipment judgment network, the working status of the target person and the wearing status of the protective equipment are obtained, the working status of the target person and the wearing status of the protective equipment are input into the loss function, and the two-dimensional human perception network and three-dimensional human shape reconstruction network are trained through back propagation.
8. The method for detecting compliance of personal protective equipment according to claim 7, characterized in that: The preprocessing includes: at least one of perturbation, flipping, and stretching of the training video.
Citation Information
Patent Citations
Three-dimensional human motion capturing method and device
CN118609206A
Protective clothing wearing identification method based on human body key points
CN118898857A
Single-target pose estimation method and apparatus, and electronic device
WO2025000278A1