Human body posture detection method and device, electronic equipment and storage medium

Through infrared depth image processing and model training, human body areas and skeletal key points are identified, and human skeleton feature vectors are generated, which solves the problem of inaccurate human posture assessment in autonomous driving and improves safety.

CN120599663APending Publication Date: 2025-09-05BEIJING JINGWEI HIRAIN TECH CO INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510756835.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

It is difficult to accurately assess human posture using existing technologies, which leads to safety risks in autonomous driving.

Method used

By acquiring infrared depth images, using pre-trained human body detection models and human body key point detection models, human body areas and bone key points are identified, and human skeleton feature vectors are generated to determine human body posture.

Benefits of technology

It achieves accurate detection of human posture and improves the safety of autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599663A_ABST
    Figure CN120599663A_ABST
Patent Text Reader

Abstract

The invention provides a human body posture detection method and device, electronic equipment and a storage medium. The human body posture detection method comprises the following steps: firstly, acquiring an infrared depth image in a target scene; and inputting the infrared depth image into a pre-constructed human body detection model for human body detection to obtain a human body region image. Inputting the human body region image into a pre-constructed human body key point detection model for key point detection to obtain coordinates and confidence of human body skeleton key points in the human body region image; and determining the type of each human skeleton key point based on the confidence coefficient. And finally, connecting each human skeleton key point based on the type to obtain a human skeleton feature vector, and determining a human posture based on the human skeleton feature vector. By using the method provided by the invention, the key points of the human skeleton can be accurately identified based on the infrared depth image, so that the feature vector of the human skeleton is obtained, the human posture is determined, and the safety of automatic driving is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and in particular to a method, device, electronic device and storage medium for detecting human posture. Background Art

[0002] With the development of science and technology, autonomous driving should become more and more widespread. The driving of autonomous vehicles still depends on the close cooperation between people and systems, and safe and reasonable human-computer interaction can be carried out through the detection of human posture.

[0003] In existing technologies, when detecting the human body, problems such as target positioning errors and errors in the detection areas of key points of the human body are prone to occur, making it difficult to accurately assess the human body posture, resulting in safety hazards in autonomous driving. Summary of the Invention

[0004] In view of this, the present application provides a method, device, electronic device and storage medium for detecting human posture to solve the problem in the existing technology that it is difficult to accurately evaluate human posture, resulting in safety hazards in autonomous driving.

[0005] To achieve the above objectives, this application provides the following technical solutions:

[0006] The first aspect of the present application discloses a method for detecting human posture, comprising:

[0007] Acquire infrared depth images of the target scene;

[0008] Inputting the infrared depth image into a pre-built human body detection model to perform human body detection to obtain a human body region image; wherein the human body detection model is trained based on infrared sample images with human body regions marked;

[0009] Inputting the human body region image into a pre-built human body key point detection model to perform key point detection, thereby obtaining coordinates and confidence levels of human skeletal key points in the human body region image; wherein the human body key point detection model is trained based on infrared sample images labeled with human skeletal key points;

[0010] Determining the type of each of the human skeleton key points based on the confidence level;

[0011] Based on the type, key points of the human skeleton are connected to obtain a human skeleton feature vector, and the human body posture is determined based on the human skeleton feature vector.

[0012] Optionally, the above method, wherein the inputting the infrared depth image into a pre-built human body detection model to perform human body detection to obtain a human body region image, includes:

[0013] When a human body is detected in the infrared depth image, the area where the human body is currently located is marked using a human body detection frame, and a human body area image and a human body area image confidence score are generated;

[0014] If there is only one human body region image corresponding to the current human body, directly output the human body region image corresponding to the current human body as the final result;

[0015] If there are multiple human body region images corresponding to the current human body, the human body region image with the highest confidence level is output as the final result.

[0016] Optionally, in the above method, the training process of the human body detection model includes:

[0017] Obtain infrared sample images with human body regions marked;

[0018] Inputting the infrared sample image into the human body detection initial model to perform human body detection, and obtaining a human body region image corresponding to the current infrared sample image;

[0019] Determine whether the human body region image corresponding to the current infrared sample image is consistent with the actually marked human body region;

[0020] If the human body region image corresponding to the current infrared sample image is consistent with the actually marked human body region, the construction of the human body detection model is completed;

[0021] If the human body area image corresponding to the current infrared sample image is inconsistent with the actual marked human body area, the loss function is calculated and the parameters of the initial human body detection model are adjusted based on the loss function until the human body area image corresponding to the current infrared sample image is consistent with the actual marked human body area, and the construction of the human body detection model is completed.

[0022] Optionally, the method described above, wherein the human body region image is input into a pre-built human body key point detection model for key point detection, and the coordinates and confidence levels of the human skeleton key points in the human body region image are obtained, comprises:

[0023] Performing pooling processing on the human body region image to obtain a first vector;

[0024] Processing the first vector through a fully connected layer to obtain a second vector;

[0025] Multiplying the second vector by N convolution kernels generated by random initialization to obtain N multiplication results; where N is a preset positive integer;

[0026] Adding the N multiplication results to obtain convolution kernel parameters, and performing depth-separable convolution processing on the human body region image using the convolution kernel parameters to obtain a processed feature image;

[0027] The feature image is divided into channels to obtain K sub-feature images and the confidence of each sub-feature image; wherein K is a preset positive integer, and the sub-feature image is used to represent the key points of the human skeleton.

[0028] Optionally, in the above method, determining the type of each of the human skeleton key points based on the confidence level includes:

[0029] If the confidence level of the current human skeleton key point is less than a preset first threshold, the type of the current human skeleton key point is the first type; wherein the types include the first type, the second type, and the third type; the first type indicates that the key point does not exist, the second type indicates that the key point exists but is not visible, and the third type indicates that the key point exists and is visible;

[0030] If the confidence level of the current human skeleton key point is between the first threshold and a preset second threshold, the type of the current human skeleton key point is the second type;

[0031] If the confidence of the current human skeleton key point is greater than a preset second threshold, the type of the current human skeleton key point is the third type; wherein the second threshold is greater than the first threshold.

[0032] Optionally, in the above method, the convolution in the human key point detection model is a lightweight dynamic convolution module; wherein, the lightweight dynamic convolution module includes a pooling layer and a fully connected layer; the lightweight dynamic convolution module adjusts the convolution parameters according to the human body area image input into the human key point detection model.

[0033] The second aspect of the present application discloses a human body posture detection device, comprising:

[0034] An acquisition unit, used for acquiring an infrared depth image of a target scene;

[0035] A first detection unit is configured to input the infrared depth image into a pre-built human body detection model to perform human body detection and obtain a human body region image; wherein the human body detection model is trained based on infrared sample images with human body regions marked;

[0036] A second detection unit is configured to input the human body region image into a pre-built human body key point detection model to perform key point detection, thereby obtaining coordinates and confidence levels of human skeletal key points in the human body region image; wherein the human body key point detection model is trained based on infrared sample images labeled with human skeletal key points;

[0037] A determination unit is used to determine the type of each of the human skeleton key points based on the confidence level; a connection unit is used to connect each of the human skeleton key points based on the type to obtain a human skeleton feature vector, and determine the human body posture based on the human skeleton feature vector.

[0038] Optionally, in the above device, the first detection unit includes:

[0039] an identification subunit, configured to, when a human body is detected in the infrared depth image, identify the area where the current human body is located by using a human body detection frame, and generate a human body area image and a human body area image confidence score;

[0040] a first output unit, configured to directly output the human body region image corresponding to the current human body as a final result if there is only one human body region image corresponding to the current human body;

[0041] The second output unit is configured to output the human body region image with the highest confidence as a final result if there are multiple human body region images corresponding to the current human body.

[0042] Optionally, in the above device, the first detection unit includes:

[0043] An acquisition subunit, used to acquire an infrared sample image with a human body region marked;

[0044] A detection subunit, configured to input the infrared sample image into an initial human body detection model to perform human body detection, and obtain a human body region image corresponding to the current infrared sample image;

[0045] A judging subunit, configured to judge whether the human body region image corresponding to the current infrared sample image is consistent with the actually marked human body region;

[0046] A construction subunit, configured to complete the construction of the human body detection model if the human body region image corresponding to the current infrared sample image is consistent with the actually marked human body region;

[0047] The parameter adjustment subunit is used to calculate the loss function if the human body area image corresponding to the current infrared sample image is inconsistent with the actual marked human body area, and adjust the parameters of the initial human body detection model based on the loss function until the human body area image corresponding to the current infrared sample image is consistent with the actual marked human body area, thereby completing the construction of the human body detection model.

[0048] A third aspect of the present application discloses an electronic device, comprising:

[0049] one or more processors;

[0050] a storage device having one or more programs stored thereon;

[0051] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of the first aspects of the present application.

[0052] A fourth aspect of the present application discloses a computer storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method described in any one of the first aspects of the present application is implemented.

[0053] As can be seen from the above technical solution, in a method for detecting human posture provided by the present application, an infrared depth image of the target scene is first obtained. The infrared depth image is then input into a pre-built human detection model for human body detection to obtain a human body region image. The human body region image is then input into a pre-built human key point detection model for key point detection to obtain the coordinates and confidence of the human skeleton key points in the human body region image. The type of each human skeleton key point is then determined based on the confidence. Finally, each human skeleton key point is connected based on the type to obtain a human skeleton feature vector, and the human posture is determined based on the human skeleton feature vector. It can be seen that using the method of the present application, it is possible to accurately perform human body detection on the infrared depth image through the human detection model to obtain a human body region image, and accurately identify the human skeleton key points in the human body region image through the human key point detection model, thereby obtaining a human skeleton feature vector and determining the human posture, thereby improving the safety of autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0055] Figure 1This is a flow chart of a method for detecting human posture disclosed in an embodiment of the present application;

[0056] Figure 2 This is a flowchart of an implementation of step S102 disclosed in another embodiment of the present application;

[0057] Figure 3 This is a flowchart of an implementation of step S103 disclosed in another embodiment of the present application;

[0058] Figure 4 This is a schematic diagram of the distribution of key points of the human skeleton disclosed in another embodiment of the present application;

[0059] Figure 5 A schematic diagram of a human body posture detection device disclosed in another embodiment of the present application;

[0060] Figure 6 This is a schematic diagram of an electronic device disclosed in another embodiment of the present application. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0062] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0063] Furthermore, in this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any actual relationship or order between these entities or operations.

[0064] As can be seen from the background technology, in the existing technology, when detecting the human body, problems such as target positioning errors and errors in the detection area of ​​key points of the human body are prone to occur, making it difficult to accurately assess the human body posture, resulting in safety hazards in autonomous driving.

[0065] In view of this, the present application provides a method, device, electronic device and storage medium for detecting human posture to solve the problem in the prior art that it is difficult to accurately assess human posture, resulting in safety hazards in autonomous driving.

[0066] The present application provides a method for detecting human body posture. Figure 1 As shown, specifically including:

[0067] S101: Acquire an infrared depth image of a target scene.

[0068] It should be noted that a TOF camera is used to obtain an infrared depth image of the target scene, where the target scene is the scene of the vehicle in motion. The TOF camera can continuously send light pulses to the human body, use a sensor to receive the light returned from the object, and obtain the distance between the human body and the camera by detecting the round-trip time of the light pulse. ToF can directly output depth information with better real-time performance. To meet the low-latency and real-time requirements in the field of smart cockpits, a ToF camera is used to directly obtain the depth information of the human body in the target scene, thereby achieving more accurate and rapid recognition of the corresponding human posture.

[0069] S102: Input the infrared depth image into a pre-built human body detection model to perform human body detection to obtain a human body region image; wherein the human body detection model is trained based on infrared sample images with human body regions marked.

[0070] It should be noted that the acquired infrared depth image is input into a pre-built human detection model for human detection. The human area is identified and marked with a human detection frame to obtain a human area image. The human detection model is trained based on infrared sample images with human area labels. These infrared sample images with human area labels can be collected using a TOF camera or obtained from the COCO (Common Objects in Context) dataset. This dataset is a large-scale image dataset that not only contains a large number of human images in various poses, but also contains human detection frame labeling information.

[0071] Optionally, in another embodiment of the present application, an implementation of step S102 may include:

[0072] When a human body is detected in an infrared depth image, the area where the current human body is located is marked by a human body detection frame, and a human body area image and a human body area image confidence score are generated.

[0073] If there is only one human body region image corresponding to the current human body, the human body region image corresponding to the current human body is directly output as the final result.

[0074] If there are multiple human body region images corresponding to the current human body, the human body region image with the highest confidence level is output as the final result.

[0075] It should be noted that in the process of human body detection by the human body detection model, when a human body is detected in the infrared depth image, the area where the current human body is located is marked by the human body detection frame, and then the human body area image selected by the human body detection frame and the human body area image confidence are generated. The human body area image confidence is used to characterize the credibility of the generated human body area image as the image of the currently detected human body. In order to avoid the situation where multiple human body area images are generated for the same human body due to misidentification, when there is only one human body area image corresponding to the current human body, the human body area image corresponding to the current human body is directly output as the final result; when there are multiple human body area images corresponding to the current human body, the human body area image with the highest human body area image confidence is output as the final result, so as to ensure that each detected human body has only one human body area image, and to ensure the uniqueness and accuracy of the output results of the subsequent target detection model.

[0076] Optionally, in another embodiment of the present application, an implementation of the training process of the human body detection model in step S102 is as follows: Figure 2 As shown, this may include:

[0077] S201: Obtain an infrared sample image with a human body region marked.

[0078] It should be noted that we first obtain infrared sample images with human body areas marked. This can be collected in real time using a TOF camera or obtained from the COCO (Common Objects in Context) dataset.

[0079] S202: Input the infrared sample image into the initial human body detection model to perform human body detection, and obtain a human body region image corresponding to the current infrared sample image.

[0080] It should be noted that the infrared sample image is input into the human body detection initial model for human body detection to obtain a human body region image corresponding to the current infrared sample image, wherein the human body detection initial model adopts a lightweight model.

[0081] S203: Determine whether the human body region image corresponding to the current infrared sample image is consistent with the actually marked human body region.

[0082] It should be noted that after the human body region image corresponding to the current infrared sample image is detected in the human body detection initial model, the human body region image corresponding to the current infrared sample image is compared with the human body region actually annotated in the current infrared sample image to determine, for example, whether the coordinates of the human body detection frame are consistent. This allows the accuracy of the human body region image corresponding to the current infrared sample image detected by the human body detection initial model to be determined.

[0083] S204: If the human body region image corresponding to the current infrared sample image is consistent with the actually marked human body region, the construction of the human body detection model is completed.

[0084] It should be noted that if the human body region image corresponding to the current infrared sample image is consistent with the actual marked human body region, it means that the detection result of the model is accurate, and the construction of the risk assessment model is completed.

[0085] S205. If the human body area image corresponding to the current infrared sample image is inconsistent with the actual marked human body area, the loss function is calculated, and the parameters of the initial human body detection model are adjusted based on the loss function until the human body area image corresponding to the current infrared sample image is consistent with the actual marked human body area, and the construction of the human body detection model is completed.

[0086] It should be noted that if the human body region image corresponding to the current infrared sample image is inconsistent with the actual annotated human body region, the model's detection results are inaccurate. An error function is then calculated, and the parameters of the human body detection model are adjusted based on the error function. A new round of detection is then performed until the human body region image corresponding to the current infrared sample image is consistent with the actual annotated human body region. This completes the construction of the human body detection model.

[0087] S103. Input the human body region image into a pre-built human body key point detection model for key point detection to obtain the coordinates and confidence levels of the human skeleton key points in the human body region image; wherein the human body key point detection model is trained based on infrared sample images marked with human skeleton key points.

[0088] It should be noted that after acquiring the human body region image, the human body region image is subjected to key point detection using a pre-built human body key point detection model to obtain the coordinates of the human skeleton key points in the human body region image and the corresponding confidence levels. The human body key point detection model is trained based on infrared sample images annotated with human skeleton key points. First, infrared images are captured using a TOF camera. Then, the captured infrared images are annotated with reference to the annotation files of the COCO dataset to obtain infrared sample images annotated with human skeleton key points. This is used to create a sample image dataset for model training. The training method can adopt the training method of the above-mentioned human body detection model, which will not be repeated here.

[0089] In this human key point detection model, the human key point detection indicator used is the key point similarity (Object Keypoint Similarity, OKS), as shown in the following formula:

[0090]

[0091] Among them, p is the number of the person, i is the number of the key point of the human body, is the Euclidean distance of the predicted key points for each person, s is the scale factor of the current person, is the normalization factor of the i-th human skeleton point, reflecting the standard deviation of the current skeleton annotation. The larger the value, the more difficult it is to annotate the point. The key point type of the i-th key point of the p-th person is 0, 1, or 2, which respectively represent: 0 means the key point does not exist in the ground truth, 1 means the key point exists but is invisible or occluded, and 2 means the key point exists and is visible.

[0092] The evaluation indicator of the human key point detection model is the mean average precision (mAP). Specifically, different values ​​are set for the artificial threshold T in the AP indicator, and then multiple precision indicators are obtained. Finally, the multiple precision indicators are averaged to obtain the average precision, as shown in the following formula:

[0093]

[0094] Optionally, in another embodiment of the present application, an implementation of step S103 is as follows: Figure 3 As shown, this may include:

[0095] S301: Perform pooling processing on the human body region image to obtain a first vector.

[0096] It should be noted that the human body region image is a feature map of size 64×48×32. The average pooling of the channel dimension is performed through the pooling layer in the human key point detection to obtain the first vector of size 1×1×32.

[0097] S302: Process the first vector through a fully connected layer to obtain a second vector.

[0098] It should be noted that the first vector is then processed through a fully connected layer with an input size of 32 and an output size of 4 in human key point detection to obtain a vector of size 1×1×4.

[0099] S303. Multiply the second vector by N convolution kernels generated by random initialization to obtain N multiplication results; where N is a preset positive integer.

[0100] It should be noted that the second vector is multiplied by N convolution kernels generated by random initialization, for example, N is 4, that is, the second vector of size 1×1×4 is multiplied by 4 convolution kernels of size 3×3×32 generated by random initialization, to obtain four multiplication results.

[0101] S304: Add the N multiplication results to obtain convolution kernel parameters, and use the convolution kernel parameters to perform depth-wise separable convolution processing on the human body region image to obtain a processed feature image.

[0102] It should be noted that the four multiplication results are added together to obtain a convolution kernel parameter of size 3×3×32. Then, the convolution kernel parameter of size 3×3×32 is used to perform a depthwise separable convolution operation on the feature image of size 64×48×32 input into the lightweight dynamic convolution module to generate a feature image of size 64×48×32.

[0103] S305 , dividing the feature image into channels to obtain K sub-feature images and the confidence level of each sub-feature image; wherein K is a preset positive integer, and the sub-feature images are used to represent key points of the human skeleton.

[0104] It should be noted that in this embodiment, K takes the value of 17, and the processed feature image is divided into channels to generate a feature image of size 64×48×17. The feature image of size 64×48×17 contains 17 sub-feature images of size 64×48, so 17 sub-feature images of size 64×48 and the confidence of each sub-feature image are obtained. These 17 sub-feature images of size 64×48 correspond to the 17 predicted key points of the human skeleton. The key points are as follows: Figure 4 As shown, they are: 0-nose, 1-left eye, 2-right eye, 3-left ear, 4-right ear, 5-left shoulder, 6-right shoulder, 7-left elbow, 8-right elbow, 9-left wrist, 10-right wrist, 11-left hip, 12-right hip, 13-left knee, 14-right knee, 15-left foot, 16-right foot.

[0105] Optionally, in another embodiment of the present application, the convolution in the above-mentioned human key point detection model is a lightweight dynamic convolution module;

[0106] The lightweight dynamic convolution module, consisting of a pooling layer and a fully connected layer, adjusts convolution parameters based on input features (i.e., the human region image fed into the human keypoint detection model). This addresses the issue of variability in human posture to a certain extent. Furthermore, compared to conventional convolution, the lightweight dynamic convolution module reduces the number of parameters and computational complexity.

[0107] S104: Determine the type of each human skeleton key point based on the confidence level.

[0108] It should be noted that the confidence level of each skeletal keypoint in a human region image can be used to determine its type. Different types of skeletal keypoints have different characteristics. For example, some keypoints may be absent, some may be present but invisible, and some may be present and visible. By determining the presence and visibility of each skeletal keypoint based on its type, accurate human posture assessment can be performed.

[0109] Optionally, in another embodiment of the present application, an implementation of the above step S104 may include:

[0110] If the confidence of the current human skeleton key point is less than a preset first threshold, the type of the current human skeleton key point is the first type; wherein the types include the first type, the second type and the third type; the first type indicates that the key point does not exist, the second type indicates that the key point exists but is not visible, and the third type indicates that the key point exists and is visible.

[0111] If the confidence level of the current human skeleton key point is between the first threshold and the preset second threshold, the type of the current human skeleton key point is the second type.

[0112] If the confidence of the current human skeleton key point is greater than a preset second threshold, the type of the current human skeleton key point is the third type; wherein the second threshold is greater than the first threshold.

[0113] It should be noted that the types of human skeleton key points are divided into three categories, namely type 1 (key point does not exist), type 2 (key point exists but is not visible) and type 3 (key point exists and is visible). The second threshold is greater than the first threshold, for example, the second threshold is 0.6 and the first threshold is 0.4. If the confidence of the current human skeleton key point is less than 0.4, the type of the current human skeleton key point is type 1, that is, the current key point does not exist; if the confidence of the current human skeleton key point is between 0.4 and 0.6, the type of the current human skeleton key point is type 2, that is, the key point exists but is not visible. If the confidence of the current human skeleton key point is greater than 0.6, the type of the current human skeleton key point is type 3, that is, the key point exists and is visible.

[0114] S105 . Connecting key points of the human skeleton based on the type to obtain a human skeleton feature vector, and determining the human body posture based on the human skeleton feature vector.

[0115] It should be noted that the human skeleton key points are connected according to their type and the positions of the joints of the human limbs to form a human skeleton feature vector. Key points that do not exist or are present but invisible are not connected, while key points that are present and visible are connected. After generating the human skeleton feature vector, the human posture can be identified using it. Autonomous vehicles can perform relevant driving operations by identifying human posture, thereby improving autonomous driving safety.

[0116] In a method for detecting human posture provided by an embodiment of the present application, an infrared depth image of a target scene is first acquired. The infrared depth image is then input into a pre-built human detection model for human body detection to obtain a human body region image. The human body region image is then input into a pre-built human key point detection model for key point detection to obtain the coordinates and confidence levels of the human skeleton key points in the human body region image. The type of each human skeleton key point is then determined based on the confidence level. Finally, each human skeleton key point is connected based on the type to obtain a human skeleton feature vector, and the human posture is determined based on the human skeleton feature vector. It can be seen that, using the method of the present application, it is possible to accurately perform human body detection on the infrared depth image through the human detection model to obtain a human body region image, and accurately identify the human skeleton key points in the human body region image through the human key point detection model, thereby obtaining a human skeleton feature vector and determining the human posture, thereby improving the safety of autonomous driving.

[0117] Another embodiment of the present application further provides a human body posture detection device, such as Figure 5 As shown, specifically including:

[0118] The acquisition unit 501 is configured to acquire an infrared depth image of a target scene.

[0119] The first detection unit 502 is used to input the infrared depth image into a pre-built human body detection model to perform human body detection and obtain a human body area image; wherein the human body detection model is trained based on infrared sample images with human body areas marked.

[0120] The second detection unit 503 is used to input the human body area image into a pre-built human body key point detection model for key point detection to obtain the coordinates and confidence levels of the human skeleton key points in the human body area image; wherein the human body key point detection model is trained based on infrared sample images marked with human skeleton key points.

[0121] The determination unit 504 is configured to determine the type of each human skeleton key point based on the confidence level.

[0122] The connection unit 505 is used to connect the key points of the human skeleton based on the type to obtain the human skeleton feature vector, and determine the human posture based on the human skeleton feature vector.

[0123] In this embodiment, the specific execution process of the acquisition unit 501, the first detection unit 502, the second detection unit 503, the determination unit 504 and the connection unit 505 can be found in the corresponding Figure 1 The content of the method embodiment will not be repeated here.

[0124] In a human posture detection device provided in an embodiment of the present application, first, an acquisition unit 501 acquires an infrared depth image of a target scene. Then, a first detection unit 502 inputs the infrared depth image into a pre-built human detection model to perform human body detection and obtain a human body region image. Then, a second detection unit 503 inputs the human body region image into a pre-built human key point detection model to perform key point detection and obtain the coordinates and confidence levels of the human skeleton key points in the human body region image. The determination unit 504 then determines the type of each human skeleton key point based on the confidence level. Finally, a connection unit 505 connects each human skeleton key point based on the type to obtain a human skeleton feature vector, and determines the human posture based on the human skeleton feature vector. It can be seen that using the method of the present application, it is possible to accurately perform human body detection on the infrared depth image through the human detection model to obtain a human body region image, and accurately identify the human skeleton key points in the human body region image through the human key point detection model, thereby obtaining a human skeleton feature vector and determining the human posture, thereby improving the safety of autonomous driving.

[0125] Optionally, in another embodiment of the present application, an implementation of the first detection unit 502 includes:

[0126] The identification subunit is used to identify the area where the current human body is located through the human body detection frame when a human body is detected in the infrared depth image, and generate a human body area image and a human body area image confidence score.

[0127] The first output unit is configured to directly output the human body region image corresponding to the current human body as a final result if there is only one human body region image corresponding to the current human body.

[0128] The second output unit is configured to output the human body region image with the highest confidence as a final result if there are multiple human body region images corresponding to the current human body.

[0129] In this embodiment, the specific execution process of the identification subunit, the first output unit, and the second output unit can be found in the corresponding method embodiments described above, and will not be repeated here.

[0130] Optionally, in another embodiment of the present application, an implementation of the first detection unit 502 includes:

[0131] The acquisition subunit is used to acquire infrared sample images with human body areas marked.

[0132] The detection subunit is used to input the infrared sample image into the human body detection initial model to perform human body detection and obtain the human body area image corresponding to the current infrared sample image.

[0133] The judging subunit is used to judge whether the human body region image corresponding to the current infrared sample image is consistent with the actually marked human body region.

[0134] The construction subunit is used to complete the construction of the human body detection model if the human body area image corresponding to the current infrared sample image is consistent with the actual marked human body area.

[0135] The parameter adjustment subunit is used to calculate the loss function if the human body area image corresponding to the current infrared sample image is inconsistent with the actual marked human body area, and adjust the parameters of the initial human body detection model based on the loss function until the human body area image corresponding to the current infrared sample image is consistent with the actual marked human body area, thereby completing the construction of the human body detection model.

[0136] In this embodiment, the specific execution process of the acquisition subunit, the detection subunit, the judgment subunit, the construction subunit and the parameter adjustment subunit can be found in the corresponding Figure 2 The content of the method embodiment will not be repeated here.

[0137] Another embodiment of the application further provides an electronic device, such as Figure 6 As shown, specifically including:

[0138] One or more processors 601 .

[0139] The storage device 602 stores one or more programs.

[0140] When one or more programs are executed by one or more processors 601 , the one or more processors 401 implement any one of the methods in the above embodiments.

[0141] Another embodiment of the present application further provides a computer storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, any one of the methods in the above embodiments is implemented.

[0142] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0143] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0144] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for detecting human body posture, characterized in that: include: Acquire infrared depth images of the target scene; Inputting the infrared depth image into a pre-built human body detection model to perform human body detection to obtain a human body region image; wherein the human body detection model is trained based on infrared sample images with human body regions marked; Inputting the human body region image into a pre-built human body key point detection model to perform key point detection, thereby obtaining coordinates and confidence levels of human skeletal key points in the human body region image; wherein the human body key point detection model is trained based on infrared sample images labeled with human skeletal key points; Determining the type of each of the human skeleton key points based on the confidence level; Based on the type, key points of the human skeleton are connected to obtain a human skeleton feature vector, and the human body posture is determined based on the human skeleton feature vector.

2. The method according to claim 1, characterized in that The step of inputting the infrared depth image into a pre-built human body detection model to perform human body detection to obtain a human body region image comprises: When a human body is detected in the infrared depth image, the area where the human body is currently located is marked using a human body detection frame, and a human body area image and a human body area image confidence score are generated; If there is only one human body region image corresponding to the current human body, directly output the human body region image corresponding to the current human body as the final result; If there are multiple human body region images corresponding to the current human body, the human body region image with the highest confidence level is output as the final result.

3. The method according to claim 1, characterized in that The training process of the human body detection model includes: Obtain infrared sample images with human body regions marked; Inputting the infrared sample image into the human body detection initial model to perform human body detection, and obtaining a human body region image corresponding to the current infrared sample image; Determine whether the human body region image corresponding to the current infrared sample image is consistent with the actually marked human body region; If the human body region image corresponding to the current infrared sample image is consistent with the actually marked human body region, the construction of the human body detection model is completed; If the human body area image corresponding to the current infrared sample image is inconsistent with the actual marked human body area, the loss function is calculated and the parameters of the initial human body detection model are adjusted based on the loss function until the human body area image corresponding to the current infrared sample image is consistent with the actual marked human body area, and the construction of the human body detection model is completed.

4. The method according to claim 1, wherein The step of inputting the human body region image into a pre-built human body key point detection model to perform key point detection, and obtaining coordinates and confidence levels of human skeleton key points in the human body region image, comprises: Performing pooling processing on the human body region image to obtain a first vector; Processing the first vector through a fully connected layer to obtain a second vector; Multiplying the second vector by N convolution kernels generated by random initialization to obtain N multiplication results; where N is a preset positive integer; Adding the N multiplication results to obtain convolution kernel parameters, and performing depth-separable convolution processing on the human body region image using the convolution kernel parameters to obtain a processed feature image; The feature image is divided into channels to obtain K sub-feature images and the confidence of each sub-feature image; wherein K is a preset positive integer, and the sub-feature image is used to represent the key points of the human skeleton.

5. The method according to claim 1, wherein Determining the type of each of the human skeleton key points based on the confidence level includes: If the confidence level of the current human skeleton key point is less than a preset first threshold, the type of the current human skeleton key point is the first type; wherein the types include the first type, the second type, and the third type; the first type indicates that the key point does not exist, the second type indicates that the key point exists but is not visible, and the third type indicates that the key point exists and is visible; If the confidence level of the current human skeleton key point is between the first threshold and a preset second threshold, the type of the current human skeleton key point is the second type; If the confidence of the current human skeleton key point is greater than a preset second threshold, the type of the current human skeleton key point is the third type; wherein the second threshold is greater than the first threshold.

6. The method according to claim 4, characterized in that The convolution in the human key point detection model is a lightweight dynamic convolution module; wherein, the lightweight dynamic convolution module includes a pooling layer and a fully connected layer; the lightweight dynamic convolution module adjusts the convolution parameters according to the human body area image input into the human key point detection model.

7. A human body posture detection device, characterized in that: include: An acquisition unit, used for acquiring an infrared depth image of a target scene; A first detection unit is configured to input the infrared depth image into a pre-built human body detection model to perform human body detection and obtain a human body region image; wherein the human body detection model is trained based on infrared sample images with human body regions marked; A second detection unit is configured to input the human body region image into a pre-built human body key point detection model to perform key point detection, thereby obtaining coordinates and confidence levels of human skeletal key points in the human body region image; wherein the human body key point detection model is trained based on infrared sample images labeled with human skeletal key points; A determination unit, configured to determine the type of each of the human skeleton key points based on the confidence level; The connection unit is used to connect the key points of the human skeleton based on the type to obtain a human skeleton feature vector, and determine the human body posture based on the human skeleton feature vector.

8. The device according to claim 7, characterized in that The first detection unit includes: An acquisition subunit, used to acquire an infrared sample image with a human body region marked; A detection subunit, configured to input the infrared sample image into an initial human body detection model to perform human body detection, and obtain a human body region image corresponding to the current infrared sample image; A judging subunit, configured to judge whether the human body region image corresponding to the current infrared sample image is consistent with the actually marked human body region; A construction subunit, configured to complete the construction of the human body detection model if the human body region image corresponding to the current infrared sample image is consistent with the actually marked human body region; The parameter adjustment subunit is used to calculate the loss function if the human body area image corresponding to the current infrared sample image is inconsistent with the actual marked human body area, and adjust the parameters of the initial human body detection model based on the loss function until the human body area image corresponding to the current infrared sample image is consistent with the actual marked human body area, thereby completing the construction of the human body detection model.

9. An electronic device, characterized in that: include: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 6.

10. A computer storage medium, characterized in that A computer program is stored thereon, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Construction site image recognition method and system based on multifunctional safety helmet

    CN121170602A