Skeleton key point detection method, electronic device, and computer-readable storage medium

CN122761418APending Publication Date: 2026-09-15XUANCHENG LUXSHARE PRECISION IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610966605.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-15

AI Technical Summary

Technical Problem

[0004]本发明实施例的主要目的在于提出一种骨骼关键点检测方法、电子设备及计算机可读存储介质,旨在解决遮挡场景下的骨骼关键点检测误差大的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122761418A_ABST
    Figure CN122761418A_ABST
Patent Text Reader

Abstract

The application provides a skeleton key point detection method, an electronic device and a computer readable storage medium. The method comprises the steps of: acquiring an image to be recognized, and performing feature extraction on the image to be recognized to obtain image features; acquiring individual hierarchical vector features, and performing feature extraction on the image features based on the individual hierarchical vector features to obtain skeleton key point features, wherein the individual hierarchical vector features indicate a hierarchical structure of individuals, parts and skeleton key points; and performing skeleton key point prediction based on the skeleton key point features to obtain target skeleton key points corresponding to the image to be recognized. By constructing individual hierarchical vector features based on the hierarchical structure of individuals, parts and skeleton key points, the natural structured association of individual skeletons can be used to construct constraints of skeleton key points, so that the accuracy of skeleton key point detection in a shielding scene can be improved based on the constraint relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision, and more particularly to a method for detecting skeletal key points, an electronic device, and a computer-readable storage medium. Background Technology

[0002] In existing human skeletal key point detection tasks, the mainstream technical solutions are mainly divided into two categories: bottom-up and top-down.

[0003] The bottom-up approach first detects all skeletal keypoints globally, then uses an association algorithm to cluster skeletal keypoints belonging to the same person into a complete pose. Its typical technical architecture includes two core components: a skeletal keypoint detection module and an instance aggregation module. The top-down approach first uses an object detection algorithm to determine the bounding box of each human body in the image, then performs skeletal keypoint detection on each bounding box individually to achieve "one person, one bounding box, one pose". Its typical technical architecture includes a human body detection module and a single instance skeletal keypoint detection module. However, regardless of whether it is bottom-up or top-down, the detection of skeletal keypoints will produce large errors when there is occlusion or noise in the scene. Summary of the Invention

[0004] The main objective of this invention is to propose a skeletal key point detection method, electronic device, and computer-readable storage medium, aiming to solve the problem of large skeletal key point detection errors in occluded scenes.

[0005] This invention provides a method for detecting skeletal key points, the method comprising the following steps: Acquire the image to be identified, and extract features from the image to be identified to obtain image features; Individual hierarchical vector features are obtained, and skeletal keypoint features are extracted from the image features based on the individual hierarchical vector features, wherein the individual hierarchical vector features indicate the hierarchical structure of individual-part-skeletal keypoints; Based on the skeletal key point features, skeletal key point prediction is performed to obtain the target skeletal key points corresponding to the image to be identified.

[0006] Optionally, the step of extracting image features from the image to be identified includes: Obtain the trained recognition model, wherein the recognition model includes an encoder; The image to be identified is input into the trained recognition model so that the encoder can extract the image features.

[0007] Optionally, obtaining individual hierarchical vector features includes: Obtain the trained recognition model, wherein the recognition model includes a self-attention module, and the self-attention module includes a series of individual self-attention layers, part self-attention layers and skeletal key point self-attention layers; Obtain a preset query vector and input the preset query vector into the individual self-attention layer, wherein the preset query vector indicates a preset structure of individual-part-skeletal key points, and the individual self-attention layer is used to extract the association between individuals; The individual extraction vector output by the individual self-attention layer is input into the part self-attention layer, wherein the part self-attention layer is used to extract the association between parts in the same individual; The part extraction vector output by the part self-attention layer is input to the skeletal keypoint self-attention layer, wherein the skeletal keypoint self-attention layer is used to extract the association between skeletal keypoints in the same part. Obtain the individual-level vector features output by the self-attention layer of the skeletal keypoints.

[0008] Optionally, the step of extracting skeletal keypoint features from the image features based on the individual hierarchical vector features includes: Obtain the trained recognition model, wherein the recognition model includes a cross-attention layer; The individual hierarchical vector features and the image features are input into the cross-attention layer; Obtain the skeletal keypoint features output by the cross-attention layer.

[0009] Optionally, obtaining individual-level vector features and extracting skeletal keypoint features from the image features based on the individual-level vector features includes: Obtain the trained recognition model, wherein the recognition model includes multiple cascaded attention modules, each attention module including a self-attention module and a cross-attention layer, and the self-attention module including cascaded individual self-attention layers, location self-attention layers, and skeletal keypoint self-attention layers; the output of the cross-attention layer in the preceding attention module is connected to the input of the individual self-attention layer in the subsequent attention module; wherein: Obtain a preset query vector and input the preset query vector into the individual self-attention layer of the first attention module; The image features are input into the cross-attention layer in each attention module; Obtain the individual-level vector features of the skeletal keypoints in the last attention module output from the attention layer; The individual hierarchical vector features are input into the cross-attention layer in the last attention module; Obtain the skeletal keypoint features output by the cross-attention layer in the last attention module.

[0010] Optionally, the step of predicting skeletal key points based on the skeletal key point features to obtain the target skeletal key points corresponding to the image to be identified includes: Obtain the trained recognition model, wherein the recognition model includes a prediction head, and the prediction head includes an individual prediction head and a skeletal keypoint prediction head; The skeletal keypoint features are respectively input into the individual prediction head and the skeletal keypoint prediction head; Obtain the individual existence probability output by the individual prediction head, and the bone point coordinates and bone point confidence output by the skeletal keypoint prediction head; The existence of an individual is determined based on the probability of its existence, and the coordinates of the existing bone points contained within the existing individual are determined from the bone point coordinates. The target bone key points are determined based on the confidence level of the existing bone point coordinates.

[0011] Optionally, the method further includes: The initial recognition model is trained based on the target loss function to obtain the trained recognition model; wherein: The target loss function is constructed based on individual differences, differences in body parts within individuals, and differences in skeletal key points within body parts.

[0012] This invention also provides a method for detecting skeletal key points, the method comprising: An image to be identified is acquired and input into a trained recognition model, so that the trained recognition model can extract features from the image to be identified to obtain image features, and extract skeletal key point features from the image features based on individual hierarchical vector features, wherein the individual hierarchical vector features indicate the hierarchical structure of individual-part-skeletal key point, and skeletal key point prediction is performed based on the skeletal key point features to obtain the target skeletal key points corresponding to the image to be identified; Obtain the target skeleton key points output by the trained recognition model.

[0013] This invention also provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the skeletal keypoint detection method as described above.

[0014] This invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the skeletal keypoint detection method described above.

[0015] This invention proposes a skeletal keypoint detection method, electronic device, and computer-readable storage medium. The method involves acquiring an image to be identified and extracting image features from the image. Individual-level vector features are then acquired, and skeletal keypoint features are extracted from the image features based on these individual-part-skeletal keypoint features. The individual-level vector features indicate a hierarchical structure of individual-part-skeletal keypoints. Skeletal keypoint prediction is performed based on these features to obtain the target skeletal keypoints corresponding to the image to be identified. By constructing individual-level vector features based on a hierarchical structure of individual-part-skeletal keypoints, the inherent structured relationships of individual skeletons can be used to construct constraints on skeletal keypoints. This constraint relationship improves the accuracy of skeletal keypoint detection in occluded scenarios. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the first embodiment of the skeletal key point detection method of the present invention; Figure 2 This is a schematic diagram of the skeletal key points in the skeletal key point detection method of this invention. Figure 3 This is a schematic diagram of the skeletal key point detection method according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the overall structure of the recognition model in the skeletal key point detection method of this invention. Figure 5 This is a schematic diagram of the overall processing of the recognition model in the skeletal key point detection method of this invention. Figure 6 This is a schematic diagram of the module structure of an electronic device in an embodiment of the present invention. Detailed Implementation

[0019] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0020] This invention provides a method for detecting skeletal key points, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the skeletal key point detection method of the present invention. The method includes the following steps: Step S10: Obtain the image to be identified, and extract features from the image to be identified to obtain image features; The image to be identified is the image for which skeletal keypoint detection is required. This image can be acquired using an image acquisition device set up in a specific scene. The acquisition scene can be set based on actual needs, such as acquiring images of scenes containing individuals with skeletal structures. The type of individual can be set based on actual needs, such as human, animal, or robot; the following explanation uses human as an example. For example, acquiring images of scenes containing human figures such as workshop workers, drivers, or pedestrians means that the image to be identified contains images of complete or partial individuals. If the person is completely within the acquisition range of the image acquisition device and there are no other objects obstructing the person, an image containing a complete individual can be acquired. If part of the person is outside the acquisition range of the image acquisition device, or if other objects obstruct the person, an image containing a partial individual will be acquired. One image to be identified can contain multiple individual images. For example, if the acquisition range of the image acquisition device contains multiple individuals, these multiple individual images can simultaneously contain images of complete individuals and partial individuals.

[0021] The specific method of feature extraction can be set according to actual needs. For example, CNN (Convolutional Neural Network) can be used to extract the corresponding image features from the image to be recognized.

[0022] Step S20: Obtain individual hierarchical vector features, and extract skeletal keypoint features from the image features based on the individual hierarchical vector features, wherein the individual hierarchical vector features indicate the hierarchical structure of individual-part-skeletal keypoints; Individual hierarchical vector features indicate the hierarchical structure of individual-location-skeletal keypoints; where: Individual hierarchy indicates the relationship between different individuals. For example, if the image to be identified contains images corresponding to three human bodies, then the relationship between the three independent human bodies can be indicated by the individual hierarchy.

[0023] Location hierarchy indicates the relationships between different locations within the same individual; skeletal key points within an individual can be classified as belonging to different locations, as shown in [reference needed]. Figure 2 The skeletal key points include 16 points: 0-right ankle, 1-right knee, 2-right hip, 3-left hip, 4-left knee, 5-left ankle, 6-pelvis, 7-chest, 8-upper neck, 9-top of head, 10-right wrist, 11-right elbow, 12-right shoulder, 13-left shoulder, 14-left elbow, and 15-left wrist. In this embodiment, the 16 skeletal key points are divided into 5 parts, namely: The left upper limb includes three key skeletal points: 13-left shoulder, 14-left elbow, and 15-left wrist. The right upper limb includes three key skeletal points: 12-right shoulder, 11-right elbow, and 10-right wrist. The left lower limb includes three key skeletal points: 3-left hip, 4-left knee, and 5-left ankle. The right lower limb includes three key skeletal points: 2-right hip, 1-right knee, and 0-right ankle. The trunk includes four key skeletal points: 9-top of the head, 8-upper neck, 7-chest, and 6-pelvis. The part hierarchy can indicate the relationship between different parts of the same individual.

[0024] Understandably, based on the above division, the key skeletal points corresponding to each part can be used to construct features in sequence, thereby clarifying the location of each key skeletal point, as shown in [reference]. Figure 3 ,: The left upper limb line L1 is formed by connecting the following lines in sequence: 13 - left shoulder, 14 - left elbow, and 15 - left wrist. The right upper limb line L2 is formed by connecting the lines 12-right shoulder, 11-right elbow, and 10-right wrist in sequence; The left lower limb line L3 is formed by connecting the left hip, the left knee, and the left ankle in sequence. The right lower limb line L4 is formed by connecting the right hip, the right knee, and the right ankle in sequence. The midline of the torso, L5, is formed by connecting the following lines in sequence: 9-top of the head, 8-upper neck, 7-chest, and 6-pelvis.

[0025] In practical applications, the location or key skeletal feature lines within a location can be used to construct the location hierarchy. The following explanation will use feature lines as an example.

[0026] Skeletal keypoint hierarchy indicates the relationships between different skeletal keypoints in the same region.

[0027] The above 16 skeletal key points and 5 body parts are examples. In practical applications, the position and number of skeletal key points, as well as the division and number of body parts, can be set according to actual needs. The following explanation will use the above 16 skeletal key points and 5 body parts as examples. Other setting methods can be implemented by analogy and will not be elaborated further.

[0028] In this embodiment, the individual-level vector features are set as a hierarchical structure of individual-part-skeletal keypoints. This allows for feature extraction from the image based on this hierarchical structure. First, different individuals in the image features are identified. Then, for each individual, the parts they contain are identified, and for each part, the skeletal keypoints they contain are identified. Therefore, by using individual-level vector features to extract image features, the feature extraction operation is based on the naturally existing skeletal structure of the individual, ensuring that the obtained skeletal keypoint features conform to the individual's structural characteristics. In this way, even if the individual in the image to be identified is occluded, the skeletal keypoints can still be accurately identified based on the skeletal chain constraints of individual, part, and skeletal keypoints, thereby reducing the false negative rate and ensuring the consistency of the pose determined based on the skeletal keypoints.

[0029] Skeletal keypoint features are features obtained after extracting features from the image; the specific feature extraction method can be set according to actual needs, such as feature extraction through cross attention.

[0030] Step S30: Based on the skeletal key point features, perform skeletal key point prediction to obtain the target skeletal key points corresponding to the image to be identified.

[0031] After obtaining the skeletal keypoint features, the skeletal keypoints can be predicted based on the keypoint features, thereby obtaining the target skeletal keypoints corresponding to the final output image to be identified.

[0032] The target skeletal key points are the skeletal key points obtained by detecting the image to be recognized.

[0033] After obtaining the target skeleton key points, further analysis can be performed based on them. For example, the posture of the person can be determined based on the target skeleton key points, or the joint angles, distances between skeleton key points, and positions of skeleton key points can be calculated based on the target skeleton key points. Then, it can be used to judge whether the person's movements conform to standard movements and whether the person is within a safe area. Specific subsequent applications can be set according to actual needs and are not limited here.

[0034] This embodiment constructs individual-level vector features based on a hierarchical structure of individual-part-skeletal key points, enabling the use of the naturally existing structured associations of individual skeletons to construct constraints on skeletal key points. In occluded scenarios, this constraint relationship is used to improve the accuracy of skeletal key point detection.

[0035] Furthermore, in the second embodiment of the skeletal key point detection method of the present invention based on the first embodiment of the present invention, step S10 includes the following steps: Step S11: Obtain the trained recognition model, wherein the recognition model includes an encoder; Step S12: Input the image to be recognized into the trained recognition model so that the encoder can extract the image features from the image to be recognized.

[0036] The recognition model is used to identify the target skeletal key points in the image to be recognized.

[0037] The encoder is used to extract features from the image to be recognized in order to obtain the corresponding image features.

[0038] The specific structure of the encoder can be set according to actual needs. For example, the encoder includes a backbone network 100 and a feature fusion network 200. The backbone network 100 is used to extract multi-scale image features of the image to be recognized, and the feature fusion network 200 is used to fuse the multi-scale image features to obtain image features.

[0039] The specific type of the backbone network 100 can be set according to actual needs. For example, if the backbone network adopts a convolutional neural network, the specific type of the convolutional neural network can be set according to actual needs, such as ResNet or MobileNet.

[0040] The specific type of the feature fusion network 200 can be set according to actual needs, such as using FPN (Feature Pyramid Network).

[0041] In this embodiment, by setting a recognition model that includes an encoder, it is possible to extract image features.

[0042] Further, see Figure 5 In the third embodiment of the skeletal key point detection method of the present invention based on the first embodiment of the present invention, step S20 includes the following steps: Step S21: Obtain the trained recognition model, wherein the recognition model includes a self-attention module, and the self-attention module includes a series of individual self-attention layers, part self-attention layers and skeletal key point self-attention layers. Step S22: Obtain a preset query vector and input the preset query vector into the individual self-attention layer, wherein the preset query vector indicates a preset structure of individual-part-skeletal key points, and the individual self-attention layer is used to extract the association between individuals; Step S23: Input the individual extraction vector output by the individual self-attention layer into the part self-attention layer, wherein the part self-attention layer is used to extract the association between parts in the same individual; Step S24: Input the part extraction vector output by the part self-attention layer to the skeletal keypoint self-attention layer, wherein the skeletal keypoint self-attention layer is used to extract the association between skeletal keypoints in the same part. Step S25: Obtain the individual-level vector features output by the self-attention layer of the skeletal key points.

[0043] The self-attention module is used to capture the relationships between specific elements within a preset query vector.

[0044] In this embodiment, the self-attention module specifically includes an individual self-attention layer, a part self-attention layer, and a skeletal key point self-attention layer, which respectively capture the correlation between individuals, parts, and skeletal key points.

[0045] The preset query vector is a pre-defined query vector used to represent the basic structure of an individual's skeletal chain; the preset query vector is as follows:

[0046] Where Query is the preset query vector; R indicates the real tensor; D ins The number of individuals; D line The number of parts / feature lines in a single individual; D pt The number of skeletal keypoints in a single region / feature line; D f For feature dimensions.

[0047] The preset query vector indicates that for the recognition model, it supports a maximum of D elements in the input image. ins There are 10 individuals, and each individual has at most D. line There are _ ... pt There are 1 skeletal key points, with a feature dimension of D. fUnderstandably, at this point, the number of individuals, feature lines, and skeletal keypoints that can be supported is at most. When an individual is occluded, the actual part or skeletal keypoints displayed in the image will be even fewer. However, based on the construction of the preset query vector, it is possible to predict the occluded part.

[0048] If the preset query vector is:

[0049] This means that the recognition model can simultaneously detect up to 30 people in an image, and detect 5 feature lines for each person. Each feature line contains 4 skeletal key points, and the feature dimension is 256.

[0050] Individual self-attention layers are used to learn the global associations between different individuals in image features, such as the positional relationship between different individuals and the occlusion priority, so as to distinguish different individuals and avoid confusion in multi-person scenes.

[0051] In the individual self-attention layer, multiple individual queries are set based on the number of individuals in the preset query vector, such as setting N. max =30, meaning a maximum of 30 individual queries are supported; the queried individual instances Qins = {Qins1, Qins2, ..., Qins30}.

[0052] It is understandable that the individual self-attention layer learns the global relationships between individuals. Therefore, in order to focus on individuals, the preset query vector is transformed into a two-dimensional feature with the number of individuals as the main component, such as:

[0053] At this point, the location, skeletal key points, and feature dimensions are merged into a single feature, while the individual remains an independent feature, thereby strengthening the relationship between individuals.

[0054] The individual self-attention layer outputs the processed vector, which is the vector extracted by the individual and sent to the part self-attention layer.

[0055] The part-based attention layer is used to learn the relationships between different parts of the same individual, thereby ensuring the rationality of the given parts in terms of posture and proportion. For example, there is a positional constraint between the left shoulder and chest between the left upper limb line and the driving midline, thus ensuring that the given parts conform to the individual's structural characteristics.

[0056] In the part-based self-attention layer, multiple part queries are set based on the number of parts in the preset query vector, such as setting 5 part queries for each individual; the queried part query Qline={Qline 1 1. Qline 1 2, ..., Qline 305}, where the superscript of Qline indicates the corresponding individual, and the subscript indicates the change in part within the individual. That is, each individual contains 5 parts, and 30 individuals contain a total of 30×5=150 parts.

[0057] Understandably, the part-based self-attention layer learns the relationships between parts within the same individual. Therefore, in order to focus on parts, the preset query vector is transformed into a two-dimensional feature with the number of parts as the main component, such as:

[0058] At this point, the individual, skeletal key points, and feature dimensions are merged into a single feature, while the parts are treated as independent features, thereby strengthening the relationship between the parts.

[0059] The processed vector output by the part self-attention layer is the part extraction vector to the bone keypoint self-attention layer.

[0060] The skeletal keypoint self-attention layer is used to learn the sequential associations between different skeletal keypoints in the same body part. For example, for the left upper limb line, the connection order of the corresponding skeletal keypoints is left shoulder, left elbow, and left wrist. Based on the sequential association, the sequential structure between skeletal keypoints is guaranteed to meet the characteristics of the body part.

[0061] In the skeletal keypoint self-attention layer, multiple skeletal keypoint queries are set based on the number of skeletal keypoints in the preset query vector. For example, four skeletal keypoint queries are set for the midline of the torso, and three skeletal keypoint queries are set for other feature lines. Understandably, the skeletal keypoints learn the relationships between skeletal keypoints within the same body part through the self-attention layer. Therefore, in order to focus on skeletal keypoints, the preset query vector is transformed into a two-dimensional feature with the number of skeletal keypoints as the main component, such as:

[0062] At this point, the individual, part, and feature dimensions are merged into a single feature, while the skeletal keypoints are treated as independent features. This strengthens the relationships between skeletal keypoints, allowing the model to learn the connections between skeletal keypoints along a line.

[0063] The preset query vector learns the relationships between individuals, between parts within the same individual, and between skeletal feature points within the same part by passing through the individual self-attention layer, the part self-attention layer, and the skeletal feature point self-attention layer in sequence. This achieves the learning of hierarchical relationships between individuals, parts, and skeletal feature points, and yields individual hierarchical vector features.

[0064] Furthermore, in the fourth embodiment of the skeletal key point detection method of the present invention based on the first embodiment of the present invention, step S20 includes the following steps: Step S26: Obtain the trained recognition model, wherein the recognition model includes a cross-attention layer; Step S27: Input the individual hierarchical vector features and the image features into the cross-attention layer; Step S28: Obtain the skeletal keypoint features output by the cross-attention layer.

[0065] Cross-attention layers are used to establish a relationship between two input features.

[0066] Individual-level vector features indicate the hierarchical features of individuals, locations, and skeletal key points. By using a cross-attention layer, skeletal key point features are extracted from image features based on these hierarchical features. This allows the extraction of skeletal key point features to be constrained by the hierarchical features of individuals, locations, and skeletal key points, thereby extracting skeletal key point features from image features based on the skeletal chain hierarchy.

[0067] Skeletal keypoint features include skeletal keypoint features, which in turn contain structural features of the skeletal chain. For example, a skeletal keypoint is the left palm skeletal keypoint on the left upper limb line of the first individual in the image.

[0068] Furthermore, in the fifth embodiment of the skeletal key point detection method of the present invention based on the first embodiment of the present invention, step S20 includes the following steps: Step S29: Obtain the trained recognition model, wherein the recognition model includes multiple cascaded attention modules, each attention module including a self-attention module and a cross-attention layer. The self-attention module includes cascaded individual self-attention layers, location self-attention layers, and skeletal keypoint self-attention layers; the output of the cross-attention layer in the preceding attention module is connected to the input of the individual self-attention layer in the subsequent attention module; wherein: Step S210: Obtain a preset query vector and input the preset query vector into the individual self-attention layer of the first attention module; Step S211: Input the image features into the cross-attention layer in each attention module; Step S212: Obtain the individual-level vector features of the skeletal keypoints output from the attention layer in the last attention module; Step S213: Input the individual hierarchical vector features into the cross-attention layer in the last attention module; Step S214: Obtain the skeletal keypoint features output by the cross-attention layer in the last attention module.

[0069] In this embodiment, a multi-layer attention module is set up, which enables multiple rounds of fusion and optimization of query vectors and image features, thereby improving the accuracy of skeletal key point feature extraction. The specific number of attention modules can be set according to actual needs. It is understood that too few attention modules will lead to insufficient feature fusion, while too many attention modules will lead to overfitting. Therefore, the number of attention modules can be determined based on the actual scenario requirements, such as setting the number of attention modules to 6.

[0070] For the first attention module, a preset query vector is input to the individual self-attention layer. The skeletal keypoint self-attention layer outputs the first volume-level vector feature to the cross-attention layer. The cross-attention layer extracts features from the image features based on the first volume-level vector feature to obtain the first skeletal keypoint feature. The first skeletal keypoint feature is then used as the input to the individual self-attention layer of the second attention module. The execution between the individual self-attention layer and the skeletal keypoint self-attention layer is described in the aforementioned embodiment and will not be repeated here.

[0071] The second attention module to the (N-1)th attention module: the individual self-attention layer extracts features based on the input skeletal keypoint features, the skeletal keypoint self-attention layer outputs individual-level vector features to the cross attention layer, the cross attention layer extracts features from the image features based on the individual-level vector features to obtain skeletal keypoint features, and uses the skeletal keypoint features as the input of the individual self-attention layer of the next attention module; The Nth attention module, which is the last attention module, has an individual self-attention layer that extracts features based on the input skeletal keypoint features. The skeletal keypoint self-attention layer outputs individual-level vector features to the cross-attention layer. The cross-attention layer extracts features from the image features based on the individual-level vector features to obtain skeletal keypoint features. At this point, the skeletal keypoint features are the final output for skeletal keypoint prediction.

[0072] Furthermore, in the sixth embodiment of the skeletal key point detection method of the present invention based on the first embodiment of the present invention, step S30 includes the following steps: Step S31: Obtain the trained recognition model, wherein the recognition model includes a prediction head, and the prediction head includes an individual prediction head and a skeletal key point prediction head; Step S32: Input the skeletal keypoint features into the individual prediction head and the skeletal keypoint prediction head respectively; Step S33: Obtain the individual existence probability output by the individual prediction head, and the bone point coordinates and bone point confidence output by the skeletal keypoint prediction head. Step S34: Determine the existence of an individual based on the existence probability of the individual, and determine the coordinates of the existing bone points contained within the existing individual in the bone point coordinates; Step S35: Determine the target bone key points based on the confidence level of the existing bone point coordinates.

[0073] The prediction header is used to output relevant information about the prediction results.

[0074] Individual prediction heads are used to output prediction information related to individuals; the specific structure of individual prediction heads can be set according to actual needs, such as using a two-layer FC (Fully Connected Network).

[0075] The individual prediction head outputs the existence probability of each individual in the skeletal keypoint features; the existence probability of an individual indicates the probability that the individual actually exists; a preset individual probability threshold can be set, and when the existence probability of an individual is greater than or equal to the preset individual probability threshold, the individual is considered to exist.

[0076] The skeleton keypoint prediction head is used to output prediction information related to skeleton keypoints; the specific structure of the skeleton keypoint prediction head can be set according to actual needs, such as using two layers of FC.

[0077] The skeletal keypoint prediction head outputs the corresponding skeletal point confidence score for each skeletal keypoint in the skeletal keypoint features. The skeletal point confidence score indicates the probability that the skeletal keypoint actually exists. A preset skeletal keypoint threshold can be set. When the skeletal point confidence score is greater than or equal to the preset skeletal keypoint threshold, the skeletal keypoint is considered to exist.

[0078] We can first determine the existence of individuals based on their probability of existence, and for individuals that do not exist, we can determine that the skeletal key points they contain do not exist. For existing skeletal keypoints in an individual, the existence of a skeletal keypoint is further determined based on the skeletal point confidence score. If the existence of a skeletal keypoint is determined based on the skeletal point confidence score, it is taken as the target skeletal keypoint, and the coordinates of the existing skeletal point of the target skeletal keypoint are output.

[0079] The existence of skeletal point coordinates indicates the location of the target skeletal key points in the image to be identified.

[0080] Furthermore, in the seventh embodiment of the skeletal key point detection method of the present invention based on the first embodiment, the method further includes: Step S40: Train the initial recognition model based on the target loss function to obtain the trained recognition model; wherein: The target loss function is constructed based on individual differences, differences in body parts within individuals, and differences in skeletal key points within body parts.

[0081] The initial recognition model is the recognition model before training.

[0082] In order to enable the recognition model to further learn the hierarchical structure of individuals, parts, and skeletal key points, this embodiment constructs a loss function based on the differences of individuals, parts, and skeletal key points.

[0083] First, for the training samples, each training sample contains sample images. Each individual is labeled within the sample images, and within each individual, the coordinates of skeletal keypoints are labeled. Feature lines corresponding to different body parts are constructed by connecting these keypoints. For example, the label set corresponding to each individual includes GT_L={GT_L1, GT_L2, GT_L3, GT_L4, GT_L5}, where GT_L represents the individual. Different individuals can be distinguished using subscripts, such as GT_L... 1 For the first individual, GT_L 2 For the second individual; GT_L i The i-th feature line is indicated by GT_L1, which is the left upper limb line, GT_L2, which is the right upper limb line, GT_L3, which is the left lower limb line, GT_L4, which is the right lower limb line, and GT_L5, which is the trunk midline.

[0084] The training samples are input into the initial recognition model, and the initial recognition model outputs the prediction results corresponding to the training samples. The prediction results are matched with the training samples to obtain the differences between the two. In this embodiment, matching is performed for individuals, parts, and skeletal key points respectively.

[0085] The specific method of individual matching can be set according to actual needs. For example, the Hungarian algorithm can be used to calculate the matching loss between the predicted result and the individuals in the training sample. Specifically, it can include confidence loss + loss of all skeletal keypoints within the individual; for example, the confidence loss can be constructed using BCE (binary cross-entropy loss).

[0086] Among them, L ins_conf y represents the confidence loss of an individual; M represents the total number of individuals; m The label for the actual existence of individual m; m This represents the probability of the individual's existence in the training results.

[0087] The confidence loss function is used to optimize the prediction accuracy of whether an individual exists, and to filter out empty instances.

[0088] For individuals that are successfully matched, their constituent body parts are matched sequentially. Body part matching can be based on human anatomical constraints, calculating the distance error between key skeletal points along feature lines within the same individual, such as the physiological distance between the left upper limb line (13 - left shoulder) and the trunk midline (7 - chest).

[0089] Among them, L line_assoc The feature line association loss for the region; P assoc This represents the set of skeletal key point association pairs between feature lines, such as the left upper limb line 13-left shoulder and the trunk midline 7-chest represented as (13, 7); d pred (•) represents the Euclidean distance between point i and point j in the prediction result; d GT (•) represents the Euclidean distance between point i and point j in the training sample.

[0090] For successfully matched parts, their skeletal key points are matched sequentially. In this embodiment, multiple loss mechanisms are used to construct the skeletal key points, including coordinate loss, confidence loss, and feature line angle loss. The coordinate loss can be constructed using L1 loss to optimize the accuracy of the coordinates of successfully matched skeletal keypoints:

[0091] Among them, L p2p For coordinate loss; K is the number of valid skeletal keypoints, i.e., the number after filtering out empty instances and skeletal keypoints with a confidence level less than a threshold; (u k v k ) indicates the coordinates of the k-th skeletal keypoint, the pred subscript indicates the skeletal keypoint in the prediction result, and the GT subscript indicates the skeletal keypoint in the training sample.

[0092] Confidence loss L pt_conf Focal Loss can be used to optimize the prediction accuracy of skeletal keypoint confidence.

[0093] The feature line angle loss is used to supervise the direction and angle of the connection between adjacent skeletal keypoints within the same feature line:

[0094]

[0095]

[0096] Among them, L line_dir For feature line angle loss; Nl N represents the number of feature lines contained in the current individual. v The number of skeletal keypoints in the current feature line; edge pred_i,j edge is the edge vector of the j-th segment of the i-th feature line in the prediction result. GT_i,j Let (j+1) mod N be the edge vector of the j-th feature line in the training samples. v Indicates to repeatedly select points.

[0097] After determining all loss functions, a weighted combination of the loss functions is performed to obtain the final loss function L. total :

[0098] Wherein, λ1~λ5 are the weights corresponding to each loss function, and the specific values ​​can be set according to actual needs.

[0099] This invention provides a method for detecting skeletal key points. In an eighth embodiment of this method, the method includes the following steps: Step S50: Obtain the image to be identified and input the image to be identified into the trained recognition model so that the trained recognition model can extract features from the image to be identified to obtain image features, and extract skeletal key point features from the image features based on individual hierarchical vector features. The individual hierarchical vector features indicate the hierarchical structure of individual-part-skeletal key point. Based on the skeletal key point features, skeletal key point prediction is performed to obtain the target skeletal key points corresponding to the image to be identified. Step S60: Obtain the target skeleton key points output by the trained recognition model.

[0100] In this embodiment, an end-to-end recognition model architecture is constructed. After the image to be recognized is input into the trained recognition model, the recognition model outputs the target skeletal key points corresponding to the image to be recognized. This eliminates the intermediate error propagation in the multi-stage process and makes the training and inference process simpler.

[0101] The recognition model specifically includes an encoder, an attention module, and a prediction head. The specific structure and function of the encoder, attention module, and prediction head are described in the foregoing embodiments and will not be repeated here.

[0102] This embodiment constructs individual-level vector features based on a hierarchical structure of individual-part-skeletal key points, enabling the use of the naturally existing structured associations of individual skeletons to construct constraints on skeletal key points. In occluded scenarios, this constraint relationship is used to improve the accuracy of skeletal key point detection.

[0103] The overall implementation principle of this application is explained below: This application allows for flexible setting of skeletal keypoints based on actual business needs, and does not mandate the use of all 16 skeletal keypoints from the MPII human pose dataset, nor does it require the detection of irrelevant parts. Two examples are provided below: Intelligent cockpit driver status monitoring: The system collects images of the driver's upper body through the built-in camera in the cockpit, selects only the key skeletal points of the upper limbs and torso, and constructs three feature lines: the left upper limb line, the right upper limb line, and the torso. The recognition model outputs the coordinates of the key skeletal points of the driver's upper body end-to-end. The backend calculates and realizes real-time monitoring of hand status, head posture, and dangerous driving behaviors.

[0104] Workstation / Workshop Personnel Safety Monitoring: In industrial workstation and workshop safety monitoring scenarios, images of the operator's upper body are collected by fixed cameras at the workstation. Only the key points of the upper limbs and torso are detected, and three feature lines are constructed: the left upper limb line, the right upper limb line, and the torso. The model outputs the coordinates of the key points of the upper body skeleton, which are used to judge abnormal postures and behaviors such as raising hands beyond the boundary, violating regulations, and leaving the post.

[0105] The 16 standard skeletal keypoints in the MPII human pose dataset are only used to illustrate technical principles, verify model performance, and define human structural standards. In actual product applications, skeletal keypoints can be customized according to the scenario, without being limited to 16, and there is no need to detect irrelevant parts in the scene, such as the lower limbs in the driving and workshop scenarios mentioned above.

[0106] Explanation of Attention Mechanism for Hierarchical Association Learning: The attention module (transformer decoder) takes two inputs: image features extracted by the CNN and a predefined query vector (Query). The three-level self-attention mechanism—instance (i.e., individual), feature lines (i.e., body parts), and skeletal keypoints—only applies to the Query, aiming to teach it the hierarchical relationship between the person, limbs, and skeletal keypoints. Assume the Query vector is:

[0107] Among them, D ins D represents the number of instances; line D represents the number of individual feature lines to be detected; pt D represents the number of skeletal keypoints corresponding to each feature line; f This represents the feature dimension. For example, if D ins =30, D line =5,D pt =4,D f=256. This means that our trained model can simultaneously detect up to 30 people in an image, and detect 5 feature lines for each person. Each feature line contains 4 skeletal keypoints, and the feature dimension is 256.

[0108] Instance-level self-attention: This layer learns "relationships between people," shifting the dimension of the query from D... ins ×D line ×D pt ×D f The four-dimensional features are transformed into D ins ×(D line *D pt *D f The two-dimensional features of a single person are combined into a single overall feature. Using instances as objects, the model learns the positional and distance relationships between people. Linear self-attention: This layer learns the relationships between limbs within the same person, thus shifting the dimension of the query from D... ins ×D line ×D pt ×D f The four-dimensional features are transformed into D line ×(D ins *D pt *D f The two-dimensional features are calculated independently for each individual. Using feature lines as objects, the model learns the relative positional relationships between the left arm, right arm, torso, left leg, and right leg of the same person. Point-level self-attention: This layer learns the relationships between skeletal keypoints along the same line, thus expanding the Query dimension from D... ins ×D line ×D pt ×D f The four-dimensional features are transformed into D pt ×(D ins *D line *D f The two-dimensional features of the model are used to learn the connection relationship between the skeletal key points on a single line. Explanation of model detection output results: The current model only outputs the coordinates of skeletal feature points for all people in the image, without directly outputting pose conclusions. For specific applications, such as pose detection, subsequent calculations of joint angles, distances between skeletal keypoints, and keypoint positions can be used to determine whether the action conforms to standard movements and whether it is within a safe zone.

[0109] Function Description: The MPII dataset contains 16 human skeletal keypoints covering crucial parts of the human body from the head to the limbs, which are of great significance for studying human posture and motion analysis. This invention uses the MPII human skeletal keypoint dataset as an example; the distribution of human skeletal keypoints is as follows: Figure 2 As shown, from the front, it includes: 0-right ankle, 1-right knee, 2-right hip, 3-left hip, 4-left knee, 5-left ankle, 6-pelvis, 7-chest, 8-upper neck, 9-top of head, 10-right wrist, 11-right elbow, 12-right shoulder, 13-left shoulder, 14-left elbow, 15-left wrist.

[0110] The technical solution of this invention revolves around the core of "three-level structure association modeling" and has the following technical definitions: Individual skeletal feature lines are defined as follows: Based on 16 standard human skeletal key points from the MPII dataset, and combined with the skeletal connections and motion topology in human anatomy, they are abstracted into 5 feature lines with fixed associations, denoted as L1~L5, such as... Figure 3 As shown. Each line contains a fixed number of skeletal keypoints, denoted as Pk, where k is the keypoint number: L1 (Left Upper Limb Line): Left shoulder 13 → Left elbow 14 → Left wrist P5, including 3 key skeletal points; L2 (Right Upper Limb Line): Right shoulder 12 → Right elbow 11 → Right wrist 10, including 3 key skeletal points; L3 (Left Lower Limb Line): Left hip 3 → Left knee 4 → Left ankle 5, containing 3 key skeletal points; L4 (Right Lower Limb Line): Right Hip 2 → Right Knee 1 → Right Ankle 0, containing 3 key skeletal points; L5 (Trunk Midline): Top of head 9 → Upper neck 8 → Chest 7 → Pelvis 6, containing 4 key skeletal points; Multi-person scene definition: The input image contains 1 to N human instances (N≤30, covering most real-world scenes). Each instance corresponds to 5 independent feature lines. The network model distinguishes the feature line affiliation of different instances.

[0111] The working principle and method of this embodiment: Training sample construction: Based on the MPII dataset, the 16 skeletal keypoints of each human instance were transformed into "line annotations": For each instance, according to the definitions of L1 to L5, extract the coordinates of the corresponding skeletal keypoints to form a GT set of 5 feature lines, denoted as GT_L={GT_L1, GT_L2, GT_L3, GT_L4, GT_L5}; assign a unique instance ID to each instance, and group the feature lines of different instances by ID (e.g., the GT_L of instance 1 is denoted as GT_L). 1 Example 2 is denoted as GT_L2 ); Data augmentation: Random flipping, scaling, rotation, and brightness adjustment are applied to the training samples to ensure network robustness.

[0112] Network Architecture Design: This embodiment adopts an end-to-end architecture of "convolutional encoder + Transformer decoder". The core innovation lies in the multi-level hierarchical associative modeling component of "instance (i.e., viewpoint) - line (i.e., feature line) - point". This enables structured feature learning from "global human instance" to "local skeletal key points". The overall architecture is as follows: Figure 4 As shown, it includes the following core modules: The image feature encoder includes a backbone network 100 and a feature fusion network 200. Backbone Network 100: Lightweight convolutional backbone networks, such as ResNet and MobileNet series, are used to extract multi-scale image features; Feature Fusion Network 200: Multi-scale image features are fused through Feature Fusion Network (FPN) to obtain a global-local collaborative feature map.

[0113] Multi-level hierarchical association modeling, i.e., attention module 300: Through a 6-layer Transformer decoder, i.e., the attention module, self-attention and cross-attention are performed sequentially at the "instance level → line level → point level" to achieve hierarchical association learning. See Figure 5 Self-Attention: Each decoder layer contains three self-attention layers: First layer: Instance-Level Self-Attention, i.e., individual self-attention layer 311: Set N max =30 instance queries (denoted as Qins={Qins 1 Qins 2 Qins 30 Each instance query corresponds to a human instance. Instance-level self-attention learns the global relationships between different human instances, such as positional relationships and occlusion priority, solving the instance confusion problem in multi-person scenes; The second layer: Line-Level Self-Attention, i.e., the location self-attention layer 312: Each instance query is associated with 5 line queries, denoted as Qline={Qline 1 1. Qline 1 2, ..., Qline 305}, a total of 30 x 5 = 150 line queries. Line-level self-attention learns the relationship between 5 lines within the same human body instance to ensure the rationality of posture proportions. For example, L1 (left upper limb) and L5 (torso midline) capture the positional constraints of "left shoulder and chest" through self-attention. The third layer: Point-Level Self-Attention, i.e., the skeletal keypoint self-attention layer 313: The number of point queries is allocated according to the type of line: L1~L4 correspond to 3 point queries per line, and L5 corresponds to 4 point queries; Point-level self-attention learns the sequential association of skeletal keypoints within a single line, such as the connection logic of "left shoulder → left elbow → left palm", which improves the robustness of point prediction.

[0114] Cross-Attention Layer 321: Multi-level queries interact with image feature maps through cross-attention, extracting skeletal key point features of the corresponding lines of instances from the feature maps. For example, in the L1 query of instance 1, the features of the left upper limb region of instance 1 are focused on the image. The prediction head 400 includes the instance prediction head 410, the skeletal keypoint coordinate prediction head 420, and the confidence prediction head 430.

[0115] Instance prediction head 410 (Instance conf branch): Each instance outputs its confidence score (conf) through a Layer 2 fully connected network (FC). ins (0~1), conf_ins represents the probability of the existence of a human instance; Skeletal keypoint coordinate prediction head 420 (Pts conf branch) and confidence prediction head 430 (Ptsbranch): Each point query outputs the coordinates loc_pts and confidence (conf_pt, 0~1) of the corresponding skeletal keypoint through a 2-layer fully connected network (FC). conf_pts represents the coordinates of the skeletal keypoint and the detection reliability.

[0116] Training optimization section: We designed a training strategy of "hierarchical matching + multiple loss functions" to improve learning efficiency. 3.1 Hierarchical Matching The prediction results need to be paired with the labeled ground truth (GT) of the training samples. This embodiment uses a hierarchical matching method, which is implemented in three steps: "instance matching → line matching → point matching". Instance matching: The Hungarian Algorithm is used to calculate the matching loss between the predicted instance and the GT instance. The loss = confidence loss + loss of all points within the instance. The optimal instance matching relationship is determined, such as predicting instance 1 to match GT instance 2. Line matching: For successfully matched instances, line matching is performed in a fixed order of L1 to L5, such as predicting the L1 of instance 1 to match the L1 of GT instance 2. Point matching: For lines that are successfully matched, point matching is performed according to the fixed sequence number of the skeletal key points, such as matching the predicted point 1 of L1 with point 13 of L1 in GT.

[0117] Multiple loss calculation: Total loss L total It consists of five types of weighted losses, balancing objectives such as instance reliability, line association constraints, point positioning accuracy, and attitude rationality;

[0118] Wherein, λ1~λ5 are the weights corresponding to each loss function, and the specific values ​​can be set according to actual needs.

[0119] The loss terms are defined as follows: Instance confidence loss Lins conf The binary cross-entropy loss (BCE) is used to optimize the prediction accuracy of whether a human instance exists, filtering out empty instances. The formula is as follows:

[0120] Among them, L ins_conf y represents the confidence loss of an individual; M represents the total number of individuals; m The label for the actual existence of individual m; m This represents the probability of the individual's existence in the training results.

[0121] Line correlation loss Lline assoc Based on human anatomical constraints, the distance error between key skeletal points along the same intra-line is calculated, such as the physiological distance between the left shoulder (13) on the left upper limb line and the chest (7) on the trunk midline. The formula is as follows:

[0122] Among them, L line_assoc The feature line association loss for the region; P assoc This represents the set of skeletal key point association pairs between feature lines, such as the left upper limb line 13-left shoulder and the trunk midline 7-chest represented as (13, 7); d pred (•) represents the Euclidean distance between point i and point j in the prediction result; d GT (•) represents the Euclidean distance between point i and point j in the training sample.

[0123] Point-to-point loss L p2p The L1 loss function is used to optimize the accuracy of the coordinates of successfully matched skeletal keypoints. The formula is as follows:

[0124] Among them, L p2p For coordinate loss; K is the number of valid skeletal keypoints, i.e., the number after filtering out empty instances and skeletal keypoints with a confidence level less than a threshold; (u k v k ) indicates the coordinates of the k-th skeletal keypoint, the pred subscript indicates the skeletal keypoint in the prediction result, and the GT subscript indicates the skeletal keypoint in the training sample.

[0125] Point confidence loss L pt_conf FocalLoss is used to optimize the prediction accuracy of skeletal keypoint confidence. Linear angle loss L line_dir The direction and angle of the line connecting adjacent skeletal keypoints within the supervision line are given by the following formula:

[0126]

[0127]

[0128] Among them, L line_dir For feature line angle loss; N l N represents the number of feature lines contained in the current individual. v The number of skeletal keypoints in the current feature line; edge pred_i,j edge is the edge vector of the j-th segment of the i-th feature line in the prediction result. GT_i,j Let (j+1) mod N be the edge vector of the j-th feature line in the training samples. v Indicates to repeatedly select points.

[0129] Deploy a skeletal keypoint detection model: To deploy the model in a real-world product for real-time individual skeletal keypoint detection, the model needs to be converted to a format supported by the target platform. The following example demonstrates the model deployment steps using TDA4VL.

[0130] Model Training: Train the skeletal keypoint detection model using an open-source neural network training framework. Open-source frameworks include, but are not limited to, Caffe, PyTorch, Tensorflow, and MXNet. Model conversion: Convert the recognition model trained by the open-source framework into ONNX format; Model optimization and quantization: The model was optimized and quantized using Texas Instruments (TI) model conversion tools. Model Deployment: The final result is a model file that can run on TDA4VL. By deploying this model file to the target platform, real-time skeletal keypoint detection is achieved.

[0131] The improvements of this embodiment compared to existing technologies are as follows: This embodiment addresses the shortcomings of existing individual skeletal keypoint detection technologies, such as complex multi-stage processes, lack of structured association, insufficient robustness to occlusion, and poor hardware adaptability. It achieves breakthroughs through four major innovations: end-to-end architecture design, hierarchical association modeling, anatomical loss enhancement, and lightweight deployment optimization. The details are as follows: End-to-end unified optimization: Simplify processes and improve overall collaboration; This embodiment proposes an end-to-end convolutional encoder + Transformer decoder architecture. Taking the image to be recognized as input, it outputs the skeletal key points of multiple human instances in one step, achieving end-to-end joint optimization from image to skeletal key points. This design eliminates intermediate error propagation in the multi-stage process, making the training and inference process simpler.

[0132] Instance-line-point hierarchical association modeling: enhancing occlusion robustness and pose consistency; Existing technologies treat skeletal keypoints as isolated individual targets, failing to utilize the naturally occurring structured associations of the human skeleton, such as the "shoulder-elbow-wrist" upper limb skeletal chain constraints. This results in high false negative rates and poor pose consistency in occluded scenarios. This embodiment innovatively designs an instance-level self-attention → line-level self-attention → point-level self-attention mechanism, systematically solving the above problems through progressive association learning from global to local.

[0133] Anatomical constraint loss: Utilizing prior knowledge of human anatomy to improve postural rationality; This embodiment specifically designs two types of anatomical constraint losses: line correlation loss L. line_assoc Linear angle loss L line_dir The model constrains the physiological distance range of key skeletal points between lines within the instance, using anatomical priors to improve the rationality of localization; at the same time, it supervises the direction of the connection between adjacent key skeletal points within the line to ensure that it conforms to the human movement mechanism; the introduction of anatomical loss enables the model to actively fit the human structural rules, improving the detection accuracy under complex postures.

[0134] Lightweight and universal optimization: Supports efficient deployment across multiple hardware devices; Existing high-precision skeletal keypoint detection models have high computational complexity, making them difficult to run in real time on edge devices and limiting the application of the technology. This embodiment adopts a lightweight convolutional architecture and a Transformer decoder, resulting in low computational cost. Furthermore, the trained model can be converted to ONNX format and further optimized and quantized using the Texas Instruments (TI) toolchain, adapting it to edge chips such as the TDA4L and expanding the practical application scenarios of the technology.

[0135] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0136] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0137] Reference Figure 6 In terms of hardware structure, the electronic device may include components such as a communication module 10, a memory 20, and a processor 30. In the electronic device, the processor 30 is connected to both the memory 20 and the communication module 10. The memory 20 stores a computer program, which is executed by the processor 30. When the computer program is executed, it implements the steps of the above-described method embodiments.

[0138] The communication module 10 can connect to external communication devices via a network. The communication module 10 can receive requests from the external communication devices and can also send requests, instructions, and information to the external communication devices. The external communication devices can be other electronic devices, servers, or IoT devices, such as televisions, etc.

[0139] The memory 20 can be used to store software programs and various data. The memory 20 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as acquiring an image to be recognized and extracting image features from the image to be recognized), etc.; the data storage area may include a database, and may store data or information created based on system usage. Furthermore, the memory 20 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0140] The processor 30 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 20, and by calling data stored in the memory 20, it performs various functions and processes data, thereby providing overall monitoring of the electronic device. The processor 30 may include one or more processing units; optionally, the processor 30 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 30.

[0141] although Figure 6 Not shown, but the above electronic device may also include a circuit control module for connecting to a power supply to ensure the normal operation of other components. Those skilled in the art will understand that... Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0142] This invention also proposes a computer-readable storage medium on which a computer program is stored. The computer-readable storage medium may be... Figure 6 The memory 20 in the electronic device may be at least one of ROM (Read-Only Memory) / RAM (Random Access Memory), magnetic disk, optical disk, etc. The computer-readable storage medium includes a number of instructions to cause a terminal device with a processor (which may be a television, automobile, mobile phone, computer, server, terminal, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0143] In this invention, the terms "first," "second," "third," "fourth," and "fifth" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0144] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0145] Although embodiments of the present invention have been shown and described above, the scope of protection of the present invention is not limited thereto. It is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, and substitutions to the above embodiments within the scope of the present invention, and such changes, modifications, and substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for detecting skeletal key points, characterized in that, The skeletal key point detection method includes: Acquire the image to be identified, and extract features from the image to be identified to obtain image features; Individual hierarchical vector features are obtained, and skeletal keypoint features are extracted from the image features based on the individual hierarchical vector features, wherein the individual hierarchical vector features indicate the hierarchical structure of individual-part-skeletal keypoints; Based on the skeletal key point features, skeletal key point prediction is performed to obtain the target skeletal key points corresponding to the image to be identified.

2. The skeletal key point detection method as described in claim 1, characterized in that, The image features obtained by extracting features from the image to be identified include: Obtain the trained recognition model, wherein the recognition model includes an encoder; The image to be identified is input into the trained recognition model so that the encoder can extract the image features.

3. The skeletal key point detection method as described in claim 1, characterized in that, The acquisition of individual hierarchical vector features includes: Obtain the trained recognition model, wherein the recognition model includes a self-attention module, and the self-attention module includes a series of individual self-attention layers, part self-attention layers and skeletal key point self-attention layers; Obtain a preset query vector and input the preset query vector into the individual self-attention layer, wherein the preset query vector indicates a preset structure of individual-part-skeletal key points, and the individual self-attention layer is used to extract the association between individuals; The individual extraction vector output by the individual self-attention layer is input into the part self-attention layer, wherein the part self-attention layer is used to extract the association between parts in the same individual; The part extraction vector output by the part self-attention layer is input to the skeletal keypoint self-attention layer, wherein the skeletal keypoint self-attention layer is used to extract the association between skeletal keypoints in the same part. Obtain the individual-level vector features output by the self-attention layer of the skeletal keypoints.

4. The skeletal key point detection method as described in claim 1, characterized in that, The process of extracting skeletal keypoint features from the image features based on the individual hierarchical vector features includes: Obtain the trained recognition model, wherein the recognition model includes a cross-attention layer; The individual hierarchical vector features and the image features are input into the cross-attention layer; Obtain the skeletal keypoint features output by the cross-attention layer.

5. The skeletal key point detection method as described in claim 1, characterized in that, The step of obtaining individual hierarchical vector features and extracting skeletal keypoint features from the image features based on the individual hierarchical vector features includes: Obtain the trained recognition model, wherein the recognition model includes multiple cascaded attention modules, each attention module including a self-attention module and a cross-attention layer, and the self-attention module including cascaded individual self-attention layers, location self-attention layers, and skeletal keypoint self-attention layers; the output of the cross-attention layer in the preceding attention module is connected to the input of the individual self-attention layer in the subsequent attention module; wherein: Obtain a preset query vector and input the preset query vector into the individual self-attention layer of the first attention module; The image features are input into the cross-attention layer in each attention module; Obtain the individual-level vector features of the skeletal keypoints in the last attention module output from the attention layer; The individual hierarchical vector features are input into the cross-attention layer in the last attention module; Obtain the skeletal keypoint features output by the cross-attention layer in the last attention module.

6. The skeletal key point detection method as described in claim 1, characterized in that, The step of predicting skeletal key points based on the skeletal key point features to obtain the target skeletal key points corresponding to the image to be identified includes: Obtain the trained recognition model, wherein the recognition model includes a prediction head, and the prediction head includes an individual prediction head and a skeletal keypoint prediction head; The skeletal keypoint features are respectively input into the individual prediction head and the skeletal keypoint prediction head; Obtain the individual existence probability output by the individual prediction head, and the bone point coordinates and bone point confidence output by the skeletal keypoint prediction head; The existence of an individual is determined based on the probability of its existence, and the coordinates of the existing bone points contained within the existing individual are determined from the bone point coordinates. The target bone key points are determined based on the confidence level of the existing bone point coordinates.

7. The skeletal key point detection method according to any one of claims 2 to 6, characterized in that, The method further includes: The initial recognition model is trained based on the target loss function to obtain the trained recognition model; wherein: The target loss function is constructed based on individual differences, differences in body parts within individuals, and differences in skeletal key points within body parts.

8. A method for detecting skeletal key points, characterized in that, The skeletal key point detection method includes: An image to be identified is acquired and input into a trained recognition model. The trained recognition model extracts features from the image to obtain image features and extracts skeletal keypoint features from the image features based on individual hierarchical vector features. The individual hierarchical vector features indicate the hierarchical structure of individual-part-skeletal keypoints. Based on the skeletal keypoint features, skeletal keypoint prediction is performed to obtain the target skeletal keypoints corresponding to the image to be identified. Obtain the target skeleton key points output by the trained recognition model.

9. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the skeletal keypoint detection method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the skeletal keypoint detection method as described in any one of claims 1 to 8.