Pose Recognition Method, Device, Storage Medium, and Electronic Device

The pose recognition model extracts image feature data, predicts the locations of the center point and the adaptive point, and uses the adaptive point to predict the key point set, solving the problem of low pose recognition accuracy in the single-stage pose regression method, achieving efficient pose recognition effect.

CN114139630BActive Publication Date: 2025-07-18BEIJING HORIZON INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111463757.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-02
Publication Date
2025-07-18
Estimated Expiration
2041-12-02

AI Technical Summary

Technical Problem

The single-stage pose regression method cannot fully encode pose information of different scales and deformations because it only uses the characteristics of the center point, resulting in low pose recognition accuracy.

Method used

Image feature data is extracted through the pose recognition model, the positions of the center point and the adaptive point are predicted, the key point set is used to predict, the target pose is determined by combining the center point and the key point set, and the parallel processing of multiple network branches is used to improve the accuracy of pose recognition.

Benefits of technology

It realizes fine-grained representation of poses of different scales and deformations, improves the accuracy of pose recognition, and avoids the burden of computing and storage in the post-processing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114139630B_ABST
    Figure CN114139630B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a posture recognition method, apparatus, storage medium, and electronic device. The method includes: extracting first feature data of an image including an object to be recognized by using a posture recognition model; predicting the position of the center point of the object to be recognized and the positions of adaptive points respectively corresponding to each part based on the first feature data, where the center point represents the imaging point of the center point part of the object to be recognized; predicting a key point set respectively corresponding to each part based on the first feature data and the positions of the adaptive points respectively corresponding to each part; determining the target posture of the object to be recognized based on the position of the center point and the key point sets respectively corresponding to each part. Through each adaptive point, postures of different scales and deformations can be characterized in a fine-grained manner, and thus the association between the key points and the object to be recognized can be clarified, and the accuracy of posture recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to computer vision technology, and in particular to a method, device, storage medium, and electronic device for pose recognition. Background Art

[0002] In the field of computer vision, pose recognition is used to locate the key point positions of an object to be recognized in an image, and to represent the pose of the object to be recognized based on the key point positions, such as human pose recognition. With the application of deep learning technology, great progress has been made in this field and it has promoted the development of fields such as human-computer interaction and behavior recognition.

[0003] In related technologies, a single-stage pose regression method usually first predicts the center point of the human body, and then predicts multiple key points based on the center point to obtain the pose of the human body. Summary of the Invention

[0004] To solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a method, device, storage medium, and electronic device for human pose recognition.

[0005] According to one aspect of the embodiments of the present disclosure, a method for human pose recognition is provided. The method includes: extracting first feature data of an image containing an object to be recognized by using a pose recognition model; based on the first feature data, predicting the position of the center point of the object to be recognized and the positions of adaptive points corresponding to each part, where the center point represents the imaging point of the center point part of the object to be recognized; based on the first feature data and the positions of the adaptive points corresponding to each part, predicting key point sets corresponding to each part; and based on the position of the center point and the key point sets corresponding to each part, determining the target pose of the object to be recognized.

[0006] According to another aspect of the embodiments of the present disclosure, a method for training a pose recognition model is provided. The method includes: obtaining a training set, where the training set includes sample images with labeled sample labels, and the sample labels include the position of the sample center point of the object to be recognized, the positions of sample key points, and the sample center point heat map and sample key point heat maps corresponding to the sample images; processing the sample feature data based on the initial pose regression sub-network of the initial pose recognition model to obtain the predicted center point confidence of each pixel point and the positions of the corresponding predicted key points; processing the sample feature data based on the key point heat map network of the initial pose recognition model to generate the predicted key point heat map of the sample image; determining a first loss function based on the predicted center point confidence of each pixel point and the sample center point heat map; determining a second loss function based on the positions of the sample key points and the positions of the predicted key points corresponding to the reference pixel points, where the position of the reference pixel point is the same as the position of the sample center point; determining the value of a third loss function based on the predicted key point heat map and the sample key point heat map; adjusting the parameters of the initial pose recognition model based on the first loss function, the second loss function, and the third loss function, and when the termination condition is satisfied, deleting the key point heat map network to obtain the pose recognition model.

[0007] According to another aspect of the embodiments of the present disclosure, a human pose determination device is provided, including: a feature extraction unit configured to extract first feature data of an image including an object to be recognized by using a pose recognition model; a first prediction unit configured to predict the position of the center point of the object to be recognized and the positions of adaptive points corresponding to each part based on the first feature data, where the center point represents the imaging point of the center point part of the object to be recognized; a second prediction unit configured to predict a set of key points corresponding to each part based on the first feature data and the positions of the adaptive points corresponding to each part; a pose determination unit configured to determine the target pose of the object to be recognized based on the position of the center point and the set of key points corresponding to each part.

[0008] According to another aspect of the embodiments of the present disclosure, there is provided an apparatus for training a pose recognition model, including: a sample acquisition unit configured to acquire a training set, the training set including sample images with labeled sample labels, and the sample labels including the position of the sample center point of the object to be recognized, the positions of the sample key points, and the sample center point heat map and the sample key point heat map corresponding to the sample images; a feature extraction unit configured to process the sample images in the training set based on the initial backbone network of the pre-constructed initial pose recognition model to obtain sample feature data; a pose prediction unit configured to process the sample feature data based on the initial pose regression sub-network of the initial pose recognition model to obtain the predicted center point confidence of each pixel point and the positions of the corresponding predicted key points; a heat map prediction unit configured to process the sample feature data based on the key point heat map network of the initial pose recognition model to generate the predicted key point heat map of the sample image; a first loss unit configured to determine a first loss function based on the predicted center point confidence of each pixel point and the sample center point heat map; a second loss unit configured to determine a second loss function based on the positions of the sample key points and the positions of the predicted key points corresponding to the reference pixel points, and the position of the reference pixel point is the same as the position of the sample center point; a third loss unit configured to determine a third loss function based on the predicted key point heat map and the sample key point heat map; and a model training unit configured to adjust the parameters of the initial pose recognition model based on the first loss function, the second loss function, and the third loss function until the termination condition is satisfied, and delete the key point heat map network to obtain the pose recognition model.

[0009] Based on the human pose determination method provided in the above embodiments of the present disclosure, the position of the center point of the object to be recognized and the positions of the adaptive points corresponding to each part can be predicted using the first feature data. Then, based on the first feature data and the positions of the adaptive points corresponding to each part, the key point set corresponding to each part can be predicted, and the target pose of the object to be recognized can be determined based on the position of the center point and the key point sets corresponding to each part. Through each adaptive point, the pose with different scales and deformations can be characterized in a fine-grained manner, thereby clarifying the association between the key points and the object to be recognized, and improving the accuracy of pose recognition.

[0010] The technical solutions of the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0011] By describing the embodiments of the present disclosure in more detail with reference to the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation to the present disclosure. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0012] Figure 1(a) is a schematic diagram of a system architecture of the posture recognition method of the present disclosure;

[0013] Figure 1(b) is a schematic diagram of a target posture in an embodiment of the posture recognition method of the present disclosure;

[0014] Figure 2 It is a flowchart of an embodiment of the posture recognition method of the present disclosure.

[0015] Figure 3 It is a flowchart of predicting the positions of the prediction center point and the adaptive point in an embodiment of the posture recognition method of the present disclosure;

[0016] Figure 4 It is a flowchart of predicting a key point set in an embodiment of the posture recognition method of the present disclosure;

[0017] Figure 5 It is a flowchart of predicting the position of a candidate adaptive point in an embodiment of the posture recognition method of the present disclosure;

[0018] Figure 6 It is a flowchart of predicting a candidate key point set in an embodiment of the posture recognition method of the present disclosure;

[0019] Figure 7 It is a flowchart of an embodiment of the method for training a posture recognition model of the present disclosure;

[0020] Figure 8 It is a schematic structural diagram of an embodiment of the posture recognition device of the present disclosure;

[0021] Figure 9 It is a schematic structural diagram of an embodiment of the device for training a posture recognition model of the present disclosure;

[0022] Figure 10 It is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed implementation manners

[0023] Next, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.

[0024] It should be noted that: Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0025] Those skilled in the art can understand that the terms "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.

[0026] It should also be understood that in the embodiments of the present disclosure, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.

[0027] It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, without clear limitation or contrary indication in the context, it can generally be understood as one or more.

[0028] In addition, the term "and / or" in the present disclosure is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.

[0029] It should also be understood that the description of each embodiment in the present disclosure emphasizes the differences between the embodiments, and their similarities can be referred to each other. For the sake of brevity, they will not be elaborated one by one.

[0030] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0031] The following description of at least one exemplary embodiment is actually only illustrative and in no way a limitation on the present disclosure and its application or use.

[0032] Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods and devices should be regarded as part of the specification.

[0033] It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0034] Embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate together with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.

[0035] Terminal devices, computer systems, servers and other electronic devices can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules may include routines, programs, target programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.

[0036] Overview of the Application

[0037] In the process of implementing the present disclosure, the inventors found that in the process of predicting key points based on the center point by a single-stage pose regression method, since only the features of the center point are utilized, it is impossible to fully encode the pose information of different scales and deformations, resulting in the inability to finely characterize the poses of different scales and deformations, and the accuracy of pose recognition is relatively low.

[0038] Exemplary System

[0039] The following provides an exemplary description of the posture recognition method of the present disclosure with reference to Fig. 1(a). Fig. 1(a) is a schematic diagram of a system architecture of the posture recognition method of the present disclosure. As shown in Fig. 1, the system may include a posture recognition model and a max pooling kernel 170. Among them, the posture recognition model may include a backbone network 110, a key point regression network branch 120, a region perception network branch 130, and a center point perception network branch 140. The execution entity may be a terminal device or a server loaded with computer instructions of the posture recognition model. When the execution entity obtains an image containing an object to be recognized, it may extract first feature data from the image through the backbone network 110 of the posture recognition model (for example, a convolutional neural network such as ResNet or HRNet). Then, the center point perception network branch 140 predicts the center point confidence of each pixel point based on the first feature data, and the max pooling kernel 170 screens out the center point of the object to be recognized according to the center point confidence. The key point regression network branch 120 extracts key point regression features from the first feature data, and the region perception network branch 130 predicts the positions 150 of the adaptive points corresponding to each part of the object to be recognized based on the key point regression features. Then, the key point regression network branch 120 predicts the key point sets 160 corresponding to each part based on the positions of the respective adaptive points and the key point regression features. Finally, the position of the object to be recognized is determined according to the position of the center point, and the key point sets belonging to the object to be recognized are determined, so as to determine the target posture of the object to be recognized.

[0040] As shown in Fig. 1(b), when the object to be recognized is a human body, the target posture may include key point sets corresponding to 7 parts respectively. The local posture of each part and the relative position of the part in the human body instance can be characterized by the key point sets.

[0041] Exemplary Method

[0042] In the embodiments of the present disclosure, "candidate" means to be determined. For example, a candidate adaptive point means a point that is to be determined and has a certain probability of becoming an adaptive point. When it is determined through judgment that the candidate adaptive point meets the preset conditions (for example, the pixel point corresponding to the candidate adaptive point is determined as the center point), the candidate adaptive point is correspondingly determined as an adaptive point. During this process, the attributes of the point (such as position, semantic information, etc.) will not change.

[0043] Figure 2 It is a flowchart of the posture recognition method provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device, such as Figure 2 As shown, the following steps are included:

[0044] Step 210, using the posture recognition model to extract first feature data of an image containing an object to be recognized.

[0045] In this embodiment, the object to be recognized may be, for example, a human body, an animal, or other objects that can recognize postures. The first feature data may include, but is not limited to, the texture features of the image, as well as the semantic information, boundary information, position information, etc. of the pixel points. The first feature data may be, for example, a multi-dimensional matrix.

[0046] With reference to FIG. 1(a) for exemplary illustration, the backbone network in the pose recognition model can be used to extract features from the image to obtain the first feature data.

[0047] Step 220: Based on the first feature data, predict the position of the center point of the object to be recognized and the positions of the adaptive points corresponding to each part.

[0048] Among them, the center point represents the imaging point of the center point part of the object to be recognized.

[0049] In this embodiment, the adaptive points correspond to each part of the object to be recognized, and the relative position between the part and the center of the object to be recognized can be represented by the relative position between the adaptive point and the center point.

[0050] As an example, the convolutional layer or fully connected layer in the pose recognition model can be used to process the first feature data to predict the position of the center point of the object to be recognized and the positions of the adaptive points corresponding to each part.

[0051] In a specific example, the execution entity can adopt a per-pixel processing method. First, each pixel point is assumed to be the center point, and then the pose recognition model is used to preset the center point confidence of each pixel point, and, predict the candidate adaptive points corresponding to each part when the pixel point is the center point; then, all pixel points can be screened by the center point confidence. When the confidence of the pixel point meets the preset conditions (for example, the confidence is greater than a predetermined value or the confidence is a local maximum), the pixel point is determined to be the center point, and the candidate adaptive points corresponding to the pixel point are the adaptive points. Thus, the center point of the object to be processed and the adaptive points corresponding to each part can be obtained.

[0052] Step 230: Based on the first feature data and the positions of the adaptive points corresponding to each part, predict the key point sets corresponding to each part.

[0053] In this embodiment, each part corresponds to a key point set, and the key point set may include one or more key points.

[0054] As shown in FIG. 1(b), in the example of human pose recognition, each part can be determined according to the joint structure of the human body, and the key points can represent the imaging points of the joints of the human body. For example, the head region may include 5 key points, and other parts may include two key points.

[0055] In a specific example, the execution entity can extract the feature data corresponding to the positions of each adaptive point from the first feature data by means of bilinear interpolation, and then use a convolutional layer or a fully connected layer to process the extracted feature data to predict one or more key points corresponding to each adaptive point, so as to obtain the key point set corresponding to each part.

[0056] Step 240: Determine the target pose of the object to be recognized based on the position of the center point and the key point sets respectively corresponding to each part.

[0057] In an example of multi-person pose recognition, the execution entity can determine the positions of each human body instance in the image according to the position of the center point, and then determine the key point sets respectively corresponding to each part of each human body instance according to the key point sets corresponding to each center point, and determine the target poses of each human body in the image according to the multiple key point sets.

[0058] The pose recognition method provided in this embodiment can use the first feature data to predict the position of the center point of the object to be recognized and the positions of the adaptive points respectively corresponding to each part, and then predict the key point set corresponding to each part according to the first feature data and the positions of the adaptive points respectively corresponding to each part, and determine the target pose of the object to be recognized according to the position of the center point and the key point sets respectively corresponding to each part. Through each adaptive point, the pose of different scales and deformations can be characterized in a fine-grained manner, thereby clarifying the association between the key points and the object to be recognized, and thus improving the accuracy of pose recognition.

[0059] Any of the pose recognition methods provided in the embodiments of the present disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices, servers, etc. Alternatively, any of the pose recognition methods provided in the embodiments of the present disclosure can be executed by a processor. For example, the processor executes any of the pose recognition methods mentioned in the embodiments of the present disclosure by calling the corresponding instructions stored in the memory. This will not be elaborated below.

[0060] Next, refer to Figure 3 , Figure 3 shows a flowchart for predicting the positions of the center point and the adaptive points in an embodiment of the pose recognition method of the present disclosure. As Figure 3 shown, the above step 220 may further include the following steps:

[0061] Step 310: Extract features from the first feature data based on the key point regression network branch of the pose recognition model to obtain second feature data.

[0062] As an example, the key point regression network branch can extract features from the first feature data through a convolutional layer to obtain second feature data.

[0063] Step 320: Process the second feature data based on the region perception network branch of the pose recognition model, and predict the positions of the candidate adaptive points corresponding to each part for each pixel point in the image.

[0064] In this embodiment, before determining the center point, each pixel point in the image has a certain probability of becoming the center point, and this probability is the confidence of the center point. Based on this, the execution entity can first assume each pixel point as the center point, and thereby predict the positions of the candidate adaptive points corresponding to each part for each pixel point. The candidate adaptive points represent the to-be-determined adaptive points, and when the pixel point is determined as the center point, the candidate adaptive points are the adaptive points.

[0065] As an example, the region perception network branch can use a convolutional layer or a fully connected layer to process the second feature data, and predict the positions of the candidate adaptive points corresponding to each part for each pixel point.

[0066] Step 330: Extract features from the first feature data based on the center point perception network branch of the pose recognition model to obtain the third feature data.

[0067] Step 340: Extract the center regression features corresponding to each pixel point from the third image feature based on the positions of the candidate adaptive points corresponding to each part for each pixel point.

[0068] The feature data in this disclosure (such as including the first feature data, the second feature data, the third feature data, the center regression features, the key point regression features, and other feature data) can be in the form of a multi-dimensional matrix, and the corresponding part of the feature data can be extracted from the feature map based on the position of the pixel point in the image (usually referring to the pixel coordinates).

[0069] Step 350: Predict the center point confidence of each pixel point based on the center regression features corresponding to each pixel point.

[0070] In a specific example, the execution entity can use the first convolutional layer in the center point perception network branch to extract the third feature data from the first feature data, and then extract the feature data corresponding to the positions of the respective candidate adaptive points from the third feature data; then splice the extracted feature data to obtain the center regression features corresponding to each pixel point; and then use the second convolutional layer or the fully connected layer to process the center regression features to predict the center point confidence of each pixel point.

[0071] It should be noted that the center point perception network branch and the key point regression network in this embodiment can be processed in parallel, and the present disclosure does not limit their order.

[0072] Step 360: Use a maximum pooling kernel to determine the position of the central point of the object to be recognized as the position of the central point of the pixel points whose central point confidence is greater than the preset threshold, and determine the positions of the candidate adaptive points corresponding to each part of the pixel point as the positions of the adaptive points corresponding to each part of the object to be recognized.

[0073] In this embodiment, the central point confidence represents the probability that a pixel point is the central point. The higher the central confidence, the higher the matching degree between the pixel point and the central point. Correspondingly, the accuracy of representing the central point of the object to be recognized by this pixel point is higher.

[0074] In Figure 3 In the illustrated embodiment, a per-pixel processing method is adopted. The pose recognition model is used to predict the central point confidence of each pixel point and its corresponding candidate adaptive points. Then, the matching degree between the pixel point and the central point is evaluated through the central point confidence. The pixel points with a central point confidence greater than the preset threshold are determined as the central points, and the candidate adaptive points corresponding to these pixel points are the adaptive points corresponding to each part. On the one hand, it can improve the accuracy of predicting the positions of the central point and the adaptive points. On the other hand, through the parallel processing method of each network branch in the pose recognition model, the operation efficiency can be improved.

[0075] Further referring to Figure 4 , Figure 4 shows a flowchart for predicting a key point set in an embodiment of the pose recognition method of the present disclosure. Based on the embodiments shown in Figure 3 and Figure 2 , the above step 230 may further include:

[0076] Step 410: Use the key point regression network branch to extract the key point regression features corresponding to each part of each pixel point from the second feature data based on the positions of the candidate adaptive points corresponding to each part of each pixel point.

[0077] As an example, the key point regression network branch can be used to extract the feature data corresponding to the positions of each candidate adaptive point from the second feature data by means of bilinear interpolation, and then splice the extracted feature data, and use the spliced feature data as the key point regression features corresponding to each part of the pixel point.

[0078] Step 420: Based on the key point regression features corresponding to each part of each pixel point and the positions of its corresponding candidate adaptive points, predict the candidate key point sets corresponding to each part of each pixel point.

[0079] In this embodiment, the execution entity may use the key point regression network branch to predict one or more candidate key points corresponding to the positions of each candidate adaptive point, so as to obtain a set of candidate key points for each part corresponding to each pixel point.

[0080] Step 430: After determining the position of the pixel point with the center point confidence greater than the preset threshold as the position of the center point of the object to be recognized by using the maximum pooling kernel, determine the set of candidate key points for each part corresponding to the pixel point as the set of key points corresponding to each of the parts.

[0081] In this embodiment, the candidate adaptive point represents an undetermined adaptive point, and the set of candidate key points represents an undetermined set of key points. By evaluating the pixel points through the maximum pooling kernel, when a pixel point is determined as the center point, the corresponding candidate adaptive point and the set of candidate key points are correspondingly referred to as the adaptive point and the set of key points.

[0082] From Figure 4 It can be seen that Figure 4 the process reflects that the set of candidate key points corresponding to each candidate adaptive point is predicted based on the pose recognition model. When a pixel point is determined as the center point, the corresponding candidate adaptive point and candidate key points of the pixel point can be synchronously determined as the adaptive point and the set of key points. Compared with the top-down and bottom-up two-stage recognition methods in the related art, the pose recognition method in this embodiment can determine the position of the center point of the object to be recognized and the set of key points without post-processing, realizing single-stage pose recognition, and can avoid the computing burden and storage burden in the post-processing process, improving the recognition efficiency.

[0083] Next, referring to Figure 5 , Figure 5 shows a flowchart for predicting the position of candidate adaptive points in an embodiment of the pose recognition method of the present disclosure. As shown in Figure 5 shown, based on the embodiments shown in Figure 4 and Figure 3 , the above step 320 may further include the following steps:

[0084] Step 510: Based on the second feature data, predict the first offset corresponding to each pixel point for each part.

[0085] In this embodiment, the first offset is a vector pointing from the pixel point to the adaptive point corresponding to each part, and can represent the relative position of the part of the object to be recognized and the center point.

[0086] As an example, the region perception branch in the pose recognition model may use a convolutional layer or a fully connected layer to perform prediction based on the second feature data to obtain one or more first offsets of each pixel point.

[0087] Step 520: Based on the position of each pixel point and its first offset corresponding to each part, determine the positions of the candidate adaptive points corresponding to each part for each pixel point.

[0088] In Figure 5 the shown process, the first offset of the pixel point corresponding to each part can be predicted through the second feature data, and then combined with the position of the pixel point to determine the position of the candidate adaptive point. Through the position of the adaptive point, the local features can be perceived more accurately.

[0089] Then refer to Figure 6 , Figure 6 which shows the flowchart of predicting the candidate key point set in an embodiment of the posture recognition method of the present disclosure. As Figure 6 shown, based on the embodiments shown in Figure 4 and Figure 5 the above step 420 includes:

[0090] Step 610: Based on the key point regression features of each part corresponding to each pixel point respectively, predict one or more second offsets corresponding to each candidate adaptive point.

[0091] In this embodiment, the second offset is a vector pointing from the candidate adaptive point to the candidate key point.

[0092] As an example, the key point regression network branch in the posture recognition model can use a convolutional layer or a fully connected layer to predict the key point regression features of each part corresponding to each pixel point respectively, and determine one or more second offsets corresponding to each candidate adaptive point.

[0093] Step 620: Based on the positions of the candidate adaptive points of each part corresponding to each pixel point respectively, and one or more second offsets corresponding to each candidate adaptive point, predict the positions of one or more candidate key points of each part corresponding to each pixel point respectively, so as to generate a candidate key point set of each part corresponding to each pixel point.

[0094] In the process of implementing the present disclosure, the inventors also found that related technologies usually predict the position of the downstream key point based on the position of the upstream key point. For example, when predicting the position of the wrist joint key point, first the position from the center point to the shoulder joint key point needs to be predicted, then based on the position of the shoulder joint key point, the position of the elbow joint key point is predicted, and then based on the position of the elbow joint key point, the position of the wrist joint key point is predicted. This results in the errors of the shoulder joint key point and the elbow joint key point accumulating into the error of the wrist joint key point, so the accuracy of key point prediction is relatively low.

[0095] From Figure 6 it can be seen that Figure 6The process shown embodies the steps of "predicting the second offset based on the key point regression feature and then determining the position of the candidate key point in combination with the position of the candidate adaptive point". Since the position of the candidate adaptive point is not predefined but predicted based on the feature data, the cumulative error in the process of predicting the candidate key point can be reduced.

[0096] Next, refer to Figure 7 , Figure 7 which shows a flowchart of an embodiment of the method for training a pose recognition model according to the present disclosure. As Figure 7 shown, the process includes the following steps:

[0097] Step 710, obtain a training set.

[0098] Among them, the training set includes sample images with labeled sample labels. The sample labels include the position of the sample center point of the object to be recognized, the position of the sample key points, as well as the sample center point heat map and the sample key point heat map of the sample image.

[0099] In this embodiment, the sample center point heat map can represent the reference confidence of each pixel point as the center point, and the sample key point heat map can represent the probability of each pixel point as the key point.

[0100] Step 720, process the sample images in the training set based on the initial backbone network of the pre-constructed initial pose recognition model to obtain sample feature data.

[0101] Step 730, process the sample feature data based on the initial pose regression sub-network of the initial pose recognition model to obtain the predicted center point confidence of each pixel point and the corresponding position of the predicted key point.

[0102] As an example, the initial pose regression sub-network can include an initial center point perception network branch, an initial region perception network branch, and an initial key point regression network branch. The initial region perception network can predict the position of the candidate adaptive point of each pixel point, and the initial center point perception network can predict the predicted center point confidence of each pixel point; the initial region perception network branch can predict the position of the candidate adaptive point corresponding to each pixel point; the key point regression network branch can predict the position of the predicted key point corresponding to each candidate adaptive point, so as to obtain the position of the predicted key point corresponding to each pixel point.

[0103] Step 740, process the sample feature data based on the key point heat map sub-network of the initial pose recognition model to generate the predicted key point heat map of the sample image.

[0104] In this embodiment, the key point heatmap network predicts the confidence of multiple types of key points corresponding to each pixel point based on the sample feature data, and generates a multi-channel predicted key point heatmap based on the confidence of multiple types of key points corresponding to each pixel point. Each type corresponds to one channel. The type of the key point can, for example, represent the joint type of the object to be recognized represented by the key point. For example, the key points corresponding to the shoulder joint belong to the same type of key points.

[0105] Step 750: Determine the first loss function based on the predicted center point confidence of each pixel point and the sample center point heatmap.

[0106] As an example, the execution entity can first determine the reference confidence corresponding to the position from the sample center point heatmap according to the position of the pixel point, then determine the difference between the predicted center point confidence of each pixel point and the reference confidence, and thereby determine the first loss function.

[0107] Step 760: Determine the second loss function based on the position of the sample key points and the position of the predicted key points corresponding to the reference pixel points.

[0108] Wherein, the position of the reference pixel point is the same as the position of the sample center point.

[0109] As an example, after the execution entity obtains the predicted key points of each pixel point through the initial pose regression network, it can determine the position of the reference pixel point according to the position of the sample center point, and then obtain the position of the predicted key points corresponding to the reference pixel point. Then, it determines the predicted offset between the position of the predicted key points and the position of the reference pixel point, and determines the sample offset between the position of the sample key points and the position of the sample center point. After that, it can determine the second loss value according to the difference between the predicted offset and the sample offset.

[0110] Step 770: Determine the third loss function based on the predicted key point heatmap and the sample key point heatmap.

[0111] As an example, the execution entity can first determine the difference between the pixel values of the pixel points at the same position in the predicted key point heatmap and the sample label, and then determine the value of the loss function based on the differences of all pixel points.

[0112] Step 780: Adjust the parameters of the initial pose recognition model based on the first loss function, the second loss function, and the third loss function until the termination condition is met, delete the key point heatmap network, and obtain the pose recognition model.

[0113] In Figure 7In the illustrated embodiment, a first loss function can be used to constrain the process of predicting the center point confidence in the pose recognition model, and a second loss function can be used to constrain the process of predicting the positions of key points in the pose recognition model. At the same time, a sample key point heat map and a key point heat map network branch can be used to assist the backbone network in learning the extraction strategy of the structured pose information of the object to be recognized, which can improve the training efficiency.

[0114] Exemplary Device

[0115] Reference is made below to Figure 8 , Figure 8 which shows a schematic structural diagram of an embodiment of the pose recognition device of the present disclosure. As Figure 8 shown, the device includes: a feature extraction unit 810 configured to extract first feature data of an image including an object to be recognized by using a pose recognition model; a first prediction unit 820 configured to predict the position of the center point of the object to be recognized and the positions of adaptive points corresponding to each part based on the first feature data, where the center point represents the imaging point of the center point part of the object to be recognized; a second prediction unit 830 configured to predict a set of key points corresponding to each part based on the first feature data and the positions of the adaptive points corresponding to each part; and a pose determination unit 840 configured to determine the target pose of the object to be recognized based on the position of the center point and the set of key points corresponding to each part.

[0116] In one embodiment, the first prediction unit 820 further includes: a first extraction module configured to extract features from the first feature data based on the key point regression network branch of the pose recognition model to obtain second feature data; a first prediction module configured to process the second feature data based on the region perception network branch of the pose recognition model to predict the positions of candidate adaptive points corresponding to each part for each pixel point in the image; a second extraction module configured to extract features from the first feature data based on the center point perception network branch of the pose recognition model to obtain third feature data; a third extraction module configured to extract the center regression feature corresponding to each pixel point from the third image feature based on the positions of the candidate adaptive points corresponding to each part for each pixel point; a second prediction module configured to predict the center point confidence of each pixel point based on the center regression feature corresponding to each pixel point; and a first determination module configured to use a maximum pooling kernel to determine the position of the pixel point with a center point confidence greater than a preset threshold as the position of the center point of the object to be recognized, and determine the positions of the candidate adaptive points corresponding to each part for the pixel point as the positions of the adaptive points corresponding to each part of the object to be recognized.

[0117] In one embodiment, the second prediction unit 830 further includes: a fourth extraction module configured to extract, by using a key point regression network branch, key point regression features of each part corresponding to each pixel point from the second feature data based on the positions of candidate adaptive points of each part corresponding to each pixel point; a third prediction module configured to predict a set of candidate key points of each part corresponding to each pixel point based on the key point regression features of each part corresponding to each pixel point and the positions of the corresponding candidate adaptive points; a second determination module configured to, after determining the position of a pixel point with a center point confidence greater than a preset threshold as the position of the center point of the object to be recognized by using a maximum pooling kernel, determine the set of candidate key points of each part corresponding to the pixel point as the set of key points corresponding to each part.

[0118] In one embodiment, the first prediction module further includes: a first offset sub-module configured to predict a first offset amount of each pixel point corresponding to each part based on the second feature data; a first position sub-module configured to determine the positions of candidate adaptive points of each pixel point corresponding to each part based on the position of each pixel point and its first offset amount corresponding to each part.

[0119] In one embodiment, the third prediction module further includes: a second offset sub-module configured to predict one or more second offset amounts corresponding to each candidate adaptive point based on the key point regression features of each part corresponding to each pixel point; a second position sub-module configured to predict the positions of one or more candidate key points of each part corresponding to each pixel point based on the positions of the candidate adaptive points of each part corresponding to each pixel point and the one or more second offset amounts corresponding to each candidate adaptive point, so as to generate a set of candidate key points of each part corresponding to each pixel point.

[0120] Next, referring to Figure 9 , Figure 9 shows a schematic structural diagram of an embodiment of the apparatus for training a pose recognition model according to the present disclosure, as Figure 9As shown in the figure, the device includes: a sample acquisition unit 910 configured to acquire a training set, the training set including sample images with labeled sample tags, the sample tags including the position of the sample center point of the object to be recognized, the positions of sample key points, and the sample center point heat map and sample key point heat map of the sample image; a feature extraction unit 920 configured to process the sample images in the training set based on the initial backbone network of the pre-constructed initial pose recognition model to obtain sample feature data; a pose prediction unit 930 configured to process the sample feature data based on the initial pose regression sub-network of the initial pose recognition model to obtain the predicted center point confidence of each pixel point and the positions of its corresponding predicted key points; a heat map prediction unit 940 configured to process the sample feature data based on the key point heat map network of the initial pose recognition model to generate the predicted key point heat map of the sample image; a first loss unit 950 configured to determine a first loss function based on the predicted center point confidence of each pixel point and the sample center point confidence; a second loss unit 960 configured to determine a second loss function based on the positions of the sample key points and the positions of the predicted key points corresponding to the reference pixel points, the position of the reference pixel point being the same as the position of the sample center point; a third loss unit 970 configured to determine a third loss function based on the predicted key point heat map and the sample key point heat map; and a model training unit 980 configured to adjust the parameters of the initial pose recognition model based on the first loss function, the second loss function, and the third loss function until the termination condition is met, delete the key point heat map network, and obtain the pose recognition model.

[0121] Exemplary Electronic Device

[0122] Next, with reference to Figure 10 the electronic device according to an embodiment of the present disclosure will be described. The electronic device may be any one or both of the first device 100 and the second device 200, or a stand-alone device independent of them, and the stand-alone device may communicate with the first device and the second device to receive the input signals collected from them.

[0123] Figure 10 The block diagram of the electronic device according to an embodiment of the present disclosure is illustrated.

[0124] As Figure 10 shown, the electronic device 10 includes one or more processors 11 and a memory 12.

[0125] The processor 11 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.

[0126] The memory 12 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 11 may run the program instructions to implement the gesture recognition method and / or the method of training a gesture recognition model according to various embodiments of the present disclosure described above and / or other desired functions. Various contents such as input signals, signal components, noise components, etc. may also be stored in the computer-readable storage media.

[0127] In one example, the electronic device 10 may further include: an input device 13 and an output device 14, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0128] For example, when the electronic device is the first device 100 or the second device 200, the input device 13 may be the above-mentioned microphone or microphone array for capturing the input signal of the sound source. When the electronic device is a stand-alone device, the input device 13 may be a communication network connector for receiving the collected input signals from the first device 100 and the second device 200.

[0129] In addition, the input device 13 may further include, for example, a keyboard, a mouse, and so on.

[0130] The output device 14 may output various information to the outside, including the determined distance information, direction information, etc. The output device 14 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0131] Of course, for simplicity, Figure 10 only some of the components related to the present disclosure in the electronic device 10 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 10 may further include any other appropriate components.

[0132] Exemplary Computer Program Product and Computer Readable Storage Medium

[0133] In addition to the above methods and devices, the embodiments of the present disclosure may also be computer program products, which include computer program instructions that, when run by a processor, cause the processor to execute the steps in the gesture recognition method and / or the method of training a gesture recognition model according to various embodiments of the present disclosure described in the "Exemplary Method" section above of this specification.

[0134] The computer program product can be written in any combination of one or more programming languages for executing the program code of the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0135] In addition, an embodiment of the present disclosure can also be a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the posture recognition method and / or the method of training a posture recognition model according to various embodiments of the present disclosure described in the "Exemplary Method" section above of this specification.

[0136] The computer-readable storage medium can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0137] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for the purposes of illustration and easy understanding, and are not limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0138] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts between the embodiments, reference can be made to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and reference can be made to the partial description of the method embodiment for the relevant parts.

[0139] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the phrase "and / or", and can be used interchangeably with it, unless the context clearly indicates otherwise. The phrase "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with it.

[0140] The methods and apparatuses of this disclosure can be implemented in many ways. For example, the methods and apparatuses of this disclosure can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the methods is for illustration only, and the steps of the methods of this disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, this disclosure can also be implemented as a program recorded on a recording medium, and these programs include machine-readable instructions for implementing the methods according to this disclosure. Therefore, this disclosure also covers a recording medium storing a program for executing the methods according to this disclosure.

[0141] It should also be noted that in the apparatuses, equipment, and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.

[0142] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0143] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.

Claims

1. A gesture recognition method, comprising: extracting first feature data of an image containing an object to be recognized by using a gesture recognition model, where the first feature data includes at least one of the texture feature of the image and the semantic information, boundary information, and position information of pixel points; predicting the position of the center point of the object to be recognized and the positions of adaptive points corresponding to each part based on the first feature data by using the gesture recognition model, where the center point represents the imaging point of the center point part of the object to be recognized, and the adaptive points correspond to each part of the object to be recognized; predicting a key point set corresponding to each part based on the first feature data and the positions of the adaptive points corresponding to each part; determining the target gesture of the object to be recognized based on the position of the center point and the key point sets corresponding to each part; wherein, predicting the key point set corresponding to each part based on the first feature data and the positions of the adaptive points corresponding to each part includes: extracting the feature data corresponding to the positions of the adaptive points from the first feature data; predicting the key points corresponding to the adaptive points by using the extracted feature data to obtain the key point sets corresponding to each part; 2. The method according to claim 1, wherein, predicting the position of the center point of the object to be recognized and the positions of the adaptive points corresponding to each part based on the first feature data includes: performing feature extraction on the first feature data by using the key point regression network branch of the gesture recognition model to obtain second feature data; processing the second feature data by using the region perception network branch of the gesture recognition model to predict the positions of candidate adaptive points corresponding to each pixel point in the image for each part; performing feature extraction on the first feature data by using the center point perception network branch of the gesture recognition model to obtain third feature data; extracting the center regression feature corresponding to each pixel point from the third feature data based on the positions of the candidate adaptive points corresponding to each pixel point for each part; predicting the center point confidence of each pixel point based on the center regression feature corresponding to each pixel point; using a maximum pooling kernel to determine the position of the pixel point with the center point confidence greater than a preset threshold as the position of the center point of the object to be recognized, and determining the positions of the candidate adaptive points corresponding to each part of the object to be recognized for this pixel point as the positions of the adaptive points corresponding to each part of the object to be recognized; 3. The method according to claim 2, wherein predicting the key point set corresponding to each part based on the first feature data and the positions of the adaptive points corresponding to each part includes: using the key point regression network branch to extract the key point regression features corresponding to each part for each pixel point from the second feature data based on the positions of the candidate adaptive points corresponding to each part for each pixel point; Predict the candidate key point sets of each part corresponding to each pixel point based on the key point regression features of each part corresponding to each pixel point respectively and the positions of the corresponding candidate adaptive points thereof. After using the maximum pooling kernel to determine the position of the pixel point with the center point confidence greater than the preset threshold as the position of the center point of the object to be recognized, determine the candidate key point sets of each part corresponding to the pixel point as the key point sets corresponding to each part respectively.

4. The method according to claim 3, wherein, Based on the region perception network branch of the pose recognition model, process the second feature data to predict the positions of the candidate adaptive points corresponding to each pixel point in the image for each part, including: Based on the second feature data, predict the first offset corresponding to each pixel point for each part. Based on the position of each pixel point and the first offset corresponding to each part thereof, determine the positions of the candidate adaptive points corresponding to each pixel point for each part.

5. The method according to claim 4, wherein Predict the candidate key point sets of each part corresponding to each pixel point based on the key point regression features of each part corresponding to each pixel point respectively and the positions of the corresponding candidate adaptive points thereof, including: Based on the key point regression features of each part corresponding to each pixel point respectively, predict one or more second offsets corresponding to each candidate adaptive point. Based on the positions of the candidate adaptive points of each part corresponding to each pixel point respectively and one or more second offsets corresponding to each candidate adaptive point, predict the positions of one or more candidate key points of each part corresponding to each pixel point to generate the candidate key point sets of each part corresponding to each pixel point.

6. A method for training a pose recognition model, including: Obtain a training set, where the training set includes sample images with labeled sample labels, and the sample labels include the position of the sample center point of the object to be recognized, the positions of the sample key points, and the sample center point heat map and sample key point heat map corresponding to the sample image. Process the sample images in the training set based on the initial backbone network of the pre-constructed initial pose recognition model to obtain sample feature data. Process the sample feature data based on the initial pose regression sub-network of the initial pose recognition model to obtain the predicted center point confidence of each pixel point and the positions of the corresponding predicted key points. Process the sample feature data based on the key point heat map network of the initial pose recognition model to generate the predicted key point heat map of the sample image. Based on the predicted center point confidence of each pixel point and the sample center point heat map, determine the first loss function. Based on the positions of the sample key points and the positions of the predicted key points corresponding to the reference pixel points, determine the second loss function, where the position of the reference pixel point is the same as the position of the sample center point. Based on the predicted key point heat map and the sample key point heat map, determine the third loss function. Based on the first loss function, the second loss function, and the third loss function, adjust the parameters of the initial pose recognition model until the termination condition is satisfied, delete the keypoint heatmap network, and obtain the pose recognition model.

7. A pose recognition device, comprising: A feature extraction unit configured to extract first feature data of an image including an object to be recognized by using a pose recognition model, where the first feature data includes at least one of a texture feature of the image and semantic information, boundary information, and position information of a pixel point; A first prediction unit configured to predict the position of the center point of the object to be recognized and the positions of adaptive points corresponding to each part respectively based on the first feature data by using the pose recognition model, where the center point represents an imaging point of the center point part of the object to be recognized, and the adaptive points correspond to each part of the object to be recognized; A second prediction unit configured to predict a set of keypoints corresponding to each part respectively based on the first feature data and the positions of the adaptive points corresponding to each part respectively; the second prediction unit is further configured to extract feature data corresponding to the positions of the adaptive points from the first feature data, and predict the keypoints corresponding to the adaptive points by using the extracted feature data to obtain the set of keypoints corresponding to each part respectively; A pose determination unit configured to determine the target pose of the object to be recognized based on the position of the center point and the set of keypoints corresponding to each part respectively.

8. A device for training a pose recognition model, comprising: A sample acquisition unit configured to acquire a training set, where the training set includes sample images with labeled sample tags, and the sample tags include the position of the sample center point of the object to be recognized, the positions of the sample keypoints, and the sample center point heatmap and sample keypoint heatmap corresponding to the sample images; A feature extraction unit configured to process the sample images in the training set based on an initial backbone network of a pre-constructed initial pose recognition model to obtain sample feature data; A pose prediction unit configured to process the sample feature data based on an initial pose regression sub-network of the initial pose recognition model to obtain the predicted center point confidence of each pixel point and the positions of the corresponding predicted keypoints; A heatmap prediction unit configured to process the sample feature data based on the keypoint heatmap network of the initial pose recognition model to generate a predicted keypoint heatmap of the sample image; A first loss unit configured to determine a first loss function based on the predicted center point confidence of each pixel point and the sample center point heatmap; A second loss function unit configured to determine a second loss function based on the positions of the sample keypoints and the positions of the predicted keypoints corresponding to the reference pixel points, where the position of the reference pixel point is the same as the position of the sample center point; A third loss unit configured to determine a third loss function based on the predicted keypoint heatmap and the sample keypoint heatmap; The model training unit is configured to adjust the parameters of the initial pose recognition model based on the first loss function, the second loss function, and the third loss function until a termination condition is satisfied, and delete the key point heat map network to obtain the pose recognition model.

9. A computer-readable storage medium storing a computer program for executing the method according to any one of claims 1-7 above.

10. An electronic device, comprising: a processor; a memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of claims 1-7 above.

Citation Information

Patent Citations

  • Attitude recognition method and system based on thermodynamic diagram and offset vector and storage medium

    CN111191622A

  • Attitude determination method, device and equipment, and storage medium

    CN112241731A