Posture Recognition Method, Device, Storage Medium and Electronic Device

By extracting image feature data and predicting the query position of key points, the problem of inaccurate pose scores in the existing pose recognition methods is solved, and a higher pose recognition accuracy is achieved.

CN114170439BActive Publication Date: 2025-06-17BEIJING HORIZON INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111463800.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-02
Publication Date
2025-06-17
Estimated Expiration
2041-12-02

AI Technical Summary

Technical Problem

The existing pose recognition method is based on the 2D Gaussian core and the discrete {0,1} estimated pose score, resulting in the pose score not accurately representing the quality of pose regression.

Method used

By extracting the feature data of the image, predict the key point query position corresponding to each pixel point, and predict the pose quality score and key point position based on the feature data and key point query position, and determine the target pose of the object to be identified.

Benefits of technology

The correlation between the pose quality score and the target pose is improved and the accuracy of pose recognition is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114170439B_ABST
    Figure CN114170439B_ABST
Patent Text Reader

Abstract

An embodiment of the present disclosure discloses a gesture recognition method, device, storage medium, and electronic device. The method includes: extracting feature data of an image including an object to be recognized; predicting a key point query position corresponding to each pixel point from the feature data based on a preset prediction method, where the pixel point represents an imaging point of a candidate center point part of the object to be recognized, and the key point corresponding to the pixel point represents an imaging point of a candidate key part of the object to be recognized; predicting a gesture quality score corresponding to each pixel point and the position of the key point based on the feature data and the key point query position corresponding to each pixel point; determining pixel points whose gesture quality scores meet a preset condition as target pixel points, and determining the position of the key point corresponding to the target pixel point as the position of the target key point; and determining the target gesture of the object to be recognized based on the position of the target pixel point and the position of the target key point. The accuracy of gesture recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to computer vision technology, and in particular, to a pose recognition method, apparatus, storage medium, and electronic device. Background Art

[0002] In the field of computer vision, pose recognition is used to locate the key point positions of an object to be recognized in an image, and the pose of the object to be recognized is characterized based on the key point positions, such as human pose recognition. With the application of deep learning technology, great progress has been made in this field and it has promoted the development of fields such as human-computer interaction and behavior recognition.

[0003] In related technologies, pose recognition methods usually determine the pose score of an object instance to be recognized based on a two-dimensional Gaussian kernel and discrete {1, 0}, so as to characterize the quality of pose recognition. Summary of the Invention

[0004] To solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a pose recognition method, apparatus, storage medium, and electronic device.

[0005] According to one aspect of the embodiments of the present disclosure, a pose recognition method is provided, including: extracting feature data of an image including an object to be recognized; predicting a key point query position corresponding to each pixel point from the feature data based on a preset prediction method, where the pixel point represents an imaging point of a candidate center point part of the object to be recognized, and the key point represents an imaging point of a candidate key part of the object to be recognized; predicting a pose quality score corresponding to each pixel point and the position of the key point based on the feature data and the key point query position corresponding to each pixel point; determining a target pixel point for a pixel point whose pose quality score meets a preset condition, and determining the position of the key point corresponding to the target pixel point as the position of the target key point; determining the target pose of the object to be recognized based on the position of the target pixel point and the position of the target key point.

[0006] According to another aspect of the embodiments of the present disclosure, a method for training a pose recognition model is provided, including: obtaining a sample image marked with a reference quality score of a pixel point and a sample pose of an object to be recognized, where the sample pose includes the position of a sample target pixel point representing an imaging point of a center point part of the object to be recognized and the position of a sample key point representing an imaging point of a key part of the object to be recognized; processing the sample image by using a pre-constructed initial pose recognition model to obtain a predicted quality score corresponding to each pixel point and a predicted pose of the object to be recognized; determining a first loss function based on the predicted pose and the sample pose; determining a second loss function based on the predicted quality score corresponding to each pixel point and a preset reference quality score; adjusting the initial pose recognition model based on the first loss value and the second loss value until a training stop condition is met to obtain a pose recognition model.

[0007] According to another aspect of the embodiments of the present disclosure, a gesture recognition device is provided, including: a feature extraction unit configured to extract feature data of an image including an object to be recognized; a first prediction unit configured to predict a key point query position corresponding to each pixel point from the feature data based on a preset prediction method, where the pixel point represents an imaging point of a candidate center point part of the object to be recognized, and the key point corresponding to the pixel point represents an imaging point of a candidate key part of the object to be recognized; a second prediction unit configured to predict a gesture quality score corresponding to each pixel point and the position of the key point based on the feature data and the key point query position corresponding to each pixel point; a target determination unit configured to determine a target pixel point for the pixel points whose gesture quality scores meet a preset condition, and determine the position of the key point corresponding to the target pixel point as the position of the target key point; a gesture determination unit configured to determine the target gesture of the object to be recognized based on the position of the target pixel point and the position of the target key point.

[0008] According to another aspect of the embodiments of the present disclosure, a device for training a gesture recognition model is provided, including: a sample acquisition unit configured to acquire a sample image marked with a sample gesture of an object to be recognized, where the sample gesture includes the position of a sample target pixel point representing an imaging point of a center point part of the object to be recognized and the position of a sample key point representing an imaging point of a key part of the object to be recognized; a model prediction unit configured to process the sample image by using a pre-constructed initial gesture recognition model to obtain a predicted quality score corresponding to each pixel point and the predicted gesture of the object to be recognized; a first loss unit configured to determine a first loss function based on the predicted gesture and the sample gesture; a second loss unit configured to determine a second loss function based on the predicted quality score corresponding to each pixel point and a preset reference quality score; a model training unit configured to adjust the initial gesture recognition model based on the first loss value and the second loss value until a training stop condition is met to obtain a gesture recognition model.

[0009] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, and the storage medium stores a computer program for executing the method in any of the above embodiments.

[0010] According to another aspect of the embodiments of the present disclosure, an electronic device is provided, and the electronic device includes: a processor; a memory for storing executable instructions that can be executed by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method in any of the above embodiments.

[0011] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product including computer programs / instructions which, when executed by a processor, implement the method in any of the above embodiments.

[0012] Based on the pose recognition method provided in the above embodiments of the present disclosure, the query position of the key point corresponding to each pixel point can be predicted using the feature data of the image, and then based on the feature data and the query position of the key point, the pose quality score corresponding to each pixel point and the position of the key point can be predicted, where the pixel point represents the imaging point of the candidate center point part of the object to be recognized, and the key point corresponding to the pixel point represents the imaging point of the candidate key part of the object to be recognized; then the pixel points whose pose quality scores meet the preset conditions are determined as target pixel points, and the positions of the key points corresponding to the target pixel points are determined as the positions of the target key points; finally, based on the positions of the target pixel points and the positions of the target key points, the target pose of the object to be recognized is determined. By predicting the pose quality score of the pixel point through the feature data and the query position of the key point, and using this to represent the accuracy of pose recognition, the correlation between the pose quality score and the target pose is improved, which helps to improve the accuracy of pose recognition.

[0013] The technical solution of the present disclosure will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings

[0014] By describing the embodiments of the present disclosure in more detail with reference to the drawings, the above and other objects, features, and advantages of the present disclosure will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure, and do not constitute a limitation to the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.

[0015] Figure 1 is an exemplary architecture diagram of a deep network model applicable to the pose recognition method of the present disclosure;

[0016] Figure 2 is a flowchart of the pose recognition method provided in an exemplary embodiment of the present disclosure;

[0017] Figure 3 is a flowchart of predicting the query position of the key point in an embodiment of the pose recognition method of the present disclosure;

[0018] Figure 4 is a flowchart of predicting the position of the key point in an embodiment of the pose recognition method of the present disclosure;

[0019] Figure 5 is a flowchart of generating the semantic feature of the key point in an embodiment of the pose recognition method of the present disclosure;

[0020] Figure 6 It is a flowchart for predicting the pose quality score in an embodiment of the pose recognition method of the present disclosure;

[0021] Figure 7 It is a flowchart of an embodiment of the method for training a pose recognition model of the present disclosure;

[0022] Figure 8 It is a flowchart for determining the second loss function value in an embodiment of the method for training a pose recognition model of the present disclosure;

[0023] Figure 9 It is a flowchart for generating a sample score map in an embodiment of the method for training a pose recognition model of the present disclosure;

[0024] Figure 10 It is a schematic structural diagram of an embodiment of the pose recognition device of the present disclosure;

[0025] Figure 11 It is a schematic structural diagram of an embodiment of the device for training a pose recognition model of the present disclosure;

[0026] Figure 12 It is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed implementation manners

[0027] Next, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.

[0028] It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.

[0029] Those skilled in the art can understand that terms such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.

[0030] It should also be understood that in the embodiments of the present disclosure, "a plurality" may refer to two or more, and "at least one" may refer to one, two or more.

[0031] It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, unless clearly defined or given a contrary indication in the context, it can generally be understood as one or more.

[0032] In addition, the term "and / or" in this disclosure is merely a description of the relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, both A and B exist simultaneously, and B exists alone. In addition, the character " / " in this disclosure generally represents an "or" relationship between the associated objects before and after.

[0033] It should also be understood that the descriptions of the various embodiments in this disclosure emphasize the differences between the various embodiments, and their similarities or similarities can be referred to each other. For the sake of brevity, they will not be elaborated one by one.

[0034] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn in actual proportional relationships.

[0035] The following description of at least one exemplary embodiment is actually merely illustrative and in no way limits this disclosure or its application or use.

[0036] Well-known technologies, methods, and devices for those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods, and devices should be regarded as part of the specification.

[0037] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0038] The embodiments of this disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate together with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, small computer systems, large computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.

[0039] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, target programs, components, logics, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.

[0040] Overview of the Application

[0041] In the process of implementing the present disclosure, the inventors found that single-stage pose recognition methods usually use a 2D Gaussian kernel and discrete {0, 1} to estimate the pose score of the object to be recognized, and use this to characterize the quality of pose recognition. This method results in the pose score not being able to accurately characterize the quality of pose regression.

[0042] Exemplary System

[0043] Next, refer to Figure 1 , Figure 1 shows an architecture diagram of a pose recognition model that can implement the pose recognition method of the present disclosure. As Figure 1 shown, the execution entity can be, for example, a terminal device or a server, on which computer instructions of the pose recognition model can be pre-loaded. The pose recognition model can be, for example, a deep network model constructed based on convolutional neural networks such as ResNet and HRNet.

[0044] When the execution entity obtains an image containing the object to be recognized, the pose recognition model can extract the feature data Rg of the image through the backbone network 110 (which can be a convolutional neural network such as ResNet and HRNet). The first network branch 120 predicts the key point query position 140 corresponding to each pixel point based on the feature data Rg, and then predicts the position 150 of the key point corresponding to each pixel point according to the key point query position and the feature data. The second network branch 130 determines the pose feature 150 corresponding to each pixel point according to the key point query position 140 and the feature data Rg, and predicts the pose quality score 170 corresponding to each pixel point from the pose feature 150. Finally, the method of non-maximum suppression is performed through the pooling kernel 180, and the target pixel points are screened out from all pixel points based on the pose quality score. 190 is the set of target pixel points and their corresponding target key points. Then, based on the positions of the target pixel points and the positions of their corresponding target key points, the target pose 191 of the object to be recognized is determined.

[0045] Exemplary Method

[0046] Figure 2 is a schematic flow chart of a pose recognition method provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device, such as Figure 2 as shown, including the following steps:

[0047] Step 210: Extract feature data of an image including an object to be recognized.

[0048] In this embodiment, the object to be recognized can be, for example, a human body, an animal, or other objects that can recognize poses. The feature data can include, but is not limited to, the texture features of the image and the semantic information, boundary information, position information, etc. of the pixel points. The feature data can be, for example, a multi-dimensional matrix.

[0049] As an example, the convolutional layer or encoder-decoder in a deep network can be used to extract feature data from the image.

[0050] Step 220: Predict the key point query position corresponding to each pixel point from the feature data based on a preset prediction method.

[0051] Among them, the pixel point represents the imaging point of the candidate center point part of the object to be recognized, and the key point corresponding to the pixel point represents the imaging point of the candidate key part of the object to be recognized.

[0052] In this embodiment, the key point query position represents the position encoding the feature information (including but not limited to semantic information and position information) of the key point. The key point query position can be a continuous pixel position key point query position obtained by prediction.

[0053] As an example, the prediction method can be implemented using the convolutional layer or fully connected layer in a deep network model. The execution entity can traverse the pixel points in the image, extract the corresponding features from the feature data according to the positions of the pixel points, and then predict the key point query position corresponding to the pixel point based on the extracted features.

[0054] It should be noted that each pixel point can correspond to one or more key point query positions, and the present disclosure does not limit this.

[0055] Step 230: Predict the pose quality score corresponding to each pixel point and the position of the key point based on the feature data and the key point query position corresponding to each pixel point.

[0056] In this embodiment, the pose quality score represents the matching degree between the predicted pose of the pixel point and the true pose of the object to be recognized.

[0057] As an example, the execution entity can extract the features of the key point query position corresponding to the pixel point from the feature data based on the key point query position, and then predict the key points and pose quality scores corresponding to the pixel point based on the features of the key point query position respectively. When the pixel coordinates of the key point query position are integers, the feature data can be directly extracted from the feature data based on the pixel coordinates; when the pixel coordinates of the key point query position include non-integers, bilinear interpolation can be used to extract the feature data from the feature data.

[0058] It should be noted that each key point query position can correspond to one or more key points, and the position of a pixel point or a key point generally refers to the pixel coordinates of the point in the image.

[0059] Step 240: Determine the pixel points whose pose quality scores meet the preset conditions as target pixel points, and determine the positions of the key points corresponding to the target pixel points as the positions of the target key points.

[0060] The preset conditions are used to evaluate the pose quality. For example, it can be greater than a preset score threshold, or it can be the local score maximum. The score threshold can be set according to experience. In a specific example, a large number of pose recognition results can be obtained first, and then the pixel points in the pose recognition results (such as the imaging points corresponding to the center part and key parts of the object to be recognized) can be scored according to the matching degree between the pose recognition results and the pose. Finally, the score threshold can be determined through statistical analysis.

[0061] In this embodiment, the target pixel points and their corresponding target key points represent the pixel points that meet the pose quality requirements. Through step 230, the execution entity can determine the pose quality scores of each pixel point and the positions of the corresponding key points (which can include one or more), and regard each pixel point and the position of its corresponding key point as candidate data for determining the target pose of the object to be recognized. Then, by comparing the pose quality scores with the preset conditions, the target pixel points that meet the preset conditions are determined from all pixel points. Correspondingly, the positions of the key points corresponding to the target pixel points can be used as the positions of the target key points.

[0062] Step 250: Determine the target pose of the object to be recognized based on the positions of the target pixel points and the positions of the target key points.

[0063] Generally, the pose of the object to be recognized can be characterized by the central position of the object to be recognized and the relative positions of multiple key parts. Furthermore, the central position and key parts of the object to be recognized can be abstracted as points, and thus the pose of the object to be recognized can be characterized by the positions of the points. Taking human pose recognition as an example, the target pixel points can represent the position of the human body center, and the positions of the target key points can represent the positions of the human body joint points.

[0064] In this embodiment, the execution subject may determine the position of the target pixel point as the center position of the object to be recognized, and determine the positions of the target key points as the positions of the respective key parts of the object to be recognized, so as to obtain the target pose of the object to be recognized.

[0065] The pose recognition method provided in this embodiment can use the feature data of the image to predict the key point query position corresponding to each pixel point, and then, based on the feature data and the key point query position, predict the pose quality score corresponding to each pixel point and the position of the key point, where the pixel point represents the imaging point of the candidate center point part of the object to be recognized, and the key point corresponding to the pixel point represents the imaging point of the candidate key part of the object to be recognized; then, determine the pixel points whose pose quality scores meet the preset conditions as target pixel points, and determine the positions of the key points corresponding to the target pixel points as the positions of the target key points; finally, based on the positions of the target pixel points and the positions of the target key points, determine the target pose of the object to be recognized. By predicting the pose quality score of the pixel point through the feature data and the key point query position, and using this to represent the accuracy of pose recognition, the correlation degree between the pose quality score and the target pose is improved, which helps to improve the accuracy of pose recognition.

[0066] Any of the pose recognition methods provided in the embodiments of the present disclosure may be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices, servers, etc. Alternatively, any of the pose recognition methods provided in the embodiments of the present disclosure may be executed by a processor, for example, the processor executes any of the pose recognition methods mentioned in the embodiments of the present disclosure by calling the corresponding instructions stored in the memory. Details will not be described below.

[0067] Next, refer to Figure 3 , Figure 3 shows a flowchart for predicting the key point query position in an embodiment of the pose recognition method of the present disclosure. In some optional implementation manners of the above Figure 2 shown embodiment, step 220 may include the following steps:

[0068] Step 310: The number of key points is a first preset number. Based on the position of each pixel point, perform feature extraction of the feature data for the first preset number of times to obtain the first preset number of key point regression features corresponding to each pixel point.

[0069] Combined with Figure 1For exemplary illustration, the execution subject can use the first network branch in the pose recognition model to extract the first preset number of key point regression features corresponding to each pixel point from the feature data. Specifically, the first network branch can include the first preset number of independent sub-branches, and each sub-branch can extract the features corresponding to the position of the pixel point from the feature data through a 1X1 convolutional layer to obtain the key point regression features corresponding to the sub-branch. Through the parallel processing of each sub-branch, the first preset number of key point regression features corresponding to each pixel point can be obtained.

[0070] It should be noted that the first preset number can be one or more.

[0071] Step 320: Based on the first preset number of key point regression features corresponding to each pixel point, predict the first preset number of first offsets corresponding to each pixel point.

[0072] In this embodiment, the first offset represents the vector pointing from the pixel point to the key point query position, and each key point regression feature corresponds to a first offset. Continuing with the Figure 1 For exemplary illustration, each sub-branch of the first network branch can predict a first offset.

[0073] Step 330: Based on the position of each pixel point and the first preset number of first offsets corresponding thereto, determine the first preset number of key point query positions corresponding to each pixel point.

[0074] As an example, if the position of the pixel point is (a, b) and the first offsets are (1, 2), (-2, 3), then the key point query positions can be (a + 1, b + 2), (a - 2, b + 3).

[0075] From Figure 3 It can be seen that Figure 3 The shown process reflects extracting multiple independent key point regression features from the feature data, then predicting a first offset based on each key point regression feature, and thus obtaining multiple key point query positions, which can avoid the correlation error between multiple key point query positions caused by feature sharing and help improve the accuracy.

[0076] Further referring to Figure 4 , Figure 4 shows a flowchart of predicting the position of key points in an embodiment of the pose recognition method of the present disclosure. As shown in Figure 4 , on the basis of the embodiment shown in Figure 3 , the above step 230 can include the following steps:

[0077] Step 410: Generate the first preset number of key-point semantic features corresponding to each pixel based on the first preset number of key-point regression features and the first preset number of key-point query positions corresponding to each pixel point.

[0078] Combined with Figure 1 For exemplary illustration, a sub-branch in the first network branch may extract the features corresponding to the key-point query positions from the key-point regression features, and then encode the extracted features into key-point semantic features through a convolutional layer or an encoder, so as to obtain the first preset number of key-point semantic features corresponding to each pixel point.

[0079] Step 420: Predict the first preset number of second offsets corresponding to each pixel point based on each key-point semantic feature.

[0080] In this embodiment, the second offset represents the vector pointing from the key-point query position to the key point.

[0081] Combined with Figure 1 For exemplary illustration, a sub-branch in the first network branch may map the key-point semantic features to the second offsets by using a convolutional layer or a fully connected layer, so as to obtain the first preset number of second offsets corresponding to each pixel point.

[0082] Step 430: Determine the positions of the first preset number of key points corresponding to each pixel point based on the first preset number of second offsets and the key-point query positions corresponding to each pixel point.

[0083] In related technologies, in some pose recognition methods, when calculating key points, the positions of key points are often determined by using the relative distances between key points.

[0084] In Figure 4 In the shown process, the key-point semantic features can be extracted from the key-point regression features through the key-point query positions, then the second offsets are predicted based on the key-point semantic features, and finally the positions of the key points are determined according to the key-point query positions and the second offsets. Compared with determining the positions of key points by using the relative positions between key points, by using the first offset and the second offset to represent the relative positions of pixel points and each key point, and thus determining the positions of each key point, the cumulative error can be avoided and the accuracy can be improved.

[0085] Further referring to Figure 5 , Figure 5 shows a flowchart of generating key-point semantic features in an embodiment of the pose recognition method of the present disclosure. As Figure 5 shown, in some optional implementation manners of the embodiment shown in Figure 4 the above step 410 may include the following steps:

[0086] Step 510: Extract the key-point query position features corresponding to the key-point regression features from each key-point regression feature.

[0087] Step 520: Based on each key-point query position feature, predict a second preset number of third offsets corresponding to each key-point query position.

[0088] Step 530: Based on each key-point query position and the second preset number of third offsets corresponding thereto, determine the positions of the second preset number of enhanced pixel points corresponding to each key-point query position.

[0089] Step 540: Based on the positions of the second preset number of enhanced pixel points corresponding to each key-point query position, extract the second preset number of enhanced pixel point features from the feature data.

[0090] Step 550: Fuse each key-point query position feature and the second preset number of enhanced pixel point features corresponding thereto to generate a key-point semantic feature corresponding to each key-point query position, and obtain the first preset number of key-point semantic features corresponding to each pixel point.

[0091] In a specific example, the execution subject may first extract the key-point query position features from the key-point regression features, and then predict the positions of N enhanced pixel points based on the key-point query position features, where N is a preset positive integer; then, extract N enhanced pixel point features from the feature data; thereafter, fuse the N enhanced pixel point features with the key-point query position features to obtain the key-point semantic feature corresponding to the key-point query position, and the key-point semantic feature may include the features of (N + 1) positions.

[0092] Figure 5 The illustrated embodiment reflects the steps of predicting the positions of enhanced pixel points based on the key-point query position features, and then fusing the enhanced pixel point features with the key-point query position features to generate the key-point semantic features, so that the key-point semantic features can include the features of a larger number of pixel points, which can increase the amount of information in the key-point semantic features and help improve the accuracy.

[0093] Next, refer to Figure 6 Figure 6 which shows a flowchart of predicting the pose quality score in an embodiment of the pose recognition method of the present disclosure, and the above step 230 may include the following steps:

[0094] Step 610: Perform feature extraction on the feature data to obtain the instance features of the image.

[0095] Continue to combine with Figure 1 ​For exemplary illustration, the execution entity may use the convolutional layer in the second network branch to extract features from the feature data to obtain the instance features of the image.

[0096] Step 620: Based on the query positions of the first preset number of key points corresponding to each pixel point, extract the key point instance features corresponding to each pixel point from the instance features.

[0097] Step 630: Based on the key point instance features corresponding to each pixel point, generate the pose features corresponding to each pixel point.

[0098] As an example, the execution entity may splice the first preset number of key point instance features corresponding to the same pixel point using the second network branch, and use the spliced features as the pose features corresponding to the pixel point.

[0099] Step 640: Based on the pose features corresponding to each pixel point, predict the pose quality score corresponding to each pixel point.

[0100] As an example, the execution entity may input the pose features into the convolutional layer or the fully connected layer in the second network branch, and the convolutional layer or the fully connected layer maps the pose features to the pose quality score.

[0101] In Figure 6 the illustrated embodiment, the pose features of the pixel points are generated based on the instance features at the key point query positions, so that the pose features have a correlation with the key points. Furthermore, the pose quality score obtained from the pose features can more accurately characterize the matching degree between the pixel points and their key points, and the pose of the object to be recognized.

[0102] Next, referring to Figure 7 , Figure 7 shows a flowchart of an embodiment of the method for training a pose recognition model according to the present disclosure. The process includes the following steps:

[0103] Step 710: Obtain the reference quality score of the labeled pixel points and the sample image of the sample pose of the object to be recognized.

[0104] Among them, the sample pose includes the position of the sample target pixel point representing the imaging point of the center point part of the object to be recognized, and the position of the sample key points representing the imaging points of the key parts of the object to be recognized. For example, the position of the sample key points can be represented by the offset of the sample target pixel point.

[0105] As an example, the reference quality score of the pixel points may adopt the data form of a grayscale image. Among them, the higher the reference quality score, the higher the pixel value in the grayscale image.

[0106] Step 720: Process the sample image using the pre-constructed initial pose recognition model to obtain the predicted quality score corresponding to each pixel point and the predicted pose of the object to be recognized.

[0107] The pose recognition model in this embodiment is used to implement the pose recognition method in any of the foregoing embodiments.

[0108] As an example, the initial pose recognition model can adopt Figure 1 the architecture shown in the figure. The initial backbone network extracts sample feature data from the sample image, and then the first initial network branch predicts the key point query position corresponding to each pixel point, and based on the key point query position and the sample feature data, predicts the sample key points corresponding to each pixel point; the second initial network branch predicts the predicted quality score corresponding to each pixel point according to the key point query position and the image features; finally, the maximum pooling kernel filters out the sample target pixel points from all pixel points based on the predicted quality score, and then determines the predicted pose of the object to be recognized based on the position of the sample target pixel points and the position of the corresponding sample key points.

[0109] Step 730: Determine the first loss function value based on the predicted pose and the sample pose.

[0110] In this embodiment, the first loss function is used to constrain the prediction process of the key points in the initial pose recognition model, so that the initial pose recognition model learns the prediction strategy of the key points.

[0111] In a specific example, the execution subject can determine the first loss function value according to the difference between each point in the predicted pose and each point in the sample pose.

[0112] Step 740: Determine the second loss function value based on the predicted quality score corresponding to each pixel point and the reference quality score.

[0113] In this embodiment, the second loss function is used to constrain the prediction process of the pose quality score in the initial pose recognition model, so that the initial pose recognition model learns the prediction strategy of the pose quality score.

[0114] As an example, the execution subject can first determine the difference between the predicted quality score of each pixel point and the reference quality score, and then determine the second loss function value based on the differences corresponding to all pixel points.

[0115] Step 750: Adjust the initial pose recognition model based on the first loss function value and the second loss function value until the training stop condition is met, and obtain the pose recognition model.

[0116] The execution entity can utilize the backpropagation characteristic of the deep network model to take the derivatives of the first loss function value and the second loss function value, and adjust the initial pose recognition model according to the derivative results. When the training stop condition is met (for example, the number of iterations reaches the preset number, or the first loss function value and the second loss function value converge simultaneously), it indicates that the accuracy of the current initial pose recognition model has reached the requirement. At this time, the training can be terminated, and the current initial pose recognition model is determined as the pose recognition model.

[0117] The method for training a pose recognition model provided in this embodiment constrains the pose prediction process and the quality score prediction process of the pose recognition model through the first loss function and the second loss function respectively. The pose recognition model trained in this way can more accurately predict the pose of the object to be recognized.

[0118] Next, refer to Figure 8 , Figure 8 which shows a flowchart for determining the second loss function value in an embodiment of the method for training a pose recognition model of the present disclosure. The process includes the following steps:

[0119] Step 810: Map the position of each pixel point in the sample image to the predicted sample score map to obtain the mapped point of each pixel point in the sample score map.

[0120] In a specific example, the execution entity can map the coordinates of the pixel point in the sample image to the sample score map according to the ratio of the size of the sample score map to the size of the sample image, so as to obtain the coordinates of the corresponding mapped point in the sample score map. For example, if the size of the sample score map is 1 / 4 of the size of the sample image, and the coordinates of the pixel point in the sample image are (m, n), then the coordinates of this pixel point in the sample score map are (m / 4, n / 4).

[0121] Step 820: Determine the pixel value of each mapped point as the reference quality score corresponding to each pixel point.

[0122] In this embodiment, the sample score map is a single-channel grayscale map constructed based on the reference quality score of each pixel point. In the sample score map, the pixel value (i.e., the grayscale value) of each pixel point is the reference quality score of this pixel point.

[0123] In some optional implementation manners of this embodiment, the sample score map can be generated through the process shown in Figure 9 As shown in Figure 9 which shows, the process includes the following steps:

[0124] Step 910: Determine the sample candidate key points corresponding to each pixel point within the preset area.

[0125] As an example, the preset region can be the overall region of the object to be recognized in the sample image, or can be a central region determined with the center of the object to be recognized as the center and a preset radius.

[0126] Step 920: Determine the reference quality score of each pixel point in the region of the object to be recognized based on the similarity between the sample candidate key points and the sample key points.

[0127] Step 930: Map the pixel points in the sample image to a single-channel image, use the reference pose quality score as the pixel value of the pixel points in the preset region in the single-channel image, and determine the pixel value of the pixel points outside the preset region in the single-channel image to be 0, to obtain a sample score map.

[0128] In this implementation, the higher the reference quality score of the pixel points in the sample image, the higher the pixel value of the pixel points in the sample score map.

[0129] In a specific example, pixel point A is located within the preset region, and the similarity between the candidate key point of pixel point A and the labeled sample key point is 0.8, then the pixel value of the mapped point of pixel point A in the sample score map is 0.8; pixel point B is located outside the preset region, then the pixel value of the mapped point of pixel point B in the sample score map is 0.

[0130] Compared with using discrete {0, 1} to represent the quality score of pixel points in the Gaussian heat map in the related art, Figure 9 in the shown implementation, by determining the reference quality score through the similarity between the sample candidate key points and the sample key points, and using the reference quality score as the pixel value in the sample score map, a continuous numerical interval can be used to represent the reference quality score of the pixel points in the sample image, and the predicted pose quality score of different pixel points can be more accurately represented.

[0131] Continue to refer to Figure 8 Step 830: Determine the second loss function value based on the predicted quality score and the reference quality score corresponding to each pixel point.

[0132] From Figure 8 it can be seen that Figure 8 the shown embodiment reflects determining the reference quality score of pixel points according to the sample score map, and further determining the second loss function value. Since the sample score map can use a continuous numerical interval to represent the reference quality score of pixel points, it can more accurately represent the quality of the description of the pose of the object to be recognized by pixel points, thereby improving the training effect of the pose recognition model.

[0133] Exemplary Apparatus

[0134] Then refer to Figure 10 Figure 10The structural schematic diagram of an embodiment of the posture recognition device of the present disclosure is shown. As Figure 10 shown, the device includes: a feature extraction unit 1010 configured to extract feature data of an image including an object to be recognized; a first prediction unit 1020 configured to predict a key point query position corresponding to each pixel point from the feature data based on a preset prediction method, where the pixel point represents an imaging point of the candidate center point part of the object to be recognized, and the key point corresponding to the pixel point represents an imaging point of the candidate key part of the object to be recognized; a second prediction unit 1030 configured to predict a posture quality score corresponding to each pixel point and the position of the key point based on the feature data and the key point query position corresponding to each pixel point; a target determination unit 1040 configured to determine a pixel point with a posture quality score meeting a preset condition as a target pixel point, and determine the position of the key point corresponding to the target pixel point as the position of the target key point; a posture determination unit 1050 configured to determine the target posture of the object to be recognized based on the position of the target pixel point and the position of the target key point.

[0135] In one embodiment, the first prediction unit 1020 further includes: a first extraction module configured to perform a first preset number of feature extractions on the feature data respectively based on the position of each pixel point to obtain a first preset number of key point regression features corresponding to each pixel point; a first prediction module configured to predict a first preset number of first offsets corresponding to each pixel point based on the first preset number of key point regression features corresponding to each pixel point; a first determination module configured to determine a first preset number of key point query positions corresponding to each pixel point based on the position of each pixel point and its corresponding first preset number of first offsets.

[0136] In one embodiment, the second prediction unit 1030 includes: a semantic feature module configured to generate a first preset number of key point semantic features corresponding to each pixel point based on the first preset number of key point regression features and the key point query position corresponding to each pixel point; a second prediction module configured to predict a first preset number of second offsets corresponding to each pixel point based on each key point semantic feature; a position determination module configured to determine the position of a first preset number of key points corresponding to each pixel point based on the first preset number of second offsets and the key point query position corresponding to each pixel point.

[0137] In one embodiment, the semantic feature module further includes: an associated feature sub-module configured to extract the key point query position feature corresponding to the key point regression feature from each key point regression feature; a third prediction sub-module configured to predict a second preset number of third offsets corresponding to each key point query position based on each key point query position feature; a position determination sub-module configured to determine the positions of a second preset number of enhanced pixel points corresponding to each key point query position based on each key point query position and its corresponding second preset number of third offsets; an enhanced feature sub-module configured to extract a second preset number of enhanced pixel point features from the feature data based on the positions of the second preset number of enhanced pixel points corresponding to each key point query position; and a feature generation sub-module configured to fuse each key point query position feature and its corresponding second preset number of enhanced pixel point features to generate the key point semantic feature corresponding to each key point query position, obtaining the first preset number of key point semantic features corresponding to each pixel point.

[0138] In one embodiment, the second prediction unit 1030 further includes: an instance feature module configured to perform feature extraction on the feature data to obtain the instance feature of the image; a second extraction module configured to extract the key point instance feature corresponding to each pixel point from the instance feature based on the first preset number of key point query positions corresponding to each pixel point; a feature generation module configured to generate the pose feature corresponding to each pixel point based on the key point instance feature corresponding to each pixel point; and a score prediction module configured to predict the pose quality score corresponding to each pixel point based on the pose feature corresponding to each pixel point.

[0139] Next, referring to Figure 11 , Figure 11 shows a schematic structural diagram of an embodiment of the apparatus for training a pose recognition model of the present disclosure, as Figure 11As shown, the device includes: a sample acquisition unit 1110 configured to acquire a sample image marked with a reference quality score of pixel points and a sample pose of an object to be recognized, where the sample pose includes the position of a sample target pixel point representing the imaging point of the center point part of the object to be recognized and the position of a sample key point representing the imaging point of the key part of the object to be recognized; a model prediction unit 1120 configured to process the sample image using a pre-constructed initial pose recognition model to obtain a predicted quality score corresponding to each pixel point and a predicted pose of the object to be recognized; a first loss unit 1130 configured to determine a first loss function based on the predicted pose and the sample pose; a second loss unit 1140 configured to determine a second loss function based on the predicted quality score corresponding to each pixel point and a preset reference quality score; and a model training unit 1150 configured to adjust the initial pose recognition model based on the first loss value and the second loss value until a training stop condition is met to obtain a pose recognition model.

[0140] In one embodiment, the second loss unit 1140 further includes: a mapping module configured to map the position of each pixel point in the sample image to a predicted sample score map to obtain a mapped point of each pixel point in the sample score map; a score determination module configured to determine the pixel value of each mapped point as the reference quality score corresponding to each pixel point; and a loss determination module configured to determine a second loss function based on the predicted quality score and the reference quality score corresponding to each pixel point.

[0141] The device further includes a score map construction unit, including: a first determination module configured to determine a sample candidate key point corresponding to each pixel point within the area of the object to be recognized; a reference score module configured to determine a reference quality score for each pixel point within the area of the object to be recognized based on the similarity between the sample candidate key point and the sample key point; and an image construction module configured to map the pixel points in the sample image to a single-channel image, use the reference pose quality score as the pixel value of the pixel points within the area of the sample object in the single-channel image, and determine the pixel value of the pixel points outside the area of the sample object in the single-channel image as 0 to obtain a sample score map.

[0142] Exemplary Electronic Device

[0143] Next, refer to Figure 12 to describe an electronic device according to an embodiment of the present disclosure. Figure 12 FIG. illustrates a block diagram of an electronic device according to an embodiment of the present disclosure.

[0144] As Figure 12 shown, the electronic device 1200 includes one or more processors 1210 and a memory 1220.

[0145] The processor 1210 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 1200 to perform desired functions.

[0146] The memory 1220 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 1210 can run the program instructions to implement the gesture recognition method and / or the method of training a gesture recognition model in the various embodiments of the present disclosure described above and / or other desired functions. Various contents such as input signals, signal components, noise components, etc. can also be stored in the computer-readable storage media.

[0147] In one example, the electronic device 1200 can further include: an input device 1230 and an output device 1240, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0148] For example, the input device 1230 can be the above-mentioned microphone or microphone array for capturing the input signal of the sound source. The input device 1230 can be a communication network connector for receiving the collected input signal.

[0149] In addition, the input device 1230 can further include, for example, a keyboard, a mouse, and so on.

[0150] The output device 1240 can output various information to the outside, including the determined distance information, direction information, etc. The output device 1240 can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and so on.

[0151] Of course, for simplicity, Figure 12 only some of the components related to the present disclosure in the electronic device 1200 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 1200 can further include any other appropriate components.

[0152] Exemplary Computer Program Product and Computer Readable Storage Medium

[0153] In addition to the above methods and devices, embodiments of the present disclosure may also be computer program products, which include computer program instructions that, when run on a processor, cause the processor to execute the steps in the gesture recognition methods and / or the methods for training a gesture recognition model according to various embodiments of the present disclosure described in the "Exemplary Methods" section above of this specification.

[0154] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0155] Furthermore, embodiments of the present disclosure may also be computer-readable storage media, on which computer program instructions are stored, and when the computer program instructions are run on a processor, the processor is caused to execute the steps in the gesture recognition methods and / or the methods for training a gesture recognition model according to various embodiments of the present disclosure described in the "Exemplary Methods" section above of this specification.

[0156] The computer-readable storage media may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0157] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-disclosed specific details are only for illustrative purposes and for ease of understanding, and are not limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0158] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For system embodiments, since they basically correspond to method embodiments, they are described relatively simply. For relevant parts, reference can be made to the partial description of method embodiments.

[0159] The block diagrams of devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.

[0160] The methods and apparatuses of this disclosure can be implemented in many ways. For example, the methods and apparatuses of this disclosure can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the method is only for illustration, and the steps of the method of this disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, this disclosure can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the method according to this disclosure. Therefore, this disclosure also covers a recording medium storing a program for executing the method according to this disclosure.

[0161] It should also be noted that in the apparatuses, equipment, and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.

[0162] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0163] The foregoing description has been presented for purposes of illustration and description. In addition, this description is not intended to limit embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A gesture recognition method, comprising: Extract the feature data of the image including the object to be recognized; Predict the key point query position corresponding to each pixel point from the feature data based on a preset prediction method, where the pixel point represents the imaging point of the candidate center point part of the object to be recognized, and the key point represents the imaging point of the candidate key part of the object to be recognized; Predict the pose quality score corresponding to each pixel point and the position of the key point based on the feature data and the key point query position corresponding to each pixel point; Determine the pixel points with pose quality scores meeting the preset conditions as target pixel points, and determine the positions of the key points corresponding to the target pixel points as the positions of the target key points; Determine the target pose of the object to be recognized based on the positions of the target pixel points and the positions of the target key points.

2. The method according to claim 1, wherein, The number of key points is a first preset number. The predicting the key point query position corresponding to each pixel point from the feature data based on a preset prediction method includes: Perform feature extraction of the first preset number on the feature data respectively based on the position of each pixel point to obtain the first preset number of key point regression features corresponding to each pixel point; Predict the first preset number of first offsets corresponding to each pixel point based on the first preset number of the key point regression features corresponding to each pixel point; Determine the first preset number of key point query positions corresponding to each pixel point based on the position of each pixel point and the first preset number of the first offsets corresponding thereto.

3. The method according to claim 2, wherein, The predicting the position of the key point corresponding to each pixel point based on the feature data and the key point query position corresponding to each pixel point includes: Generate the first preset number of key point semantic features corresponding to each pixel point based on the first preset number of the key point regression features and the key point query position corresponding to each pixel point; Predict the first preset number of second offsets corresponding to each pixel point based on the first preset number of the key point semantic features corresponding to each pixel point; Determine the first preset number of key point positions corresponding to each pixel point based on the first preset number of the second offsets and the key point query position corresponding to each pixel point.

4. The method according to claim 3, wherein, The generating the first preset number of key point semantic features corresponding to each pixel point based on the first preset number of the key point regression features and the key point query position corresponding to each pixel point includes: Extract each key point query position feature from each key point regression feature; Predict the second preset number of third offsets corresponding to each key point query position based on each key point query position feature; Determine the positions of the second preset number of enhanced pixel points corresponding to each key point query position based on each key point query position and the second preset number of the third offsets corresponding thereto; Extract the second preset number of enhanced pixel point features from the feature data based on the positions of the second preset number of the enhanced pixel points corresponding to each key point query position; Fuse the position feature of each of the key points and the corresponding second preset number of the enhanced pixel point features to generate the key point semantic feature corresponding to each key point query position, and obtain the first preset number of key point semantic features corresponding to each pixel point.

5. The method according to any one of claims 2 to 4, wherein, Predict the pose quality score corresponding to each pixel point based on the feature data and the key point query position corresponding to each pixel point, including: Extract the feature data of the image by performing feature extraction on the feature data to obtain the instance feature of the image; Based on the first preset number of key point query positions corresponding to each pixel point, extract the first preset number of key point instance features corresponding to each pixel point from the instance feature; Generate the pose feature corresponding to each pixel point based on the first preset number of the key point instance features corresponding to each pixel point; Predict the pose quality score corresponding to each pixel point based on the pose feature corresponding to each pixel point.

6. A method for training a gesture recognition model, comprising: Obtain a sample image marked with the reference quality score of the pixel point and the sample pose of the object to be recognized, where the sample pose includes the position of the sample target pixel point representing the imaging point of the center point part of the object to be recognized, and the position of the sample key point representing the imaging point of the key part of the object to be recognized; Process the sample image using a pre-constructed initial pose recognition model to obtain the predicted quality score corresponding to each pixel point and the predicted pose of the object to be recognized; Determine the first loss function value based on the predicted pose and the sample pose; Determine the second loss function value based on the predicted quality score and the reference quality score corresponding to each pixel point; Adjust the initial pose recognition model based on the first loss function value and the second loss function value until the training stop condition is satisfied to obtain the pose recognition model.

7. Based on the predicted quality score and the reference quality score corresponding to each pixel point, determining a second loss function, comprising: Map the position of each pixel point in the sample image to the predicted sample score map to obtain the mapped point of each pixel point in the sample score map; Determine the pixel value of each mapped point as the reference quality score corresponding to each pixel point; Determine the second loss function value based on the predicted quality score and the reference quality score corresponding to each pixel point; Wherein, the sample score map is obtained through the following steps: Determine the sample candidate key points corresponding to each pixel point within the preset region; Determine the reference quality score of each pixel point within the region of the object to be recognized based on the similarity between the sample candidate key points and the sample key points; Map the pixel points in the sample image to a single-channel image, use the reference pose quality score as the pixel value of the pixel points within the preset region in the single-channel image, and determine the pixel value of the pixel points outside the preset region in the single-channel image as 0 to obtain the sample score map.

8. An attitude recognition device, comprising: A feature extraction unit configured to extract the feature data of an image including the object to be recognized; A first prediction unit, configured to predict a key point query position corresponding to each pixel point from the feature data based on a preset prediction method, where the pixel point represents an imaging point of the center point part of the candidate object to be recognized, and the key point represents an imaging point of the key part of the candidate object to be recognized; A second prediction unit, configured to predict a pose quality score corresponding to each pixel point and the position of the key point based on the feature data and the key point query position corresponding to each pixel point; A target determination unit, configured to determine the pixel points whose pose quality scores meet the preset conditions as target pixel points, and determine the position of the key point corresponding to the target pixel points as the position of the target key point; A pose determination unit, configured to determine the target pose of the object to be recognized based on the positions of the target pixel points and the positions of the target key points; 9. A device for training an attitude recognition model, comprising: A sample acquisition unit, configured to acquire a sample image marked with a reference quality score of pixel points and a sample pose of the object to be recognized, where the sample pose includes the position of a sample target pixel point representing the imaging point of the center point part of the object to be recognized and the position of a sample key point representing the imaging point of the key part of the object to be recognized; A model prediction unit, configured to process the sample image by using a pre-constructed initial pose recognition model to obtain a predicted quality score corresponding to each pixel point and the predicted pose of the object to be recognized; A first loss unit, configured to determine a first loss function value based on the predicted pose and the sample pose; A second loss unit, configured to determine a second loss function value based on the predicted quality score corresponding to each pixel point and a preset reference quality score; A model training unit, configured to adjust the initial pose recognition model based on the first loss value and the second loss value until a training stop condition is met, and obtain a pose recognition model.

10. A computer-readable storage medium storing a computer program for performing the method according to any one of claims 1-7 above.

11. An electronic device, the electronic device comprising: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of claims 1-7 above.

Citation Information

Patent Citations

  • Human body posture recognition method, device and equipment and storage medium

    CN110781765A

  • Face quality evaluation method and device

    CN110837750A