Training methods, apparatus, electronic devices, and computer programs for posture estimation models

By integrating distribution functions to adjust model parameters, the method improves the accuracy and efficiency of pose estimation models, addressing the challenge of capturing image intrinsic information and reducing resource consumption.

JP2026517354APending Publication Date: 2026-05-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2024-10-08
Publication Date
2026-05-29

Smart Images

  • Figure 2026517354000001_ABST
    Figure 2026517354000001_ABST
Patent Text Reader

Abstract

This application relates to a data processing technology, specifically a training method, apparatus, electronic device, and storage medium for a pose estimation model, which can be used in various scenarios such as cloud technology, artificial intelligence, smart transportation, and driver assistance. A training sample is obtained, which includes a sample image and the first coordinate of a preset keypoint within the sample image. The initial pose estimation model performs pose estimation of the object within the sample image and obtains the predicted second coordinate of the preset keypoint within the sample image and L sets of predicted parameter values. The L sets of predicted parameter values ​​are determined for L pre-set distribution functions. Based on the predicted second coordinate and the L sets of predicted parameter values, the L distribution functions are integrated to obtain a predicted probability distribution. The predicted probability distribution represents the distribution of the probability that each pixel point in the predicted sample image is a preset keypoint. The distribution loss is determined based on the difference between the predicted probability distribution and the target probability distribution. The target probability distribution is the probability distribution determined based on the first coordinate. Based on the distribution loss, the model parameters of the initial pose estimation model are adjusted to obtain a target pose estimation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims priority to a Chinese patent application filed on October 23, 2023, with application number 202311370780.4 and invention title "Training Method, Apparatus, Electronic Device, and Storage Medium for Pose Estimation Model".

[0002] This application relates to the field of data processing technology, and particularly to a training method, apparatus, electronic device, and storage medium for a pose estimation model.

Background Art

[0003] With the development of artificial intelligence technology, it is possible to use a pose estimation model to estimate the pose of an object in a recognized image, obtain the coordinate information of each key point describing the pose of the object, and further estimate the pose of the object using the obtained coordinate information.

[0004] In related technologies, when training a pose estimation model, it is difficult to capture the inherent information of the input sample image, and the training effect of the pose estimation model cannot be ensured. Furthermore, when estimating the pose of an object in an image using a trained pose estimation model, the coordinate information of the key points of the object cannot be accurately obtained, resulting in a decrease in the pose estimation effect.

Summary of the Invention

[0005] Embodiments of this application provide a training method, apparatus, electronic device, and storage medium for a pose estimation model to improve the pose estimation accuracy of the pose estimation model.

[0006] According to an embodiment of this application, a training method for a pose estimation model is provided. This method includes a step of obtaining a training sample, where the training sample includes a sample image and a first coordinate of a preset key point in the sample image, and the preset key point is used for positioning the pose of an object in the sample image, The steps include: performing pose estimation of an object in a sample image using an initial pose estimation model, obtaining the predicted second coordinates of preset keypoints in the sample image and L sets of predicted parameter values, wherein the L sets of predicted parameter values ​​are determined for L pre-set distribution functions; A step of obtaining a predicted probability distribution by integrating L distribution functions based on a predicted second coordinate and L sets of predicted parameter values, wherein the predicted probability distribution represents the distribution of the probability that each pixel point in the predicted sample image is a preset keypoint. A step of determining the distribution loss based on the difference between the predicted probability distribution and the target probability distribution, wherein the target probability distribution is a probability distribution determined based on the first coordinate, The process includes the steps of adjusting the model parameters of the initial pose estimation model based on the distribution loss to obtain a target pose estimation model.

[0007] According to the embodiment of the present invention, a training device for a posture estimation model is provided. This device is An acquisition unit for acquiring training samples, wherein the training sample includes a sample image and the first coordinate of a preset keypoint within the sample image, and the preset keypoint is used to perform pose positioning of an object within the sample image, and the acquisition unit... The initial pose estimation model is used to estimate the pose of an object in a sample image, and the predicted second coordinates of preset keypoints in the sample image and L sets of predicted parameter values ​​are obtained, wherein the L sets of predicted parameter values ​​are determined for L predefined distribution functions. Based on the predicted second coordinate and L sets of predicted parameter values, L distribution functions are integrated to obtain a predicted probability distribution, and the distribution loss is determined based on the difference between the predicted probability distribution and the target probability distribution, wherein the predicted probability distribution represents the distribution of the probability that each pixel point in the predicted sample image is a preset keypoint, and the target probability distribution is the probability distribution determined based on the first coordinate. The system includes a training unit used to adjust the model parameters of an initial pose estimation model based on distributional loss and obtain a target pose estimation model.

[0008] According to an embodiment of the present invention, an electronic device is provided comprising memory, a processor, and a computer program stored in memory and executable by the processor. The processor executes the computer program to realize the above method.

[0009] A computer-readable storage medium is provided on which a computer program is stored. The computer program, when executed by a processor, accomplishes the above method.

[0010] We provide computer program products, including computer programs. When a computer program is executed by a processor, it achieves the above-described method. [Brief explanation of the drawing]

[0011] [Figure 1] This is a schematic diagram of possible application scenarios in the embodiments of the present invention. [Figure 2A] This is a schematic diagram of the training process for the posture estimation model in the embodiment of the present invention. [Figure 2B] This is another schematic diagram of the training process for the posture estimation model in the embodiment of the present invention. [Figure 3] This is a schematic diagram of the output results of the initial posture estimation model in the embodiment of the present invention. [Figure 4A] This is a schematic diagram of the process for determining the predictive probability distribution corresponding to the preset keypoints in the embodiment of the present invention. [Figure 4B] This is a schematic diagram showing the correspondence between the predicted probability distribution and the sample image in the embodiment of the present invention. [Figure 4C] This is a schematic diagram of the dynamic adjustment of the target probability distribution in the embodiment of the present invention. [Figure 4D] This is a schematic diagram of the process for calculating the model loss for preset keypoints in an embodiment of the present invention. [Figure 5A] It is a schematic diagram of a process for realizing business processing using the target posture estimation model in the embodiment of the present application. [Figure 5B] It is a schematic diagram of a process for sorting and acquiring the processed image in the embodiment of the present application. [Figure 5C] It is a schematic diagram of the posture estimation process in the embodiment of the present application. [Figure 6A] It is a schematic diagram of the palmprint recognition process in the embodiment of the present application. [Figure 6B] It is a schematic diagram of the processing logic when realizing action recognition using the target posture estimation model in the embodiment of the present application. [Figure 7] It is a schematic diagram of the logical structure of the training device for the posture estimation model according to the embodiment of the present application. [Figure 8] It is a schematic diagram of the hardware structure of the electronic device to which the embodiment of the present application is applied.

Embodiments for Carrying out the Invention

[0012] In order to make the objectives, technical means, and advantages of the embodiments of the present application clearer, hereinafter, while referring to the drawings in the embodiments of the present application, the technical means of the present application will be clearly and completely described. As is clear, the described embodiments are not all embodiments, but some embodiments of the technical means of the present application. All other embodiments that can be obtained by those skilled in the art without creative labor based on the embodiments described in this document of the present application all belong to the protection scope of the technical means of the present application.

[0013] Terms such as "first", "second", etc. in the specification, claims, and above-mentioned drawings of the present application are used to distinguish similar objects and are not necessarily used to explain a specific order or priority. It should be understood that data used in this way may be exchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein.

[0014] Hereinafter, to facilitate the understanding of those skilled in the art, some of the terms used in the embodiments of the present application will be described.

[0015] Human body detection: By using target detection technology to identify the area where the human body exists in an image, the human body area image is extracted from the image. The image can be replaced with a photograph. Also, the image may be an image frame of a video. The video may be pre-collected or collected in real time.

[0016] Hand detection: By using target detection technology to identify the area where the hand exists in an image, the hand area image is extracted from the image.

[0017] Key point: It may include connection points and / or joint points of the human body skeleton such as the head, shoulders, elbows, wrists, hip joints, knees, and ankles, or specific points of human body parts such as points on the face and hands.

[0018] Human body pose estimation: It is to estimate the coordinates of the key points of the human body skeleton in various poses within an image. Human body pose estimation usually includes full-body pose estimation and partial pose estimation of the human body. Human body pose estimation aims to predict the position information of predefined key points (also called preset key points) of the human body within the image. It is a basic task in computer vision and is widely used in various vision tasks. It is an important preprocessing operation for many downstream tasks (such as human motion analysis, activity recognition, motion capture, etc.). Human body pose estimation can also be defined as searching for a specific pose in a space composed of all joint poses.

[0019] Hand pose estimation: It is to estimate the coordinates of the key points of the hand skeleton in various poses within an image. Hand pose estimation aims to predict the position information of predefined key points (also called preset key points) of the hand within the image. It is a basic task in computer vision and is widely used in various vision tasks. It is an important preprocessing operation for many downstream tasks (such as gesture recognition, hand motion analysis, motion capture, etc.).

[0020] Regression-based pose estimation: In embodiments of the present invention, the coordinates of keypoints may be output by regression using a pose estimation model on an input image. The pose estimation model may be an initial pose estimation model used for model training, or it may be a target pose estimation model finally obtained through model training. Regression-based pose estimation estimates the posture of a human body by, for example, representing the positions of joint points in two-dimensional coordinates and learning the mapping from body features to the positions of body parts, thereby training the network so that the network output directly indicates the coordinates of each joint point.

[0021] Linear layer: In a neural network, this layer performs a linear transformation on the input. In deep learning, neural networks typically consist of multiple layers, each responsible for a specific task. These layers include pooling layers, linear layers, and activation function layers. A linear layer, for example, multiplies the input data by a weight matrix and adds a bias term to generate the output.

[0022] A probability distribution is a distribution whose sum is 1, where the value of each point (for example, a pixel position in an image) represents the probability of that point being represented. In other words, a probability distribution represents the probability that a random variable will take a specific value.

[0023] A Gaussian mixture model is a probability distribution model composed of a linear combination of multiple Gaussian distribution functions.

[0024] Monte Carlo estimation is a method of performing approximate numerical calculations by randomly sampling from a probability model.

[0025] Pearson correlation coefficient: Used to indicate the degree of correlation between two variables, its value ranges from -1 to 1.

[0026] Heatmap-based pose estimation: For an input image, the model outputs a corresponding heatmap and generates keypoint coordinates.

[0027] The argmax function is used to obtain the array index corresponding to the element with the largest value in the input array.

[0028] When selecting a method for implementing pose estimation, the applicant considered that using a heatmap-based pose estimation technique would require generating a high-resolution likelihood heatmap based on a feature map. In the heatmap, locations where the model considers keypoints to be most likely to appear are marked with high probability, and the remaining locations are marked with low probability. Based on the heatmap, the argmax function is used to obtain the keypoint coordinates predicted by the model. However, in the heatmap-based pose estimation method, the prediction head generates a high-resolution likelihood heatmap based on the input feature map, and since the number of heatmaps is the number of predicted keypoints, a heatmap is generated for each keypoint. This method undoubtedly consumes a lot of memory and is computationally expensive, making it difficult to use in real-time scenarios where speed is required or in IoT devices with limited computational resources. In addition, because the size of the heatmap is limited, the keypoint coordinates obtained using the argmax function often contain quantization errors, which also affects the final performance of the model.

[0029] Based on this, the applicant considered whether it might be possible to reduce the consumption of memory and computational resources in pose estimation using conventional regression-based pose estimation techniques.

[0030] Therefore, assuming processing using conventional regression-based pose estimation techniques, global mean pooling is used during processing to simplify features, the prediction head contains only multiple linear layers, and the keypoint coordinates predicted using the regression method are output directly.

[0031] However, in conventional regression-based pose estimation techniques, the directly regressed coordinate values ​​(vectors) do not have the same spatial dimension as the input image. That is, the output coordinate values ​​correspond to specific points, while the input is an image, so the two do not have the same spatial dimension. Therefore, during model training, the constraints on the coordinate values ​​become implicit and inconsistent constraints, preventing the model from accurately capturing the intrinsic information of the input image, resulting in poor training results.

[0032] Therefore, this invention proposes a training method, apparatus, electronic device, and storage medium for a posture estimation model. In the process of training an initial posture estimation model based on regression, the output results of the initial posture estimation model are adjusted so that, after regression processing, predicted coordinates corresponding to each preset keypoint are output, and in addition, L sets of predicted parameter values ​​corresponding to each predicted keypoint are output. This provides a processing base for the parameterization and integration process of L distribution functions performed for each preset keypoint.

[0033] Furthermore, in the process of establishing constraints for model training, by constructing a predictive probability distribution corresponding to each preset keypoint, the predicted coordinates of the preset keypoints can be transformed into a probability distribution on the corresponding sample image by aggregating L distribution functions performed for each preset keypoint. This makes it possible for the input image fed into the initial pose estimation model to have the same dimensions as the predictive probability distribution that formed the basis for establishing the constraints of the initial pose estimation model. Consequently, the trained target pose estimation model captures the intrinsic information of the image more accurately, improving the representational power of the target pose estimation model, leading to more accurate positioning of keypoints of the object in the image and more accurate estimation of the object's pose, as well as improving the processing performance of the target pose estimation model.

[0034] Furthermore, by combining the network characteristic that the regression-based network structure itself is lightweight, the process of training an initial pose estimation model to obtain a target pose estimation model can, on the one hand, ensure the pose estimation performance of the model and improve the accuracy of pose estimation, while on the other hand, reduce the consumption of memory and computational resources, alleviate the time burden, and improve resource utilization.

[0035] The embodiments of this application will be described below with reference to the drawings. The embodiments described herein are for illustrative and interpretive purposes only and are not intended to limit this application. Furthermore, all embodiments of this application (including those described in the claims and specification) and their features can be combined with each other, insofar as they do not contradict each other.

[0036] Figure 1 is a schematic diagram of a possible application scenario in an embodiment of the present invention. This schematic diagram of the application scenario includes a server device 110 and a client device 120.

[0037] In some feasible embodiments of the present invention, the server device 110 can train a target pose estimation model. Furthermore, the server device 110 can independently perform a pose estimation task in a specific pose estimation scenario. Alternatively, the server device 110 can transmit a trained target pose estimation model to a client device 120, and the client device 120 can perform a pose estimation task in a specific pose estimation scenario.

[0038] Alternatively, in several other feasible embodiments, the client device 120 can train a target pose recognition model to perform a pose estimation task in a specific pose estimation scenario.

[0039] The server equipment 110 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0040] The client device 120 includes, but is not limited to, mobile phones, tablet computers, laptop computers, e-readers, smart voice interaction devices, smart home appliances, in-car terminals, and aircraft. The embodiments of this application can be used in a variety of scenarios and include, but are not limited to, cloud technology, artificial intelligence, smart transportation, and driver assistance.

[0041] In feasible embodiments of the present invention, the related object can initiate a pose estimation request for an image to be processed using a target application on the client device 120, thereby enabling the processing device that performs pose estimation on the image to be processed and obtains a pose estimation result. The target application may be an applet application, a client application, or a web application, and the processing device may specifically be a server device 110 or a client device 120, but the present invention is not particularly limited thereto.

[0042] In the embodiments of this application, the server device 110 and the client device 120 can communicate via a wired or wireless network. The following description outlines the relevant processing processes, using as an example the processing device performing training of a target pose estimation model and processing of pose estimation tasks. Here, depending on the actual processing needs, the processing device may specifically be the server device 110 or the client device 120.

[0043] The following describes several possible application scenarios related to pose estimation.

[0044] Scenario 1: Identify the area to be recognized in the identity recognition process. In the application scenario of Scenario 1, after determining the recognition information to be used for identity recognition, it is possible to determine each preset keypoint that needs to be estimated in the image for pose estimation based on the recognition information used. For example, when using palm prints for identity recognition, each preset keypoint determined can identify at least the area of ​​the hand in the image. As another example, when using irises for identity recognition, each preset keypoint determined can identify at least the area of ​​the eye in the image. As yet another example, when using gestures for identity recognition, each preset keypoint determined can identify at least a different gesture. After training a target pose estimation model, the processing device can use the target pose estimation model to output coordinate information for each predicted keypoint based on the image being recognized. Furthermore, it can use this coordinate information to determine the recognized region for identity recognition and extract the recognized region from the image being recognized.

[0045] Scenario 2: Perform motion recognition in the anomaly detection process. In the application scenario of Scenario 2, the target object for motion recognition can be determined. The target object for motion recognition may be a living human or animal, or an inanimate object that exhibits various movements through mechanical motion. Furthermore, preset keypoints for pose positioning are determined for the target object, training samples are created for each target, and a target pose estimation model is trained using each training sample. Next, target detection technology is used to detect the region where the object to be recognized exists from the captured original image, and the recognized image containing the region where the object to be recognized exists is cut out from the original image. Then, using the target pose estimation model, pose estimation is performed on the object in the recognized image, predicted coordinates corresponding to each preset keypoint are determined, and based on each predicted coordinate, recognition of abnormal movements (such as falls) is achieved.

[0046] Scenario 3: Perform motion recognition during the motion instruction process. In the application scenario corresponding to Scenario 3, the target object for motion recognition can be determined. The target object for motion recognition may be a "human". Furthermore, preset keypoints for pose positioning are determined for the target object for motion recognition, training samples are created for each target, and the target pose estimation model is trained using each training sample. Next, using target detection technology, the region containing the object to be recognized is detected from the captured original image, and the recognized image containing the region containing the object is extracted from the original image. Then, using a target pose estimation model, the pose of the object in the recognized image is estimated, predictive coordinates corresponding to each preset keypoint are determined, and tasks such as dance motion recognition and dance step pattern recognition are performed based on these predictive coordinates.

[0047] Furthermore, specific embodiments of this application include the acquisition and processing of sample images and images to be processed. If the embodiments described herein are used in a specific product or technology, permission or consent from the relevant parties is required, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0048] The training process for the posture estimation model will be explained below from the perspective of the processing equipment, with reference to the diagrams.

[0049] Figure 2A is a schematic diagram of the training process for the posture estimation model in an embodiment of the present invention. As shown in Figure 2A, this process includes the following steps 101 to 105.

[0050] In step 101, the processing device can acquire a training sample. The training sample includes a sample image and the first coordinates of a preset keypoint within the sample image. The preset keypoint is used to position the object within the sample image. The "first coordinate" is also called the "sample coordinate" because it is a coordinate within the "sample image". In all embodiments of this application, the description of the "sample coordinate" is also a description of the "first coordinate". The first coordinate / sample coordinate may be entered manually.

[0051] In step 102, the processing device performs pose estimation of the object in the sample image using the initial pose estimation model, and obtains the predicted second coordinates of the preset keypoints in the sample image and L sets of predicted parameter values, the L sets of predicted parameter values ​​being determined for L pre-set distribution functions. The "predicted second coordinates" are the coordinates of the preset keypoints predicted by the initial pose estimation model and are therefore also called "predicted coordinates". In all embodiments of this application, the description of "predicted coordinates" is also a description of "predicted second coordinates".

[0052] In step 103, the processing unit integrates L distribution functions based on the predicted second coordinates and L sets of predicted parameter values ​​to obtain a predicted probability distribution, which represents the distribution of the probability that each pixel point in the predicted sample image is a preset keypoint.

[0053] In step 104, the processing device determines the distribution loss based on the difference between the predicted probability distribution and the target probability distribution, where the target probability distribution is the probability distribution determined based on the first coordinate.

[0054] In step 105, the processing unit adjusts the model parameters of the initial pose estimation model based on the distributed loss to obtain the target pose estimation model.

[0055] According to embodiments of the present invention, the integration in step 103 may include the step of obtaining a predicted probability distribution by performing parameter assignment and weighted summation on L distribution functions based on the predicted second coordinates and L sets of predicted parameter values. More specifically, the integration may include the steps of obtaining L distribution function results by assigning parameters to each corresponding distribution function among L distribution functions based on the predicted second coordinates and the predicted function parameter values ​​for each of the L sets of predicted parameter values, and performing weighted summation on the L distribution function results using the component weights of the L sets of predicted parameter values ​​as components.

[0056] If the pre-set L distribution functions are L Gaussian distributions, the integration may include the steps of: determining the mean matrix for each of the L Gaussian distributions based on the predicted second coordinate, and determining the covariance matrix and component weights of the Gaussian distributions based on a set of predicted parameter values ​​corresponding to the Gaussian distributions to obtain the Gaussian distribution results after parameter assignment; and performing a Gaussian mixture process on the L Gaussian distribution results according to the component weights determined for each of the L Gaussian distributions to obtain the predicted probability distribution.

[0057] The target probability distribution may be a Gaussian distribution determined by the following steps: determining the target mean matrix based on the first coordinate system, determining the standard deviations of the horizontal and vertical axes of the image coordinate system, and determining the target covariance matrix of the target probability distribution based on the standard deviations of the horizontal and vertical axes; and assigning parameters to a standard Gaussian distribution based on the target mean matrix and target covariance matrix to obtain the target probability distribution.

[0058] The steps for determining the standard deviations of the horizontal and vertical axes may include: determining a norm value representing the difference between the first coordinate and the predicted second coordinate based on the first coordinate and the predicted second coordinate; and, if it is determined that the norm value exceeds a set threshold, determining the norm value as the standard deviation of the horizontal and vertical axes of the target probability distribution, and, if it is determined that the norm value does not exceed a set threshold, determining the set threshold as the standard deviation of the horizontal and vertical axes of the target probability distribution, wherein the standard deviations of the horizontal and vertical axes of the target probability distribution are the same value.

[0059] Prior to step 105, the method may further include a step of calculating a positional loss based on the difference between the predicted second coordinate and the first coordinate, wherein the positional loss represents the loss of the initial pose estimation model in predicting the coordinate positions of preset keypoints in sample images. A step of adjusting the model parameters of the initial pose estimation model based on the distributional loss includes a step of adjusting the model parameters of the initial pose estimation model based on the distributional loss and the positional loss.

[0060] The initial pose estimation model includes a backbone network and a prediction head having multiple linear layers. The step of estimating the pose of an object in a sample image using the initial pose estimation model may include the steps of extracting features from the sample image using the backbone network and regressing the extracted features in linear layers to predict the second coordinate and the predicted parameter values ​​of L-set.

[0061] After obtaining a target pose estimation model, this method may further include the steps of: obtaining an image to be processed; and using the target pose estimation model, performing pose estimation on the object to be recognized in the image to be processed and obtaining coordinate information of preset keypoints in the image to be processed.

[0062] The step of acquiring the image to be processed includes the steps of acquiring the original image, performing object recognition processing on the original image and determining a target region containing the object to be recognized within the original image, wherein the object to be recognized is the target object for pose estimation, and cutting out the image content corresponding to the target region from the original image to acquire the image to be processed.

[0063] The process further includes the steps of: obtaining coordinate information of preset keypoints within the image to be processed; determining the state features of the object to be recognized within the image based on the positional relationship between each piece of coordinate information; and determining a target state that matches the object to be recognized based on the matching status between the state features and the candidate state features corresponding to each candidate state.

[0064] Figure 2B is another schematic diagram of the training process of the pose estimation model in the embodiment of the present application. Figure 2B provides a more detailed explanation of the training process of the pose estimation model shown in Figure 2A. The relevant model training processes will be described below with reference to Figure 2B.

[0065] Step 201: The processing equipment acquires each training sample.

[0066] In embodiments of the present invention, to train a target pose estimation model, a processing device acquires each training sample arranged according to the pose estimation needs. Each training sample includes a sample image and the sample coordinates within the sample image for each preset keypoint. Each preset keypoint is used for pose positioning.

[0067] In the embodiments of this application, each preset keypoint may be selected according to the needs of posture estimation. In feasible embodiments, when performing posture estimation on a "human," each preset keypoint may include common human body keypoints (also called human skeleton keypoints), or both human body keypoints and other customized keypoints. When performing posture estimation on a part of the human body, each preset keypoint may include keypoints of a localized area of ​​the human body. For example, when performing gesture recognition, each preset keypoint may include common hand keypoints (including keypoints for each finger joint, etc.). In this way, by adaptively selecting each preset keypoint to suit a specific posture estimation task, posture positioning of different objects can be effectively achieved.

[0068] Step 202: The processing device uses an initial pose estimation model to perform pose estimation on sample images included in the selected training samples, and obtains predicted coordinates and L sets of predicted parameter values ​​corresponding to each preset keypoint obtained through regression processing.

[0069] In the embodiments of this invention, an initial pose estimation model can be obtained by adjusting the output of a conventional regression-based pose estimation network. The backbone network of the initial pose estimation model may be any network that performs feature extraction in the pose estimation scenario, such as StemNet or HRNet-W48. The prediction head of the initial pose estimation model includes multiple linear layers and is used to implement computation and prediction functions.

[0070] Here, the adjustment of a conventional regression-based pose estimation network includes, for example, the processing and output of the linear layer in the prediction head. Therefore, by adjusting the processing and output of the linear layer, the initial pose estimation model can, during training, not only output predicted coordinates but also a set of predicted parameter values ​​corresponding to each of L pre-defined distribution functions. The number of outputs of the linear layer can be determined by the number of pre-defined distribution functions and the assignment requirements for each distribution function. Furthermore, the present invention allows the format of the output content to be set according to the actual processing needs.

[0071] The initial pose model is, for example, a neural network, and includes feature extraction and regression processing. Feature extraction includes, for example, extracting features such as edges, textures, and colors from a sample image. In the real world, there are many situations where there is some relationship between two or more variables, but one variable is not precise enough to determine the other. Regression is a prediction method that examines the relationship between the dependent variable (target) and independent variable (predictor) in such variables. In the regression processing, for example, linear regression is used, and by applying a linear transformation to the extracted features, the predicted coordinates of preset keypoints and predicted parameter values ​​for L distribution functions are obtained.

[0072] Assume that a set of predicted parameter values ​​includes four parameters: the standard deviation on the x-axis (also called the horizontal axis), the standard deviation on the y-axis (also called the vertical axis), the Pearson correlation coefficient, and the component weights of the distribution function. The Pearson correlation coefficient is used to represent the correlation between the standard deviation on the x-axis and the standard deviation on the y-axis. In some embodiments, by adjusting the output of the linear layer, predicted coordinates and four parameter vectors can be output for each preset keypoint. If the total number of preset distribution functions is L, each parameter vector contains L parameter values, and parameter values ​​at the same position in different parameter vectors constitute a set of predicted parameter values. Alternatively, in some other embodiments, by adjusting the output of the linear layer, predicted coordinates and L parameter vectors can be output for each preset keypoint. Each parameter vector contains four parameter values: the standard deviation on the x-axis (also called the horizontal axis), the standard deviation on the y-axis (also called the vertical axis), the correlation coefficient, and the weight coefficients of the distribution function. Alternatively, in some other embodiments, by adjusting the output of the linear layer, predicted coordinates and one parameter vector can be output for each preset keypoint. This parameter vector contains 4*L parameter values. The first four parameter values ​​can be considered as a set of predicted parameter values. The horizontal and vertical axes mentioned above are in the image coordinate system. This image coordinate system can have the lower-left corner of the image as the origin, with the length of the image as the vertical axis and the width as the horizontal axis. "Predicted parameter values" are sometimes abbreviated as "predicted parameters."

[0073] After acquiring each training sample and the constructed initial pose estimation model, the processing device can select training samples from each training sample and perform iterative training.

[0074] In the embodiments of this invention, the batch size can be determined according to the actual processing needs. In the following description, the relevant training process is described using the case where the batch size is 1 as an example. When the batch size is greater than 1, a loss value can be determined for each acquired sample image. Furthermore, the model parameters can be adjusted based on the loss values ​​determined for different sample images.

[0075] In the embodiment of the present invention, the processing device selects a training sample from each training sample to be used in the current iteration and acquires the sample image used in the current iteration. Subsequently, the processing device performs pose estimation on the selected sample image using an initial pose estimation model and obtains predicted coordinates corresponding to each preset keypoint and L sets of predicted parameter values ​​obtained through regression processing. The L sets of predicted parameter values ​​are determined for L pre-set distribution functions.

[0076] In the embodiments of this application, depending on the actual processing needs, the L types of pre-set distribution functions may be any one or a combination thereof of distribution functions having a clear probability function, such as a Gaussian distribution, a Laplace distribution, a Dirac distribution, or a multinomial distribution. This application schematically describes the case where the L pre-set distribution functions are L Gaussian distributions as an example.

[0077] To make it clear, the output of the initial pose estimation model corresponds to the type of distribution function selected. In other words, when selecting a different type of distribution function, the parameters required when assigning parameters to the distribution function will also be different. Therefore, to satisfy the distribution function assignment requirements, the output of the initial pose estimation model can be adaptively adjusted during the model construction phase.

[0078] For example, Figure 3 is a schematic diagram of the output results of the initial pose estimation model in an embodiment of the present invention. Assume that the total number of preset keypoints is n, and that the distribution function set in advance for each preset keypoint is L Gaussian distributions. As can be seen from the contents shown in Figure 3, for each preset keypoint, the initial pose estimation model can obtain the corresponding predicted coordinates and L sets of predicted parameter values. Taking the predicted parameter value group 1 corresponding to preset keypoint 1 as an example, the predicted parameter value group 1 includes the parameter value { It contains TIFF2026517354000002.tif5170. Here, TIFF2026517354000003.tif5170 has the standard deviation on the horizontal axis. TIFF2026517354000004.tif5170 has the standard deviation on the vertical axis. TIFF2026517354000005.tif5170 is the determined Pearson correlation coefficient. TIFF2026517354000006.tif5170 is the component weights determined for the corresponding distribution function (for example, Gaussian distribution 1 among L Gaussian distributions).

[0079] Standard deviation is an indicator of how much a dataset is scattered from its mean. A larger standard deviation means that most values ​​are far from the mean, while a smaller standard deviation means that these values ​​are close to the mean. (Pearson correlation coefficient) TIFF2026517354000007.tif5170 shows the standard deviation on the horizontal axis. TIFF2026517354000008.tif5170 and the standard deviation on the vertical axis This is used to represent the correlations in TIFF2026517354000009.tif5170. Since each of the L Gaussian distributions can be considered as one component, their weights are called component weights. Each Gaussian distribution is a single probability density function. The covariance matrix of a Gaussian distribution (also called a component distribution) is jointly determined by the standard deviation on the x-axis, the standard deviation on the y-axis, and the Pearson correlation coefficient corresponding to the Gaussian distribution.

[0080] Step 203: For each preset keypoint, the processing unit performs the following operations: it integrates L distribution functions based on the predicted coordinates corresponding to the preset keypoint and L sets of predicted parameter values ​​to obtain the predicted probability distribution in the sample image of the predicted keypoint, and determines the distribution loss based on the distribution difference between the predicted probability distribution and the corresponding target probability distribution.

[0081] The predicted probability distribution is used to describe the probability that each pixel point in a sample image is a preset keypoint. The predicted probability distribution of a predicted keypoint in a sample image is a probability distribution determined by considering each pixel in the sample image as a predicted keypoint of a preset keypoint. In the embodiments of this application, assuming that the L pre-set distribution functions are specifically L Gaussian distributions, the processing device performs the following operations on each Gaussian distribution in the process of determining the corresponding predicted probability distribution for each predicted keypoint. That is, it determines the mean matrix of the Gaussian distribution based on the predicted coordinates corresponding to the Gaussian distribution, and determines the covariance matrix and component weights of the Gaussian distribution based on a set of predicted parameter values ​​corresponding to the Gaussian distribution to obtain the Gaussian distribution result after parameter assignment. Then, it performs a Gaussian mixing process on the L Gaussian distribution results according to the component weights determined for each of the L Gaussian distributions to obtain the predicted probability distribution of the corresponding predicted keypoint in the sample image.

[0082] Specifically, in this application, the probability that each pixel point in the sample image is a preset keypoint (also called a Gaussian mixed representation of preset keypoints) is determined by integrating L parameterized Gaussian distributions for each preset keypoint. Here, the sample image may be a two-dimensional image depending on the actual processing needs. Thus, since the coordinate positions of the pixel points in the sample image are two-dimensional, the L Gaussian distributions used are specifically L two-dimensional Gaussian distributions. Therefore, when assigning parameters to each Gaussian distribution to materialize it, the mean matrix and covariance matrix of the Gaussian distribution can be determined for each Gaussian distribution. Here, the mean matrix is, for example, a 1x2 matrix, and the covariance matrix is, for example, a 2x2 matrix.

[0083] If L distribution functions are L Gaussian distributions, the predicted parameter values ​​obtained for each set include the horizontal standard deviation, the vertical standard deviation, the Pearson correlation coefficient, and the component weights of the distribution function. Thus, when determining the mean matrix for a Gaussian distribution, the two coordinate values ​​contained in the predicted coordinates corresponding to the Gaussian distribution can be used as the two elements of the mean matrix. That is, the predicted coordinates are the mean of the Gaussian distribution. The covariance matrix of a Gaussian distribution can be determined using the following equation (1). TIFF2026517354000010.tif17170In the formula, TIFF2026517354000011.tif5170 is the horizontal axis standard deviation included in a set of predicted parameter values ​​corresponding to a Gaussian distribution. TIFF2026517354000012.tif5170 is the vertical axis standard deviation included in a set of predicted parameter values ​​corresponding to a Gaussian distribution. TIFF2026517354000013.tif5170 is the Pearson correlation coefficient, and the range of values ​​is The filename is TIFF2026517354000014.tif5170, where C is the constructed covariance matrix.

[0084] For L Gaussian distributions corresponding to preset keypoints, construct L covariance matrices. The file is recorded as TIFF2026517354000015.tif5170. Here, the covariance matrix for each Gaussian distribution can be constructed using equation (1).

[0085] The collaborative determination of the final predicted probability distribution based on L Gaussian distributions can be understood as a Gaussian mixture distribution. The L Gaussian distributions can be understood as L Gaussian components in the Gaussian mixture distribution. The parameters of the L Gaussian components are It is recorded as TIFF2026517354000016.tif5170. For L Gaussian distributions corresponding to preset keypoints, the L Gaussian distributions have the same mean matrix. It has TIFF2026517354000017.tif4170, but has a different covariance matrix and component weights.

[0086] A Gaussian mixture distribution obtained by integrating multiple Gaussian distributions is not a standard Gaussian distribution, and therefore its standard deviation is not an accurate value. When predicting its standard deviation using an initial pose estimation model, it is difficult to describe it accurately with a specific formula. To describe the standard deviation, in the embodiments of this application, the distribution is sampled using the Monte Carlo method to obtain the shape of the distribution. The basic procedure of the Monte Carlo method is to extract the required number of samples without knowing the Gaussian mixture distribution and to make these samples fit the Gaussian mixture distribution. This process is called sampling. Sampling approximates the standard deviation of the Gaussian mixture distribution. For example, after constructing a distribution and generating a large number of random numbers as samples, selection can be made from these samples in a predetermined manner.

[0087] Furthermore, for each preset keypoint, L Gaussian distributions are assigned parameters. These assigned parameters include the mean matrix and the covariance matrix. Next, the L Gaussian distributions corresponding to the preset keypoints are weighted and merged according to the component weights contained in the corresponding L sets of predicted parameter values ​​to obtain the predicted probability distribution in the sample image of the corresponding predicted keypoint after Gaussian mixing. The relevant mixing process is shown in equation (2). TIFF2026517354000018.tif17170In the formula, TIFF2026517354000019.tif5170 is a Gaussian mixture representation (also called a predicted probability distribution) obtained for a preset keypoint (assuming preset keypoint 1), where L is the total number of each pre-configured Gaussian distribution. TIFF2026517354000020.tif5170 represents the component weights predicted by the initial pose estimation model for a Gaussian distribution i. TIFF2026517354000021.tif5170 is the covariance matrix determined for a Gaussian distribution i. TIFF2026517354000022.tif4170 is a mean matrix determined based on the predicted coordinates of preset keypoint 1, where q is a variable representing a matrix determined by the coordinates of any pixel point in the sample image.

[0088] For example, Figure 4A is a schematic diagram of the process for determining the predicted probability distribution corresponding to a preset keypoint in an embodiment of the present invention. As can be seen from the processing process shown in Figure 4A, after determining the predicted coordinates and L sets of predicted parameter values ​​corresponding to preset keypoint 1, a corresponding mean matrix can be determined based on the predicted coordinates, and a corresponding covariance matrix can be constructed based on the parameter values ​​in each set of predicted parameter values. Furthermore, a Gaussian distribution with L component weights is specifically determined based on the obtained mean matrix and L covariance matrices, and these Gaussian distributions are added together to obtain the predicted probability distribution corresponding to preset keypoint 1.

[0089] As another example, Figure 4B is a schematic diagram of the correspondence between the predicted probability distribution and the sample image in an embodiment of the present invention. As can be seen from the contents of Figure 4B, after obtaining the predicted probability distribution for preset keypoint 1, a corresponding probability value can be determined for each pixel point in the sample image. Here, the determined probability value is used to represent the probability that the pixel point is preset keypoint 1. As can be seen from the contents of Figure 4B, for a pixel point q in the sample image, a two-dimensional matrix of the pixel point can be determined based on the pixel coordinates of the pixel point q in the sample image, and then the probability value of the pixel point q can be determined using equation (2).

[0090] In this way, by using the predicted coordinates directly predicted by the initial pose estimation model and the predicted parameter values ​​of L sets, a corresponding predicted probability distribution, i.e., a Gaussian mixture representation corresponding to each preset keypoint, can be determined for each preset keypoint. By using a Gaussian mixture representation, the coordinates of the preset keypoints can be transformed into a probability distribution in image space, enabling the constraints considered during the training of the regression model and the input image to have the same spatial dimension, thereby improving the model's expressive power and supporting improved model training performance.

[0091] Furthermore, assuming that the L distribution functions are specifically L Gaussian distributions and the target probability distribution is also a Gaussian distribution, the processing device, when determining the corresponding target probability distribution for each preset keypoint, determines the target mean matrix based on the sample coordinates of the preset keypoint, determines the standard deviation of each coordinate axis in the image coordinate system of the sample image, determines the target covariance matrix of the target probability distribution based on the standard deviation of each coordinate axis, assigns parameters to the standard Gaussian distribution based on the target mean matrix and target covariance matrix, and obtains the target probability distribution. Here, the target probability distribution is the probability distribution of the sample image determined based on the sample coordinates of the preset keypoint. The target probability distribution is relative to the predicted probability distribution and is the target that the predicted probability distribution should approach. The "target mean matrix" is the mean matrix, and the "target covariance matrix" is the covariance matrix. These two are named because they are the mean matrix and covariance matrix of the "target" probability distribution, respectively.

[0092] The process for determining the target probability distribution described above will be explained below, using the example of constructing a target probability distribution for a preset keypoint (assuming preset keypoint 1).

[0093] For example, a target probability distribution can be constructed using a standard Gaussian distribution. Specifically, the standard Gaussian distribution used here is a bivariate standard Gaussian distribution where the standard deviation on the x-axis and the standard deviation on the y-axis are equal. The standard Gaussian distribution is as follows: In formula TIFF2026517354000023.tif18170, μ is the expected value (mean) and σ is the standard deviation.

[0094] When determining the target mean matrix, the x and y coordinate values ​​included in the sample coordinates of preset keypoint 1 are determined as elements of the target mean matrix. The number of coordinate dimensions of the sample coordinates is the same as the number of elements of the target mean matrix. For example, assuming that the sample coordinates of preset keypoint are (10,25), the target mean matrix determined for this preset keypoint will be [10,25]. When determining the corresponding target covariance matrix for preset keypoint 1, the standard deviations of the x and y axes can be determined individually. Depending on the actual processing needs, the standard deviation values ​​for each coordinate axis may be fixed values ​​or may change dynamically according to the difference between the predicted coordinates and sample coordinates of the preset keypoint. The standard deviation values ​​for each coordinate axis are the same. The above L Gaussian distributions can also be constructed based on bivariate standard Gaussian distributions.

[0095] Figure 4C is a schematic diagram of the dynamic adjustment of the target probability distribution in an embodiment of the present invention. As can be seen from the contents of Figure 4C, Figure 4C is a diagram that shows the variables in one dimension and intuitively shows the change of the entire distribution. A diagram that shows the variables in one dimension can be extended to a diagram that shows the variables in two dimensions. In Figure 4C, TIFF2026517354000024.tif5170 is the mean in the one-dimensional plot of the predicted probability distribution, and x g This is the mean in the one-dimensional plot of the target probability distribution.

[0096] As can be seen from the dynamic changes in the predicted probability distribution and target probability distribution shown in Figure 4C, as training progresses, The value of TIFF2026517354000025.tif5170 gradually increases by x g This approaches the target probability distribution. In the initial training phase, the target probability distribution and the Gaussian mixture representation (also called the predicted probability distribution) are made to intersect as much as possible, and the value of the standard deviation can be adjusted stepwise to obtain a dynamic target probability distribution, which plays an effective role when adjusting parameters based on the difference in distributions.

[0097] Therefore, since the target probability distribution is a Gaussian distribution, the processing device can control the standard deviation of the target Gaussian distribution by multiplying the difference between the predicted coordinates and the sample coordinates by a predetermined coefficient. For example, the processing device can determine a norm value representing the difference between the sample coordinates and the predicted coordinates based on the sample coordinates and the predicted coordinates of the predicted keypoints. If it is determined that the norm value exceeds a set threshold, the norm value is determined as the standard deviation of each coordinate axis of the target Gaussian distribution; if it is determined that the norm value does not exceed the set threshold, the set threshold is determined as the standard deviation of each coordinate axis of the target Gaussian distribution. Here, the standard deviation of each coordinate axis of the target Gaussian distribution is the same value.

[0098] For example, if we set the coefficient to α, the standard deviation of the target Gaussian distribution... TIFF2026517354000026.tif5170 is given by the following equation (3). TIFF2026517354000027.tif7170 formula, TIFF2026517354000028.tif4170 represents the predicted coordinates of a preset keypoint. TIFF2026517354000029.tif5170 represents sample coordinates of a preset keypoint. TIFF2026517354000030.tif6170 represents the process of obtaining a 2D result by subtracting the sample coordinates from the predicted coordinates, and then calculating the norm of the 2D result.

[0099] However, if the standard deviation of the target Gaussian distribution continues to change, the predicted probability distribution cannot reach the convergence target, which is unfavorable for determining convergence in the model training process. Also, in the training process, the predicted probability distribution is fitted to the shape of the target probability distribution by the L1 loss term of the distribution loss, but if the target probability distribution is constantly changing, the predicted probability distribution cannot learn useful shape information. The L1 loss term will be explained in detail later in the process of calculating the model loss. For this reason, when the target probability distribution converges to a certain state, we stop the change in the target probability distribution and keep it unchanged. If the standard deviation threshold corresponding to this state is t, the target probability distribution is given by the following equation (4). TIFF2026517354000031.tif12170 formula, TIFF2026517354000032.tif5170 is the target probability distribution in the sample image of the preset keypoint. TIFF2026517354000033.tif4170 is a predicted coordinate determined based on preset keypoints. TIFF2026517354000034.tif5170 is a sample coordinate determined based on preset keypoints. TIFF2026517354000035.tif5170 is the standard deviation determined for a preset keypoint, where the value of t is determined by the actual processing needs.

[0100] In this way, when training the model, the value of the standard deviation can be dynamically determined based on the difference between the predicted coordinates and the sample coordinates. This ensures that the target probability distribution and the predicted probability distribution intersect as much as possible, allowing the distribution loss determined based on the target and predicted probability distributions to play a more effective role in the model training process.

[0101] After determining the standard deviations for the x and y axes, the target covariance matrix is ​​obtained using the following equation (5). TIFF2026517354000036.tif15170 formula, TIFF2026517354000037.tif5170 is the calculated target covariance matrix. TIFF2026517354000038.tif5170 is the standard deviation on the horizontal axis (also called the horizontal axis). TIFF2026517354000039.tif6170 represents the standard deviation on the vertical axis (also called the y-axis).

[0102] Furthermore, after assigning parameters to the standard Gaussian distribution based on the obtained target mean matrix and target covariance matrix, the target probability distribution corresponding to preset keypoint 1 is obtained.

[0103] In this way, by establishing a target probability distribution with the highest probability value at the sample coordinates based on the sample coordinates corresponding to the preset keypoints in the sample image space, a comparison criterion is provided between the model prediction results and the created predicted probability distribution.

[0104] Furthermore, after determining the corresponding predicted probability distribution and target probability distribution for each preset keypoint, the processing device can calculate the corresponding distribution loss using the following formula. TIFF2026517354000040.tif6170, TIFF2026517354000041.tif5170 is the distribution loss, TIFF2026517354000042.tif5170 is the target probability distribution. TIFF2026517354000043.tif5170 is a predicted probability distribution, TIFF2026517354000044.tif6170 represents the KL divergence between the predicted probability distribution and the target probability distribution. TIFF2026517354000045.tif6170 represents the L1 loss for the target probability distribution and the predicted probability distribution. TIFF2026517354000046.tif4170 is a smoothing coefficient for smoothing the two loss terms, and its specific value is set according to the actual processing needs.

[0105] In the embodiments of this application, the KL divergence value is unstable and shows some fluctuation, especially when the probability density of the distribution is zero. Therefore, when determining the distribution loss, an L1 loss term can be added in addition to the calculation of the KL divergence.

[0106] The processing device can perform the operation of calculating the position (pixel position) loss for each preset keypoint based on the difference between the predicted coordinates and sample coordinates corresponding to each preset keypoint, before adjusting the model parameters of the initial pose estimation model based on the distribution loss determined for each preset keypoint.

[0107] Specifically, the positional loss can be calculated using the following equation (7). TIFF2026517354000047.tif6170 formula, TIFF2026517354000048.tif5170 is a position loss determined in accordance with preset keypoints. TIFF2026517354000049.tif4170 represents the predicted coordinates of a preset keypoint. TIFF2026517354000050.tif5170 represents sample coordinates of a preset keypoint. TIFF2026517354000051.tif5170 represents the calculation of the L1 loss.

[0108] In this way, by calculating the position loss based on the difference between the predicted coordinates and the sample coordinates, the influence of the regression loss of the regression model itself can be retained in the calculated model loss, thereby constraining the coordinate regression values.

[0109] Step 204: The processing equipment adjusts the model parameters of the initial pose estimation model based on each distributed loss.

[0110] In a feasible embodiment of the present invention, when performing step 204, the processing equipment can adjust the model parameters of the initial pose estimation model based on the distributed loss determined for each preset keypoint.

[0111] In several other feasible embodiments, when positional losses are introduced, the model parameters of the initial pose estimation model can be adjusted based on each distributional loss and each positional loss when tuning the model parameters. The final loss function is given by equation (8). TIFF2026517354000052.tif6170 and TIFF2026517354000053.tif4170 are the loss values ​​ultimately determined for the preset keypoints. TIFF2026517354000054.tif5170 is the distributed loss calculated for preset keypoints. TIFF2026517354000055.tif5170 is the position loss calculated for preset keypoints (also known as regression loss, which represents the loss of the initial pose estimation model when predicting the position of preset keypoints in the sample image), TIFF2026517354000056.tif4170 is the coefficient for positional loss. This prevents the two distributions from becoming misaligned in the early stages of model training, where the predicted probability distribution deviates significantly from the target probability distribution. Furthermore, it prevents the distribution loss from remaining unchanged due to the misalignment of the two distributions. Introducing positional loss also ensures that the model adapts to the initial convergence state as quickly as possible, thereby improving the training efficiency of the initial pose estimation model.

[0112] Figure 4D is a schematic diagram of the process for calculating the model loss for a preset keypoint in an embodiment of the present invention. As can be seen from the contents shown in Figure 4D, when a sample image is input to the initial pose estimation model, the model obtains the predicted coordinates output for preset keypoint 1 and L sets of predicted parameter values. Furthermore, L parameters of a pre-set Gaussian distribution are assigned to each preset keypoint 1, and a Gaussian mixture process is performed to obtain the corresponding predicted probability distribution. For intuitive understanding, the Gaussian distribution diagram shown in Figure 4D is a schematic diagram of a one-dimensional variable. Furthermore, the distribution loss is determined based on the difference between the predicted probability distribution obtained for preset keypoint 1 and the corresponding target probability distribution.

[0113] Step 205: The processing unit determines whether the model convergence condition is met. If it is met, it executes step 206; otherwise, it executes step 202.

[0114] In the embodiments of this invention, the pre-set convergence conditions may be that the total number of training iterations reaches a first threshold, or that the number of times the model loss calculated in multiple model training iterations is consecutively lower than the second threshold reaches a third threshold. Here, the values ​​of the first threshold, the second threshold, and the third threshold are set according to the actual processing needs.

[0115] Step 206: The processing unit outputs the trained target pose estimation model.

[0116] Specifically, the processing device iteratively performs the training process shown in steps 202-204 on the initial pose estimation model until a predetermined convergence condition is met, thereby obtaining a trained target pose estimation model.

[0117] Therefore, the processing equipment can perform tasks in different operational scenarios based on the acquired target attitude estimation model.

[0118] Figure 5A is a schematic diagram of the process for implementing business processing using the target pose estimation model in the embodiment of the present invention. The business processing process executed using the target pose estimation model will be described below with reference to Figure 5A.

[0119] Step 501: The processing device acquires the image to be processed.

[0120] In possible realizations of the present invention, the processing device may acquire the image to be processed collected by an image acquisition device, or it may acquire the image to be processed selected by a party from a client device. The image to be processed is the image to be used for pose estimation.

[0121] In several other possible implementations, the acquired raw image can be cropped to obtain the image to be processed in order to reduce the processing load on the target pose estimation model. For example, after acquiring the raw image, the processing device performs object recognition processing on the raw image to determine the target region containing the object to be recognized within the raw image. Here, the object to be recognized is the target object for pose estimation. Subsequently, the processing device crops the raw image to obtain the image content corresponding to the target region and obtains the image to be processed.

[0122] For example, Figure 5B is a schematic diagram of the process for organizing and acquiring images to be processed in an embodiment of the present invention. As can be seen from the contents shown in Figure 5B, when performing pose estimation on a "human," the processing device acquires the original image, then performs target detection on the original image according to the actual pose estimation needs, and can identify the human body region or human body part region that is the target of pose estimation in the form of a target detection frame. The detection method used for target detection may be a normal human body region detection method (such as the YOLO algorithm) or a human body part region detection method. Subsequently, the processing device cuts out the region identified by the target detection frame (denoted as the ROI region) from the original image and uses it as input to the target pose estimation model. In the embodiment of the present invention, the process of cutting out the original image to acquire the image to be processed can also be applied to the generation of sample images in the model training stage. In this way, by cutting out the region of interest from the original image, interference from background content can be minimized in the acquired image to be processed, and the pose estimation effect of the specified target can be ensured.

[0123] Step 502: The processing device uses a target pose estimation model to perform pose estimation on the object to be recognized in the image to be processed and obtains the coordinate information of each preset keypoint in the image to be processed.

[0124] The processing device inputs the image to be processed into a target pose estimation model, performs pose estimation, and then obtains predicted coordinate information corresponding to each preset keypoint.

[0125] For example, Figure 5C is a schematic diagram of the posture estimation process in an embodiment of the present invention. As can be seen from the contents shown in Figure 5C, when the processing device inputs the image to be processed into the target posture estimation model, it obtains the predicted coordinates corresponding to each preset keypoint output from the target posture model.

[0126] Thus, when performing a specific pose estimation task, it is not necessary to individually determine the predicted probability distribution for each preset keypoint, and the process of determining the predicted probability distribution can remain in the model training phase. The function of calculating the predicted probability distribution based on the model output can be considered a plugin connected to the pose estimation model. Therefore, when processing based on a trained target pose estimation model, the relevant plugin can be directly removed without contributing to the overall time consumption. This improves the training effect of the model, eliminates the burden of resource consumption when applying the model, and ensures the efficiency of pose estimation.

[0127] Furthermore, after acquiring coordinate information of each preset keypoint within the image to be processed, the processing device can determine the state characteristics of the object to be recognized within the image based on the positional relationship between the coordinate information, and then determine the target state that matches the object to be recognized based on the matching status between the state characteristics and the candidate state features corresponding to each candidate state.

[0128] In the embodiments of the present invention, corresponding candidate state features can be pre-stored for each candidate state. Here, each candidate state can be selected from states such as various body postures, various gestures, or identity authentication states for various objects. When the candidate state is a various body posture or various gestures, the pre-stored candidate state features can represent the relative positions of each preset keypoint in the corresponding posture or gesture. When the candidate state is an identity authentication state, the candidate state features may specifically be features for realizing identity authentication, such as palm print features or iris features.

[0129] Therefore, the processing device can determine the state characteristics of the object to be recognized based on the positional relationship between each preset keypoint, and determine the target state corresponding to the object to be recognized based on the state characteristics. Here, the object to be recognized is the target object for pose estimation within the image to be processed.

[0130] In this way, the pose estimation results enable the determination of the state of the object being recognized in various application scenarios.

[0131] The following describes the relevant business process using several business processes that utilize the target attitude estimation model, with reference to the diagrams.

[0132] Figure 6A is a schematic diagram of the palm print recognition process in an embodiment of the present invention. The process for achieving palm print recognition based on the target pose estimation model will be described below with reference to Figure 6A.

[0133] As privacy concerns gradually increase, palm print recognition will be used more widely in real-world application scenarios such as payments and identity verification. The human body posture estimation technology according to the embodiment of this application can be used on a human hand to position the palm region by detecting key points of the hand from images collected in real time.

[0134] Specifically, before recognizing each user's palm print, a palm detection model and a palm keypoint detection model using human body pose estimation technology (i.e., a target pose estimation model) are combined to form a palm recognition assembly, and the user's palm can be registered in the backend registration library.

[0135] Subsequently, in each recognition process, after acquiring a captured image, palm detection and hand pose estimation are performed on the image to finally determine the hand region within the image, and the hand region is cropped from the image. Next, palm print recognition is performed by comparing the hand region image with each photograph in the registered library to recognize the user's identity and perform identity authentication. Because palm detection roughly determines the hand region from the image, and the hand pose estimation process determines the position of each preset keypoint of the hand, accurate identification of the hand region becomes possible.

[0136] Figure 6B is a schematic diagram of the processing logic when realizing motion recognition using the target pose estimation model in the embodiment of the present invention. The human body pose estimation proposed in the present invention can be used for recognizing motions, gestures, and step patterns, such as determining fall states and disease signals, and for automated instruction in fitness, sports, and dance. As can be seen from the contents shown in Figure 6B, the processing logic when performing motion recognition, gesture recognition, and step pattern recognition involves capturing an image, identifying a region of the human body or hand by target detection, realizing human body pose estimation using the target pose estimation model, identifying preset key points of the human body or hand, extracting a region of interest (ROI) based on each determined preset key point, and then recognizing subsequent motions, gestures, and step patterns based on the determined region of interest.

[0137] Furthermore, the applicant compared the posture estimation method conceived during the conceptual stage of the invention with the posture estimation method proposed in this application and obtained the following comparison results.

[0138] Specifically, please refer to Table 1, which compares the model test results of the embodiments of this application. The applicant tested the processing effectiveness of the pose estimation method proposed in this application with other possible pose estimation methods in the validation set of the public dataset MSCOCO. The metrics for evaluating processing effectiveness include the number of parameters, GFLOPs (Giga Floating-Point Operations Per Second), and mAP (mean Average Precision). The number of parameters and GFLOPs represent the model processing speed; the smaller the number of parameters and GFLOPs, the higher the model processing speed. mAP represents the model prediction accuracy; the higher the mAP, the higher the model prediction accuracy.

[0139] TIFF2026517354000057.tif173170

[0140] ResNet-50 and StemNet are small-scale backbone networks, while ResNet-152 and HRNet are large-scale backbone networks. The larger the W coefficient of HRNet, the deeper and wider the network layers become, resulting in a larger model. ResNet-50 is larger than StemNet, while HRNet-W32 and ResNet-152 are almost the same size.

[0141] In summary, when compared with other pose estimation methods using a backbone network of similar scale, the target pose estimation model trained using the model training method proposed in this application outperforms all other methods and can maintain a narrower range for the number of parameters and GFLOPs. While SimpleBaselines is a heatmap model, the processing performance of this application is superior to that of the heatmap model. Therefore, the target pose estimation model trained using the training method proposed in this application far surpasses other currently available methods and offers a clear processing advantage.

[0142] Thus, the pose estimation model training method proposed in this application can achieve performance comparable to a heatmap model by training a regression model (i.e., an initial pose estimation model) while minimizing time consumption. Furthermore, this application is extremely time-intensive and applicable to real-time human body pose estimation scenarios. In summary, this application proposes an innovative training method that uses Gaussian mixture processing to represent the positions of preset keypoints and trains the model by minimizing the difference between the predicted probability distribution and the target probability distribution using Monte Carlo estimation.

[0143] Based on a similar inventive concept, Figure 7 shows a schematic diagram of the logical structure of a training device for a posture estimation model according to an embodiment of the present application. The posture estimation model training device 700 comprises an acquisition unit 701 and a training unit 702.

[0144] According to one embodiment of the present invention, the acquisition unit 701 is used to acquire a training sample. The training sample includes a sample image and the first coordinates of a preset keypoint within the sample image, the preset keypoint being used to position the pose of an object within the sample image.

[0145] The training unit 702 performs pose estimation of an object in a sample image using an initial pose estimation model, obtains the predicted second coordinates of preset keypoints in the sample image and L sets of predicted parameter values, wherein the L sets of predicted parameter values ​​are determined for L pre-set distribution functions, integrates the L distribution functions based on the predicted second coordinates and L sets of predicted parameter values ​​to obtain a predicted probability distribution, and determines the distribution loss based on the difference between the predicted probability distribution and the target probability distribution, wherein the predicted probability distribution represents the distribution of the probability that each pixel point in the predicted sample image is a preset keypoint, and the target probability distribution is the probability distribution determined based on the first coordinate, and adjusts the model parameters of the initial pose estimation model based on the distribution loss to obtain a target pose estimation model.

[0146] The training unit 702 is used to obtain the predicted probability distribution by performing parameter assignment and weighted summation on L distribution functions based on the predicted second coordinate and L sets of predicted parameter values.

[0147] The training unit 702 is used to obtain L distribution function results by assigning parameters to each corresponding distribution function among L distribution functions based on the predicted second coordinate and the predicted function parameter values ​​for each of the L predicted parameter values, and to obtain the predicted probability distribution by performing a weighted sum on the L distribution function results using the component weights of the L predicted parameter values ​​as components.

[0148] If the pre-set L distribution functions are L Gaussian distributions, the training unit is used to determine the mean matrix of each of the L Gaussian distributions based on the predicted second coordinate, and to determine the covariance matrix and component weights of the Gaussian distributions based on a set of predicted parameter values ​​corresponding to the Gaussian distributions, thereby obtaining the Gaussian distribution results after parameter assignment. It is also used to perform Gaussian mixture processing on the L Gaussian distribution results according to the component weights determined for each of the L Gaussian distributions, thereby obtaining the predicted probability distribution.

[0149] The target probability distribution may be a Gaussian distribution determined by the following steps: determining the target mean matrix based on the first coordinate system, determining the standard deviations of the horizontal and vertical axes of the image coordinate system, and determining the target covariance matrix of the target probability distribution based on the standard deviations of the horizontal and vertical axes; and assigning parameters to a standard Gaussian distribution based on the target mean matrix and target covariance matrix to obtain the target probability distribution.

[0150] When determining the standard deviation of each corresponding coordinate axis, the training unit 702 determines a norm value representing the difference between the first coordinate and the predicted second coordinate based on the first coordinate and the predicted second coordinate; if it determines that the norm value exceeds a set threshold, it determines the norm value as the standard deviation of the horizontal and vertical axes of the target probability distribution; and if it determines that the norm value does not exceed a set threshold, it determines the set threshold as the standard deviation of the horizontal and vertical axes of the target probability distribution, so that the standard deviations of the horizontal and vertical axes of the target probability distribution are the same value.

[0151] According to another embodiment of the present invention, the acquisition unit 701 is used to acquire each training sample. One training sample includes one sample image and the sample coordinates within the one sample image for each preset keypoint. Each preset keypoint is used for pose positioning.

[0152] Training unit 702 is used to perform multiple iterative training on the initial pose estimation model based on each training sample to obtain the target pose estimation model. In one iterative training, the following operations are performed:

[0153] Pose estimation is performed on sample images included in the selected training samples, and the predicted coordinates corresponding to each preset keypoint and L sets of predicted parameter values ​​are obtained through regression processing. The L sets of predicted parameter values ​​are determined for L predefined distribution functions.

[0154] For each preset keypoint, the following operations are performed: The training model 702 integrates L distribution functions based on the predicted coordinates and L sets of predicted parameter values ​​corresponding to the preset keypoint to obtain the predicted probability distribution in the sample image for the corresponding predicted keypoint, and determines the distribution loss based on the distribution difference between the predicted probability distribution and the corresponding target probability distribution. The target probability distribution is the probability distribution of the sample image determined based on the first coordinate of the preset keypoint. Based on each distribution loss, the training model 702 adjusts the model parameters of the initial pose estimation model.

[0155] If the pre-set L distribution functions are L Gaussian distributions, and the L distribution functions are integrated based on the corresponding predicted coordinates and L sets of predicted parameter values ​​to obtain the predicted probability distribution in the sample image of the corresponding predicted keypoint, then the training unit 702 will: Based on the corresponding predicted coordinates, the mean matrix of the Gaussian distribution is determined, and based on the corresponding set of predicted parameter values, the covariance matrix and component weights corresponding to the Gaussian distribution are determined, and the Gaussian distribution result after parameter assignment is obtained. This process is used to perform a Gaussian mixing operation on L Gaussian distributions according to the component weights determined for each distribution, thereby obtaining the predicted probability distribution within the sample image of the corresponding predicted keypoint.

[0156] The target probability distribution is, Based on the sample coordinates of the corresponding preset keypoints, the target mean matrix is ​​determined, along with the standard deviation of each corresponding coordinate axis. Based on the standard deviation of each coordinate axis, the target covariance matrix corresponding to the target probability distribution is determined. Based on the target mean matrix and target covariance matrix, parameters are assigned to the standard Gaussian distribution to obtain the target probability distribution. It may also be a Gaussian distribution determined by [the relevant factor].

[0157] When determining the standard deviation of each corresponding coordinate axis, the training unit 702, Based on the sample coordinates and the predicted coordinates of the predicted keypoints, a norm value representing the difference between the sample coordinates and the predicted coordinates is determined, This is used to determine whether the norm value exceeds a set threshold, to determine the norm value as the standard deviation of each axis of the target probability distribution, and whether the norm value does not exceed the set threshold, to determine the set threshold as the standard deviation of each axis of the target probability distribution, such that the standard deviations of each axis of the target probability distribution are the same.

[0158] Before adjusting the model parameters of the initial pose estimation model based on each distribution loss, the training unit 702 further: For each preset keypoint, it is used to perform the operation of calculating the position loss based on the difference between the corresponding predicted coordinates and sample coordinates. Adjusting the model parameters of the initial pose estimation model based on each distribution loss is possible. This involves adjusting the model parameters of the initial pose estimation model based on each distribution loss and each position loss.

[0159] After acquiring the target attitude estimation model, the apparatus may further include a processing unit 703. The processing unit 703 is: Acquiring the image to be processed, It is used to perform pose estimation on the object to be recognized within the image being processed using a target pose estimation model, and to obtain the coordinate information of each preset keypoint within the image being processed.

[0160] When acquiring the image to be processed, the processing unit 703 performs the following: Obtaining the original image, This involves performing object recognition processing on the original image and determining the target region within the original image that includes the object to be recognized, wherein the object to be recognized is the target object for pose estimation. It is used to extract the image content corresponding to the target region from the original image and obtain the image to be processed.

[0161] After obtaining the coordinate information of each preset keypoint within the image to be processed, the processing unit 703 further: Based on the positional relationships between each coordinate information, the state characteristics of the object to be recognized within the image to be processed are determined, It is used to determine the target state that matches the object to be recognized, based on the matching status between state features and candidate state features corresponding to each candidate state.

[0162] Details of the training apparatus for the posture estimation model in the embodiment of this application can be found in the description of the embodiment of the method.

[0163] After introducing a training method and apparatus for a posture estimation model according to an exemplary embodiment of the present application, an electronic device according to another exemplary embodiment of the present application will be introduced.

[0164] As those skilled in the art will understand, various embodiments of the present application may be implemented as systems, methods, or program products. Therefore, various embodiments of the present application may specifically be implemented as complete hardware embodiments, complete software embodiments (including firmware, microcode, etc.), or embodiments combining hardware and software. Hereinafter, these may be collectively referred to as “circuits,” “modules,” or “systems.”

[0165] Based on an inventive concept similar to that of the embodiments of the above method, when the electronic device in the embodiment of the present application corresponds to a processing device, Figure 8 shows a schematic diagram of the hardware structure of the electronic device to which the embodiment of the present application is applied. The electronic device 800 may include at least a processor 801 and a memory 802. A computer program is stored in the memory 802. When the computer program is executed by the processor 801, the processor 801 performs one of the steps of training the pose estimation model described above.

[0166] In some possible embodiments, the electronic device according to the present application may include at least one processor and at least one memory in which a computer program is stored. Once the computer program is executed by the processor, the processor performs steps of training a pose estimation model according to the various exemplary embodiments of the present application described herein. For example, the processor may perform the steps shown in Figures 2A and 2B.

[0167] Based on an inventive concept similar to that of the embodiments of the method described above, various embodiments of training the attitude estimation model according to the present application may be implemented as a program product including program code. When the program product is executed on an electronic device, the program code is used to cause the electronic device to perform the steps of training the attitude estimation model according to the various exemplary embodiments of the present application described herein. For example, the electronic device may perform the steps shown in Figures 2A and 2B.

[0168] A program product may use any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0169] While embodiments of this application have been described, those skilled in the art, upon understanding the basic creative concepts, may make further changes and modifications to these embodiments. Therefore, the attached claims are intended to be construed as encompassing the embodiments and all changes and modifications that fall within the scope of this application.

[0170] Clearly, a person skilled in the art can make various modifications and variations of this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations of this application fall within the scope of the claims of this application and the equivalent art, this application is intended to include such modifications and variations as well.

Claims

1. A method for training a posture estimation model, A step of obtaining a training sample, wherein the training sample includes a sample image and a first coordinate of a preset keypoint within the sample image, and the preset keypoint is used to position an object in the sample image. A step of performing an initial pose estimation model to estimate the pose of the object in the sample image, and obtaining the predicted second coordinates of the preset keypoints in the sample image and L sets of predicted parameter values, wherein the L sets of predicted parameter values ​​are determined for L predetermined distribution functions. A step of obtaining a predicted probability distribution by integrating the L distribution functions based on the predicted second coordinates and the L predicted parameter values, wherein the predicted probability distribution represents the distribution of the probability that each pixel point in the predicted sample image is the preset key point. A step of determining the distribution loss based on the difference between the predicted probability distribution and the target probability distribution, wherein the target probability distribution is a probability distribution determined based on the first coordinate, The steps include adjusting the model parameters of the initial attitude estimation model based on the distribution loss and obtaining a target attitude estimation model, A method for training a posture estimation model, characterized by including the following:

2. The aforementioned integration A step of obtaining the predicted probability distribution by performing parameter assignment and weighted summation on the L distribution functions based on the predicted second coordinates and the L predicted parameter values, The method according to claim 1, characterized by including

3. The aforementioned integration The steps include: assigning parameters to corresponding distribution functions among the L distribution functions based on the predicted second coordinates and the predicted function parameter values ​​for each of the L predicted parameter values, and obtaining L distribution function results; The steps include: obtaining the predicted probability distribution by performing a weighted sum on the L distribution function results, using the L distribution function results as components and the component weights of the L sets of predicted parameter values; The method according to claim 1 or 2, characterized by including

4. The aforementioned L pre-set distribution functions are L Gaussian distributions, The aforementioned integration The steps include determining the mean matrix of each of the L Gaussian distributions based on the predicted second coordinates, and determining the covariance matrix and component weights of the Gaussian distributions based on a set of predicted parameter values ​​corresponding to the Gaussian distributions, thereby obtaining the Gaussian distribution results after parameter assignment. The steps include: performing a Gaussian mixing process on the L Gaussian distribution results according to the component weights determined for each of the L Gaussian distributions to obtain the predicted probability distribution; The method according to any one of claims 1 to 3, characterized by including

5. The aforementioned target probability distribution is The steps include determining the target mean matrix based on the first coordinate system, determining the standard deviations of the horizontal and vertical axes of the image coordinate system, and determining the target covariance matrix of the target probability distribution based on the standard deviations of the horizontal and vertical axes, The steps include assigning parameters to a standard Gaussian distribution based on the target mean matrix and the target covariance matrix, and obtaining the target probability distribution, The method according to any one of claims 1 to 4, characterized in that it is a Gaussian distribution determined by

6. The step of determining the standard deviations of the horizontal and vertical axes, respectively, is: A step of determining a norm value representing the difference between the first coordinate and the predicted second coordinate based on the first coordinate and the predicted second coordinate, If it is determined that the norm value exceeds a set threshold, the norm value is determined as the standard deviation of the horizontal and vertical axes of the target probability distribution; and if it is determined that the norm value does not exceed the set threshold, the set threshold is determined as the standard deviation of the horizontal and vertical axes of the target probability distribution, wherein the standard deviations of the horizontal and vertical axes of the target probability distribution are the same value. The method according to claim 5, characterized by including

7. Before adjusting the model parameters of the initial pose estimation model based on the distribution loss, the method A step of calculating a position loss based on the difference between the predicted second coordinate and the first coordinate, wherein the position loss represents the loss of the initial pose estimation model in predicting the coordinate position of the preset keypoint in the sample image. It further includes, The step of adjusting the model parameters of the initial pose estimation model based on the distribution loss is: The step of adjusting the model parameters of the initial attitude estimation model based on the distribution loss and the position loss. The method according to any one of claims 1 to 6, characterized by including

8. The initial attitude estimation model includes a backbone network and a prediction head having multiple linear layers. The step of estimating the pose of the object in the sample image using the initial pose estimation model is: The steps include: performing feature extraction of the sample image using the backbone network; The steps include: performing regression processing on the extracted features using a linear layer to predict the second coordinate and the predicted parameter values ​​of the L-set; The method according to any one of claims 1 to 7, characterized by including

9. After obtaining the target attitude estimation model, the method is as follows: The steps include acquiring the image to be processed, The steps include: using the target pose estimation model, performing pose estimation processing on the object to be recognized in the image to be processed, and obtaining the coordinate information of the preset keypoints in the image to be processed; The method according to any one of claims 1 to 8, further comprising:

10. The step of acquiring the image to be processed is: Steps to obtain the original image, A step of performing object recognition processing on the original image and determining a target region in the original image that includes the object to be recognized, wherein the object to be recognized is the target object for pose estimation. The steps include: obtaining a processed image by cutting out the image content corresponding to the target region from the original image; The method according to claim 9, characterized by including

11. After obtaining the coordinate information of the preset keypoints in the image to be processed, The steps include determining the state characteristics of the object to be recognized within the image to be processed based on the positional relationship between each coordinate information, The step of determining a target state that matches the object to be recognized, based on the matching status between the state features and the candidate state features corresponding to each candidate state, The method according to claim 9 or 10, further comprising:

12. A training device for posture estimation models, An acquisition unit for acquiring training samples, wherein the training sample includes a sample image and a first coordinate of a preset keypoint within the sample image, and the preset keypoint is used to position the pose of an object within the sample image; The initial pose estimation model is used to estimate the pose of the object in the sample image, and the predicted second coordinates of the preset keypoints in the sample image and L sets of predicted parameter values ​​are obtained, wherein the L sets of predicted parameter values ​​are determined for L predetermined distribution functions. Based on the predicted second coordinates and the L predicted parameter values, the L distribution functions are integrated to obtain a predicted probability distribution, and the distribution loss is determined based on the difference between the predicted probability distribution and the target probability distribution, wherein the predicted probability distribution represents the distribution of the probability that each pixel point in the predicted sample image is the preset key point, and the target probability distribution is the probability distribution determined based on the first coordinates. A training unit used to adjust the model parameters of the initial posture estimation model based on the distribution loss and obtain a target posture estimation model, A training device for a posture estimation model, characterized by comprising the following:

13. The aforementioned training unit is Based on the predicted second coordinates and the predicted parameter values ​​of the L sets, parameter assignment and weighted summation are performed on the L distribution functions to obtain the predicted probability distribution. The apparatus according to claim 12, characterized in that it is used for the following.

14. The aforementioned training unit is Based on the predicted second coordinates and the predicted function parameter values ​​for each of the L predicted parameter values, parameters are assigned to each corresponding distribution function among the L distribution functions, and L distribution function results are obtained. The predicted probability distribution is obtained by taking the L distribution function results as components and performing a weighted sum on the L distribution function results using the component weights of the L sets of predicted parameter values, The apparatus according to claim 12 or 13, characterized in that it is used for the following.

15. The aforementioned L pre-set distribution functions are L Gaussian distributions, The aforementioned training unit is Based on the predicted second coordinates, the mean matrix of each of the L Gaussian distributions is determined, and based on a set of predicted parameter values ​​corresponding to the Gaussian distributions, the covariance matrix and component weights of the Gaussian distributions are determined to obtain the Gaussian distribution results after parameter assignment. The process involves performing a Gaussian mixing operation on the L Gaussian distribution results according to the component weights determined for each of the L Gaussian distributions, thereby obtaining the predicted probability distribution. The apparatus according to any one of claims 12 to 14, characterized in that it is used for the following purpose.

16. The aforementioned target probability distribution is Based on the aforementioned first coordinate system, the target mean matrix is ​​determined, and the standard deviations of the horizontal and vertical axes of the image coordinate system are determined, respectively. Based on the standard deviations of the horizontal and vertical axes, the target covariance matrix of the target probability distribution is determined. Based on the aforementioned target mean matrix and target covariance matrix, parameters are assigned to the standard Gaussian distribution to obtain the aforementioned target probability distribution. The apparatus according to any one of claims 12 to 15, characterized in that it is a Gaussian distribution determined by

17. When determining the standard deviations of the horizontal and vertical axes, the training unit, Based on the first coordinate and the predicted second coordinate, a norm value representing the difference between the first coordinate and the predicted second coordinate is determined, If it is determined that the norm value exceeds the set threshold, the norm value is determined as the standard deviation of the horizontal and vertical axes of the target probability distribution, and if it is determined that the norm value does not exceed the set threshold, the set threshold is determined as the standard deviation of the horizontal and vertical axes of the target probability distribution, wherein the standard deviations of the horizontal and vertical axes of the target probability distribution are the same value. The apparatus according to claim 16, characterized in that it is used for [a specific purpose].

18. An electronic device comprising memory, a processor, and a computer program stored in memory and executable by the processor, wherein the processor, when executing the computer program, realizes the method according to any one of claims 1 to 11.

19. A computer-readable storage medium in which a computer program is stored, wherein the computer program, when executed by a processor, implements the method described in any one of claims 1 to 11.

20. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method described in any one of claims 1 to 11.