Human body posture estimation model training method and device, storage medium and product
By using 2D key points of multi-view images to optimize 3D posture information in the human posture estimation model, the problem of low training efficiency in the prior art is solved, and efficient and high-precision human posture estimation is achieved.
Patent Information
- Application Number
- CN202510374314.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-25
AI Technical Summary
The existing human posture estimation model has low training efficiency, making it difficult to accurately and automatically extract prior information, and the optimization process is time-consuming and inefficient.
By inputting the first view image into the human body posture estimation model, 3D human body posture information is output, and 2D human body key points are extracted from multiple second view images, and the 3D human body posture information of the first view is optimized based on these key points, and the model is finally trained.
It realizes efficient training of high-precision human posture estimation model, reducing training time and improving estimation accuracy.
Smart Images

Figure CN120375466A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method and device for training a human pose estimation model, a storage medium, and a product. Background Art
[0002] Human pose estimation can well detect the poses and movements of the human body in various fields. For example, in VR and AR applications, this technology can enable users to interact with the virtual environment, enhancing the immersion and experience; in the field of agricultural machinery, it can monitor agricultural machinery drivers or agricultural machinery safety officers in real time to ensure that they maintain correct postures when driving and operating agricultural machinery, thereby improving safety and work efficiency.
[0003] Due to the various movements and shapes exhibited by the human body in real-world scenarios, human pose estimation is challenging. Currently, models trained using optimization-based methods are usually employed to estimate 3D human poses and shapes. Specifically, a parametric human model is defined or a pre-scanned 3D model is used as a template, and prior information such as key points, skeletons, contours, and RGB-D images is utilized to construct an energy function for optimizing the model. By minimizing the energy function, the predefined parametric human model is fitted to the prior information to estimate the human pose and shape.
[0004] Although a human pose estimation model can be trained using an optimization-based method, due to the complexity of the real human body, it is often difficult to accurately automatically extract prior information, and the optimization process is usually time-consuming, resulting in the problem of low training efficiency of the human pose estimation model. Summary of the Invention
[0005] The purpose of the embodiments of this application is to provide a method and device for training a human pose estimation model, a storage medium, and a product, so as to solve the problem of low training efficiency of the existing human pose estimation model.
[0006] To solve the above technical problems, this specification is implemented as follows: In a first aspect, a method for training a human pose estimation model is provided, including: Inputting a first-view image into the human pose estimation model to output 3D human pose information corresponding to the first-view image, where the first-view image corresponds to an image of the human body taken from a first perspective; Extracting 2D human key points of the human body from a plurality of second-view images respectively, where the second-view corresponds to an image of the human body taken from a second perspective, the second perspective is different from the first perspective, and the second perspectives of the plurality of second-view images are different from each other; Optimizing the 3D human pose information corresponding to the first-view image based on the 2D human key points to obtain optimized 3D human pose information; Train the human pose estimation model based on the optimized 3D human pose information.
[0007] Optionally, the 3D human pose information includes the pose parameters, shape parameters, and camera parameters of the human body.
[0008] Optionally, optimizing the 3D human pose information corresponding to the first view image based on the 2D human key points to obtain the optimized 3D human pose information includes: Determine the 3D human key points corresponding to the 3D human pose information of the first view image; Determine the total projection error function between the 2D human key points obtained by projecting the 3D human key points onto each second view and the 2D human key points extracted from the corresponding second view, where the total projection error function is determined based on the pose parameters, shape parameters, and camera parameters of the human body; Based on the total projection error function and preset constraint conditions, construct an optimization model for the 3D human pose information with the goal of minimizing the total projection error function; Perform an optimal solution for the optimization model to obtain the optimized 3D human pose information.
[0009] Optionally, determining the 3D human key points corresponding to the 3D human pose information of the first view image includes: Input the 3D human pose information corresponding to the first view image into the SMPL model to obtain the 3D human key points corresponding to the 3D human pose information.
[0010] Optionally, determining the total projection error function between the 2D human key points obtained by projecting the 3D human key points onto each second view and the 2D human key points extracted from the corresponding second view includes: Project the 3D human key points onto the 2D planes corresponding to each second view respectively to obtain the 2D human key points projected onto each second view; Based on the error sum between the projected 2D human key points corresponding to each second view and the extracted 2D human key points, with the pose parameters, shape parameters, and camera parameters of different views as variables, construct the total projection error function.
[0011] Optionally, training the human pose estimation model based on the optimized 3D human pose information includes: Determine an estimation loss function based on the optimized 3D human pose information and the 3D human pose information output by the human pose estimation model; Adjust the parameters of the human body pose estimation model until the estimation loss function is minimized, and train the human body pose estimation model.
[0012] Optionally, it further includes: Select each view image sample in turn from the view image sample set including the first view image and the multiple second view images as the first view image for training the human body pose estimation model.
[0013] In a second aspect, a training device for a human body pose estimation model is provided, including a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0014] In a third aspect, a readable storage medium is provided. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0015] In a fourth aspect, a computer program product is provided. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to cause a computer to execute the steps of the method described in the first aspect.
[0016] In the embodiments of the present application, by inputting a first view image into a human body pose estimation model to output 3D human body pose information corresponding to the first view image, the first view image corresponds to an image obtained by photographing a human body from a first perspective; 2D human body key points of the human body are respectively extracted from multiple second view images, the second view corresponds to an image obtained by photographing the human body from a second perspective, and the second perspective is different from the first perspective, and the second perspectives of the multiple second view images are different from each other; based on the 2D human body key points, the 3D human body pose information corresponding to the first view image is optimized to obtain optimized 3D human body pose information; based on the optimized 3D human body pose information, the human body pose estimation model is trained, whereby a human body pose estimation model with high estimation accuracy can be trained time-saving and effectively. Description of the Drawings
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings: Figure 1 It is a schematic flowchart of the training method of the human body pose estimation model in the embodiments of the present application.
[0018] Figure 2It is a schematic diagram of the application scenario of the training method of the human pose estimation model according to the embodiment of the present application.
[0019] Figure 3 It is a structural block diagram of the training device of the human pose estimation model according to the embodiment of the present application. Detailed implementation manners
[0020] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application. The accompanying drawing numbers in the present application are only used to distinguish each step in the solution and are not used to limit the execution order of each step. The specific execution order shall be subject to the description in the specification.
[0021] To solve the problems existing in the prior art, the embodiment of the present application provides a training method for a human pose estimation model, as Figure 1 shown, including the following steps 102 to step 108.
[0022] Step 102, input the first view image into the human pose estimation model to output the 3D human pose information corresponding to the first view image, where the first view image corresponds to an image obtained by photographing a human body from a first perspective.
[0023] Combined with Figure 2 , Figure 2 is an application scenario example of the training method of the human pose estimation model according to the embodiment of the present application. The first view image 10 is an image obtained by photographing the corresponding human body with a camera placed at a certain angle, that is, the first perspective. Here, the first view image 10 is used as the main view and is a single view.
[0024] The first view image 10 is input into the human pose estimation model 30, and the human pose estimation model 30 is used to estimate the 3D human pose information corresponding to the appearance of the human body in the first view image 10. For example, the human pose estimation model 30 is a parameter regressor, and by regressing the image parameters, the 3D human pose information can be obtained , where i represents the view image number corresponding to the first view image 10, and reg represents that the human pose estimation model 30 input with the first view image 10 numbered i is a parameter regressor.
[0025] Optionally, the 3D human pose information includes the pose parameters, shape parameters, and camera parameters of the human body.
[0026] The 3D pose information includes 3D pose parameters that describe the actions of each key point of the human body, and 3D shape parameters that describe the shape of the human body, such as height, weight, and build. The key points of the human body can be represented by the joint points of the human body, such as shoulders, knees, wrists, etc. The 3D human pose parameters can be represented by the axis angles of the 3D space X / Y / Z of the human key points, and the human shape parameters can be represented by vector values. Different vector values describe different human shapes. The camera parameters describe the translation vectors of the 3D space X / Y / Z in the external parameters of the camera.
[0027] For example, combined with Figure 2 the first view image 10 of a single view, the corresponding pose parameters (24*3), shape parameters (10), and camera parameters (3) are obtained. The vector of the pose parameters (24*3) describes the axis angles of the 3D space of 24 human joint points, the vector of the shape parameters (10) describes 10 different human shapes, and the vector of the camera parameters (3) describes the translation vectors of the 3D space X / Y / Z in the external parameters of the camera. Define the parameters regressed from the i-th first view image as where represents the pose parameters, represents the shape parameters, represents the camera parameters.
[0028] Step 104: Extract the 2D human key points of the human body from multiple second view images respectively. The second view corresponds to the image obtained by photographing the human body from the second perspective, and the second perspective is different from the first perspective. The second perspectives of the multiple second view images are different from each other.
[0029] The second view image 20 is the image obtained by photographing the corresponding human body with cameras placed at multiple angles different from the first perspective, that is, the second perspective. Multiple cameras are placed at different angles to photograph the same human body respectively, and multiple second view images 20 with different second perspectives are obtained correspondingly. Here, the second view image 20 is used as multiple views derived from the first view image 10.
[0030] The 2D human key points 40 are used to represent the human joint points in the 2D space X / Y, and the key point coordinates are in the format of (x, y). Different perspective transformations result in occlusion of the human body parts photographed in the corresponding second view image 20, so the number and type of the extracted 2D human key points 40 may vary.
[0031] As Figure 2 shown, the 2D human key points 40 extracted from multiple different second view images 20 can be fused together to obtain the fused 2D human key points 50.
[0032] Step 106: Optimize the 3D human pose information corresponding to the first view image based on the 2D human key points to obtain the optimized 3D human pose information.
[0033] In this step, the 2D human key points extracted from the second view images 20 of multiple perspectives are used to supplement the 2D human key points that do not appear in the single-view, single-perspective first view image 10, so as to optimize the 3D human pose information estimated corresponding to the first view image 10.
[0034] Based on the solution provided in the above embodiment, optionally, in the above step 106, the optimizing the 3D human pose information corresponding to the first view image based on the 2D human key points to obtain the optimized 3D human pose information includes: determining the 3D human key points corresponding to the 3D human pose information of the first view image; determining the total projection error function between the 2D human key points obtained by projecting the 3D human key points onto each second view angle and the 2D human key points extracted from the corresponding second view angle, where the total projection error function is determined based on the pose parameters, shape parameters, and camera parameters of the human body; constructing an optimization model of the 3D human pose information with the goal of minimizing the total projection error function based on the total projection error function and preset constraint conditions; and performing an optimal solution to the optimization model to obtain the optimized 3D human pose information.
[0035] In this embodiment, the optimization of the 3D human pose information estimated corresponding to the first view image 10 is achieved by constructing an error function corresponding to the 2D human key points of the first view image 10 and the 2D human key points of multiple second view images 20.
[0036] First, based on the 3D human pose information of the first view image 10 output by the human pose estimation model 30 to be trained, determine the 3D human key points corresponding to the first view image 10.
[0037] In one embodiment, the determining the 3D human key points corresponding to the 3D human pose information of the first view image includes: inputting the 3D human pose information corresponding to the first view image into the SMPL model to obtain the 3D human key points corresponding to the 3D human pose information.
[0038] The Skinned Multi-Person Linear Model (SMPL) is a parametric model of human shape and pose, which can detect the key points of multiple joints of the human body. Input the 3D human pose information estimated for the corresponding human body in the first view image 10 into the SMPL model to output the 3D human key points corresponding to the first view image 10, including the coordinates and categories of the key points.
[0039] Similarly, due to the perspective transformation, there are differences in the human bodies captured in the first view image 10 and each second view image 20, but there are also a certain number of human key points of the same type.
[0040] Project each 3D human key point corresponding to the first view image 10 into a 2D human key point, where the same 3D human key points are respectively projected into 2D human key points corresponding to different second perspectives. In this way, for the human key points that appear in both the first view image 10 and the corresponding second view image 20, after being projected to the same perspective, the theoretically same 2D human key points should coincide.
[0041] Based on the camera parameters in the 3D human pose information of the first view image 10, the coordinates of each 3D human key point of the first view image 10 can be converted into the coordinates of 2D human key points. Specifically, using the external parameters in the camera parameters, the position coordinates of each 3D human key point are converted into camera coordinates. Then, using perspective projection, the camera coordinates are converted into normalized image coordinates. Then, using the camera internal parameters of the camera that captured the first view image 10, the normalized image coordinates are converted into the pixel coordinates of the corresponding 2D human key point coordinates. Thus, the coordinates of each 3D human key point corresponding to the first view image 10 can be converted one by one into the pixel coordinates of 2D human key points.
[0042] The 2D human key points obtained by projecting the 3D human key points corresponding to the first view image 10 into different second perspectives are different, and then a total projection error function is constructed between the 2D human key points obtained by projecting the 3D human key points into each second perspective and the 2D human key points extracted from the corresponding second perspective.
[0043] Optionally, determining the total projection error function between the 2D human key points obtained by projecting the 3D human key points into each second perspective and the 2D human key points extracted from the corresponding second perspective includes: obtaining the 2D human key points projected into each second perspective by projecting the 3D human key points onto the 2D planes corresponding to each second perspective; constructing the total projection error function based on the sum of the errors between the projected 2D human key points corresponding to each second perspective and the extracted 2D human key points, with the pose parameters, shape parameters of the human body, and camera parameters of different perspectives as variables.
[0044] The 3D human body key points corresponding to the first view image 10, after being projected onto the second view angle corresponding to the target second view image 20, should theoretically coincide with the 2D human body key points obtained by corresponding projection and the 2D human body key points extracted from the target second view image 20. However, since the human body pose estimation model is a model that needs to be trained, the accuracy of the 3D human body pose information output by the human body pose estimation model may not be high enough. Correspondingly, there is a deviation between the 2D human body key points corresponding to the projection of the first view image 10 and the 2D human body key points extracted from the second view image 20 corresponding to the second view angle. Summing the deviations corresponding to multiple second view images 20, the total projection error function is obtained.
[0045] The total projection error function takes the parameters corresponding to the output of the human body pose estimation model, that is, the pose parameters, shape parameters, and camera parameters of the human body, as variables. The optimization goal is to minimize the value of the total projection error. In this way, the smaller the total projection error, the better the solutions obtained for the pose parameters, shape parameters, and camera parameters of the human body, which are the variables, that is, the optimized 3D human body pose information. the optimized 3D human body pose information The corresponding view image 60 of Figure 2 is shown as
[0046] In one embodiment, the fused 2D human body key points 50 can be implemented based on SMPLify to optimize the 3D human body pose information of the first view image 10.
[0047] Step 108, train the human body pose estimation model based on the optimized 3D human body pose information.
[0048] Specifically, the training of the human body pose estimation model based on the optimized 3D human body pose information includes: determining an estimation loss function based on the optimized 3D human body pose information and the 3D human body pose information output by the human body pose estimation model; adjusting the parameters of the human body pose estimation model until the estimation loss function is minimized, and training to obtain the human body pose estimation model.
[0049] By combining the 2D human body key points extracted from multiple second view images 20, the 3D human body pose information estimated by the human body pose estimation model 30 can be optimized and compared with the 3D human body pose information originally estimated by the human body pose estimation model 30 to determine whether the current human body pose estimation model 30 can achieve high-precision 3D human body pose information estimation.
[0050] Combining Figure 2 the optimized 3D human body pose information and the 3D human body pose information output by the human body pose estimation model 30 The smaller the difference or loss, the higher the estimation accuracy of the human body pose estimation model 30. Taking the parameters of the human body pose estimation model 30 as variables and minimizing the estimation error function as the optimization goal, the human body pose estimation model is trained. Adjust the parameters of the human body pose estimation model 30 and loop in sequence until the estimation loss function is minimized, that is, the training of the human body pose estimation model with the current first view image 10 as a sample is completed.
[0051] In order to further improve the estimation accuracy of the human body pose estimation model, in one embodiment, the method further includes: sequentially selecting each view image sample from the view image sample set including the first view image and the plurality of second view images as the first view image for training the human body pose estimation model.
[0052] That is to say, the view image sample set includes a plurality of view images taken from different perspectives. Sequentially select one view image as the first view image, and at the same time select some or all of the other view images as the second view images, and train the human body pose estimation model in the manner of steps 102 to 108.
[0053] After the training of the current first view image is completed, then select any other view image different from the current first view image as the first view image, and at the same time select some or all of the other view images as the second view images, and train the human body pose estimation model again in the manner of steps 102 to 108 until all the view images in the view image sample set are trained. At this time, the finally trained human body pose estimation model is obtained.
[0054] Combined with Figure 2 , the corresponding loss function is expressed as , where N represents the number of view images in the view image sample set.
[0055] The human body pose estimation model can be a convolutional neural network (CNN). The CNN regresses the human body model parameters (i.e., 3D human pose information) from a single view image, and then simplifies the human body model parameters regressed by the CNN by fusing the 2D human key points extracted from multiple view images. The SMPLify is used to optimize the CNN, and the SMPL model is fitted to the 2D human key points predicted by the CNN from a single view image. The optimized human body model parameters are obtained from other view images. The optimized human body model parameters are used to supervise the training of the CNN, so that the CNN can regress better parameters.
[0056] By integrating CNN and SMPLify of multiple second-view images into a new training loop path, the initial human body model parameters regressed by CNN based on a single-view image are combined with the 2D key points obtained from multiple second-view images through SMPLify, and multiple training loops are performed to optimize CNN.
[0057] In the test, by comparing the 3D human body pose information estimated by the single-view and the multi-view of the embodiment of the present application, using multi-view images can provide a deeper understanding of human joints, without being restricted by limbs or environmental elements. The human body pose estimation model of the embodiment of the present application has robustness and competitive advantages, and can better reconstruct the 3D human body model under different environments and complex human body postures.
[0058] In the embodiment of the present application, by inputting the first-view image into the human body pose estimation model to output the 3D human body pose information corresponding to the first-view image, the first-view image is an image obtained by photographing the human body from the first perspective; 2D human body key points of the human body are respectively extracted from multiple second-view images, the second-view corresponds to an image obtained by photographing the human body from the second perspective, the second perspective is different from the first perspective, and the second perspectives of the multiple second-view images are different from each other; the 3D human body pose information corresponding to the first-view image is optimized based on the 2D human body key points to obtain optimized 3D human body pose information; based on the optimized 3D human body pose information, the human body pose estimation model is trained, whereby a human body pose estimation model with high estimation accuracy can be trained time-saving and effectively.
[0059] The embodiment of the present application also provides a human body pose estimation method, which can use the human body pose model trained by the training method based on the embodiment of the present application for human body pose estimation.
[0060] The human pose estimation model trained in the embodiments of the present application can be applied to the field of agricultural machinery, and can well detect the positions and movements of traditional agricultural machinery operators and autonomous driving agricultural machinery safety officers, and has good applications in automated operations, human-machine collaboration, and safety warning systems. Through pose estimation, the poses and movements of operators or safety officers can be monitored in real time to ensure that they maintain correct poses when driving and operating agricultural machinery, thereby improving safety and work efficiency. Combining pose estimation technology, agricultural machinery can better understand and adapt to the movements of operators, optimize the automated control system, and improve the accuracy and efficiency of operations. In a complex agricultural environment, pose estimation can optimize human-machine collaboration. By recognizing and predicting the intentions of operators, agricultural machinery can better cooperate with humans to complete tasks. By monitoring the states of operators and safety officers, pose estimation can issue alarms when detecting fatigue or improper poses, reducing the risk of accidents. In agricultural machinery training, pose estimation can be used to provide real-time feedback on the operations of trainees, helping them correct their movements and improve learning effects.
[0061] In addition, it can also be applied to human pose estimation and has extensive applications in virtual / augmented reality (VR / AR) and computer games, and can also be used in the film industry. In sports analysis and training: in the sports field, it can help athletes analyze their movements to optimize training effects and reduce the risk of injury; in health care, in rehabilitation therapy, this technology can monitor the movements of patients, helping doctors evaluate the progress of rehabilitation and develop personalized treatment plans; in virtual reality and augmented reality: in VR and AR applications, this technology can enable users to interact with the virtual environment, enhancing the sense of immersion and experience; in human-computer interaction: in smart homes and smart devices, pose estimation can be used for gesture recognition, thereby achieving touchless operation and improving the user experience; in security monitoring: in monitoring systems, this technology can detect suspicious behaviors and enhance public safety; in social media and entertainment: in social applications, pose estimation can be used to generate animated emojis or filters for users, enhancing the interactive fun; in autonomous driving: in autonomous driving technology, this technology can help identify pedestrians and other traffic participants to improve safety.
[0062] Optionally, as Figure 3 shown, the embodiments of the present application further provide a training device 2000 for a human pose estimation model, including a processor 2400 and a memory 2200. The memory 2200 stores a program or instruction that can run on the processor 2400. When the program or instruction is executed by the processor 2400, it implements each step of the training method embodiment of the above human pose estimation model and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0063] The embodiments of the present application further provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned embodiments of any method for training a human body pose estimation model, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the readable storage medium includes a computer-readable storage medium, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0064] The embodiments of the present application further provide a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to cause a computer to implement each process of the above-mentioned embodiments of any method for training a human body pose estimation model when executed, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0065] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.
[0066] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present application.
[0067] The above has described the embodiments of the present application in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. A training method for a human body pose estimation model, characterized in that, Including: Inputting a first view image into a human body pose estimation model to output 3D human body pose information corresponding to the first view image, where the first view image corresponds to an image of a human body taken from a first perspective; Extracting 2D human body key points of the human body from multiple second view images respectively, where the second view corresponds to an image of the human body taken from a second perspective different from the first perspective, and the second perspectives of the multiple second view images are different from each other; Optimizing the 3D human body pose information corresponding to the first view image based on the 2D human body key points to obtain optimized 3D human body pose information; Training the human body pose estimation model based on the optimized 3D human body pose information.
2. The method according to claim 1, characterized in that, The 3D human body pose information includes pose parameters, shape parameters, and camera parameters of the human body.
3. The method according to claim 2, wherein The optimizing the 3D human body pose information corresponding to the first view image based on the 2D human body key points to obtain optimized 3D human body pose information includes: Determining 3D human body key points corresponding to the 3D human body pose information of the first view image; Determining a total projection error function between the 2D human body key points obtained by projecting the 3D human body key points onto each second perspective and the 2D human body key points extracted from the corresponding second perspective, where the total projection error function is determined based on the pose parameters, shape parameters, and camera parameters of the human body; Based on the total projection error function and preset constraint conditions, constructing an optimization model of the 3D human body pose information with minimizing the total projection error function as the optimization goal; Performing an optimal solution to the optimization model to obtain the optimized 3D human body pose information.
4. The method according to claim 3, characterized in that The determining 3D human body key points corresponding to the 3D human body pose information of the first view image includes: Inputting the 3D human body pose information corresponding to the first view image into an SMPL model to obtain 3D human body key points corresponding to the 3D human body pose information.
5. The method according to claim 3, wherein The determining the total projection error function between the 2D human body key points obtained by projecting the 3D human body key points onto each second perspective and the 2D human body key points extracted from the corresponding second perspective includes: Obtaining 2D human body key points projected onto each second perspective by projecting the 3D human body key points onto the 2D planes corresponding to each second perspective respectively; Based on the pose parameters, shape parameters, and camera parameters of different perspectives as variables, constructing the total projection error function based on the sum of errors between the projected 2D human body key points corresponding to each second perspective and the extracted 2D human body key points.
6. The method according to claim 1, characterized in that, The training the human body pose estimation model based on the optimized 3D human body pose information includes: Determining an estimation loss function based on the optimized 3D human body pose information and the 3D human body pose information output by the human body pose estimation model; Adjusting the parameters of the human body pose estimation model until the estimation loss function is minimized, and training to obtain the human body pose estimation model.
7. The method according to claim 6, characterized in that, Also including: From the view image sample set including the first view image and the plurality of second view images, each view image sample is sequentially selected as the first view image for training the human pose estimation model.
8. A training device for a human body pose estimation model, characterized in that, Comprising a processor and a memory, the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method according to any one of claims 1-7 are implemented.
9. A readable storage medium, characterized in that, Programs or instructions are stored on the readable storage medium, and when the programs or instructions are executed by a processor, the steps of the method according to any one of claims 1-7 are implemented.
10. A computer program product, characterized in that, The computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps of the method according to any one of claims 1-7.