Human motion capture model training method and remote human body rendering method

By combining the Gaussian sphere learning model and the human body rendering model, the Gaussian sphere parameters are automatically learned, which solves the problem of inaccurate control point coordinates in the human body 3D motion capture model, realizes the generation of high-quality human body motion rendering images, and improves the effect of immersive meetings.

CN121120684APending Publication Date: 2025-12-12GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410748258.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-11
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

The control point coordinates of existing 3D motion capture models for humans lack precision, resulting in inaccurate images and affecting the effectiveness of immersive meetings.

Method used

By employing a Gaussian sphere learning model and a human body rendering model, and by acquiring target video data and human body control information, the Gaussian sphere parameters are automatically learned to constrain the human body rendering model, thereby generating high-quality human motion rendering images.

Benefits of technology

It improves the accuracy and rendering effect of human motion capture models, reduces computational complexity, increases the speed and real-time performance of image generation, and enhances the interactivity of immersive meetings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120684A_ABST
    Figure CN121120684A_ABST
Patent Text Reader

Abstract

The invention discloses a training method of a human body motion capture model, a remote human body rendering method and electronic equipment. By acquiring the target video data, applying the human body motion capture model and the human body rendering model, and automatically learning the output Gaussian ball parameters by using the Gaussian ball learning model, the high-quality and accurate human body motion rendering image can be generated, and the deviation of the human body control information on the image generation is made up, so that the performance of the model is improved, and the user experience is improved. The generated image better conforms to the real human body action; the calculation complexity can be reduced, the image generation speed can be increased, and the real-time performance and the interactivity can be improved through the application of Gaussian ball parameters; by acquiring the video data of the target human body in different postures and the human body control information corresponding to the target human body action image, the human body action rendering image can be more accurately generated, the postures and actions of the human body are better captured and presented by using the human body action capturing model and the human body rendering model, and the accuracy of image generation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a human action capture model training method, a remote human rendering method and an electronic device. BACKGROUND

[0002] An immersive meeting is a remote video conference with added sense of presence and sense of immersion, which aims to provide participants with a more realistic and immersive meeting experience, as if the participants were in the same conference room. The main body of the immersive meeting is the human body, and it is important to accurately convey important action information of the participants by 3D action capture of the human body.

[0003] Since there is no biological control point in the real physical world, 3D action capture of the human body can only be performed through a model artificially constructed, which leads to certain challenges in the accuracy of the control point coordinates. In this case, even the control point coordinates labeled by humans as the "ground truth" cannot guarantee complete accuracy. Different labelers may get different results, and the same labeler may get different results at different times or when labeling different pictures. This means that the model trained based on such labeling results in the existing scheme cannot ensure the accuracy and stability of the extracted control points. For a generated rendering model, if the results of the control points are biased, the generated images will be problematic, which has a great impact on the immersive meeting. SUMMARY

[0004] The technical problem solved by the embodiments of the present application is how to improve the quality of human action capture images and the rendering effect.

[0005] To solve the above technical problems, one technical solution adopted by the embodiments of the present application is to provide a human action capture model training method, comprising: obtaining target video data, the target video data comprising video data of a target human body in different postures, the video data comprising a plurality of original human action images, the human action capture model comprising a Gaussian sphere learning model and a human body rendering model; obtaining human body control information corresponding to a target human action image, the target human action image being an original human action image in a training state; obtaining Gaussian sphere parameters output by the Gaussian sphere learning model when the target human body image is used as a training input sample, the Gaussian sphere parameters being used to constrain a Gaussian sphere of the human body rendering model; inputting the Gaussian sphere parameters and the human body control information into the human body rendering model to obtain a human action rendering image; determining image difference information between the human action rendering image and the target human action image; and training the human action capture model according to the image difference information.

[0006] The method can generate high-quality and accurate human motion rendering images by obtaining target video data, applying a human motion capture model and a human rendering model, and automatically learning Gaussian sphere parameters output by a Gaussian sphere learning model. This automatic learning and parameter constraint method can compensate for the deviation of human control information on image generation, thereby improving the performance of the model and making the generated images more consistent with real human motion. Moreover, the application of Gaussian sphere parameters can reduce computational complexity, speed up image generation, and improve real-time performance and interactivity. In addition, by obtaining video data of a target human in different poses and human control information corresponding to the target human motion image, the method can more accurately generate human motion rendering images. By using a human motion capture model and a human rendering model, the pose and motion of the human body can be better captured and presented, improving the accuracy of image generation.

[0007] In some embodiments, the Gaussian sphere parameters include Gaussian sphere rendering parameters and position offset parameters, the Gaussian sphere learning model includes a neural network model, the human rendering model includes an SMPL human model and a 3D Gaussian rendering model, and the Gaussian sphere parameters output by the Gaussian sphere learning model when the target human image is used as a training input sample include: determining the vertex features of the vertices in the SMPL human model that are used to associate with the pose of the target human when the target human motion image is used as a training input sample; inputting the vertex features of the vertices into the neural network model to obtain Gaussian sphere rendering parameters and position offset parameters, the Gaussian sphere rendering parameters being used to represent the rendering attributes of the Gaussian sphere of the 3D Gaussian rendering model, and the position offset parameters being used to indicate the position offset of the Gaussian sphere of the 3D Gaussian rendering model. By using the neural network model to obtain the Gaussian sphere parameters output by the Gaussian sphere learning model, including the Gaussian sphere rendering parameters and the position offset parameters, accurate pose association, flexible rendering attribute control, and enhanced rendering effect can be achieved.

[0008] In some embodiments, the determination of the vertex features of the vertices in the SMPL human model that are used to associate with the pose of the target human includes: obtaining control points of the SMPL human model, the control points being key points of specific joint positions in the SMPL human model; performing pose transformation and morphological adjustment on the key points to obtain vertex positions of the vertices in the SMPL human model that are used to associate with the pose of the target human; and performing feature sampling operations on the vertex positions of the vertices in a pre-set sampling space to obtain the vertex features of the vertices. This method determines the vertex features of the vertices in the SMPL human model that are used to associate with the pose of the target human, thereby more accurately capturing and representing the pose information of the target human.

[0009] In some embodiments, the inputting the vertex feature of the vertex into the neural network model to obtain the Gaussian sphere rendering parameter and the position offset parameter comprises: inputting the vertex feature of the vertex into the neural network model to output the Gaussian sphere rendering parameter composed of spherical harmonics, opacity, a rotation matrix and a scale vector, and the position offset parameter composed of an offset vector of a Gaussian sphere center point and a weight value of the vertex. The Gaussian sphere rendering parameter can more accurately describe and control the lighting, transparency and deformation effect of the human body surface, and improve the realism and detail presentation of the rendering.

[0010] In some embodiments, the inputting the Gaussian sphere parameter and the human body control information into the human body rendering model to obtain the human body action rendering image comprises: determining a target position of a target vertex after a change of a target human body according to the position offset parameter and the human body control information, the target vertex being a vertex in the target human body action image; inputting the target position of each target vertex and the Gaussian sphere rendering parameter into the SMPL human body model to obtain the rendering parameter of a Gaussian sphere corresponding to each target vertex; and inputting the rendering parameter of the Gaussian sphere corresponding to each target vertex into the 3D Gaussian rendering model to obtain the human body action rendering image. By inputting the Gaussian sphere parameter and the human body control information into the human body rendering model, the position of the target vertex can be determined, the rendering parameter of the Gaussian sphere can be calculated, and finally the human body action rendering image can be generated. This method can improve the accuracy and realism of the rendering image, and make the generated image more consistent with the posture and action of the target human body.

[0011] In some embodiments, the determining the target position of the target vertex after the change of the target human body according to the position offset parameter and the human body control information comprises: updating the vertex position of the target vertex according to the position offset parameter to obtain a static position of the target vertex before the change of the action of the target human body; and moving the static position according to the human body control information to obtain the target position of the target vertex after the change of the action of the target human body. By updating the vertex position of the target vertex according to the position offset parameter and moving the static position according to the human body control information, the effect of retaining the static position and capturing the target position after the change of the action can be achieved. This method can ensure that the rendering result is consistent with the change of the action of the target human body, and improve the realism and accuracy of the rendering.

[0012] In some embodiments, the position offset parameter comprises an offset vector of a Gaussian sphere center point, and the updating the vertex position of the target vertex according to the position offset parameter to obtain the static position of the target vertex before the action of the target human body changes comprises: adding the coordinate corresponding to the vertex position of the target vertex and the offset vector of the Gaussian sphere center point to obtain the static position of the target vertex before the action of the target human body changes. In this way, the initial position of the target vertex can be retained, and the correction of the Gaussian sphere to the position is considered, so that the rendering result is more accurate and realistic.

[0013] In some embodiments, the human body control information comprises a camera pose when the target human body action image in the target video data is acquired, and a key point associated with a specific joint position of a human body in the target human body action image; the position offset parameter comprises a weight value of a vertex; and the moving the static position according to the human body control information to obtain the target position of the target vertex after the action of the target human body changes comprises: updating the position of the target vertex of the static position according to the camera pose, the coordinates of the key point, and the weight value of the vertex to obtain the target position of the target vertex after the action of the target human body changes. In this way, the action change of the target human body can be accurately captured by combining the camera perspective, the key point information, and the vertex weight, and the corresponding position update can be generated to achieve a more realistic and accurate rendering effect.

[0014] In some embodiments, the image difference information comprises a loss value, and the training the human body action capture model according to the image difference information comprises: optimizing the neural network parameters in the neural network model and the parameters of the SMPL human body model by using a gradient descent algorithm according to the loss value. By optimizing the neural network parameters in the neural network model and the parameters of the SMPL human body model by using the gradient descent algorithm, the model can be better fitted to the training data, and the performance and generalization ability of the model can be improved.

[0015] To solve the above technical problems, one technical scheme adopted by the embodiment of the present application is to provide a remote human body rendering method applied to an electronic device, comprising: receiving real-time human body control information of a target user sent by a target device when the electronic device establishes a conference connection with the target device, the target user being a user participating in the conference through the target device; inputting the real-time human body control information into a human body motion capture model corresponding to the target user, so that the human body motion capture model generates a human body motion rendering image of the target user; and controlling a conference interface of the electronic device to present the human body motion rendering image; wherein the human body motion capture model is obtained by training the human body motion capture model training method as described above. The method can realize real-time rendering of the human body motion of the target user in the remote conference by receiving real-time human body control information, inputting the human body motion capture model, generating a human body motion rendering image, and presenting it on the conference interface, thereby improving the interactivity and communication effect of the conference, and enabling the participants to more intuitively perceive the motion state of the target user. Meanwhile, by using the human body motion capture model as described above, the accuracy and realism of the rendering can be improved.

[0016] To solve the above technical problems, one technical scheme adopted by the embodiment of the present application is to provide an electronic device, comprising: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the human body motion capture model training method as described above, or execute the remote human body rendering method as described above. The electronic device has the same beneficial effects as the above method. BRIEF DESCRIPTION OF DRAWINGS

[0017] One or more embodiments are illustrated by way of example with reference to the accompanying drawings, which are schematic and not intended to be limiting of the embodiments, and in which like reference numerals designate similar elements in the figures and wherein the use of "for example", "e.g.", "of the example", "for instance", "e.g." and "in an example" herein indicates that various embodiments can comprise similar features, and each of the similar features is not necessarily the same. The figures are not necessarily to scale and the emphasis is on the functional operation of each element.

[0018] Figure 1 A system architecture schematic diagram of a conference system provided by the embodiment of the present application;

[0019] Figure 2 A structure schematic diagram of modules for completing 3D human body image related to the electronic device or the target device provided by the embodiment of the present application;

[0020] Figure 3 An interaction schematic diagram between the electronic device and the target device provided by the embodiment of the present application;

[0021] Figure 4An interaction diagram in which an electronic device provided by an embodiment of the present application and a target device exchange a human motion capture model;

[0022] Figure 5a A diagram in which an electronic device, a target device and a server provided by an embodiment of the present application participate in a conference;

[0023] Figure 5b A conference interface diagram provided by an embodiment of the present application;

[0024] Figure 6 A system architecture diagram of a conference system provided by another embodiment of the present application;

[0025] Figure 7 A flowchart of a training method of a human motion capture model provided by an embodiment of the present application;

[0026] Figure 8 A diagram of positions of 24 key points of a standard SMPL human model shown in an embodiment of the present application;

[0027] Figure 9 A process diagram of training a human motion capture model provided by an embodiment of the present application;

[0028] Figure 10 A diagram of learning Gaussian sphere parameters when a neural network model is a multi-layer MLP neural network provided by an embodiment of the present application;

[0029] Figure 11 A diagram of learning Gaussian sphere parameters when a neural network model is a plurality of multi-layer MLP neural networks provided by an embodiment of the present application;

[0030] Figure 12 A flowchart of a remote human rendering method provided by an embodiment of the present application;

[0031] Figure 13 A hardware structure diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0032] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0033] It should be noted that, unless otherwise specified, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device schematic diagram or the order in the flowchart.

[0034] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application.

[0035] SMPL (Skinned Multi-Person Linear) is a commonly used human shape and pose model. It's a parametric model based on linear algebra used to represent the shape and pose of the human body. SMPL describes the human's pose and shape using a set of parameters, including global rotation, global scale, joint angles, and deformation weights. The core idea of ​​SMPL is to represent the human body as a skeletal structure connected by joints, and to map the surface mesh onto this skeletal structure using linear weights. By adjusting the joint angles and deformation weights, human models with different poses and shapes can be generated. The parametric representation of the SMPL model allows it to adapt well to variations in the shape and pose of different human bodies.

[0036] Control points in a Smart Joint Model (SMPL) are key points associated with specific joint positions in the human body, used to describe posture and movement. The accuracy of the control point coordinates directly affects the transmission of important human movement information. Since biologically accurate control points do not exist in the real physical world, these control points are merely points on an artificially constructed SMPL model. Therefore, they do not possess completely accurate coordinates in a biologically anatomical sense. Even manually annotated control point coordinates, considered the "gold standard" (Gold Standard), cannot guarantee complete accuracy. Different annotators may obtain different results, and the same annotator may obtain different results at different times or when annotating different images. This means that existing models trained based on such annotation results cannot guarantee the accuracy and stability of the extracted control points. For generative rendering models, deviations in the control point results will lead to problems in the generated images, which has a significant impact on applications such as immersive meetings, virtual character animation, and virtual try-on.

[0037] To address the aforementioned issues, multiple cameras can be used to capture images simultaneously. Analyzing the data from these cameras yields more stable and accurate results, such as the Light Stage solution using dozens or even hundreds of synchronized cameras. However, this approach is difficult to implement in immersive meeting scenarios. Both cost and scenario limitations may prevent the deployment of multi-camera synchronization equipment like Light Stage. The Light Stage solution places a person within a spherical platform surrounded by dozens or even hundreds of cameras. These cameras trigger simultaneously to capture images of the person from different angles, while a lighting system illuminates the person under varying conditions to obtain images from different lighting environments.

[0038] Due to the high cost and large scale of the equipment, as well as the requirements for site and technical support, the relevant technologies cannot effectively solve the problems of inaccurate human motion capture and poor image rendering.

[0039] Therefore, this application provides a training method for a human motion capture model. This method acquires target video data, applies a human motion capture model and a human rendering model, and automatically learns the output Gaussian sphere parameters using a Gaussian sphere learning model. This enables the generation of high-quality, accurate human motion rendering images. This automated learning and parameter constraint method can compensate for deviations in image generation caused by human control information, thereby improving model performance and making the generated images more consistent with real human movements. Furthermore, the application of Gaussian sphere parameters reduces computational complexity, accelerates image generation, and improves real-time performance and interactivity. In addition, by acquiring video data of the target human body in different postures and the human control information corresponding to the target human motion images, this method can generate human motion rendering images more accurately. By using a human motion capture model and a human rendering model, it can better capture and present human postures and movements, improving the accuracy of image generation.

[0040] Please see Figure 1 This application provides a conference system 100, which includes an electronic device 200 and a target device 300, wherein the electronic device 200 and the target device 300 establish a conference connection. The electronic device 200 and the target device 300 can join the same remote conference to participate in the meeting. The conference interface of the electronic device 200 can display the human image of user A operating the target device 300, and similarly, the conference interface of the target device 300 can display the human image of user B operating the electronic device 200.

[0041] The human body image presented by electronic device 200 and the human body image presented by target device 300 are both 3D human body images. Since 3D human body images are more three-dimensional than 2D human body images, they are easier to create an immersive experience for users, making it easier for users to immerse themselves in the meeting atmosphere.

[0042] The 3D human body image can be synthesized from multiple rendered human body images by electronic device 200 or target device 300. The rendered human body images can be generated by electronic device 200 or target device 300 using corresponding human motion capture models. It is understood that both electronic device 200 and target device 300 can train their respective human motion capture models using the training methods for human motion capture models described in the various embodiments below.

[0043] Please refer to the following: Figure 2 and Figure 3 Both the electronic device 200 and the target device 300 include an N-channel camera module 41, an ISP module 42, a video decoding module 43, a 3D model generation module 44, a model transmission module 45, a 3D rendering module 46, a 3D compositing module 47, a 3D playback module 48, and a SOC processing module 49.

[0044] The camera module 41 is used to collect video data of the user's body in different postures and send the video data to the ISP module 42. The ISP module 42 processes the video data to improve image clarity, color reproduction, dynamic range, etc., and enhance the image's resolution in low light conditions.

[0045] The video decoding module 43 is used to restore the processed video data to obtain the original video image and audio signal. The 3D model generation module 44 is used to train a human motion capture model based on the human motion capture model described in the various embodiments below. The model transmission module 45 is used to send the human motion capture model to a designated device, which may be the device of the other party's participant or a server.

[0046] The 3D rendering module 46 uses a human motion capture model to process user feature information, thereby outputting a rendered human image. The 3D compositing module 47 combines multiple rendered human images into a stereoscopic human image. The 3D playback module 48 combines stereoscopic images and audio signals into a video stream. The SOC processing module 49 processes the video stream to present stereoscopic images and audio signals.

[0047] It is understood that the camera module 41 can be any type of camera, the ISP module 42 can be a separate ISP chip, or the ISP module 42 and the video decoding module 43 can be integrated on the same video processing chip. The 3D model generation module 44 can be a 3D graphics card processor, etc. The model transmission module 45 can be a wired communication module or a wireless communication module. The 3D rendering module 46, the 3D compositing module 47, the 3D playback module 48, and the 3D playback module 49 can be various types of processors, such as DSP processors. The SOC processing module 49 can be a SOC processing chip.

[0048] In some embodiments, the electronic device 200 establishes an end-to-end communication connection with the target device 300. Based on this end-to-end communication connection, the electronic device 200 and the target device 300 exchange human motion capture models for local storage. That is, the electronic device 200 can send its local human motion capture model to the target device 300 for local storage via the end-to-end communication connection, and the target device 300 can send its local human motion capture model to the electronic device 200 for local storage via the same connection.

[0049] For example, please see Figure 4 User A joins remote conference A using electronic device 200, and User B joins remote conference A using target device 300. Electronic device 200 trains and obtains a motion capture model J1 for User A, and target device 300 trains and obtains a motion capture model J2 for User B. Electronic device 200 and target device 300 can communicate end-to-end. Electronic device 200 sends User A's motion capture model J1 to target device 300, and target device 300 stores User A's motion capture model J1 locally. Target device 300 sends User B's motion capture model J2 to electronic device 200, and electronic device 200 stores User B's motion capture model J2 locally.

[0050] During the meeting, electronic device 200 captures real-time images of user A's body using a camera, obtaining a real-time human body image X1. Electronic device 200 then determines user A's human body control information based on the real-time human body image X1 and sends this information to target device 300. Target device 300 uses human motion capture model J1 to process user A's human body control information, thereby obtaining a baseline human body rendering image of user A. Target device 300 generates multiple parallax human body rendering images based on the baseline human body rendering image, and then synthesizes the baseline human body rendering image and the multiple parallax human body rendering images into a 3D human body image. Finally, target device 300 can display the 3D human body image of user A on its local meeting interface.

[0051] Similarly, target device 300 captures real-time images of user B's body using a camera, obtaining a real-time human body image X2. Then, target device 300 determines user B's human body control information based on the real-time human body image X2 and sends this information to electronic device 200. Electronic device 200 uses human motion capture model J2 to process user B's human body control information, thereby obtaining a baseline human body rendering image of user B. Finally, electronic device 200 generates a 3D human body image of user B according to the above procedure and displays the 3D human body image of user B on the local conference interface.

[0052] In related technologies, to display a 3D human image of a target device 300 on an electronic device 200, the technology directly controls the target device 300 to transmit N video streams to the electronic device 200. The electronic device 200 then processes the N video streams into a 3D human image for display. However, based on the transmission requirements of a video stream resolution of 1920*1080 and a frame rate of 30fps, the bandwidth required to transmit one video stream is: 1920*1080*3*8 / 1024 / 1024 / 1024*30 frames*1 stream = 1.39Gbps. If N=4, meaning the target device 300 transmits four video streams to the electronic device 200, then the technology requires 5.56Gbps of bandwidth. Therefore, the technology requires a large bandwidth to achieve the purpose of displaying a 3D human image. However, many data transmission scenarios cannot support such a high bandwidth data transmission, which significantly limits the application scope of the technology.

[0053] In this embodiment of the application, the amount of human body control information is relatively small, usually less than 1 Mbps. Therefore, the method provided in this embodiment of the application occupies less bandwidth, which is beneficial to broadening the application scope of the method provided in this embodiment of the application.

[0054] In some embodiments, please refer to the following: Figure 5a and Figure 5b Electronic device 200 and target device 300 establish communication connections with server 400 respectively, and server 400 can forward the human motion capture model of the other party to the other party's device.

[0055] Please continue reading. Figure 5a and Figure 5b When electronic device 200 and target device 300 join the same remote conference, server 400 can send the other party's human motion capture model to the other party's device.

[0056] Electronic device 200 sends user A's human motion capture model J1 to server 400 for storage, and target device 300 sends user B's human motion capture model J2 to server 400 for storage.

[0057] When server 400 detects that user A's user information and user B's user information appear on the same remote conference participant list, where user A's user information is bound to electronic device 200 and user B's user information is bound to target device 300, server 400 sends human motion capture model J1 to target device 300 and human motion capture model J2 to electronic device 200.

[0058] During the meeting, electronic device 200 captures real-time images of user A's body using a camera, obtaining a real-time human body image X1. Then, electronic device 200 determines user A's human body control information based on real-time human body image X1 and sends this information to server 400. Similarly, target device 300 captures real-time images of user B's body using a camera, obtaining a real-time human body image X2. Then, target device 300 determines user B's human body control information based on real-time human body image X2 and sends this information to server 400.

[0059] Server 400 forwards user B's human body control information to electronic device 200, and forwards user A's human body control information to target device 300.

[0060] The target device 300 calls the human motion capture model J1 to process the human control information of user A, and synthesizes a 3D human image according to the above method, and presents the 3D human image of user A on the local conference interface. Similarly, the electronic device 200 calls the human motion capture model J2 to process the human control information of user B, and generates a 3D human image of user B according to the above method, and presents the 3D human image of user B on the local conference interface.

[0061] It is understood that in some embodiments, the human motion capture model can be generated before the meeting begins. As mentioned above, electronic device 200 generates human motion capture model J1 for user A in advance, and target device 300 generates human motion capture model J2 for user B in advance. Then, before the meeting begins, electronic device 200 sends human motion capture model J1 to target device 300 for local storage, and target device 300 sends human motion capture model J2 to electronic device 200 for local storage. Alternatively, before the meeting begins, electronic device 200 sends human motion capture model J1 to server 400, target device 300 sends human motion capture model J2 to server 400, server 400 forwards human motion capture model J1 to target device 300 for local storage, and forwards human motion capture model J2 to electronic device 200 for local storage.

[0062] It is also understood that, in some embodiments, the human motion capture model can be generated at the start of the meeting. As mentioned earlier, at the start of the meeting, User A uses the camera of electronic device 200 to capture User A's body and records a first video data of a preset duration. User B uses the camera of target device 300 to capture User B's body and records a second video data of a preset duration. Electronic device 200 generates a human motion capture model J1 based on the first video data and sends the human motion capture model J1 directly to target device 300 for local storage or to server 400, which then forwards it to target device 300 for local storage. Simultaneously, target device 300 generates a human motion capture model J2 based on the second video data and sends the human motion capture model J2 directly to electronic device 200 for local storage or to server 400, which then forwards it to electronic device 200 for local storage.

[0063] In some embodiments, the video data used to generate the human motion capture model can be acquired by the camera carried by the electronic device 200 or the Mubao device 300 itself.

[0064] In some embodiments, the video data used to generate the human motion capture model can be acquired by cameras set up at the conference venue. Please refer to [link to relevant documentation]. Figure 6 The conference system 100 includes a first camera module 500 and a second camera module 600. The first camera module 500 includes N first cameras, which are communicatively connected to the electronic device 200. The second camera module 600 includes M second cameras, which are communicatively connected to the target device 300, where N and M are both positive integers.

[0065] User A and User B are conducting a remote video conference. When User A enters the conference room, the first camera module 500 captures video data of User A's body in different postures, and then transmits this video data to the electronic device 200. The electronic device 200 generates a human motion capture model J1 based on the video data. Similarly, when User B enters the conference room, the second camera module 600 captures video data of User B's body in different postures, and then transmits this video data to the target device 300. The target device 300 generates a human motion capture model J2 based on the video data.

[0066] Please see Figure 7 , Figure 7 This is a flowchart illustrating a training method for a human motion capture model provided in an embodiment of this application. The method includes the following steps:

[0067] S71. Acquire target video data, which includes video data of the target human body in different poses. The video data includes multiple original human motion images. The human motion capture model includes a Gaussian ball learning model and a human rendering model.

[0068] The target human body is the user's body portion from which motion rendering images are generated, and the original motion images are images of the target human body in the corresponding poses. Multiple original motion images can be obtained from a single camera or from multiple cameras.

[0069] In a current application scenario, a single-channel camera can be set up to capture the movements of a target human body. The camera's position and angle should cover all pose changes of the target human body. Video recording begins using the single-channel camera, ensuring that the target human body's movements in different poses are captured. During this process, the target human body can be prompted to perform specific actions or poses to obtain diverse data. The recorded video can be segmented to extract segments containing the target human body's movements. This process can be done using video editing software or automated human detection and tracking algorithms to ensure that each segment covers a specific pose or movement. Key frames are selected from each segment as raw human motion images. These frames should accurately represent the appearance and form of the target human body in that pose or movement. Frame selection can be based on time intervals or according to specific criteria. The selected raw human motion images are saved to a preset folder or database, and then organized and annotated to ensure that each image is associated with a corresponding pose or movement for subsequent data processing and analysis.

[0070] When a single-channel camera collects human motion data, it can keep the camera's intrinsic parameters (such as camera focal length and field of view) unchanged. The camera's aperture, shutter speed, ISO and other parameters also remain unchanged, which helps to ensure the consistency and reliability of video data.

[0071] After obtaining the original human motion images, camera position and pose, as well as human segmentation images, can be extracted. Camera position and pose refer to the position and orientation of a single camera in three-dimensional space, i.e., its position and orientation. Camera pose estimation algorithms can be used to match and triangulate feature points in the image to estimate the position and orientation of a single camera. These feature points can be human keypoints, salient features in the scene, or other detectable features. Human segmentation is the process of separating the human body region from the background in an image. Deep learning techniques can be used for human segmentation. For example, semantic segmentation models such as U-Net and Mask R-CNN can be used to segment the human body contour and key parts from the background at the pixel level, generating human segmentation images. Through the above steps, camera position and pose and human segmentation images can be extracted from each frame of the original human motion image, resulting in a camera position and pose sequence and a human segmentation image sequence.

[0072] The target video data can be acquired before the meeting starts, at the start of the meeting, or during the meeting.

[0073] S72. Obtain human control information corresponding to the target human motion image, wherein the target human motion image is the original human motion image in the training state.

[0074] Human control information includes the camera position and pose obtained in the above steps, specifically the camera pose when acquiring the target human motion image in the target video data, i.e., the camera's position and orientation. Human control information also includes key points in the target human motion image that are associated with specific joint positions of the human body, such as... Figure 8 As shown, Figure 8 The diagram shows the locations of 24 key points on a standard SMPL human body model. Specifically, the key points in this embodiment can be... Figure 8 The 24 key points shown.

[0075] For raw human motion images in training mode, target human motion images can be labeled and prepared for training machine learning models. These images contain samples of the human body in different actions or poses. For example, target human motion images include labeled keypoints, which are points at specific joint positions of the human body, such as the body, shoulder, elbow, and knee. The positional information of these keypoints can be used to train pose estimation, motion analysis, etc.; they also include action category labels, each target human motion image can be assigned a corresponding action category label, which describes the specific action performed by the human body in the image, such as walking, running, raising an arm, etc.; they can also include temporal series information of the images. If the training data is a continuous video sequence, each target human motion image can belong to a time series, and there is temporal continuity between the images; they can also include other image attributes, such as image resolution, human body scale variation information, etc.

[0076] S73. Obtain the Gaussian sphere parameters output by the Gaussian sphere learning model when the target human image is used as a training input sample. The Gaussian sphere parameters are used to constrain the Gaussian sphere of the human rendering model.

[0077] In this step, the Gaussian sphere learning model is used to learn and output Gaussian sphere parameters. The Gaussian sphere learning model can employ any suitable model architecture, such as a neural network model or other custom learning model. The human body rendering model is used to render and output human motion rendering images. The human body rendering model is connected to the Gaussian sphere learning model, and the Gaussian sphere parameters output by the Gaussian sphere learning model can be used as an input to the human body rendering model.

[0078] S74. Input the Gaussian sphere parameters and the human body control information into the human body rendering model to obtain a human body motion rendering image.

[0079] Specifically, the target human motion image, along with its associated position, pose, and key points, is input into a pre-defined human motion capture model. The model processes and analyzes the input data, extracting the human's pose and motion information, and rendering a human motion image. Before inputting the target human image into the motion capture model, preprocessing can be performed, including image normalization and resizing. Key points can also undergo coordinate normalization or other processing to adapt to the model's input requirements.

[0080] In this embodiment of the application, the human motion capture model includes a Gaussian sphere learning model and a human rendering model. The Gaussian sphere learning model includes a neural network model, and the human rendering model includes an SMPL human model and a 3D Gaussian rendering model. The Gaussian sphere parameters include Gaussian sphere rendering parameters and position offset parameters. The above step S74 specifically includes:

[0081] S741. When the target human motion image is used as a training input sample, determine the vertex features of the vertices in the SMPL human model used to associate with the pose of the target human.

[0082] Specifically, control points are obtained from the SMPL human body model. These control points are key points at specific joint locations within the SMPL human body model (e.g., key points at joint locations). Figure 8 (As shown); pose transformation and shape adjustment are performed on the key points to obtain the vertex positions of the vertices in the SMPL human model used to associate with the pose of the target human body; feature sampling is performed in a preset sampling space based on the vertex positions to obtain the vertex features. The SMPL human model is a parameterized model based on key points. By performing pose transformation and shape adjustment on the key points, the corresponding vertex positions can be calculated. These vertex positions represent the shape of the human model under a given pose and shape. Specifically, the SMPL human model uses the Linear Blend Skin (LBS) method to calculate vertex positions. The LBS method maps the positions of control points to vertex positions by performing a weighted linear combination of control points. These weights are calculated based on the pose and shape parameters of the human model. Therefore, the vertex positions are calculated from the control points using the LBS method.

[0083] S742. Input the vertex features of the vertices into the neural network model to obtain Gaussian sphere rendering parameters and position offset parameters. The Gaussian sphere rendering parameters represent the rendering attributes of the Gaussian sphere in the 3D Gaussian rendering model, and the position offset parameters indicate the position offset of the Gaussian sphere in the 3D Gaussian rendering model. Specifically, inputting the vertex features of the vertices into the neural network model outputs Gaussian sphere rendering parameters composed of spherical harmonic functions, opacity, rotation matrices, and scale vectors, and position offset parameters composed of the offset vector of the Gaussian sphere's center point and the vertex weight values.

[0084] Gaussian sphere rendering parameters and position offset parameters can be used to control the rendering properties and position offset of the Gaussian sphere in the 3D Gaussian rendering model, thereby finely controlling the appearance, lighting and other properties of the generated human motion rendering image.

[0085] S743. Determine the target position of the target vertex after the target human body changes, based on the position offset parameter and the human body control information. The target vertex is a vertex in the target human body motion image. Specifically, the vertex position of the target vertex is updated according to the position offset parameter to obtain the static position of the target vertex before the target human body's motion changes; the static position is moved according to the human body control information to obtain the target position of the target vertex after the target human body's motion changes.

[0086] The position offset parameter includes the offset vector of the center point of the Gaussian sphere. Updating the vertex position of the target vertex according to the position offset parameter to obtain the static position of the target vertex before the action of the target human body changes includes: adding the coordinates corresponding to the vertex position of the target vertex to the offset vector of the center point of the Gaussian sphere to obtain the static position of the target vertex before the action of the target human body changes.

[0087] The human body control information includes the camera pose when acquiring the target human body motion image in the target video data, and key points in the target human body motion image associated with specific joint positions of the human body; the position offset parameter includes the weight value of the vertex; the step of moving the static position according to the human body control information to obtain the target position of the target vertex after the target human body's motion changes includes: updating the position of the target vertex in the static position according to the camera pose, the coordinates of the key points and the weight value of the vertex, to obtain the target position of the target vertex after the target human body's motion changes.

[0088] This step determines the target vertex's position after changes in the target human body based on position offset parameters and human body control information. By combining camera pose, keypoint coordinates, and vertex weights, the target vertex's position after changes in the target human body's movement can be accurately calculated, which helps maintain the accuracy and consistency of the rendered image.

[0089] S744. Input the target position and Gaussian sphere rendering parameters of each target vertex into the SMPL human body model to obtain the Gaussian sphere rendering parameters corresponding to each target vertex.

[0090] The rendering parameters of a Gaussian sphere can describe the rendering properties of the Gaussian sphere on the target human body, providing necessary information for the subsequent rendering process.

[0091] S745. Input the rendering parameters of the Gaussian sphere corresponding to each target vertex into the 3D Gaussian rendering model to obtain a human motion rendering image.

[0092] Among these methods, rendering can be performed using Gaussian sphere parameters and human body control information, which can make the generated images more realistic and lifelike. At the same time, the appearance and lighting effects of the image can be adjusted according to different Gaussian sphere parameters.

[0093] In this embodiment, an SMPL human body model, a neural network model, and a 3D Gaussian rendering model are combined. By calculating and inputting vertex features, Gaussian sphere rendering parameters, and position offset parameters, accurate control and generation of rendered human motion images are achieved. This method can improve the accuracy, realism, and personalization of rendered images, providing an effective approach for generating realistic rendered human motion images.

[0094] S75. Determine the image difference information between the rendered human motion image and the target human motion image.

[0095] The image difference information includes the loss value. Specifically, the loss function can consider three parts: the acquired real image (img)... t and rendering images t L1, L ssim and L vgg The loss function is:

[0096] loss=λ1L1+λ2L ssim +λ1L vgg

[0097] in L ssim and L vgg These are the metrics used to construct the loss function, which measure image similarity and feature loss. SSIM (Structural Similarity Index) is a metric for measuring image structural similarity, considering similarity in brightness, contrast, and structure. The SSIM loss function measures the similarity between the original and generated images by calculating the structural similarity index. VGG Loss is a feature loss based on the VGG convolutional neural network. VGG Loss measures the difference between the generated and real images by comparing their feature representations in the VGG network. λ1, λ2, and λ3 are pre-defined hyperparameters used to train the MLP neural network.

[0098] S76. Train the human motion capture model based on the image difference information.

[0099] Specifically, based on the loss value, the gradient descent algorithm is used to optimize the neural network parameters (specifically the hyperparameters of the MLP neural network) in the neural network model, as well as the parameters of the SMPL human body model.

[0100] The optimization formula for the parameters of the SMPL human body model is as follows: It refers to the parameter learning rate, and AC is the motion capture module in the SMPL human model.

[0101]

[0102] The above-mentioned optimization of the parameters of the neural network model and the SMPL human body model using the gradient descent algorithm can improve model performance, optimize network structure, improve the accuracy of the human body model, and accelerate the convergence speed of the model, thereby generating higher quality, more realistic and lifelike human motion rendering images.

[0103] To illustrate in detail the training method of the human motion capture model provided in the embodiments of this application, the embodiments of this application are combined with Figure 9 The following is a detailed explanation of this:

[0104] Please see Figure 9 This application embodiment acquires multiple original human motion images frames. t The camera pose corresponding to each original human motion image is extracted using a pre-trained visual model. t Human body segmentation image (img) t The COLMAP tool can be used to extract the camera position and pose (pose1, pose2, ... pose) for each frame. t Human body segmentation models or image matting tools can be used to obtain the corresponding human body segmentation images (img1, img2, ..., img) for each frame of the image. t .

[0105] Furthermore, a motion capture neural network module AC is introduced to obtain the three-dimensional coordinates G1, G2, ... G of the control points in the SMPL model space for each frame of the original human motion image. t G t ∈R 24×3 G t =AC(frame) t The standard SMPL model contains 24 control points, distributed as follows: Figure 8 As shown.

[0106] During training, the embodiments of this application use human body segmentation images (img) t The target human motion image was used as the basis for obtaining the Gaussian sphere learning model 91 in the human segmentation image (img). t The Gaussian sphere parameters output as training input samples will frame the original human motion image. t camera pose tThe key points and Gaussian sphere parameters of specific joint positions in the human body model are input into the human body rendering model 92 to obtain a human motion rendering image. Next, this embodiment of the application calculates the image difference information between the human motion rendering image and the target human motion image, and trains the human motion capture model based on the image difference information. Since the image difference information has not yet converged, this embodiment of the application needs to continue training the human motion capture model.

[0107] The training steps described above are analogous and will not be repeated here. When the image difference information converges to a preset threshold, this embodiment of the application stops training the human motion capture model and outputs the trained human motion capture model.

[0108] In summary, this application embodiment utilizes a Gaussian sphere learning model to automatically learn the output Gaussian sphere parameters, compensating for the deviations caused by human control information in image generation. This improves the performance of the human motion capture model. Furthermore, by using the image difference information between the rendered human motion image and the target human motion image as supervisory information to train the human motion capture model, it is beneficial to improve the accuracy and reliability of the human motion capture model, enabling it to output rendered human motion images with a high degree of matching to the original human motion images. In addition, this application embodiment only requires target video data collected by a single camera to help complete the training of the human motion capture model, eliminating the need for multiple cameras, which is beneficial to improving the generation efficiency and application scope of the human motion capture model.

[0109] The Gaussian sphere parameters include Gaussian sphere rendering parameters and position offset parameters. The Gaussian sphere rendering parameters are used to represent the rendering properties of the Gaussian sphere in the 3D Gaussian rendering model, and the position offset parameters are used to indicate the position offset of the Gaussian sphere in the 3D Gaussian rendering model.

[0110] The Gaussian sphere learning model includes a neural network model, which is used to learn the rendering parameters and position offset parameters of the Gaussian sphere. This neural network model can be a Multi-Layer Perceptron (MLP) network, a Convolutional Neural Network, a Transformer Neural Network, etc. Understandably, the architecture and number of neural network models can be customized by the designer according to business needs.

[0111] The human body rendering model includes the SMPL human body model and the 3D Gaussian rendering model. The SMPL human body model is used to construct the human body model, and the 3D Gaussian rendering model is used to render the human body model to output a human motion rendering image. The Gaussian sphere rendering parameters include the spherical harmonic function color, opacity o, rotation matrix R, and scale vector S. The position offset parameters include the offset vector Δu of the center point of the Gaussian sphere and the weight values ​​W of the vertices associated with the pose of the target human body. These weight values ​​W can be the default weight values ​​of the SMPL human body model.

[0112] Here, vertices are the location points in the SMPL human model used to associate with the pose of the target human body, such as... Figure 8 As shown, the SMPL human body model is divided into 24 control points. Based on these control points, pose transformations and shape adjustments are performed to obtain the vertex positions of the vertices in the SMPL human body model used to associate with the pose of the target human body. Each vertex can be configured with a weight value W to represent the degree of influence of each joint on that vertex. During deformation, based on the vertex positions and corresponding weight values ​​of the vertices before deformation, this embodiment can calculate the vertex positions of the deformed vertices. Vertex features are used to characterize vertices; it is understood that this embodiment can use any feature representation method to generate vertex features.

[0113] For example, in this embodiment of the application, based on the vertex position, a random sampling algorithm, a distance-based sampling algorithm, a distribution-based sampling algorithm, or a sampling algorithm based on an adversarial network model can be used to randomly sample in a preset sampling space to obtain the vertex features of the vertex.

[0114] In the process of inputting the vertex features of the vertices into the neural network model to obtain the Gaussian sphere rendering parameters and position offset parameters, please refer to [link to relevant documentation]. Figure 10 The neural network model is a multi-layer MLP neural network. A multi-layer MLP neural network can automatically learn and output the spherical harmonic function color, opacity o, rotation matrix R, scale vector S, offset vector Δu, and weight values ​​W. For example, the expression for a multi-layer MLP neural network is:

[0115] W, color, o, Δu, R, S = mlp(f).

[0116] In some embodiments, please refer to Figure 11 The neural network model consists of six multi-layer MLP neural networks. These six MLP neural networks can automatically learn to output the spherical harmonic function color, opacity o, rotation matrix R, scale vector S, offset vector Δu, and weight values ​​W, respectively. For example, the expression for the six multi-layer MLP neural networks is:

[0117] W = mlp1(f)

[0118] color = mlp2(f)

[0119] o = mlp3(f)

[0120] Δu=mlp4(f)

[0121] R = mlp5(f)

[0122] S = mlp6(f)

[0123] This application embodiment uses the SMPL human body model as the baseline human body model, and then integrates Gaussian sphere rendering parameters and position offset parameters to create human motion capture images with a higher degree of matching. Furthermore, this application embodiment automatically learns and outputs Gaussian sphere parameters through a neural network model. This not only reduces the dependence on hardware—for example, while related technologies require 16 or hundreds of cameras—this application embodiment only requires 1 or 2 cameras, or a small number of cameras, to achieve the goal of training a high-performance human motion capture model. Moreover, the Gaussian sphere parameters automatically learned and output by the neural network model can compensate for the bias caused by manual annotation, decoupling human factors and improving the performance of the human motion capture model. The trained human motion capture model has higher robustness.

[0124] The process of inputting the Gaussian sphere parameters and the human body control information into the human body rendering model to obtain a human body motion rendering image includes: determining the target position of the target vertex after the target human body changes based on the position offset parameters and the human body control information, wherein the target vertex is a vertex in the target human body motion image; inputting the target position of each target vertex and the Gaussian sphere rendering parameters into the SMPL human body model to obtain the Gaussian sphere rendering parameters corresponding to each target vertex; and inputting the Gaussian sphere rendering parameters corresponding to each target vertex into the 3D Gaussian rendering model to obtain a human body motion rendering image.

[0125] Specifically, the vertex position of the target vertex is updated according to the position offset parameter to obtain the static position of the target vertex before the action of the target human body changes; the static position is moved according to the human body control information to obtain the target position of the target vertex after the action of the target human body changes.

[0126] For example, in this embodiment of the application, the static position of the target vertex before the change in the target human body is calculated according to the following formula:

[0127] (x″ t ,y″ t , z″ t ) = LBS(Gt W t , {x′ t y′ t , z′ t})

[0128] Among them, (x″ t ,y″ t , z″ t ) represents the static position of the target vertex before the target human body changes, LBS() is the LBS function of the preset linear hybrid skinning algorithm, and G t W represents the key points of the human body corresponding to the target vertex. t The weight value of the target vertex.

[0129] This application embodiment first determines the static position of the target vertex before the target human body changes, and then determines the target position of the target vertex after the target human body changes based on the static position, thus obtaining human motion rendering images with high efficiency.

[0130] It should be noted that in the above embodiments, there is no necessarily a certain order between the steps. Those skilled in the art can understand from the description of the embodiments of this application that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.

[0131] This application provides a remote human body rendering method, applied to electronic devices. Please refer to... Figure 12 The remote human body rendering method includes the following steps:

[0132] S81. When the electronic device establishes a conference connection with the target device, it receives real-time human body control information of the target user sent by the target device, wherein the target user is a user participating in the conference through the target device.

[0133] S82. Input the real-time human body control information into the human motion capture model corresponding to the target user, so that the human motion capture model generates a human motion rendering image of the target user.

[0134] S83. The conference interface of the electronic device displays the rendered image of the human body motion.

[0135] The human motion capture model is trained using the training method for the human motion capture model described in the above embodiments.

[0136] This method receives real-time human control information, inputs a human motion capture model, generates rendered human motion images, and presents them on the meeting interface. This enables real-time rendering of the target user's human movements in remote meetings, thereby enhancing the interactivity and communication effectiveness of the meeting and allowing participants to more intuitively perceive the target user's motion state. Furthermore, using the aforementioned human motion capture model improves the accuracy and realism of the rendering.

[0137] This application embodiment also provides a training device for a human motion capture model, specifically including: a target video data acquisition module, used to acquire target video data, the target video data including video data of a target human body in different postures, the video data including multiple original human motion images, the human motion capture model including a Gaussian sphere learning model and a human rendering model; a human control information acquisition module, used to acquire human control information corresponding to the target human motion image, the target human motion image being an original human motion image in a training state; a Gaussian sphere parameter acquisition module, used to acquire Gaussian sphere parameters output by the Gaussian sphere learning model when the target human image is used as a training input sample, the Gaussian sphere parameters being used to constrain the Gaussian sphere of the human rendering model; a human motion rendering image generation module, used to input the Gaussian sphere parameters and the human control information into the human rendering model to obtain a human motion rendering image; a difference information determination module, used to determine the image difference information between the human motion rendering image and the target human motion image; and a model training module, used to train the human motion capture model based on the image difference information.

[0138] It should be noted that the training device for the human motion capture model described above can execute the training method for the human motion capture model provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in the embodiments of the training device for the human motion capture model can be found in the training method for the human motion capture model provided in the embodiments of this application.

[0139] This application also provides a remote human body rendering device applied to an electronic device, specifically including: a human body control information receiving module, used to receive real-time human body control information of a target user sent by the target device when the electronic device establishes a conference connection with a target device, wherein the target user is a user participating in the conference through the target device; a human body motion rendering image generation module, used to input the real-time human body control information into a human body motion capture model corresponding to the target user, so that the human body motion capture model generates a human body motion rendering image of the target user; and a rendering image presentation module, used to control the conference interface of the electronic device to present the human body motion rendering image. The human body motion capture model is trained using the training method for the human body motion capture model in the above embodiments.

[0140] It should be noted that the aforementioned remote human body rendering device can execute the remote human body rendering method provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in the embodiments of the remote human body rendering device can be found in the remote human body rendering method provided in the embodiments of this application.

[0141] Figure 13 This is a schematic diagram of the hardware structure of the electronic device 10 provided in this application embodiment. This electronic device can be used to execute the training method for the human motion capture model in the above embodiments, and can also be used to execute the remote human rendering method in the above embodiments, such as... Figure 13 As shown, the electronic device 10 includes: one or more processors 11 and a memory 12. Figure 13 Taking a processor 11 as an example, the processor 11 and the memory 12 can be connected via a bus or other means. Figure 13 Taking the example of a connection between China and Israel via a bus.

[0142] The memory 12, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the human motion capture model training method or the remote human rendering method in the embodiments of this application. The processor 11 executes various functional applications and data processing of the electronic device by running the non-volatile software programs, instructions, and modules stored in the memory 12, that is, it implements the human motion capture model training method or the remote human rendering method in the above method embodiments.

[0143] The memory 12 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function. The data storage area may store data created based on the use of the human motion capture model training device or remote human rendering device. Furthermore, the memory 12 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 12 may optionally include memory remotely located relative to the processor 11, which can be connected to the human motion capture model training device or remote human rendering device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0144] The one or more modules are stored in the memory 12, and when executed by the one or more processors 11, they execute the training method of the human motion capture model or the remote human rendering method in any of the above method embodiments.

[0145] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.

[0146] The electronic devices in this application can exist in various forms, including but not limited to: ultra-mobile personal computer devices, smart displays or all-in-one machines, servers or server clusters, etc.

[0147] This application provides a non-volatile computer-readable storage medium storing computer-executable instructions that are executed by one or more processors, for example... Figure 13 One of the processors 11 can enable the above one or more processors to execute the training method of the human motion capture model or the remote human rendering method in any of the above method embodiments.

[0148] This application provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions, which, when executed by the electronic device, enable the electronic device to execute the human motion capture model training method or the remote human rendering method in any of the above method embodiments.

[0149] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0150] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software and a general-purpose hardware platform, or of course, using hardware. Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of this application as described above, which are not provided in detail for the sake of brevity; although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A training method for a human motion capture model, characterized in that, include: Acquire target video data, which includes video data of the target human body in different poses. The video data includes multiple original human motion images. The human motion capture model includes a Gaussian ball learning model and a human rendering model. Obtain human control information corresponding to the target human motion image, wherein the target human motion image is the original human motion image in the training state; Obtain the Gaussian sphere parameters output by the Gaussian sphere learning model when the target human image is used as a training input sample; the Gaussian sphere parameters are used to constrain the Gaussian sphere of the human rendering model. The Gaussian sphere parameters and the human body control information are input into the human body rendering model to obtain a human body motion rendering image. Determine the image difference information between the rendered human motion image and the target human motion image; The human motion capture model is trained based on the image difference information.

2. The training method according to claim 1, characterized in that, The Gaussian sphere parameters include Gaussian sphere rendering parameters and position offset parameters. The Gaussian sphere learning model includes a neural network model. The human body rendering model includes an SMPL human body model and a 3D Gaussian rendering model. Obtaining the Gaussian sphere parameters output by the Gaussian sphere learning model when the target human body image is used as a training input sample includes: When the target human motion image is used as a training input sample, the vertex features of the vertices in the SMPL human model used to associate with the pose of the target human are determined. The vertex features of the vertex are input into the neural network model to obtain Gaussian sphere rendering parameters and position offset parameters. The Gaussian sphere rendering parameters are used to represent the rendering properties of the Gaussian sphere in the 3D Gaussian rendering model, and the position offset parameters are used to indicate the position offset of the Gaussian sphere in the 3D Gaussian rendering model.

3. The training method according to claim 2, characterized in that, The vertex features used to associate the vertices in the SMPL human model with the pose of the target human body include: Obtain the control points of the SMPL human body model, wherein the control points are key points at specific joint positions in the SMPL human body model; The key points are subjected to pose transformation and shape adjustment to obtain the vertex positions of the vertices in the SMPL human body model used to associate with the pose of the target human body. Based on the vertex position, a feature sampling operation is performed in a preset sampling space to obtain the vertex features.

4. The training method according to claim 2, characterized in that, The step of inputting the vertex features of the vertex into the neural network model to obtain the Gaussian sphere rendering parameters and position offset parameters includes: The vertex features of the vertex are input into the neural network model, and the output is a Gaussian sphere rendering parameter composed of spherical harmonic function, opacity, rotation matrix and scale vector, and a position offset parameter composed of offset vector of Gaussian sphere center point and vertex weight value.

5. The training method according to claim 2, characterized in that, The step of inputting the Gaussian sphere parameters and the human body control information into the human body rendering model to obtain a human motion rendering image includes: The target position of the target vertex after the target human body changes is determined based on the position offset parameter and the human body control information, wherein the target vertex is a vertex in the target human body motion image; The target position and Gaussian sphere rendering parameters of each target vertex are input into the SMPL human body model to obtain the Gaussian sphere rendering parameters corresponding to each target vertex. The rendering parameters of the Gaussian sphere corresponding to each target vertex are input into the 3D Gaussian rendering model to obtain a human motion rendering image.

6. The training method according to claim 5, characterized in that, Determining the target vertex's position after the target human body changes based on the position offset parameter and the human body control information includes: The vertex position of the target vertex is updated according to the position offset parameter to obtain the static position of the target vertex before the action of the target human body changes; The static position is moved according to the human body control information to obtain the target position of the target vertex after the target human body's movement changes.

7. The training method according to claim 6, characterized in that, The position offset parameter includes the offset vector of the center point of the Gaussian sphere. Updating the vertex position of the target vertex according to the position offset parameter to obtain the static position of the target vertex before the change in the target human's movement includes: The coordinates corresponding to the vertex position of the target vertex are added to the offset vector of the center point of the Gaussian sphere to obtain the static position of the target vertex before the action of the target human body changes.

8. The training method according to claim 6, characterized in that, The human body control information includes the camera pose when acquiring the target human body motion image in the target video data, and key points in the target human body motion image that are associated with specific joint positions of the human body; The position offset parameter includes the weight value of the vertex; The step of moving the static position according to the human body control information to obtain the target position of the target vertex after the target human body's action changes includes: Based on the camera pose, the coordinates of the key points, and the weight values ​​of the vertices, the position of the target vertex at the static position is updated to obtain the target position of the target vertex after the action of the target human body changes.

9. The training method according to any one of claims 2 to 8, characterized in that, The image difference information includes a loss value, and training the human motion capture model based on the image difference information includes: Based on the loss value, the parameters of the neural network in the neural network model and the parameters of the SMPL human body model are optimized using the gradient descent algorithm.

10. A remote human body rendering method, applied to electronic devices, characterized in that, include: When the electronic device establishes a conference connection with the target device, it receives real-time human body control information of the target user sent by the target device. The target user is the user participating in the conference through the target device. The real-time human control information is input into the human motion capture model corresponding to the target user, so that the human motion capture model generates a human motion rendering image of the target user. The conference interface controlling the electronic device displays the rendered image of the human body movement; The human motion capture model is trained by the training method of any one of claims 1 to 8.

11. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the training method for the human motion capture model according to any one of claims 1 to 9, or to perform the remote human rendering method according to claim 10.