Three-dimensional image presentation method, conference equipment and conference system
By combining 3D camera technology and feature point-driven rendering technology with light field display, the problems of high bandwidth requirements and image deviation in immersive video conferencing are solved, enabling fast and real-time stereoscopic image presentation under low bandwidth, thus improving immersion and interactivity.
Patent Information
- Application Number
- CN202410748262.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-11
- Publication Date
- 2025-12-12
AI Technical Summary
Existing immersive video conferencing requires high bandwidth to transmit multiple video streams, which limits its application scope. Furthermore, existing technologies suffer from image bias and device complexity issues when generating rendering models, resulting in insufficient immersion and interactivity.
3D reconstruction is performed using 3D camera technology to generate accurate virtual avatars. Key feature points and rendering parameters are transmitted using feature point driving and rendering technology. Combined with light field display technology, multi-person, multi-view stereoscopic image presentation is achieved, reducing bandwidth requirements.
It enables fast, real-time rendering of stereoscopic images under low bandwidth conditions, expanding the scope of applications and improving immersion and interactivity, allowing each participant to have a clear and realistic image experience.
Smart Images

Figure CN121125931A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of remote conferencing technology, and in particular to a method for presenting stereoscopic images, conferencing equipment, and conferencing system. Background Technology
[0002] Immersive video conferencing is a new form of remote video conferencing that displays participants' heads in 3D on the meeting interface, enhancing the sense of presence and immersion. The technology involves using multiple cameras to capture multiple video streams, transmitting them to the conferencing equipment, which processes these streams to produce a 3D user image. This 3D user image is then presented to create an immersive meeting environment. However, this technology requires significant bandwidth to transmit multiple video streams, and most immersive video conferencing systems have limited bandwidth, thus restricting its application. Summary of the Invention
[0003] One objective of this application is to provide a method for presenting stereoscopic images, a conference device, and a conference system to solve the technical problem that related technologies require a large bandwidth to transmit data in immersive video conferencing.
[0004] In a first aspect, embodiments of this application provide a method for presenting a stereoscopic image, applied to a first conference device, comprising:
[0005] When the first conference device establishes a conference connection with the second conference device, it acquires the real-time user image of the first user, and the second conference device saves the 3D model data of the first rendering model corresponding to the first user.
[0006] Extract user feature data from real-time user images;
[0007] Based on the conference connection, the user feature data is sent to the second conference device, so that the second conference device inputs the user feature data into the first rendering model to obtain a real-time rendered image, and presents a stereoscopic image of the first user based on the real-time rendered image.
[0008] Optionally, before acquiring the real-time user image of the first user, the method further includes:
[0009] Determine the 3D model data of the first user, wherein the 3D model data is used to represent the first rendered model;
[0010] The 3D model data is sent to the second conference device so that the second conference device can parse the 3D model data into the first rendered model and save the 3D model data of the first rendered model locally on the second conference device.
[0011] Optionally, determining the 3D model data of the first user includes:
[0012] Acquire 2D image data, which includes multiple images of the target user taken from different side angles;
[0013] 3D model data is generated based on multiple images of the target user.
[0014] Optionally, generating 3D model data based on multiple target user images includes:
[0015] Extract the person region image from multiple target user images;
[0016] Based on the image of the person's region, 3D sub-model data of the person's character is generated, and the 3D sub-model data of the person's character is used to represent the 3D character rendering model of the first user.
[0017] The 3D model data is determined based on the character's 3D sub-model data.
[0018] Optionally, determining the 3D model data based on the character 3D sub-model data includes:
[0019] Extract environmental region images from multiple target user images;
[0020] Target environment 3D sub-model data is generated based on the environmental area image, and the target environment 3D sub-model data is used to represent the environment rendering model of the first user.
[0021] The 3D model data is determined based on the target environment 3D sub-model data and the character 3D sub-model data.
[0022] or,
[0023] Obtain preset environment 3D sub-model data, which is used to represent the environment rendering model of the first user;
[0024] The 3D model data is determined based on the preset environment 3D sub-model data and the character 3D sub-model data.
[0025] In a second aspect, embodiments of this application provide a method for presenting a stereoscopic image, applied to a second conference device, comprising:
[0026] When the second conference device establishes a conference connection with the first conference device, it acquires user feature data sent by the first conference device. The user feature data is obtained by the first conference device extracting the real-time user image of the first user.
[0027] Obtain the 3D model data of the first user;
[0028] The first user's 3D model data is parsed into the first user's first rendering model;
[0029] The user feature data is input into the first rendering model to obtain a real-time rendered image;
[0030] Generate a stereoscopic image of the first user based on the real-time rendered image;
[0031] Present a 3D image of the first user.
[0032] Optionally, generating the stereoscopic image of the first user based on the real-time rendered image includes:
[0033] Acquire preset multi-viewpoint information;
[0034] An image sequence is generated based on the multi-viewpoint information and the real-time rendered image;
[0035] Generate a stereoscopic image of the first user based on the image sequence.
[0036] Optionally, before acquiring the user characteristic data sent by the first conference device, the method further includes:
[0037] Generate 3D model data for the second user's second rendered model, where the second user is the user operating the second conference device;
[0038] Send the 3D model data of the second rendered model to the first conference device so that the first conference device can save the 3D model data of the second rendered model locally.
[0039] In a third aspect, embodiments of this application provide a first conference device, including a memory and a processor. The memory is connected to the processor, and the processor is used to execute one or more computer programs stored in the memory. When the processor executes one or more computer programs, it enables the conference device to implement the above-described method for presenting stereoscopic images.
[0040] In a fourth aspect, embodiments of this application provide a second conference device, including a memory and a processor. The memory is connected to the processor, and the processor is used to execute one or more computer programs stored in the memory. When the processor executes one or more computer programs, it enables the conference device to implement the above-described method for presenting stereoscopic images.
[0041] In a fifth aspect, embodiments of this application provide a conference system, including:
[0042] The aforementioned first conference equipment;
[0043] The second conference device mentioned above can establish a conference connection with the first conference device.
[0044] Optionally, the conference system may also include:
[0045] The first camera module includes N first cameras, which are respectively connected to the first conference device for acquiring 2D image data, where N is a positive integer.
[0046] The second camera module includes M second cameras, which are each connected to the second conference equipment for acquiring 2D image data.
[0047] In a sixth aspect, embodiments of this application provide a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, the computer program including program instructions, which, when executed by a processor, cause the processor to perform the stereoscopic image presentation method as described in any one of claims 1-5 or the stereoscopic image presentation method described above.
[0048] The embodiments of this application can achieve the following technical effects: In the stereoscopic image presentation method provided in the embodiments of this application, firstly, when the first conference device and the second conference device establish a conference connection, the real-time user image of the first user is acquired, and the second conference device saves the 3D model data of the first rendering model corresponding to the first user. Then, user feature data of the real-time user image is extracted. The user feature data is relatively small in size compared to the video stream of the first user. Therefore, based on the conference connection, the embodiments of this application can send a small amount of user feature data to the second conference device, so that the second conference device can input the user feature data into the first rendering model corresponding to the first user, so that the conference interface of the second conference device presents the stereoscopic image of the first user. In immersive video conferencing, the embodiments of this application do not need to transmit the entire video stream of the first user to the second conference device for 3D presentation, nor do they need to occupy too much bandwidth to transmit the data used to prompt the second conference device to perform 3D presentation. The embodiments of this application only require a small amount of user feature data to enable the second conference device to present the stereoscopic image of the first user, thus improving the application scope of the method provided in the embodiments of this application. In addition, since the user feature data is small in size, the embodiments of this application can transmit it to the second conference device without delay, which is beneficial for the second conference device to quickly and in real time present the stereoscopic image of the first user. Attached Figure Description
[0049] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1 A schematic diagram of the system architecture of a conference system provided in this application embodiment;
[0051] Figure 2 A schematic diagram of the structure between modules involved in the electronic device or target device provided in the embodiments of this application for completing 3D head images;
[0052] Figure 3 This is a schematic diagram illustrating the interaction between an electronic device and a target device provided in an embodiment of this application.
[0053] Figure 4 An interactive diagram illustrating the exchange of header rendering models between an electronic device and a target device, as provided in an embodiment of this application.
[0054] Figure 5a A schematic diagram illustrating the participation of an electronic device, a target device, and a server in a meeting, as provided in an embodiment of this application.
[0055] Figure 5b A schematic diagram of the meeting interface provided in this application embodiment;
[0056] Figure 5c A schematic diagram of a meeting interface provided for another embodiment of this application;
[0057] Figure 6 A schematic diagram of the system architecture of a conference system provided in another embodiment of this application;
[0058] Figure 7 This is a flowchart illustrating a method for presenting a stereoscopic image according to an embodiment of this application, wherein the executing entity is a first conference device;
[0059] Figure 8 A schematic diagram of facial key points of a human head FLAME model provided in an embodiment of this application;
[0060] Figure 9 A schematic diagram of key human body points in a standard FLAME human head model provided for an embodiment of this application;
[0061] Figure 10 This is a schematic diagram illustrating the process of training a head rendering model as provided in an embodiment of this application;
[0062] Figure 11A schematic diagram illustrating the learning of the first Gaussian sphere parameters when the first neural network model provided in this embodiment is a multilayer MLP neural network;
[0063] Figure 12 A schematic diagram illustrating the learning of the first Gaussian sphere parameters when the first neural network model provided in this embodiment is a plurality of multi-layer MLP neural networks;
[0064] Figure 13 This is a schematic diagram illustrating the process of training a human rendering model as provided in an embodiment of this application;
[0065] Figure 14 This is a flowchart illustrating a method for presenting a stereoscopic image according to an embodiment of this application, wherein the executing entity is a second conference device;
[0066] Figure 15 This is a flowchart illustrating a method for presenting a stereoscopic image according to an embodiment of this application, wherein the executing entity is a first conference device;
[0067] Figure 16 This is a flowchart illustrating a method for presenting a stereoscopic image according to an embodiment of this application, wherein the executing entity is a second conference device;
[0068] Figure 17 This is a schematic diagram of the structure of a conference device provided in an embodiment of this application. Detailed Implementation
[0069] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0070] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. Moreover, the terms "first," "second," and "third" used in this application do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.
[0071] Remote video conferencing struggles to provide the same efficient communication experience as face-to-face conversations because participants cannot truly perceive vital visual information such as body movements, facial expressions, and eye contact, and it lacks natural interaction methods. Currently available mainstream LCD or LED displays can only show the same 2D image in different directions within the space. This means that regardless of their location in the meeting room, participants see the same screen image, severely diminishing the immersive experience of remote video conferencing.
[0072] In remote video conferencing, body language, facial expressions, and eye contact are crucial nonverbal communication methods that convey participants' attention, interest, and intentions, facilitating effective communication and interaction. However, due to limitations in cameras and video transmission, participants struggle to accurately capture and convey body language and facial expressions, leading to incomplete information delivery and unnatural communication. Furthermore, mainstream 2D displays fail to provide a realistic sense of space and depth, limiting participants' ability to perceive object location and size. This further reduces immersion and naturalness of interaction, making it difficult for participants to truly perceive the presence and location of others in the meeting room. Additionally, participants in remote video conferencing are typically in their own work environments, making them susceptible to environmental distractions and other factors. Compared to face-to-face meetings, participants are more easily affected by their surroundings, impacting focus and effective communication. In conclusion, the lack of immersion in remote video conferencing stems from participants' inability to truly perceive visual information such as body language and facial expressions, and the limitations imposed by the spatial representation capabilities of 2D displays. These limitations affect the efficiency and naturalness of communication, necessitating further technological advancements and innovations to enhance the remote video conferencing experience.
[0073] The immersive conferencing solution provided by related technologies relies on multiple synchronous cameras to capture video of the meeting, transmitting the captured video to remote devices via high bandwidth, and then generating two-viewpoint images. The perspective of these two viewpoints depends on additional cameras capturing the results of binocular eye contact. In practice, this technology requires multiple cameras to capture video of the meeting, transmitting data to remote devices via high bandwidth. Furthermore, it supports single-person viewing, which limits the number of viewpoints for the light field display. Only participants whose eyes are captured by the cameras can normally view the content on the light field display; other participants cannot obtain a clear image. Therefore, this technology has the following problems in implementing immersive conferencing: it requires multiple synchronous cameras to capture video of the meeting, which increases the complexity of equipment and wiring, and places additional demands on meeting room setup and costs; to transmit video data captured by multiple cameras, the immersive conferencing solution requires a high-bandwidth network connection; the light field display in the immersive conferencing solution can provide multi-viewpoint images, but the number of viewpoints is limited by the number of participants whose eyes are captured, meaning that only a few participants can obtain a clear, realistic image, while others cannot have the same experience.
[0074] The relevant technology can use a human head FLAME model combined with multiple user head region images to be trained to train the rendering model. The multiple user head region images to be trained have been manually annotated with multiple facial key points.
[0075] In the real physical world, the head does not possess biologically significant facial landmarks. The facial landmarks mentioned above are merely a concept proposed for artificially constructed human head FLAME models. Therefore, for a head region image, the facial landmarks do not possess completely accurate biological locations. Even manual annotation of facial landmarks used as Ground Truth cannot guarantee complete accuracy; different designers annotating the same facial landmark location can easily yield different annotation results. Even the same designer annotating at different times can easily obtain different annotation results.
[0076] Even the gold standard Ground Truth annotation results cannot guarantee the accuracy and consistency of facial key points. Furthermore, the rendering model relies on the accuracy of facial key points. Therefore, the rendering model trained using facial key points obtained from Ground Truth is unreliable. The head rendering image output by the trained rendering model will have image deviation relative to the target user image. This image deviation will cause severe distortion in the head rendering image output by the rendering model, which will greatly reduce the user's experience.
[0077] In another related technology, to improve the accuracy of manual facial landmark annotation, multiple cameras are set up in front of the user. These cameras can simultaneously capture multiple images of the user's head area from different angles, such as using 16 or even hundreds of cameras. Facial landmarks are then annotated on these multiple images of the user's head area, and the annotated facial landmarks are relatively stable and reliable. However, this method has a problem: when building a rendering model for each user, it requires setting up dozens or even hundreds of cameras in front of each user to simultaneously capture their head for training. However, most scenarios lack the resources to set up multiple cameras for each participant, especially in remote conferencing. Therefore, this technology has significant limitations, a limited application scope, and low efficiency in generating rendering models.
[0078] This application embodiment utilizes 3D camera technology to perform 3D reconstruction of meeting participants, generating accurate virtual avatars. This reduces the need for multiple cameras and lowers the complexity of equipment and wiring. Based on the 3D avatar reconstruction results, feature point-driven rendering technology can be used to transmit participants' body movements and facial expressions. By capturing key feature points such as joint positions and facial expressions, virtual body movements and facial expressions can be generated in real time and transmitted to remote devices for rendering. Using feature point-driven rendering significantly reduces bandwidth requirements. Compared to transmitting video data from multiple cameras, only key feature points and rendering parameters need to be transmitted, effectively reducing data volume and bandwidth requirements. Finally, combined with light field display technology, it enables simultaneous viewing by multiple people from multiple perspectives. The light field display provides more viewpoints and generates corresponding images in real time based on the observer's position and gaze direction, ensuring each participant receives a clear and realistic image, enhancing immersion and experience.
[0079] This application provides a conference system; please refer to [link / reference]. Figure 1 The conference system 100 includes a first conference device 200 and a second conference device 300, wherein the first conference device 200 and the second conference device 300 establish a conference connection. The first conference device 200 and the second conference device 300 can join the same remote conference to participate in the meeting. The conference interface of the first conference device 200 can display the image of user A operating the second conference device 300, and similarly, the conference interface of the second conference device 300 can display the image of user B operating the first conference device 200.
[0080] The images of people displayed by the first conference device 200 and the images of people displayed by the second conference device 300 are three-dimensional images. Since three-dimensional images have a greater sense of depth than 2D rendered images, they are easier to create an immersive experience for users, making it easier for users to immerse themselves in the conference atmosphere.
[0081] The stereoscopic image can be synthesized by the first conference device 200 or the second conference device 300 from multiple 2D rendered images, wherein the 2D rendered image can be generated by the first conference device 200 or the second conference device 300 through a corresponding rendering model.
[0082] Please refer to the following: Figure 2 and Figure 3 Both the first conference equipment 200 and the second conference equipment 300 include an N-channel camera module 41, an ISP module 42, a video decoding module 43, a 3D model generation module 44, a model transmission module 45, a 3D rendering module 46, a 3D compositing module 47, a 3D playback module 48, and a SOC processing module 49.
[0083] The camera module 41 is used to acquire video data of the user's head in different postures and send the video data to the ISP module 42. The ISP module 42 processes the video data to improve image clarity, color reproduction, dynamic range, etc., and enhance the image's resolution in low light conditions.
[0084] The video decoding module 43 is used to restore the processed video data to obtain the original video image and audio signal. The 3D model generation module 44 is used to extract user feature data and generate a rendering model. The model transmission module 45 is used to send the rendering model to a designated device, which can be the device of the other participant or a server.
[0085] The 3D rendering module 46 processes user feature information using a rendering model to output a rendered image. The 3D compositing module 47 combines multiple rendered images into a stereoscopic image. The 3D playback module 48 combines the stereoscopic image and audio signal into a video stream. The SOC processing module 49 processes the video stream to present the stereoscopic image and audio signal.
[0086] It is understood that the camera module 41 can be any type of camera, and the ISP module 42 can be a separate ISP chip, or the ISP module 42 and the camera module 41 can be integrated on the same chip. It is also understood that in some embodiments, the number of ISP modules 42 is consistent with the number of camera modules 41, that is, N camera modules 41 correspond to N ISP modules 42, and one camera module 41 corresponds to one ISP module 42. In some embodiments, N camera modules 41 correspond to one ISP module 42, that is, the video streams of N camera modules 41 are aggregated into the same ISP module 42 for processing. The video decoding module 43 can be a video decoding chip. The 3D model generation module 44 can be a 3D graphics card processor, etc. The model transmission module 45 can be a wired communication module or a wireless communication module. The 3D rendering module 46, the 3D compositing module 47, the 3D playback module 48, and the 3D playback module 49 can be various types of processors, such as DSP processors. The SOC processing module 49 can be a SOC processing chip.
[0087] In some embodiments, the first conference device 200 and the second conference device 300 establish an end-to-end communication connection. Based on this connection, they exchange rendered models for local storage. Specifically, the first conference device 200 can send the 3D model data corresponding to its local rendered model to the second conference device 300 for local storage and parsing / reconstruction via the end-to-end communication connection, while the second conference device 300 can send the 3D model data corresponding to its local rendered model to the first conference device 200 for local storage and parsing / reconstruction via the same connection.
[0088] For example, please see Figure 4 User A joins remote conference A using the first conference device 200, and User B joins remote conference A using the second conference device 300. The first conference device 200 trains and obtains 3D model data of User A's first rendered model J1, and the second conference device 300 trains and obtains 3D model data of User B's second rendered model J2. The first conference device 200 and the second conference device 300 can communicate end-to-end. The first conference device 200 sends the 3D model data of User A's first rendered model J1 to the second conference device 300, and the second conference device 300 stores the 3D model data of User A's first rendered model J1 locally. The second conference device 300 sends the 3D model data of User B's second rendered model J2 to the first conference device 200, and the first conference device 200 stores the 3D model data of User B's second rendered model J2 locally.
[0089] During the meeting, the first conference device 200 captures real-time images of user A using a camera, obtaining a real-time user image X1. Then, the first conference device 200 determines user feature data based on the real-time user image X1 and sends this data to the second conference device 300. The second conference device 300 uses the first rendering model J1 to process the user feature data, thereby obtaining a real-time rendered image of user A. The second conference device 300 then generates multiple reference rendered images with parallax based on the real-time rendered image, and finally combines the real-time rendered image and the multiple reference rendered images to create a stereoscopic image. Finally, the second conference device 300 can display the stereoscopic image of user A on its local conference interface.
[0090] Similarly, the second conference device 300 captures real-time images of user B using a camera, obtaining a real-time user image X2. Then, the second conference device 300 determines user feature data based on the real-time user image X2 and sends this data to the first conference device 200. The first conference device 200 uses the second rendering model J2 to process the user feature data, thereby obtaining a real-time rendered image of user B. Finally, the first conference device 200 generates a stereoscopic image of user B using the above method and displays this stereoscopic image on its local conference interface.
[0091] In related technologies, to display a stereoscopic image from a second conference device 300 on a first conference device 200, the technology directly controls the second conference device 300 to transmit N video streams to the first conference device 200. The first conference device 200 then processes the N video streams into a stereoscopic image for display. However, based on the transmission requirements of a video stream resolution of 1920*1080 and a frame rate of 30fps, the bandwidth required to transmit one video stream is: 1920*1080*3*8 / 1024 / 1024 / 1024*30 frames*1 stream = 1.39Gbps. If N=4, meaning the second conference device 300 transmits four video streams to the first conference device 200, then the technology requires 5.56Gbps of bandwidth. Therefore, the technology requires a large bandwidth to achieve the purpose of displaying a stereoscopic image. However, many data transmission scenarios cannot support such a high bandwidth data transmission, which significantly limits the application scope of the technology.
[0092] In this embodiment of the application, the amount of user feature data is relatively small, usually less than 1 Mbps. Therefore, the method provided in this embodiment of the application occupies less bandwidth, which is beneficial to broadening the application scope of the method provided in this embodiment of the application.
[0093] In some embodiments, please refer to the following: Figure 5a , Figure 5b and Figure 5cThe first conference device 200 and the second conference device 300 establish communication connections with the server 400 respectively, and the server 400 can forward the 3D model data of the rendered model of the other party to the other party's device.
[0094] Please continue reading. Figure 5a , Figure 5b and Figure 5c When the first conference device 200 and the second conference device 300 join the same remote conference, the server 400 can send the 3D model data of the other party's rendered model to the other party's device.
[0095] The first conference device 200 sends the 3D model data of user A's first rendered model J1 to the server 400 for storage, and the second conference device 300 sends the 3D model data of user B's second rendered model J2 to the server 400 for storage.
[0096] When server 400 detects that user A's user information and user B's user information appear on the same remote conference participant list, where user A's user information is bound to the first conference device 200 and user B's user information is bound to the second conference device 300, server 400 sends the 3D model data of the first rendering model J1 to the second conference device 300 and sends the 3D model data of the second rendering model J2 to the first conference device 200.
[0097] During the meeting, the first conference device 200 captures real-time images of user A using a camera, obtaining a real-time user image X1. Then, the first conference device 200 determines user A's characteristic data based on the real-time user image X1 and sends this data to the server 400. Similarly, the second conference device 300 captures real-time images of user B using a camera, obtaining a real-time user image X2. Then, the second conference device 300 determines user B's characteristic data based on the real-time user image X2 and sends this data to the server 400.
[0098] Server 400 forwards user B's user characteristic data to the first conference device 200, and forwards user A's user characteristic data to the second conference device 300.
[0099] like Figure 5c As shown, the second conference device 300 calls the first rendering model J1 to process the user feature data of user A, and synthesizes a stereoscopic image according to the above method, and presents the stereoscopic image 5c1 of user A on the local conference interface. Similarly, the first conference device 200 calls the second rendering model J2 to process the user feature data of user B, and generates a stereoscopic image of user B according to the above method, and presents the stereoscopic image of user B on the local conference interface.
[0100] Understandably, in some embodiments, the rendered models can be generated before the meeting begins. For example, the first meeting device 200 generates 3D model data for user A's first rendered model J1 in advance, and the second meeting device 300 generates 3D model data for user B's second rendered model J2 in advance. Then, before the meeting begins, the first meeting device 200 sends the 3D model data of the first rendered model J1 to the second meeting device 300 for local storage, and the second meeting device 300 sends the 3D model data of the second rendered model J2 to the first meeting device 200 for local storage. Alternatively, before the meeting begins, the first meeting device 200 sends the 3D model data of the first rendered model J1 to the server 400, the second meeting device 300 sends the 3D model data of the second rendered model J2 to the server 400, and the server 400 forwards the 3D model data of the first rendered model J1 to the second meeting device 300 for local storage, and forwards the 3D model data of the second rendered model J2 to the first meeting device 200 for local storage.
[0101] It is also understood that, in some embodiments, the generation time of the rendered model can be at the start of the meeting. As previously mentioned, at the start of the meeting, User A uses the camera of the first conference device 200 to photograph User A's head and records first video data of a preset duration, wherein the first video data includes multiple target user images of User A. User B uses the camera of the second conference device 300 to photograph User B's head and records second video data of a preset duration, wherein the second video data includes multiple target user images of User B. The first conference device 200 generates 3D model data of the first rendered model J1 based on the first video data, and sends the 3D model data of the first rendered model J1 directly to the second conference device 300 for local storage or to the server 400, which forwards it to the second conference device 300 for local storage. At the same time, the second conference device 300 generates 3D model data of the second rendered model J2 based on the second video data, and sends the 3D model data of the second rendered model J2 directly to the first conference device 200 for local storage or to the server 400, which forwards it to the first conference device 200 for local storage.
[0102] In some embodiments, the video data used to generate the rendered model may be acquired by a camera carried by the first conference device 200 or the second conference device 300.
[0103] In some embodiments, the video data used to generate the rendered model can be captured by cameras set up at the conference venue. See also Figure 6The conference system 100 includes a first camera module 500 and a second camera module 600. The first camera module 500 includes N first cameras, which are communicatively connected to the first conference device 200. The second camera module 600 includes M second cameras, which are communicatively connected to the second conference device 300, where N and M are both positive integers, for example, N = M = 4 or N = M = 1.
[0104] User A and User B are conducting a remote video conference. When User A enters the conference room, the first camera module 500 captures video data of User A in different postures, and then transmits this video data to the first conferencing device 200. The first conferencing device 200 generates 3D model data for the first rendering model J1 based on the video data. Similarly, when User B enters the conference room, the second camera module 600 captures video data of User B in different postures, and then transmits this video data to the second conferencing device 300. The second conferencing device 300 generates 3D model data for the second rendering model J2 based on the video data.
[0105] As another aspect of this application, this application provides a method for presenting a stereoscopic image. Please refer to... Figure 7 The method for presenting a stereoscopic image includes the following steps:
[0106] S71: When the first conference device establishes a conference connection with the second conference device, it acquires the real-time user image of the first user, and the second conference device saves the 3D model data of the first rendering model corresponding to the first user.
[0107] In this step, in some embodiments, the first and second conference devices establish a conference connection based on an end-to-end communication connection. In some embodiments, the server adds the first and second conference devices to the same conference participant list, and the first conference device establishes a conference connection with the second conference device through the intermediary role of the server. The first and second conference devices can be the same type of electronic device, such as both being a conference tablet, mobile phone, desktop computer, tablet computer, or other electronic product. Alternatively, the first and second conference devices can be different types of electronic devices, such as the first conference device being a conference tablet and the second conference device being a mobile phone.
[0108] The real-time user image is a user image captured in real time by the first user. In some embodiments, the real-time user image may be obtained by capturing the first user's image using a camera carried by the first conference equipment. In some embodiments, the real-time user image may be obtained by capturing the first user's image using a camera deployed at the conference venue.
[0109] The first rendering model is a model used to output the rendered image corresponding to the first user.
[0110] Understandably, in some embodiments, the 3D model data of the first rendered model may be generated by a first conference device, which then sends the 3D model data to a second conference device for local storage. In some embodiments, the 3D model data of the first rendered model may be generated by the first conference device, which then sends the 3D model data to a server, which in turn sends the 3D model data to the second conference device for local storage.
[0111] It is also understood that, in some embodiments, the 3D model data of the first rendered model may be generated before the first conference device establishes a conference connection with the second conference device, and the first conference device generates the 3D model data of the first rendered model before the conference begins. In some embodiments, the 3D model data of the first rendered model may be generated at the start of the conference, and the first conference device generates the 3D model data of the first rendered model at the start of the conference.
[0112] S72: Extract user feature data from real-time user images.
[0113] In this step, user feature data is used to represent the features of a real-time user image, wherein the features of the real-time user image can characterize the features of a first user corresponding to the real-time user image. In some embodiments, user feature data includes head feature data, which includes pose, shape, and expression. In some embodiments, user feature data includes body feature data, which includes pose and shape. In some embodiments, user feature data includes both head feature data and body feature data.
[0114] Head feature data includes camera pose and head key points. Camera pose indicates the posture of the head captured by the camera, and multiple head key points can be effectively combined to represent the shape and expression of the head. Please refer to [link / reference]. Figure 8 The human head FLAME model is configured with k head key points Q, which can effectively combine to represent the shape and expression of the head.
[0115] Human feature data includes camera pose and key points on the human body associated with specific joint positions. Please refer to [link / reference]. Figure 9 The SMPL human body model is configured with Y human body keypoints P, which can be effectively combined to represent the shape of the human body. The SMPL (Skinned Multi-Person Linear) model is a commonly used human body shape and pose model. It is a parametric model based on linear algebra used to represent the shape and pose of the human body.
[0116] When user feature data includes head feature data, the user feature data extracted from real-time user images includes: determining the camera pose of the real-time user image and multiple head key points of the real-time user image based on a preset visual model, wherein the multiple head key points meet the requirements of the human head FLAME model.
[0117] When user feature data includes human feature data, the user feature data extracted from real-time user images includes: determining the camera pose of the real-time user image and multiple human key points of the real-time user image based on a preset visual model, wherein the multiple human key points meet the requirements of the SMPL model.
[0118] The default visual model can be either a GraphMap model or a MediaPipe model. The GraphMap model is used for 3D reconstruction from a series of 2D images, supporting features such as feature detection and matching, incremental SfM, multi-view stereo matching (MVS), and optimization. It can recover the geometry of a 3D scene and the camera pose for each image from an unordered or ordered set of 2D images. The MediaPipe model is a graph-based data processing pipeline used to build machine learning applications that utilize various data sources, such as video, audio, sensor data, and any time-series data.
[0119] S73: Based on the conference connection, send user feature data to the second conference device so that the second conference device inputs the user feature data into the first rendering model to obtain a real-time rendered image, and presents a stereoscopic image of the first user based on the real-time rendered image.
[0120] In this step, in some embodiments, the first conference device and the second conference device establish an end-to-end communication connection, and the first conference device sends user feature data to the second conference device based on the conference connection. In some embodiments, the first conference device sends user feature data to a server based on the conference connection, and the server then sends the user feature data to the second conference device based on the conference connection.
[0121] After receiving user feature data, the second conferencing device calls the first rendering model to process the user feature data, thereby obtaining a real-time rendered image. Then, the second conferencing device presents a stereoscopic image of the first user based on the real-time rendered image. In immersive video conferencing, this embodiment of the application eliminates the need to transmit the entire video stream of the first user to the second conferencing device for 3D presentation, and avoids consuming excessive bandwidth to transmit data that prompts the second conferencing device to perform 3D presentation. This embodiment of the application only requires a small amount of user feature data to enable the second conferencing device to present a stereoscopic image of the first user, thus improving the application scope of the method provided in this embodiment of the application. In addition, the data size of the first rendering model is usually 100MB. After a single transmission, the first rendering model will not be transmitted again. The data size of the user feature data is small, usually less than 1MB. The user feature data needs to be extracted and transmitted in real time. As mentioned above, compared with the method of directly transmitting video streams, such as one channel with a resolution of 1920*1080 and 30fps, the content bandwidth will occupy 1.39Gbps (1920*1080*3*8 / 1024 / 1024 / 1024*30 frames*1 channel = 1.39Gbps). If four channels are used, it will reach 5.56Gbps. Therefore, the embodiments of this application can transmit the second conference device without delay, which is beneficial for the second conference device to quickly and in real time present the stereoscopic image of the first user.
[0122] Presenting a stereoscopic image of the first user based on a real-time rendered image includes the following steps: obtaining preset multi-viewpoint information, generating an image sequence based on the multi-viewpoint information and the real-time rendered image, generating a stereoscopic image of the first user based on the image sequence, and presenting the stereoscopic image of the first user.
[0123] Multi-view information includes pose parameters of multiple virtual cameras, which represent different viewpoints and are used to generate a sequence of rendered images with parallax.
[0124] The acquisition of preset multi-viewpoint information includes: determining a preset number of virtual camera poses based on the set field of view and binocular parallax angle. These preset number of virtual camera poses constitute the multi-viewpoint information. For example, if the field of view is set to a horizontal viewing angle of ±30 degrees, the total field of view is 60 degrees, and the binocular parallax angle is set to 0.5 degrees, according to the formula (30*2) / 0.5 = 120, 120 virtual camera poses need to be preset, with one virtual camera set every 0.5 degrees from -30 degrees to +30 degrees. The pose matrix of these 120 virtual cameras contains the position and orientation information of each virtual camera, and the 120 virtual camera pose matrices constitute the multi-viewpoint information.
[0125] The image sequence comprises multiple reference rendered images, with the real-time rendered image being one of these reference rendered images. Each reference rendered image is rendered from a preset virtual camera pose. Parallax exists between any two reference rendered images due to differences in camera position and orientation; this parallax effect simulates the binocular parallax of the human eye. The parallax between any two adjacent reference rendered images is equal because the spacing between adjacent virtual camera positions is constant (e.g., 0.5 degrees), therefore the parallax between adjacent reference rendered images is also constant.
[0126] The above design enables naked-eye 3D effect that can be achieved for multiple people to view from multiple perspectives: Since each participant is in a different position and facing a different direction, the sequence of images they see will also be different. However, since the parallax between adjacent images is constant, each participant can feel the stereoscopic effect. In other words, even if multiple people watch at the same time, naked-eye 3D effect can be achieved.
[0127] The image sequence rendered in this embodiment allows participants to have a more immersive and engaging experience during the meeting. Each person can view the 3D meeting scene from their own perspective, achieving a truly multi-person, glasses-free 3D effect.
[0128] The multi-viewpoint information includes a preset number of virtual camera poses (e.g., 120 virtual camera poses). Generating an image sequence based on the multi-viewpoint information and real-time rendered images includes: adjusting the pose of a first user in the real-time rendered image according to the pose of the virtual camera in each virtual camera pose to generate a preset number of reference rendered images corresponding to the first user. Each reference rendered image contains a viewpoint associated with a virtual camera pose. Specifically, the first rendering model generates a real-time rendered image based on real-time user feature data (e.g., facial expressions, actions). For each preset virtual camera pose, the position and angle of the first user in the real-time rendered image can be adjusted according to the position and orientation of the virtual camera, thereby generating reference rendered images corresponding to each virtual camera viewpoint. Each reference rendered image contains a viewpoint associated with a virtual camera pose. Through these steps, for example, 120 reference rendered images corresponding to 120 virtual camera poses can be obtained, and these reference rendered images constitute the final image sequence.
[0129] By generating image sequences in this way, the preset multi-viewpoint information can be fully utilized to render the second attendee's image with parallax from different perspectives. In this way, the first attendee corresponding to the second conference device can observe the second attendee from different perspectives, producing a realistic naked-eye 3D effect.
[0130] Stereoscopic images refer to images that can present a realistic three-dimensional spatial effect. They possess parallax effects, creating a sense of three-dimensionality and utilizing the binocular parallax characteristic of the human eye. By inputting preset multi-viewpoint information and real-time user feature data into the first rendering model, the resulting image sequence possesses this stereoscopic effect. When viewers view these images, they can perceive the positional relationship and dynamic changes of the first user in three-dimensional space, generating an immersive experience.
[0131] Generating a stereoscopic image of the first user based on an image sequence includes the following steps: obtaining a preset multi-viewpoint subpixel mapping table and the pixel layout of the display interface of the second conferencing device; determining the size and pixel layout of the stereoscopic image to be synthesized based on the multi-viewpoint subpixel mapping table and the pixel layout; for each pixel position in the stereoscopic image to be synthesized, extracting the pixel value of the corresponding pixel position from the image sequence based on the multi-viewpoint subpixel mapping table; filling the extracted pixel value into the corresponding pixel position in the stereoscopic image to be synthesized to generate a stereoscopic image corresponding to the first user, wherein the filling position and pixel value match the multi-viewpoint subpixel mapping table and the pixel layout of the display interface.
[0132] This application embodiment obtains a preset multi-viewpoint subpixel mapping table, which defines the corresponding position of each pixel in the image sequence. It provides a mapping relationship to match each pixel position in the stereoscopic image to be synthesized with the corresponding position in the image sequence. Through this multi-viewpoint subpixel mapping table, it is possible to determine which pixel values to extract from the image sequence to synthesize the final stereoscopic image.
[0133] Simultaneously, the pixel layout of the display interface of the second conference device is acquired, that is, the pixel size and layout of the display interface are determined. The pixel layout of the display interface includes pixel size and layout method. Pixel size refers to the number of pixels and resolution of the display interface, such as the number of pixels in width and height. Layout method describes how the pixels are arranged on the display interface, such as how many pixels are in each row, how many pixels are in each column, and the rules governing their arrangement. By determining the pixel size and layout, it can be ensured that the generated stereoscopic image can be displayed correctly on the second conference device and that its position and layout are consistent with other images or elements on the interface, thereby providing a better viewing experience and enabling the generated stereoscopic image to be accurately presented to the second participant.
[0134] This application embodiment can determine the width and height of the stereoscopic image to be synthesized based on the pixel size information in the pixel layout. For example, the size of the stereoscopic image can match or adapt to the size of the display interface of the second conference device. Then, based on the arrangement information in the pixel layout, the arrangement of pixels in the stereoscopic image to be synthesized can be determined, such as the number of pixels in each row and column, and the arrangement rules between them. By ensuring that the size and pixel layout of the stereoscopic image to be synthesized match the multi-viewpoint subpixel mapping table, that is, ensuring that each pixel position has a corresponding position in the image sequence, correct matching is possible when extracting pixel values from the image sequence.
[0135] The embodiments of this application ensure that the stereoscopic image to be synthesized is compatible with the display interface of the second conference device and conforms to the preset multi-viewpoint subpixel mapping table, thereby ensuring that the generated stereoscopic image can be correctly displayed on the display device and remains consistent with the position and layout of other images or elements on the interface.
[0136] This embodiment of the application can traverse each pixel position of the stereoscopic image to be synthesized, for example, starting from the top left corner and traversing each pixel position in the stereoscopic image to be synthesized row by row and column by column. During the traversal, for the current pixel position, the position in the image sequence corresponding to the pixel position is found according to the multi-viewpoint subpixel mapping table, where the multi-viewpoint subpixel mapping table can provide position information in the image sequence corresponding to each pixel position in the stereoscopic image to be synthesized. Next, based on the position obtained from the image sequence, the pixel value of the corresponding position is extracted from the image sequence. The extracted pixel value can be the pixel value of a single image or multiple images, depending on the number of viewpoints defined in the multi-viewpoint subpixel mapping table. Then, the extracted pixel value is assigned to the corresponding pixel position in the stereoscopic image to be synthesized, thus each pixel position obtains the pixel value of the corresponding viewpoint.
[0137] In the process of generating stereoscopic images, this application determines the correct filling position based on a multi-viewpoint subpixel mapping table and correctly arranges the corresponding pixel values according to the pixel layout of the display interface to ensure that the generated stereoscopic image is consistent with the preset requirements. This enables the generation of high-quality stereoscopic images while maintaining accuracy and consistency, providing a multi-viewpoint experience that adapts to different display devices, thereby enhancing the user's viewing experience and meeting the demand for stereoscopic images. Furthermore, the precise pre-construction of the rendering model, instead of real-time camera acquisition and real-time 3D portrait modeling, allows for flexible selection of 2D views with different numbers of viewpoints during the later rendering process without affecting the clarity and accuracy of the output view, and also saves significant computational power for real-time 3D modeling.
[0138] In some embodiments, before acquiring the real-time user image of the first user, the method further includes the following steps: determining the 3D model data of the first user, the 3D model data being used to represent a first rendered model, sending the 3D model data to a second conference device so that the second conference device can parse the 3D model data into the first rendered model, and saving the 3D model data of the first rendered model locally on the second conference device.
[0139] Determining the 3D model data of the first user includes: acquiring 2D image data, which includes multiple target user images taken from different side angles, and generating 3D model data based on the multiple target user images.
[0140] It is understandable that the target user image is the original image directly captured by the camera, or the target user image is an image obtained by cutting out the video stream captured by the camera.
[0141] It is also understood that when the first rendering model is a 3D character rendering model, the target user image can be the character region image of the first user, or the target user image can be the character region image extracted by the first conferencing device from the image captured by the first user. It is also understood that when the first rendering model includes both a 3D character rendering model and an environment rendering model, the target user image includes the character region image of the first user and the corresponding environmental background image.
[0142] In some embodiments, the first rendering model includes a 3D character rendering model, the 3D character rendering model includes a first head rendering model, the user feature data includes first head feature data, the real-time rendering image includes a first head rendering image, and inputting the user feature data into the first rendering model to obtain the real-time rendering image includes the following steps: inputting the first head feature data into the first head rendering model to obtain the first head rendering image.
[0143] In some embodiments, the 3D character rendering model includes a first human body rendering model, the user feature data includes first human body feature data, and the real-time rendering image includes a first human body rendering image. Inputting the user feature data into the first rendering model to obtain the real-time rendering image includes the following steps: inputting the first human body feature data into the first human body rendering model to obtain the first human body rendering image.
[0144] In some embodiments, the 3D character rendering model includes a first head rendering model and a first human body rendering model, the user feature data includes first head feature data and first human body feature data, and the real-time rendering image includes a first head rendering image and a first human body rendering image. Inputting the user feature data into the first rendering model to obtain the real-time rendering image includes the following steps: inputting the head feature data into the first head rendering model to obtain the first head rendering image, inputting the first human body feature data into the first human body rendering model to obtain the first human body rendering image, and combining the first head rendering image and the first human body rendering image for rendering to obtain the final real-time rendering image.
[0145] The above method utilizes a first head rendering model to process first head feature data and a first human body rendering model to process first human body feature data. Then, the first head rendering image and the first human body rendering image are combined for rendering; that is, the rendering results of the first user's head and body are superimposed to obtain the final real-time rendered image of the first user. This method models and renders head expressions and the human body separately, which can better capture the real-time features of the first user, and finally generate a complete real-time rendered image of the first user through combined rendering. Therefore, it can fully utilize the technical advantages of 3D head reconstruction and human body capture, and can more accurately reflect the real-time state and movement of the first user in the meeting.
[0146] Generating 3D model data from multiple target user images includes: extracting person region images from multiple target user images, generating 3D sub-model data of the person based on the person region images, using the 3D sub-model data of the person to represent the 3D person rendering model of the first user, and determining the 3D model data based on the 3D sub-model data of the person.
[0147] In some embodiments, determining 3D model data based on character 3D sub-model data includes: extracting environmental region images from multiple target user images; generating target environment 3D sub-model data based on the environmental region images; the target environment 3D sub-model data being used to represent the environment rendering model of the first user; and determining 3D model data based on the target environment 3D sub-model data and the character 3D sub-model data. In this embodiment, the character and background environment captured in the actual scene are segmented, and separate character rendering models and environment rendering models are established. Then, the character rendering model and environment rendering model are merged to obtain the first rendering model.
[0148] In some embodiments, determining 3D model data based on character 3D sub-model data includes: acquiring preset environment 3D sub-model data, whereby the preset environment 3D sub-model data represents the environment rendering model of the first user; and determining 3D model data based on the preset environment 3D sub-model data and the character 3D sub-model data. This application embodiment segments character information from actual scene footage to create a separate character rendering model, while the environment rendering model uses a pre-set realistic background, thus enhancing the user's experience of customizing the environment background. Overall, this application embodiment achieves the segmentation of characters and background environment in scene footage, and the creation and reconstruction of 3D models, resulting in a more flexible and vivid construction of the overall environment for immersive meetings and a more three-dimensional 3D character imaging effect.
[0149] It is understood that, in this embodiment, the 3D character sub-model data can be directly sent as 3D model data to the second conference device, which then parses it into a 3D character rendering model. It is also understood that, in this embodiment, the 3D character sub-model data and environment 3D sub-model data can be packaged into 3D model data and sent to the second conference device, which then parses it into a 3D character rendering model and an environment rendering model. Furthermore, it is understood that, in this embodiment, the 3D character sub-model data and environment 3D sub-model data can be fused to obtain fused 3D model data, which is then sent to the second conference device, which parses it into a 3D character rendering model and an environment rendering model. The environment 3D sub-model data can be generated based on the first user's real-time environment data, or it can be generated based on preset environment data.
[0150] The person region image includes a head region image and / or a body region image, wherein the head region image is used to represent the head of the first user and the body region image is used to represent the body of the first user.
[0151] The 3D character sub-model data includes 3D head data and 3D human body data. The 3D character rendering model includes a head rendering model and a human body rendering model. The 3D head data is used to represent the head rendering model, and the 3D human body data is used to represent the human body rendering model. The 3D character sub-model data generated from the character region image includes: 3D head data trained from the head region image, and 3D human body data trained from the human body region image.
[0152] The head rendering model includes a first Gaussian sphere learning model and a head rendering model. The 3D head data is trained based on the head region image and includes the following steps:
[0153] S101: Determine the head control information corresponding to the target head region image.
[0154] S102: Obtain the first Gaussian sphere parameters output by the first Gaussian sphere learning model when the target head region image is used as the training input sample. The first Gaussian sphere parameters are used to constrain the Gaussian sphere of the head rendering model.
[0155] S103: Input the first Gaussian sphere parameters and head control information into the head rendering model to obtain the head rendering image.
[0156] S104: Determine the first image difference information between the head rendering image and the target head region image.
[0157] S105: Train the head rendering model based on the difference information of the first image to obtain 3D head data.
[0158] In S101, head control information is used to represent the head features of the target head corresponding to the target head region image. The head features include the target head's pose, shape, and expression. The target head's pose can be represented by the camera pose captured by the camera, and the shape and expression can be represented by a combination of multiple head key points (i.e., control points). Therefore, in some embodiments, the head control information includes the camera pose and the head key points of the target head region image. The camera pose is used to indicate the pose of the target head captured by the camera, and the multiple head key points can effectively combine to represent the head's shape and expression.
[0159] The target head region image is a head region image in a state of being to be trained. When the target head region image is used to complete the training operation in this embodiment of the application, the training state of the target head region image is marked as a trained state.
[0160] In S102, determining the head control information corresponding to the target head region image includes the following steps: determining the camera pose of the target head region image and multiple head key points of the target head region image based on a preset visual model, wherein the multiple head key points meet the requirements of the human head FLAME model. The preset visual model can be a Colmap model or a MediaPipe model.
[0161] The first Gaussian sphere learning model is a model used to learn and output the parameters of the first Gaussian sphere. This model can employ any suitable model architecture, such as a neural network model or other custom learning model. The head rendering model is a model used to render and output the head rendering image. The head rendering model is connected to the first Gaussian sphere learning model, and the first Gaussian sphere parameters output by the first Gaussian sphere learning model can be used as an input to the head rendering model.
[0162] In S103, the head rendering model is a head model used to generate a head rendering image, wherein the head rendering image is a 2D rendering image. This embodiment of the application can synthesize multiple head rendering images into a 3D head region image. The head rendering model supports various head model algorithms and various rendering algorithms, such as the head FLAME algorithm and the 3D Gaussian Splatting rendering algorithm.
[0163] In S104, the first image difference information is used to represent the degree of matching between the head rendering image and the target head region image. The first image difference information can be represented by loss value, histogram similarity, or feature matching degree.
[0164] The first image difference information includes a first total loss value. Determining the first image difference information between the head-rendered image and the target head region image includes the following steps: determining the feature loss value, image similarity loss value, and image difference loss value between the head-rendered image and the target head region image; multiplying the feature loss value by a first weight coefficient to obtain a first weighted loss value; multiplying the image similarity loss value by a second weight coefficient to obtain a second weighted loss value; multiplying the image difference loss value by a third weight coefficient to obtain a third weighted loss value; and adding the first weighted loss value, the second weighted loss value, and the third weighted loss value together to obtain the first total loss value.
[0165] For example, in embodiments of this application, the first total loss is calculated according to the following formula, where the formula is: L0=λ1L1+λ2L ssim +λ3L vgg , Where L0 is the total first loss value, L1 is the feature loss value, and L ssim The image similarity loss value (SSIM) is L. vgg The image difference loss value (VGGLoss) is given by λ1, where λ1 is the first weighting coefficient, λ2 is the second weighting coefficient, and λ3 is the third weighting coefficient. t For the t-th original head region image, render t This is the head rendering image corresponding to the t-th original head region image.
[0166] In S105, when the first image difference information is the first total loss value, this embodiment of the application optimizes the network parameters of the head rendering model using a preset gradient descent algorithm based on the first total loss value to obtain 3D head data.
[0167] After training the head rendering model based on image difference information in this embodiment, the training state of the target head region image is marked as trained. Next, this embodiment detects whether the head rendering model meets the training stop condition. If it does, the training operation is stopped; otherwise, one original head region image is selected from at least one original head region image in the untrained state as the new target head region image. The process returns to the step of determining the head control information corresponding to the target head region image, so as to continue training the head rendering model.
[0168] To illustrate in detail the training method of the head rendering model provided in the embodiments of this application, the embodiments of this application combine... Figure 10 The following is a detailed explanation of this matter:
[0169] Please see Figure 10 In this embodiment, multiple raw head region images (img) are acquired, and a pre-trained visual model is used to extract the camera poses (pose1, pose2, ..., pose) corresponding to each raw head region image. t and the k facial key points G that satisfy the human head FLAME model t ∈R k×3 Among them, the original head region image (img) j Head control information includes camera pose. j and facial key points G j ∈R k×3 .
[0170] During training, this embodiment of the application uses the original head region image img1 as the target head region image, and obtains the Gaussian sphere parameters output by the Gaussian sphere learning model 101 when the original head region image img1 is used as the training input sample. The camera pose (pose1) and facial key points (G1∈R) of the original head region image (img1) are determined. k×3 and Gaussian sphere parameters The head rendering model 102 is input into the head rendering model. The head rendering model 102 transmits head rendering parameters to the first 3D Gaussian rendering model 103 to obtain the head rendering image render1. Next, in this embodiment, the image difference information L between the head rendering image render1 and the original head region image img1 is calculated. 01 Based on image difference information L 01 Train the head rendering model. Due to image difference information L... 01 The model has not yet converged; therefore, this embodiment of the application requires continued training of the head rendering model.
[0171] Next, in this embodiment of the application, the original head region image img2 is used as the target head region image, and the Gaussian sphere parameters output by the Gaussian sphere learning model 101 when the original head region image img2 is used as the training input sample are obtained. The camera pose pose2 and facial key points G2∈R of the original head region image img2 are obtained. k×3 and Gaussian sphere parameters The head rendering model 102 is input as a head rendering model. The head rendering model 102 transmits head rendering parameters to the 3D Gaussian rendering model 103 to obtain the head rendering image render2. Next, in this embodiment, the image difference information L between the head rendering image render2 and the original head region image img2 is calculated. 02 Based on image difference information L 02 Train the head rendering model. Due to image difference information L... 02 The convergence has not yet occurred; therefore, this embodiment requires continued training of the head rendering model. The same logic applies, which will not be elaborated upon here. When the image difference information L... 02 When the training converges to a preset threshold, the embodiment of this application stops training the head rendering model and outputs the trained head rendering model.
[0172] In summary, this application embodiment utilizes a first Gaussian sphere learning model to automatically learn the output parameters of the first Gaussian sphere, compensating for the deviations caused by head control information in image generation. This improves the performance of the head rendering model. Furthermore, by using the image difference information between the rendered head image and the target head region image as supervisory information to train the head rendering model, it is beneficial to improve the accuracy and reliability of the head rendering model, enabling it to output a head rendering image with a high degree of matching to the original head region image. In addition, this application embodiment can complete the training of the head rendering model using only target video data collected by a single camera, eliminating the need for multiple cameras, thus improving the generation efficiency and application scope of the head rendering model.
[0173] The first Gaussian sphere parameters include the first Gaussian sphere rendering parameters and the first position offset parameters. The first Gaussian sphere rendering parameters are used to represent the rendering attributes of the Gaussian sphere in the 3D Gaussian rendering model, and the first position offset parameters are used to indicate the position offset of the Gaussian sphere in the 3D Gaussian rendering model.
[0174] The first Gaussian sphere learning model includes a first neural network model, which is used to learn the first Gaussian sphere rendering parameters and the first position offset parameters. The first neural network model can be a Multi-Layer Perceptron (MLP) network, a Convolutional Neural Network, a Transformer Neural Network, etc. It is understood that the architecture and number of elements in the first neural network model can be customized by the designer according to business needs.
[0175] The head rendering model includes a human head FLAME model and a first 3D Gaussian rendering model. The human head FLAME model is used to construct the head model, and the first 3D Gaussian rendering model is used to render the head model to output a head rendering image. The first Gaussian sphere rendering parameters include the spherical harmonic function color, opacity ο, rotation matrix R, and scale vector S. The first position offset parameters include the offset vector Δμ of the Gaussian sphere's center point and the weight values W of the vertices associated with the pose of the target head.
[0176] The first image difference information includes a first total loss value. Training the head rendering model based on the first image difference information includes: optimizing the network parameters of the first Gaussian sphere learning model using a preset gradient descent algorithm based on the first total loss value to obtain an adjusted head rendering model.
[0177] The steps to obtain the first Gaussian sphere parameters output by the first Gaussian sphere learning model when the target head region image is used as the training input sample include the following: when the target head region image is used as the training input sample, determine the vertex features of the vertices in the human head FLAME model that are used to associate with the pose of the target head, input the vertex features of the vertices into the first neural network model, and obtain the first Gaussian sphere rendering parameters and the first position offset parameters.
[0178] Vertices are the location points in the human head FLAME model used to associate with the pose of the target head, such as... Figure 8 As shown in the left-hand head image, the human head FLAME model is divided into multiple meshes, each with vertices. A standard human head FLAME model contains 5023 vertices, each assigned a weight value W to represent the degree of influence of each joint on that vertex. During deformation, based on the vertex positions and corresponding weight values before deformation, this embodiment can calculate the deformed vertex positions. Vertex features are used to characterize vertices; it is understood that this embodiment can use any feature representation method to generate vertex features.
[0179] Determining the vertex features of vertices in a head FLAME model used to associate with the pose of a target head includes the following steps: obtaining the vertex positions of the vertices in the head FLAME model used to associate with the pose of the target head; performing feature sampling operations in a preset sampling space based on the vertex positions to obtain the vertex features. For example, in this embodiment, based on the vertex positions, a random sampling algorithm, a distance-based sampling algorithm, a distribution-based sampling algorithm, or a sampling algorithm based on an adversarial network model can be used to randomly sample in the preset sampling space to obtain the vertex features.
[0180] For example, a human head FLAME model is configured with a three-dimensional coordinate system {x c ,y c ,z c The position of the i-th vertex in the three-dimensional coordinate system is (x... i ,y i ,z i In this embodiment of the application, the vertex position (x) of the i-th vertex is used as the basis for the calculation. i ,y i ,z i Perform feature sampling operations in the preset sampling space to obtain the vertex features of the i-th vertex.
[0181] In some embodiments, please refer to Figure 11 The first neural network model is a multi-layer MLP neural network. A multi-layer MLP neural network can automatically learn and output the spherical harmonic function color, opacity ο, rotation matrix R, scale vector S, offset vector Δμ, and weight values W. For example, the expression for a multi-layer MLP neural network is:
[0182] W, color, o, Δu, R, S = mlp(f).
[0183] In some embodiments, please refer to Figure 12 The first neural network model consists of six multi-layer MLP neural networks. These six MLP neural networks can automatically learn to output the spherical harmonic function color, opacity o, rotation matrix R, scale vector S, offset vector Δμ, and weight values W, respectively. For example, the expression for the six multi-layer MLP neural networks is:
[0184] W = mlp1(f)
[0185] color = mlp2(f)
[0186] o = mlp3(f)
[0187] Δu=mlp4(f)
[0188] R = mlp5(f)
[0189] S = mlp6(f)
[0190] This embodiment uses a human head FLAME model as the baseline human head model, and then fuses the first Gaussian sphere rendering parameters and the first position offset parameters to create a head rendering image with a higher degree of matching. Furthermore, this embodiment automatically learns and outputs Gaussian sphere parameters through a first neural network model. This not only reduces the dependence on hardware—for example, while related technologies require 16 or hundreds of cameras—this embodiment only requires one or two cameras to train a high-performance head rendering model. Moreover, the Gaussian sphere parameters automatically learned and output by the first neural network model can compensate for the bias caused by manual annotation, decoupling human factors and improving the performance of the head rendering model. The trained head rendering model has higher robustness.
[0191] The process of inputting the first Gaussian sphere parameters and head control information into the head rendering model to obtain the head rendering image includes the following steps: determining the first target position of the first target vertex after the target head changes based on the first position offset parameters and head control information. The first target vertex is a vertex in the target head region image. The first target position of each first target vertex and the first Gaussian sphere rendering parameters are input into the human head FLAME model to obtain the head rendering parameters corresponding to each first target vertex. The head rendering parameters corresponding to each first target vertex are input into the first 3D Gaussian rendering model to obtain the head rendering image. The first target vertex matches the requirements of the human head FLAME model.
[0192] Determining the first target position of the first target vertex after the target head changes based on the first position offset parameter and head control information includes: updating the vertex position of the first target vertex according to the first position offset parameter to obtain the first static position of the first target vertex before the target head changes; and moving the first static position according to the head control information to obtain the first target position of the first target vertex after the target head changes.
[0193] In some embodiments, after obtaining the first static position of the first target vertex before the change in the target head, this application embodiment can input the first static position of the first target vertex before the change in the target head into a first 3D Gaussian rendering model to obtain an intermediate rendered image of the target head in a static state. The trainer can compare the intermediate rendered image with the target head region image to verify the correctness of the head rendering model during training. Please refer to... Figure 10 The intermediate rendered image is renderc.
[0194] The first position offset parameter includes the offset vector of the center point of the Gaussian sphere. The vertex position of the first target vertex is updated according to the first position offset parameter to obtain the first static position of the first target vertex before the target head changes. This includes adding the vertex position of the first target vertex to the offset vector of the center point of the Gaussian sphere to obtain the first static position of the first target vertex before the target head changes.
[0195] For example, in this embodiment of the application, the first static position of the first target vertex before the target head changes is calculated according to the following formula:
[0196] (x' t y' t , z' t )=(x t y t , z t )+Δu
[0197] Among them, (x' t y' t , z' t (x) represents the first static position of the first target vertex before the target head changes. t y t , z t ) represents the vertex position of the first target vertex, and Δu is the offset vector of the center point of the Gaussian sphere.
[0198] The head control information includes the camera pose and head key points of the target head region image. The camera pose is used to indicate the pose of the camera when shooting the target head. The first position offset parameter includes the weight values of the vertices associated with the pose of the target head. Moving the first static position according to the head control information to obtain the first target position of the first target vertex after the target head changes includes: updating the position of the first target vertex using the first static position, camera pose, head key points and vertex weights according to a preset linear blending skinning algorithm to obtain the first target position of the first target vertex after the target head changes.
[0199] For example, in this embodiment of the application, the position of the first target after the change of the first target vertex at the target head is calculated according to the following formula:
[0200] (x″ l ,y″ l , z″ l ) = LBS(G t W t , {x' l y' l , z' l})
[0201] Among them, (x″ t ,y″ t , z″ t ) represents the first target position after the first target vertex changes in the target head, LBS() is the LBS (linear blend skinning) function of the preset linear blend skinning algorithm, G t For the facial key points corresponding to the first target vertex, W t This represents the weight value of the first target vertex.
[0202] The linear blending skinning component of the human head FLAME model can use the LBS algorithm to process the first target position and the first Gaussian sphere rendering parameters of each first target vertex, thereby obtaining the head rendering parameters corresponding to each first target vertex, where each first target vertex corresponds to a Gaussian sphere.
[0203] The first 3D Gaussian rendering model processes the head rendering parameters corresponding to each first target vertex based on the 3D Gaussian Splatting algorithm to obtain the head rendering image.
[0204] This application embodiment first determines the first static position of the first target vertex before the target head changes, and then determines the first target position of the first target vertex after the target head changes based on the first static position, thus obtaining the head rendering image efficiently.
[0205] The human body capture model includes a second Gaussian sphere learning model and a human body rendering model. The 3D human body data obtained by training based on human body region images includes the following steps:
[0206] S111: Obtain human control information corresponding to the target human body region image, where the target human body region image is a human body region image in the training state.
[0207] S112: Obtain the second Gaussian sphere parameters output by the second Gaussian sphere learning model when the target human body region image is used as the training input sample. The second Gaussian sphere parameters are used to constrain the Gaussian sphere of the human body rendering model.
[0208] S113: Input the second Gaussian sphere parameters and human body control information into the human body rendering model to obtain the human body rendering image.
[0209] S114: Determine the second image difference information between the rendered human body image and the target human body region image.
[0210] S115: Train the human body rendering model based on the difference information of the second image to obtain 3D human body data.
[0211] In S111, the human body control information includes the camera position and pose obtained in the above steps. Specifically, this refers to the camera's pose when acquiring the target human body region image in the target video data, i.e., the camera's position and orientation. The human body control information also includes human keypoints in the target human body region image that are associated with specific joint positions of the human body, such as... Figure 9 As shown, Figure 9 The image shows the locations of 24 key points on a standard SMPL human body model. Specifically, the key points in this embodiment can be... Figure 9 The 24 key points of the human body are shown.
[0212] For human body region images in training mode, the target human body region images can be labeled and prepared for training machine learning models. These images contain samples of the human body in different actions or poses. For example, target human body region images include labeled human keypoints, which are points at specific joint positions of the human body, such as the body, shoulders, elbows, and knees. The positional information of these human keypoints can be used to train pose estimation, motion analysis, etc.; they also include action category labels, where each target human body region image can be assigned a corresponding action category label. These labels describe the specific action performed by the human body in the image, such as walking, running, raising an arm, etc.; they can also include temporal series information of the images. If the training data is a continuous video sequence, each target human body region image can belong to a time series, and there is temporal continuity between the images; they can also include other image attributes, such as image resolution, human body scale variation information, etc.
[0213] In S112, the second Gaussian sphere learning model is a model used to learn and output the parameters of the second Gaussian sphere. This model can employ any suitable model architecture, such as a neural network model or other custom learning model. The human body rendering model is a model used to render and output human body rendered images. This model is connected to the second Gaussian sphere learning model, and the second Gaussian sphere parameters output by the learning model can be used as an input to the human body rendering model.
[0214] Specifically, in this embodiment, an image of a target human body region, along with its associated position, pose, and key points, is input into a preset human body rendering model. The model processes and analyzes the input data, extracting the human body's pose and motion information, and then renders a human body image. Before inputting the target human body image into the human body rendering model, preprocessing can be performed on the target human body region image, including image normalization and resizing. For key points, coordinate normalization or other processing can also be performed to adapt to the model's input requirements.
[0215] In S113, the human body rendering model is a human body model used to generate human body rendering images. These human body rendering images are 2D rendering images. This embodiment of the application can synthesize multiple human body rendering images into a 3D human body region image. The human body rendering model supports various human body model algorithms and various rendering algorithms, such as the SMPL algorithm for human body modeling and the 3D Gaussian Splatting rendering algorithm for rendering.
[0216] In S114, the second image difference information is used to represent the degree of matching between the rendered human body image and the target human body region image. The second image difference information can be represented by loss value, histogram similarity, or feature matching degree.
[0217] The second image difference information includes the second total loss value. Determining the second image difference information between the rendered human body image and the target human body region image includes the following steps: determining the feature loss value, image similarity loss value, and image difference loss value between the rendered human body image and the target human body region image; multiplying the feature loss value by the fourth weight coefficient to obtain the fourth weighted loss value; multiplying the image similarity loss value by the fifth weight coefficient to obtain the fifth weighted loss value; multiplying the image difference loss value by the sixth weight coefficient to obtain the sixth weighted loss value; and adding the fourth, fifth, and sixth weighted loss values together to obtain the second total loss value.
[0218] For example, in embodiments of this application, the second total loss is calculated according to the following formula, where the formula is: loss=λ4L4+λ5L ssim +λ6L vgg , Where loss is the total value of the second loss, L4 is the feature loss value, and L... ssim L represents the image similarity loss value. vgg The image difference loss value is represented by λ4, which is the fourth weighting coefficient, λ5, which is the fifth weighting coefficient, and λ6, which is the sixth weighting coefficient. t For the t-th original head region image, render t This is the head rendering image corresponding to the t-th original head region image.
[0219] In S115, this embodiment of the application uses the gradient descent algorithm to optimize the neural network parameters (specifically the hyperparameters of the MLP neural network) in the neural network model, as well as the parameters of the SMPL human body model.
[0220] The optimization formula for the parameters of the SMPL human body model is as follows: It refers to the parameter learning rate, and AC is the motion capture module in the SMPL human model.
[0221]
[0222] The above-mentioned optimization of the parameters of the neural network model and the SMPL human body model using the gradient descent algorithm can improve model performance, optimize network structure, improve the accuracy of the human body model, and accelerate the convergence speed of the model, thereby generating higher quality, more realistic and lifelike human body rendering images.
[0223] To illustrate in detail the training method for the human body rendering model provided in the embodiments of this application, the embodiments of this application are combined with Figure 13 The following is a detailed explanation of this matter:
[0224] Please see Figure 13 This application embodiment acquires multiple original human body images frames. t The camera pose corresponding to each original human image is extracted using a pre-trained visual model. t Human body segmentation image (img) t The COLMAP tool can be used to extract the camera position and pose (pose1, pose2, ... pose) for each frame. t Human body segmentation models or image matting tools can be used to obtain the corresponding human body segmentation images (img1, img2, ..., img) for each frame of the image. t .
[0225] Furthermore, a motion capture neural network module AC is introduced to obtain the three-dimensional coordinates G1, G2, ... G of the human body key points in the SMPL model space for each frame of the original human body image. t G t ∈R 24×3 G t =AC(frame) t The standard SMPL model contains 24 human keypoints, distributed as follows: Figure 9 As shown.
[0226] During training, the embodiments of this application use human body segmentation images (img) t The image is used as the target human body region, and the Gaussian sphere learning model 1301 is used to segment the human body image. t The Gaussian sphere parameters output as training input samples will frame the original human image. t camera pose t The key points and Gaussian sphere parameters of specific joint positions in the human body model are input into the human body rendering model 1302 to obtain a human body rendering image. Next, this embodiment of the application calculates the image difference information between the human body rendering image and the target human body region image, and trains the human body rendering model based on the image difference information. Since the image difference information has not yet converged, this embodiment of the application needs to continue training the human body rendering model.
[0227] The training steps described above are analogous and will not be repeated here. When the image difference information converges to a preset threshold, this embodiment of the application stops training the human body rendering model and outputs the trained human body rendering model.
[0228] In summary, this embodiment utilizes the second Gaussian sphere parameters automatically learned by the second Gaussian sphere learning model to compensate for the deviations caused by human body control information in image generation. This improves the performance of the human body rendering model. Furthermore, by using the image difference information between the rendered human body image and the target human body region image as supervisory information to train the human body rendering model, it is beneficial to improve the accuracy and reliability of the human body rendering model, enabling it to output rendered human body images with a high degree of matching to the original human body image. In addition, this embodiment only requires target video data collected by a single camera to help complete the training of the human body rendering model, eliminating the need for multiple cameras, which is beneficial to improving the generation efficiency and application scope of the human body rendering model.
[0229] In this embodiment of the application, the second Gaussian sphere learning model includes a second neural network model, the human body rendering model includes an SMPL human body model and a second 3D Gaussian rendering model, and the second Gaussian sphere parameters include second Gaussian sphere rendering parameters and second position offset parameters.
[0230] The process of inputting the second Gaussian sphere parameters and human body control information into the human body rendering model to obtain the human body rendering image includes the following steps:
[0231] S121. When the target human image is used as a training input sample, determine the vertex features of the vertices in the SMPL human model used to associate with the pose of the target human.
[0232] S122. Input the vertex features of the vertex into the second neural network model to obtain the second Gaussian sphere rendering parameters and the second position offset parameters.
[0233] S123. Determine the second target position of the second target vertex after the target human body changes based on the second position offset parameter and human body control information. The second target vertex is the vertex in the target human body region image.
[0234] S124. Input the second target position and second Gaussian sphere rendering parameters of each second target vertex into the SMPL human body model to obtain the Gaussian sphere rendering parameters corresponding to each second target vertex.
[0235] S125. Input the rendering parameters of the Gaussian sphere corresponding to each second target vertex into the 3D Gaussian rendering model to obtain the human body rendering image.
[0236] In S121, the key points of the SMPL human body model are obtained. These key points are the key points at specific joint locations within the SMPL human body model (e.g.,...). Figure 9 (As shown); pose transformation and shape adjustment are performed on keypoints to obtain the vertex positions of vertices in the SMPL human model used to associate with the pose of the target human body; feature sampling is performed in a preset sampling space based on the vertex positions to obtain the vertex features. The SMPL human model is a keypoint-based parametric model. By performing pose transformation and shape adjustment on keypoints, the corresponding vertex positions can be calculated. These vertex positions represent the shape of the human model under a given pose and shape. Specifically, the SMPL human model uses the Linear Blend Skin (LBS) method to calculate vertex positions. The LBS method maps the positions of human keypoints to vertex positions by performing a weighted linear combination of human keypoints. These weights are calculated based on the pose and shape parameters of the human model. Therefore, vertex positions are calculated from human keypoints using the LBS method.
[0237] In S122, the vertex features of the vertices are input into the second neural network model to obtain the second Gaussian sphere rendering parameters and the second position offset parameters. The second Gaussian sphere rendering parameters represent the rendering attributes of the Gaussian sphere in the second 3D Gaussian rendering model, and the second position offset parameters indicate the position offset of the Gaussian sphere in the second 3D Gaussian rendering model. Specifically, the vertex features of the vertices are input into the second neural network model, which outputs the second Gaussian sphere rendering parameters, composed of the spherical harmonic function, opacity, rotation matrix, and scale vector, and the second position offset parameters, composed of the offset vector of the center point of the Gaussian sphere and the weight values of the vertices.
[0238] The second Gaussian sphere rendering parameters and the second position offset parameters can be used to control the rendering attributes and position offset of the Gaussian sphere in the second 3D Gaussian rendering model, thereby finely controlling the appearance, lighting and other attributes of the generated human body rendering image.
[0239] In step S123, the target position of the second target vertex after the target human body changes is determined based on the second position offset parameter and human body control information. The second target vertex is a vertex in the target human body region image. Specifically, the vertex position of the second target vertex is updated according to the position offset parameter to obtain the second static position of the second target vertex before the target human body's action changes; the second static position is moved according to the human body control information to obtain the second target position of the second target vertex after the target human body's action changes.
[0240] The second position offset parameter includes the offset vector of the center point of the Gaussian sphere. The vertex position of the second target vertex is updated according to the second position offset parameter to obtain the second static position of the second target vertex before the action of the target human body changes. This includes adding the coordinates corresponding to the vertex position of the second target vertex to the offset vector of the center point of the Gaussian sphere to obtain the second static position of the second target vertex before the action of the target human body changes.
[0241] Human body control information includes the camera pose when acquiring the target human body region image in the target video data, and human body key points associated with specific joint positions in the target human body region image. The second position offset parameter includes the weight value of the vertex. The second static position is moved according to the human body control information to obtain the second target position of the second target vertex after the target human body's action changes. This includes updating the position of the second target vertex of the second static position according to the camera pose, the coordinates of the human body key points, and the weight value of the vertex to obtain the second target position of the second target vertex after the target human body's action changes.
[0242] This step determines the second target vertex's position after the target human body changes, based on the second position offset parameter and human body control information. By combining camera pose, human body keypoints, and vertex weights, the position of the second target vertex after the target human body changes can be accurately calculated, which helps maintain the accuracy and consistency of the rendered image.
[0243] In S124, the second target position and second Gaussian sphere rendering parameters of each second target vertex are input into the SMPL human body model to obtain the Gaussian sphere rendering parameters corresponding to each second target vertex. The Gaussian sphere rendering parameters can describe the rendering properties of the Gaussian sphere on the target human body, providing necessary information for subsequent rendering processes.
[0244] In step S125, the rendering parameters of the Gaussian sphere corresponding to each second target vertex are input into the second 3D Gaussian rendering model to obtain a human body rendering image. By using Gaussian sphere parameters and human body control information for rendering, the generated image can achieve higher realism and fidelity. Furthermore, the appearance and lighting effects of the image can be adjusted according to different Gaussian sphere parameters.
[0245] In this embodiment, the SMPL human body model, the second neural network model, and the second 3D Gaussian rendering model are combined. By calculating and inputting vertex features, second Gaussian sphere rendering parameters, and second position offset parameters, accurate control and generation of human body rendering images are achieved. This method can improve the accuracy, realism, and personalization of rendered images, providing an effective approach for generating realistic human body rendering images.
[0246] The second Gaussian sphere parameters include the second Gaussian sphere rendering parameters and the second position offset parameters. The second Gaussian sphere rendering parameters are used to represent the rendering attributes of the Gaussian sphere in the second 3D Gaussian rendering model, and the second position offset parameters are used to indicate the position offset of the Gaussian sphere in the second 3D Gaussian rendering model.
[0247] The second Gaussian sphere learning model includes a second neural network model. This second neural network model is used to learn the second Gaussian sphere rendering parameters and the second position offset parameters. The second neural network model can be a Multi-Layer Perceptron (MLP) network, a Convolutional Neural Network, a Transformer Neural Network, etc. Understandably, the architecture and number of the second neural network models can be customized by the designer according to business needs.
[0248] The human body rendering model includes an SMPL human body model and a second 3D Gaussian rendering model. The SMPL human body model is used to construct the human body model, and the second 3D Gaussian rendering model is used to render the human body model to output a human body rendering image. The second Gaussian sphere rendering parameters include the spherical harmonic function color, opacity o, rotation matrix R, and scale vector S. The second position offset parameters include the offset vector Δu of the Gaussian sphere center point and the weight values W of the vertices associated with the pose of the target human body. These weight values W can be the default weight values of the SMPL human body model.
[0249] Here, vertices are the location points in the SMPL human model used to associate with the pose of the target human body, such as... Figure 8 As shown, the SMPL human body model is divided into 24 human body keypoints. Based on these keypoints, pose transformation and shape adjustment are performed to obtain the vertex positions in the SMPL human body model used to associate with the pose of the target human body. Each vertex can be configured with a weight value W to represent the degree of influence of each joint on that vertex. During deformation, based on the vertex positions and corresponding weight values before deformation, this embodiment can calculate the vertex positions after deformation. Vertex features are used to characterize vertices; it is understood that this embodiment can use any feature representation method to generate vertex features.
[0250] For example, in this embodiment of the application, based on the vertex position, a random sampling algorithm, a distance-based sampling algorithm, a distribution-based sampling algorithm, or a sampling algorithm based on an adversarial network model can be used to randomly sample in a preset sampling space to obtain the vertex features of the vertex.
[0251] In the process of inputting the vertex features of the vertices into the second neural network model to obtain the second Gaussian sphere rendering parameters and the second position offset parameters, the neural network model is a multi-layer MLP neural network or multiple multi-layer MLP neural networks.
[0252] This application embodiment uses the SMPL human body model as the baseline human body model, and then integrates the second Gaussian sphere rendering parameters and the second position offset parameters to create a human body rendering image with a higher degree of matching. Furthermore, this application embodiment automatically learns and outputs the Gaussian sphere parameters through a second neural network model. This not only reduces the dependence on hardware—for example, while related technologies require 16 or hundreds of cameras—this application embodiment only requires one or two cameras, or a small number of cameras, to achieve the goal of training a high-performance human body rendering model. Moreover, the Gaussian sphere parameters automatically learned and output by the neural network model can compensate for the deviations caused by manual annotation, decoupling human factors and improving the performance of the human body rendering model. The trained human body rendering model has higher robustness.
[0253] The process of inputting the second Gaussian sphere parameters and human body control information into the human body rendering model to obtain the human body rendering image includes: determining the second target position of the second target vertex after the target human body changes based on the second position offset parameters and human body control information, where the second target vertex is a vertex in the target human body region image; inputting the second target position of each second target vertex and the second Gaussian sphere rendering parameters into the SMPL human body model to obtain the Gaussian sphere rendering parameters corresponding to each second target vertex; and inputting the Gaussian sphere rendering parameters corresponding to each second target vertex into the second 3D Gaussian rendering model to obtain the human body rendering image.
[0254] Specifically, the vertex position of the second target vertex is updated according to the second position offset parameter to obtain the second static position of the second target vertex before the action of the target human body changes; the second static position is moved according to the human body control information to obtain the second target position of the second target vertex after the action of the target human body changes.
[0255] For example, in this embodiment of the application, the second static position of the second target vertex before the target human body changes is calculated according to the following formula:
[0256] (x″ l ,y″ l , z″ l ) = LBS(G t W t , {x' l y' l , z' l})
[0257] (x″ t ,y″ t , z″ t) represents the second static position of the second target vertex before the target human body changes, LBS() is the LBS function of the preset linear hybrid skinning algorithm, G t W represents the key points of the human body corresponding to the second target vertex. t The weight value is the second target vertex. This embodiment first determines the second static position of the second target vertex before the target human body changes, and then determines the second target position of the second target vertex after the target human body changes based on the static position. This allows for highly efficient acquisition of the human body rendering image.
[0258] As another aspect of this application, this application provides a method for presenting a stereoscopic image, applied to a second conference device. Please refer to... Figure 14 The method for presenting a stereoscopic image includes the following steps:
[0259] S141: When the second conference device establishes a conference connection with the first conference device, it acquires the user feature data sent by the first conference device. The user feature data is obtained by the first conference device extracting the real-time user image of the first user.
[0260] S142: Obtain the 3D model data of the first user.
[0261] S143: Parse the first user's 3D model data into the first rendering model.
[0262] S144: Input user feature data into the first rendering model to obtain a real-time rendered image.
[0263] S145: Generate a stereoscopic image of the first user based on the real-time rendered image.
[0264] S146: Present a stereoscopic image of the first user.
[0265] In immersive video conferencing, this embodiment eliminates the need to transmit the entire video stream of the first user to the second conferencing device for 3D rendering. It also avoids consuming excessive bandwidth to transmit data that would enable the second conferencing device to render the 3D image. This embodiment requires only a small amount of user feature data to achieve the second conferencing device's rendering of the first user's stereoscopic image, thus expanding the application scope of the method provided. Furthermore, due to the small amount of user feature data, this embodiment can transmit it to the second conferencing device without delay, facilitating the second conferencing device's rapid and real-time rendering of the first user's stereoscopic image.
[0266] Generating a stereoscopic image of the first user based on a real-time rendered image includes the following steps: obtaining preset multi-viewpoint information, generating an image sequence based on the multi-viewpoint information and the real-time rendered image, and generating a stereoscopic image of the first user based on the image sequence.
[0267] Before acquiring the user feature data sent by the first conference device, the method further includes: generating 3D model data of a second rendered model for a second user, wherein the second user is the user operating the second conference device, and sending the 3D model data of the second rendered model to the first conference device, so that the first conference device saves the 3D model data of the second rendered model locally. The method for generating the second rendered model can be found in the methods described in the above embodiments, and will not be repeated here.
[0268] It should be noted that in the above embodiments, there is no necessarily a certain order between the steps. Those skilled in the art can understand from the description of the embodiments of this application that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.
[0269] As another aspect of the embodiments of this application, this application provides a stereoscopic image presentation device applied to a first conference device. The stereoscopic image presentation device can be a software module, which includes several instructions stored in a memory. A processor can access the memory, invoke the instructions for execution, and complete the stereoscopic image presentation method described in the above embodiments.
[0270] In some embodiments, the stereoscopic image presentation device can also be constructed from hardware devices. For example, the stereoscopic image presentation device can be constructed from one or more chips, which can work together to complete the stereoscopic image presentation method described in the various embodiments above. As another example, the stereoscopic image presentation device can also be constructed from various logic devices, such as general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), microcontrollers, ARM (Acorn RISC Machine) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0271] Please see Figure 15 The stereoscopic image presentation device 150 includes: an image acquisition module 151, a feature extraction module 152, and a data transmission module 153.
[0272] The image acquisition module 151 is used to acquire the real-time user image of the first user when the first conference device establishes a conference connection with the second conference device. The second conference device saves the 3D model data of the first rendering model corresponding to the first user. The feature extraction module 152 is used to extract user feature data from the real-time user image. The data transmission module 153 is used to send the user feature data to the second conference device based on the conference connection, so that the second conference device can input the user feature data into the first rendering model to obtain a real-time rendered image, and present a stereoscopic image of the first user based on the real-time rendered image.
[0273] In immersive video conferencing, this embodiment eliminates the need to transmit the entire video stream of the first user to the second conferencing device for 3D rendering. It also avoids consuming excessive bandwidth to transmit data that would enable the second conferencing device to render the 3D image. This embodiment requires only a small amount of user feature data to achieve the second conferencing device's rendering of the first user's stereoscopic image, thus expanding the application scope of the method provided. Furthermore, due to the small amount of user feature data, this embodiment can transmit it to the second conferencing device without delay, facilitating the second conferencing device's rapid and real-time rendering of the first user's stereoscopic image.
[0274] In some embodiments, please refer to [link / reference] before acquiring the real-time user image of the first user. Figure 15 The stereoscopic image presentation device 150 also includes a model generation module 154, which is used to acquire 3D model data of the first user, the 3D model data is used to represent the first rendering model, and send the 3D model data to the second conference device so that the second conference device can generate 3D model data of the first rendering model based on the 3D model data, and save the 3D model data of the first rendering model locally on the second conference device.
[0275] In some embodiments, the model generation module 154 is specifically used to: acquire 2D image data, the 2D image data including multiple target user images of the first user taken from different side angles, and generate 3D model data based on the multiple target user images.
[0276] In some embodiments, the model generation module 154 is specifically used to: extract a person region image from multiple target user images, generate a person 3D sub-model data from the person region image, the person 3D sub-model data is used to represent a 3D person rendering model of the first user, and determine 3D model data based on the person 3D sub-model data.
[0277] In some embodiments, determining the 3D model data based on the character 3D sub-model data includes: extracting environmental region images from multiple target user images; generating target environment 3D sub-model data based on the environmental region images, wherein the target environment 3D sub-model data is used to represent the environment rendering model of the first user; and determining the 3D model data based on the target environment 3D sub-model data and the character 3D sub-model data; or, obtaining preset environment 3D sub-model data, wherein the preset environment 3D sub-model data is used to represent the environment rendering model of the first user; and determining the 3D model data based on the preset environment 3D sub-model data and the character 3D sub-model data.
[0278] As another aspect of this application, this application provides a stereoscopic image presentation device applied to a second conference device. Please refer to... Figure 16 The stereoscopic image presentation device 160 includes a feature acquisition module 161, a data acquisition module 162, a model parsing module 163, an image rendering module 164, a stereoscopic generation module 165, and a stereoscopic presentation module 166.
[0279] The feature acquisition module 161 is used to acquire user feature data sent by the first conference device when the second conference device establishes a conference connection with the first conference device. The user feature data is obtained by the first conference device extracting the real-time user image of the first user. The data acquisition module 162 is used to acquire the 3D model data of the first user. The model parsing module 163 is used to parse the 3D model data of the first user into a first rendered model of the first user. The image rendering module 164 is used to input the user feature data into the first rendered model to obtain a real-time rendered image. The stereo generation module 165 is used to generate a stereo image of the first user based on the real-time rendered image. The stereo presentation module 166 is used to present the stereo image of the first user.
[0280] In some embodiments, the stereo generation module 165 is specifically used to: acquire preset multi-viewpoint information, generate an image sequence based on the multi-viewpoint information and real-time rendered images, and generate a stereo image of the first user based on the image sequence.
[0281] In some embodiments, please refer to [link / reference] before obtaining the user characteristic data sent by the first conference device. Figure 16 The stereoscopic image presentation device 160 also includes a model sending module 167, which is used to generate 3D model data of the second rendered model of the second user. The second user is the user who operates the second conference device. The 3D model data of the second rendered model is sent to the first conference device so that the first conference device can save the 3D model data of the second rendered model locally.
[0282] It should be noted that the stereoscopic image presentation device described above can execute the stereoscopic image presentation method provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in the embodiments of the stereoscopic image presentation device can be found in the stereoscopic image presentation method provided in the embodiments of this application.
[0283] Please see Figure 17 , Figure 17 This is a schematic diagram of a conference device provided in an embodiment of this application, wherein the conference device 170 may be a first conference device or a second conference device. The conference device 170 includes one or more processors 171 and a memory 172. The memory 172 is connected to one or more processors 171, for example, via a bus.
[0284] Processor 171 is configured to support the conferencing device in performing the corresponding functions in the methods described in the above-described method embodiments. Processor 171 may be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The aforementioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0285] Memory 172 is used to store program code, etc. Memory 172 may include volatile memory (VM), such as random access memory (RAM); memory may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); memory may also include combinations of the above types of memory.
[0286] The memory 172 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the stereoscopic image presentation method in the embodiments of this application. The processor executes the various functional applications and data processing of the stereoscopic image presentation method and the stereoscopic image presentation device by running the non-volatile software programs, instructions, and modules stored in the memory, that is, it realizes the functions of each module or unit of the stereoscopic image presentation method and the stereoscopic image presentation device provided in the above method embodiments.
[0287] The memory 172 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function. The data storage area may store data created based on the use of the stereoscopic image rendering device. In some embodiments, the memory may optionally include memory remotely located relative to the processor, which can be connected to the stereoscopic image rendering device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0288] One or more modules are stored in memory. When executed by one or more processors, they perform the stereoscopic image presentation method in any of the above method embodiments. For example, they perform the method steps described in the above method embodiments to realize the functions of the modules described in the above device embodiments.
[0289] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the method as described in the foregoing embodiments.
[0290] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0291] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A method for presenting a stereoscopic image, applied to a first conference device, characterized in that, include: When the first conference device establishes a conference connection with the second conference device, it acquires the real-time user image of the first user, and the second conference device saves the 3D model data of the first rendering model corresponding to the first user; Extract user feature data from the real-time user images; Based on the conference connection, the user feature data is sent to the second conference device, so that the second conference device inputs the user feature data into the first rendering model to obtain a real-time rendered image, and presents a stereoscopic image of the first user based on the real-time rendered image.
2. The presentation method according to claim 1, characterized in that, Before acquiring the real-time user image of the first user, the process also includes: Determine the 3D model data of the first user, wherein the 3D model data is used to represent the first rendered model; The 3D model data is sent to the second conference device so that the second conference device can parse the 3D model data into the first rendered model and save the 3D model data of the first rendered model locally on the second conference device.
3. The presentation method according to claim 2, characterized in that, The determination of the 3D model data of the first user includes: Acquire 2D image data, which includes multiple images of the target user taken from different side angles; 3D model data is generated based on multiple images of the target user.
4. The presentation method according to claim 3, characterized in that, The step of generating 3D model data based on multiple target user images includes: Extract the person region image from multiple target user images; Based on the image of the person's region, 3D sub-model data of the person's character is generated, and the 3D sub-model data of the person's character is used to represent the 3D character rendering model of the first user. The 3D model data is determined based on the character's 3D sub-model data.
5. The presentation method according to claim 4, characterized in that, The process of determining the 3D model data based on the character 3D sub-model data includes: Extract environmental region images from multiple target user images; Target environment 3D sub-model data is generated based on the environmental area image, and the target environment 3D sub-model data is used to represent the environment rendering model of the first user. The 3D model data is determined based on the target environment 3D sub-model data and the character 3D sub-model data. or, Obtain preset environment 3D sub-model data, which is used to represent the environment rendering model of the first user; The 3D model data is determined based on the preset environment 3D sub-model data and the character 3D sub-model data.
6. A method for presenting a stereoscopic image, applied to a second conference device, characterized in that, include: When the second conference device establishes a conference connection with the first conference device, it acquires user feature data sent by the first conference device. The user feature data is obtained by the first conference device extracting the real-time user image of the first user. Obtain the 3D model data of the first user; The first user's 3D model data is parsed into the first user's first rendering model; The user feature data is input into the first rendering model to obtain a real-time rendered image; Generate a stereoscopic image of the first user based on the real-time rendered image; Present a 3D image of the first user.
7. The presentation method according to claim 6, characterized in that, The step of generating the stereoscopic image of the first user based on the real-time rendered image includes: Acquire preset multi-viewpoint information; An image sequence is generated based on the multi-viewpoint information and the real-time rendered image; Generate a stereoscopic image of the first user based on the image sequence.
8. The presentation method according to claim 6, characterized in that, Before acquiring the user characteristic data sent by the first conference device, the process also includes: Generate 3D model data for the second user's second rendered model, where the second user is the user operating the second conference device; Send the 3D model data of the second rendered model to the first conference device so that the first conference device can save the 3D model data of the second rendered model locally.
9. A first conference device, characterized in that, The device includes a memory and a processor, the memory being connected to the processor, the processor being configured to execute one or more computer programs stored in the memory, and the processor, when executing the one or more computer programs, causing the conferencing device to implement the method for presenting stereoscopic images as described in any one of claims 1-5.
10. A second conference device, characterized in that, The device includes a memory and a processor, the memory being connected to the processor, the processor being configured to execute one or more computer programs stored in the memory, and the processor, in executing the one or more computer programs, causing the conferencing device to implement the method for presenting stereoscopic images as described in any one of claims 6 to 8.
11. A conference system, characterized in that, include: The first conference device as described in claim 9; The second conference device as described in claim 10 is capable of establishing a conference connection with the first conference device.
12. The conference system according to claim 11, characterized in that, Also includes: The first camera module includes N first cameras, which are respectively connected to the first conference device for acquiring 2D image data, where N is a positive integer. The second camera module includes M second cameras, which are each connected to the second conference equipment for acquiring 2D image data.
13. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program, the computer program including program instructions, which, when executed by a processor, cause the processor to perform the method for presenting a stereoscopic image as described in any one of claims 1-5 or the method for presenting a stereoscopic image as described in any one of claims 6 to 8.