Implementation method for immersive conference, and related electronic device

By utilizing 3D camera technology and light field display technology, an immersive meeting experience with multiple participants and multiple perspectives has been achieved. This solves the problems of high bandwidth, complex equipment, and insufficient immersion in existing technologies, thereby improving the immersiveness and rendering efficiency of remote meetings.

WO2025256457A1PCT designated stage Publication Date: 2025-12-18GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1

Patent Information

Application Number
PCT/CN2025/099343
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-11
Filing Date
2025-06-05
Publication Date
2025-12-18

AI Technical Summary

Technical Problem

Existing immersive conferencing solutions have high bandwidth requirements, support a limited number of video streams, can only achieve multiple perspectives for a single person, other participants cannot obtain clear images, and have complex equipment and wiring, resulting in insufficient immersion.

Method used

Using 3D camera technology for three-dimensional reconstruction to generate virtual avatars, feature point driving and rendering technology are used to convey body movements and facial expressions, combined with light field display technology to achieve a multi-person, multi-view immersive experience.

Benefits of technology

It reduces device complexity and bandwidth requirements, supports immersive experiences with multiple users and multiple perspectives, and improves rendering efficiency and immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025099343_18122025_PF_FP_ABST
    Figure CN2025099343_18122025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers. Disclosed are an implementation method for an immersive conference, and a related electronic device. In the method, a 3D stereoscopic image of a second conference participant is constructed by using a second rendering model, and instead of an original video stream, only a relatively small amount of feature data needs to be transmitted, thereby effectively reducing the network bandwidth required by transmission; each conference participant can obtain the 3D stereoscopic image of the second conference participant at his / her own angle of view, thereby supporting a multi-person and multi-angle of view mode; and by means of the pre-constructed second rendering model, the 3D stereoscopic image of the second conference participant can be reconstructed in real time on the basis of the feature data, without the need for complex real-time rendering calculation, such that the overall rendering efficiency is improved, and even a terminal device with a relatively weak performance can also smoothly present a vivid 3D visual effect. The present application reduces the network bandwidth required by transmission, enables each conference participant to obtain an immersive experience at his / her own angle of view, and also improves the overall rendering efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Implementation method of immersive conference and related electronic device

[0001] Cross Reference to Related Applications

[0002] The present application claims priority from the Chinese patent application No. 202410748260.0 and titled "Implementation method of immersive conference and related electronic device", filed on June 11, 2024, with the China Patent Office, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0003] The present application relates to the technical field of computer, and in particular, to an implementation method of immersive conference and related electronic device. BACKGROUND

[0004] Immersive conference is based on remote video conference and adds the sense of being on the scene and the sense of being brought in, which aims to make the participants obtain a more realistic and immersive conference experience, as if the participants are in the same conference room.

[0005] The current related immersive conference scheme is based on multiple synchronous cameras to collect conference scene videos, and transmits the collected videos to a remote device through high bandwidth, and then generates two-viewpoint images, the two-viewpoint angles depend on the additional camera to capture the binocular vision results.

[0006] The related art immersive conference scheme has high bandwidth requirement, limited supported video lines, and therefore limited supported switching angles; in addition, the number of view points of light field display is limited, only the participants whose binocular vision is captured by the camera can normally watch the content of the light field display screen, and other participants cannot obtain clear images. SUMMARY

[0007] An object of the embodiments of the present application is to provide an implementation method of immersive conference and related electronic device, to solve the technical problem of how to reduce bandwidth requirement and realize multi-person and multi-angle immersive conference experience.

[0008] To solve the above technical problems, one technical scheme adopted by the embodiments of the present application is to provide an implementation method of an immersive conference, applied to a first electronic device, comprising: receiving video stream data sent by a second electronic device participating in a conference together with the first electronic device, the video stream data comprising audio data and real-time user feature data, the real-time user feature data being obtained by the second electronic device performing feature extraction on a real-time collected image stream of a second conference participant, the real-time user feature data comprising expression information and action information of the second conference participant; obtaining a second rendering model of the second conference participant, wherein the second rendering model is constructed based on head posture information, expression information and action information in a plurality of frames of human body images of the second conference participant, and the head posture and facial expression of any two frames of the human body images are different at least in one; processing the real-time user feature data by using the second rendering model to obtain a stereoscopic image of the second conference participant; and presenting the stereoscopic image and the audio data. Compared with directly transmitting high-definition video, the implementation method uses a second rendering model to construct a 3D stereoscopic image of a second conference participant, only needs to transmit relatively small feature data (such as expression and action data) instead of original video stream, which effectively reduces the required network bandwidth for transmission, so that a good remote conference experience can be achieved even in a bandwidth-limited scenario. Secondly, each conference participant can obtain a 3D stereoscopic image of the second conference participant from his own perspective, thereby supporting multi-person and multi-perspective. Finally, the pre-constructed second rendering model can reconstruct the 3D stereoscopic image of the second conference participant in real time according to the feature data, without complex real-time rendering calculation, which improves the overall rendering efficiency, so that even a terminal device with weak performance can also smoothly present a lively 3D visual effect. Therefore, the implementation method uses a second rendering model to construct a stereoscopic image of a conference participant, not only greatly reduces the required network bandwidth for transmission, but also supports each conference participant to obtain an immersive experience from his own perspective, and also improves the overall rendering efficiency.

[0009] To solve the above technical problems, one technical scheme adopted by the embodiments of the present application is to provide an implementation method of an immersive conference, applied to a second electronic device, comprising: when the second electronic device establishes a conference connection with a first electronic device, acquiring an image stream of a second conference participant; performing feature extraction on the image stream of the second conference participant to obtain real-time user feature data, wherein the real-time user feature data comprises expression information and action information; parsing audio data from the image stream of the second conference participant; packaging the audio data and the real-time user feature data into video stream data and sending them to the first electronic device, so that the first electronic device inputs the real-time user feature data into a second rendering model corresponding to the second conference participant, generates a stereoscopic image, and makes the first electronic device present the stereoscopic image of the second conference participant and the audio data. The implementation method only needs to transmit small feature data (expression, action information) and audio data, instead of original high-definition video stream, which greatly reduces the required bandwidth of network transmission, so that good remote conference experience can be achieved even in a bandwidth-limited scenario; the second rendering model constructed in advance can efficiently reconstruct the 3D stereoscopic image of the second conference participant according to the feature data in real time, without complex real-time rendering calculation, which improves the overall rendering efficiency; the first electronic device can render the stereoscopic image of the second conference participant locally according to the received real-time feature data, each conference participant can render the stereoscopic image of the second conference participant according to his own perspective, and a multi-perspective immersive experience is achieved. Therefore, the implementation method of the immersive conference applied to the second electronic device not only effectively reduces the bandwidth requirement, but also supports multi-perspective interaction, and improves the rendering efficiency.

[0010] To solve the above technical problems, one technical scheme adopted by the embodiments of the present application is to provide a first electronic device, comprising: a memory and a processor, the memory being connected to the processor, the processor being used to execute one or more computer programs stored in the memory, and the processor, when executing the one or more computer programs, causing the first electronic device to implement an implementation method of an immersive conference applied to a first electronic device. The first electronic device has the beneficial effects corresponding to the implementation method of the immersive conference applied to the first electronic device.

[0011] To solve the above technical problems, one of the technical solutions adopted by the embodiments of the present application is to provide a second electronic device, comprising a memory and a processor, the memory is connected to the processor, the processor is used to execute one or more computer programs stored in the memory, and the processor, when executing the one or more computer programs, causes the second electronic device to implement an implementation method of an immersive conference applied to a second electronic device. The second electronic device has the beneficial effects corresponding to the implementation method of the immersive conference applied to the second electronic device.

[0012] To solve the above technical problems, one of the technical solutions adopted by the embodiments of the present application is to provide a computer readable storage medium, the computer readable storage medium stores computer executable instructions, when the computer executable instructions are executed by the first electronic device or the second electronic device, the first electronic device or the second electronic device executes the implementation method of the immersive conference as described above.

[0013] To solve the above technical problems, one of the technical solutions adopted by the embodiments of the present application is to provide a computer program product, the computer program product comprises a computer program stored on a computer readable storage medium, the computer program comprises program instructions, when the program instructions are executed by the first electronic device or the second electronic device, the first electronic device or the second electronic device executes the implementation method of the immersive conference as described above. BRIEF DESCRIPTION OF DRAWINGS

[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the description of the embodiments of the present application will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0015] Fig. 1 is a schematic diagram of an application scenario provided by the embodiments of the present application;

[0016] Fig. 2 is a structural block diagram of an implementation system of an immersive conference corresponding to the application scenario provided by the embodiments of the present application;

[0017] Fig. 3 is a schematic diagram of the position of a control point in a standard SMPL human body model provided by the embodiments of the present application;

[0018] Fig. 4 is a schematic diagram of the position of a control point in a standard FLAME human head model space provided by the embodiments of the present application;

[0019] Fig. 5 is a flow chart of an implementation method of an immersive conference provided by the embodiments of the present application;

[0020] FIG. 6 is a flowchart of a method for processing real-time user feature data using a second rendering model to obtain a stereoscopic image of a second conference participant according to an embodiment of the present application;

[0021] FIG. 7 is a flowchart of a method for generating a stereoscopic image of a second conference participant according to an image sequence according to an embodiment of the present application;

[0022] FIG. 8 is a flowchart of a method for training a second 3D head model according to an embodiment of the present application;

[0023] FIG. 9 is a flowchart of a method for training a second human motion capture model according to an embodiment of the present application;

[0024] FIG. 10 is a flowchart of a method for generating a first rendering model of a first conference participant participating in a conference through a first electronic device according to an embodiment of the present application;

[0025] FIG. 11 is a flowchart of a method for implementing an immersive conference according to another embodiment of the present application;

[0026] FIG. 12 is a flowchart of a method for generating a second rendering model of a second conference participant participating in a conference through a second electronic device according to an embodiment of the present application;

[0027] FIG. 13 is a structural schematic diagram of an implementation device of an immersive conference according to an embodiment of the present application;

[0028] FIG. 14 is a structural schematic diagram of an implementation device of an immersive conference according to another embodiment of the present application;

[0029] FIG. 15 is a structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the objectives, technical solutions and advantages of the present application clearer, further detailed descriptions will be given to the present application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative work fall within the scope of protection of the present application.

[0031] It should be noted that the various features of the embodiments of the present application can be combined with each other, and are within the protection scope of the present application, if there is no conflict. In addition, although the functional modules are divided in the device schematic diagram, and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Furthermore, the "first", "second", "third" and the like used in the present application do not limit the data and execution order, but only distinguish the same items or similar items with basically the same function and effect.

[0032] It should be noted that in the following various embodiments, there is no certain sequence between the following steps, and those skilled in the art can understand from the description of the embodiments of the present application that the following steps in different embodiments can have different execution sequences, that is, can be executed in parallel, can be exchanged, and the like.

[0033] Remote video conferencing is difficult to provide an efficient communication experience comparable to face-to-face conversation, because participants cannot truly perceive important visual information such as body movements, facial expressions, eye contact, and lack natural interaction methods. The mainstream liquid crystal display or LED display used in the current market can only display the same 2D effect image in different directions in space, which means that no matter which position in the conference room the participant stands, the screen image seen by the participant is the same. These limitations seriously weaken the immersion of remote video conferencing.

[0034] In remote video conferencing, body movements, facial expressions, eye contact, etc. are one of the most important non-verbal communication methods, which can convey participants' attention, interest and intention, and help to establish effective communication and interaction. However, due to the limitations of cameras and video transmission, participants have difficulty accurately capturing and conveying body movements and facial expressions, which leads to incomplete information transmission and unnatural communication. In addition, the mainstream 2D display cannot provide real spatial and depth perception, limiting the ability of participants to perceive object position and size, which further reduces the immersion and naturalness of interaction, and participants cannot truly perceive the presence and position of others in the conference room. In addition, participants in remote video conferencing are usually in their own work environment, which is easily disturbed by environmental interference and distraction factors. Compared with face-to-face meetings, participants are more easily disturbed by the surrounding environment, affecting concentration and effective communication. In summary, the lack of immersion in remote video conferencing is that participants cannot truly perceive visual information such as body movements and facial expressions, and are limited by the spatial expression ability of 2D display. These limitations affect the efficiency and naturalness of communication, and further technology and innovation are needed to improve the experience of remote video conferencing.

[0035] Currently, the related immersive conference solution is based on multiple synchronous cameras to collect conference site videos, and the collected videos are transmitted to a remote device through high bandwidth, and then two-viewpoint images are generated, the two-viewpoint angles depend on the additional camera to capture the binocular visual line results. In the specific implementation of the related technology, multiple cameras are needed to collect conference site videos, and the data is transmitted to the remote device end through high bandwidth. In addition, it supports single-person viewing, which limits the number of viewpoints of the light field display, only the participants whose binocular visual lines are captured by the camera can normally watch the content of the light field display screen, and other participants cannot obtain clear images. Therefore, the related technology has the following problems when implementing an immersive conference: multiple synchronous cameras are needed to collect conference site videos, which increases the complexity of devices and wiring, and additional requirements for conference room settings and costs; in order to transmit the video data collected by multiple cameras, the immersive conference solution requires high-bandwidth network connection; the light field display screen in the immersive conference solution can provide multi-viewpoint images, but the number of viewpoints is limited by the number of participants whose binocular visual lines are captured, which means that only a few participants can obtain clear and real images, while other participants cannot have the same experience.

[0036] Based on this, the embodiment of the present application provides an implementation method of an immersive conference, mainly by using 3D camera technology, the conference participants can be reconstructed in three dimensions, and their accurate virtual images can be generated, which can reduce the need for multiple cameras, and reduce the complexity of devices and wiring; based on the results of 3D portrait reconstruction, feature point driving and rendering technology can be used to realize the transmission of participants' body movements and facial expressions, by capturing key feature points of participants such as joint positions and facial expressions, virtual body movements and facial expressions can be generated in real time, and transmitted to a remote device for rendering; by using the feature point driving and rendering method, the transmission demand for bandwidth can be greatly reduced, compared with transmitting video data collected by multiple cameras, only key feature points and rendering parameters need to be transmitted, which effectively reduces the data volume and bandwidth requirement; finally, combined with light field display technology, the effect of multiple people watching at multiple angles at the same time can be realized, the light field display screen can provide more viewpoints, and real-time corresponding images can be generated according to the position and visual line direction of the observer, so that each participant can obtain clear and real images, and the sense of immersion and experience are improved.

[0037] The implementation method of the immersive conference proposed by the embodiment of the present application will be introduced below through specific embodiments.

[0038] Firstly, some terms related to the embodiment of the present application are introduced:

[0039] The SMPL (Skinned Multi-Person Linear) model is a commonly used human shape and pose model, which is a linear algebra-based parameterized model used to represent human shapes and poses. The SMPL model describes human poses and shapes through a set of parameters, including global rotation, global scale, joint angles, and deformation weights. The core idea of SMPL is to represent the human body as a skeletal structure connected by joints, and use linear weights to map the surface mesh to the skeletal structure. By adjusting the joint angles and deformation weights, different poses and shapes of the human body model can be generated. The parameterized representation of the SMPL model makes it well suited to adapt to different human shape and pose changes. The standard SMPL human model includes 24 control points, as shown in Figure 3.

[0040] The FLAME (Fast Lightweight Adaptive Meshes) model is a computer graphics model used for face modeling and animation. It is a highly customizable and flexible model that can reconstruct and simulate realistic three-dimensional faces from a single photo or video sequence. The design goal of the FLAME model is to maintain high accuracy and computational efficiency, and it is based on a statistical shape model and a linear expression model, combining information about face geometry, expression, skin color, and lighting. The FLAME model can generate a complete three-dimensional face mesh, including face geometry, texture, and normal information. The standard FLAME model includes k control points, as shown in Figure 4.

[0041] Light field display is a special display technology designed to simulate the propagation and optical effects of light fields in the real world, to present realistic three-dimensional images and visual experiences.

[0042] Multi-view subpixel mapping is a technology used for light field display and multi-view display, designed to improve the accuracy of mapping and rendering multiple viewpoints in display systems, involving subpixel-level processing and mapping of images for each viewpoint.

[0043] Referring to FIG. 1, FIG. 1 is a schematic diagram of an application scenario provided by an embodiment of the present application. The application scenario includes three conference participants, namely a first conference participant 10, a second conference participant 20, and a third conference participant 30, participating in a remote conference. When the three conference participants jointly participate in the conference, the three conference participants can achieve immersive conference effects of the conference.

[0044] Referring to FIG. 2, the first participant 10 includes a first electronic device 11, a camera 12, a microphone 13, and a first participant 14, the second participant 20 includes a second electronic device 21, a camera 22, a microphone 23, and a second participant 24, and the third participant 30 includes a third electronic device 31, a camera 32, a microphone 33, and a third participant 34.

[0045] The first electronic device 11, the second electronic device 12, and the third electronic device 13 are communicatively connected to each other. When the three participants jointly participate in the conference, the first electronic device 11 receives video stream data respectively sent by the second electronic device 21 and the third electronic device 31, which includes real-time user feature data of the second participant 24 collected by the camera 22 and audio data of the second participant 24 collected by the microphone 23, and also includes real-time user feature data of the third participant 34 collected by the camera 32 and audio data of the third participant 34 collected by the microphone 33. The real-time user feature data of the second participant 24 is obtained by the second electronic device 21 performing feature extraction on an image stream of the second participant 24 collected in real time, and includes expression information and action information of the second participant 24. The real-time user feature data of the third participant 34 is obtained by the third electronic device 31 performing feature extraction on an image stream of the third participant 34 collected in real time, and includes expression information and action information of the third participant 34.

[0046] The single-frame image can be extracted from the video stream data, thereby obtaining an image stream. The expression information and the action information of the participant are extracted based on the image stream. The process of extracting the expression information of the participant based on the image stream can include: performing face detection based on the image stream, extracting facial feature points, performing expression recognition according to the facial feature points, and outputting expression information corresponding to the expression recognition result. The process of extracting the action information of the participant based on the image stream can include: performing human body detection based on the image stream, determining human body key points, performing action analysis according to the key points, and outputting action information corresponding to the action analysis result. In specific implementation, a pre-trained model or tool (such as an open-source tool MediaPipe) can be used to extract the control point positions in the standard SMPL human body model (as shown in FIG. 3) and the control point positions in the standard FLAME human head model space (as shown in FIG. 4), and then based on the regression method, the SMPL-based action information and the FLAME-based expression information are obtained for the image frames.

[0047] After the first electronic device 11 receives the video stream data respectively sent by the second electronic device 21 and the third electronic device 31, the first electronic device 11 obtains the second rendering model of the second participant 24 and the third rendering model of the third participant 34 pre-stored in the first electronic device 11. Then, the second rendering model is used to process the real-time user feature data corresponding to the second participant 24 to obtain the stereoscopic image of the second participant 24, and the third rendering model is used to process the real-time user feature data corresponding to the third participant 34 to obtain the stereoscopic image of the third participant 34. Finally, the stereoscopic image and audio data of the second participant 24 and the stereoscopic image and audio data of the third participant 34 are presented on the display interface corresponding to the conference of the first electronic device 11. As shown in FIG. 1, the first participant 14 in the first conference party 10 can see the stereoscopic images of the second participant 24 and the third participant 34 and the audio data corresponding to each of them, as if the first participant 14, the second participant 24 and the third participant 34 are all in the conference scene of the first conference party 10. After the first electronic device 11 obtains the audio data, audio-video synchronization processing is performed. By analyzing the timestamp information of the video frame and the audio data, it is ensured that the motion and expression of the person in the video are completely synchronized with the sound. When rendering the stereoscopic images of the second participant 24 and the third participant 34, the corresponding audio data is played synchronously. In this way, when the first participant 14 watches the stereoscopic images, he can hear the sound that matches them, as if he is having a face-to-face conversation with a real person.

[0048] Through the above audio-video synchronization processing, when the first participant 14 watches the stereoscopic images of other participants, he can feel as if he is on the scene, thus improving the sense of presence and communication effect of the remote conference.

[0049] Similarly, the second electronic device 21 receives the video stream data respectively sent by the first electronic device 11 and the third electronic device 31. The video stream data includes the real-time user feature data of the first participant 14 collected by the camera 12 and the audio data of the first participant 14 collected by the microphone 13, and also includes the real-time user feature data of the third participant 34 collected by the camera 32 and the audio data of the third participant 34 collected by the microphone 33. The real-time user feature data of the first participant 14 is obtained by the first electronic device 11 extracting features from the image stream of the first participant 14 collected in real time, which includes the expression information and motion information of the first participant 14. The real-time user feature data of the third participant 34 is obtained by the third electronic device 31 extracting features from the image stream of the third participant 34 collected in real time, which includes the expression information and motion information of the third participant 34.

[0050] After the second electronic device 21 receives the video stream data sent by the first electronic device 11 and the third electronic device 31 respectively, the second electronic device 21 obtains the first rendering model of the first participant 14 and the third rendering model of the third participant 34 pre-stored in the second electronic device 21; then, the first rendering model is used to process the real-time user feature data corresponding to the first participant 14 to obtain the stereoscopic image of the first participant 14, and the third rendering model is used to process the real-time user feature data corresponding to the third participant 34 to obtain the stereoscopic image of the third participant 34; finally, the stereoscopic image and the audio data of the first participant 14 and the stereoscopic image and the audio data of the third participant 34 are presented on the display interface corresponding to the conference of the second electronic device 21, and the effect is also shown in FIG. 1.

[0051] Similarly, the third electronic device 31 receives the video stream data sent by the first electronic device 11 and the second electronic device 21 respectively, and the video stream data includes the real-time user feature data of the first participant 14 collected by the camera 12 and the audio data of the first participant 14 collected by the microphone 13, and also includes the real-time user feature data of the second participant 24 collected by the camera 22 and the audio data of the second participant 24 collected by the microphone 23. The real-time user feature data of the first participant 14 is obtained by the first electronic device 11 extracting features from the image stream of the first participant 14 collected in real time, and includes the expression information and the action information of the first participant 14. The real-time user feature data of the second participant 24 is obtained by the second electronic device 21 extracting features from the image stream of the second participant 24 collected in real time, and includes the expression information and the action information of the second participant 24.

[0052] After the third electronic device 31 receives the video stream data sent by the first electronic device 11 and the second electronic device 21 respectively, the third electronic device 31 obtains the first rendering model of the first participant 14 and the second rendering model of the second participant 24 pre-stored in the third electronic device 31; then, the first rendering model is used to process the real-time user feature data corresponding to the first participant 14 to obtain the stereoscopic image of the first participant 14, and the second rendering model is used to process the real-time user feature data corresponding to the second participant 24 to obtain the stereoscopic image of the second participant 24; finally, the stereoscopic image and the audio data of the first participant 14 and the stereoscopic image and the audio data of the second participant 24 are presented on the display interface corresponding to the conference of the third electronic device 31, and the effect is also shown in FIG. 1.

[0053] The three co-participating conference participants collect corresponding video stream data in real time, which is transmitted to the electronic devices of the other two participants respectively. Each electronic device stores the rendering model of the other participants in advance. After receiving the video stream data of the other participants, the real-time feature data is processed by the pre-prepared rendering model to generate a stereoscopic image of the participants. At the same time, the system aligns the time stamp of the received video and audio data to ensure that the sound and picture are synchronized. Each participant's electronic device can display the stereoscopic images and synchronized audio of the other two participants, which provides a visual and auditory experience as if they were in the same physical conference room, giving a sense of being there. Therefore, in the application scenarios shown in FIG. 1 and FIG. 2, the three conference participants can experience immersive remote conferencing on their respective devices. This solution that combines real-time video and audio, rendering models, and audio and video synchronization effectively enhances the sense of presence and communication effect in a distributed collaboration scenario.

[0054] The camera 12, the camera 22, and the camera 32 can each be a single-lens camera, which has a lower hardware cost than a multi-lens camera, can reduce the deployment cost of the entire system, and can reduce the design complexity. Of course, the camera 12, the camera 22, and the camera 32 can also be multi-lens cameras, each of which includes a camera module composed of multiple cameras.

[0055] The microphone 13, the microphone 23, and the microphone 33 can each be selected according to the specific conference room environment and audio quality requirements to ensure that the voice of each participant can be captured and transmitted with high quality.

[0056] The first participant 14, the second participant 24, and the third participant 34 can each be a single participant or multiple participants. When a participant includes multiple participants, for each participant including multiple participants, the system can use video synthesis technology to combine the images of multiple participants into one image, which can display the images of all participants in a limited display space. In addition, for each participant including multiple participants, the system can mix the voice signals of multiple participants to form a synthesized audio, which can ensure that the other two participants can clearly hear the speech of all participants. Furthermore, the system can automatically switch to the best camera angle and image according to the position of the current speaker, which can enable the other participants to better focus on the current speaker. Moreover, the system can provide participant management functions to allow the host or administrator to control the operations of the participants in the multi-participant conference room, such as going on the microphone and muting, which helps to ensure the orderly progress of the conference.

[0057] It should be noted that Fig. 1 shows a conference with three parties participating, and in other application scenarios, there can be fewer two parties participating in the conference, or more four parties, five parties, etc. participating in the conference, and the implementation principle is similar to that of Fig. 1.

[0058] Please refer to Fig. 5, which is a flowchart of an implementation method of an immersive conference provided by an embodiment of the present application. The method is applied to a first electronic device and includes the following steps:

[0059] S11, receiving video stream data sent by a second electronic device participating in a conference with the first electronic device, the video stream data including audio data and real-time user feature data, the real-time user feature data being obtained by the second electronic device performing feature extraction on a real-time image stream of a second conference participant, the real-time user feature data including expression information and action information of the second conference participant.

[0060] Wherein, the real-time feature extraction of the second electronic device on the image stream of the second conference participant can include the following steps: face detection and tracking, using computer vision algorithms to detect and track the face position of the second conference participant in real time; facial key point positioning, positioning a series of key facial feature points such as eyebrows, eyes, nose, and lips in the detected face region; performing expression recognition, i.e., based on the position and deformation of the facial key points, using a pre-trained expression recognition model to detect the facial expression of the second conference participant in real time, such as happy, angry, surprised, etc., and these expression information constitutes part of the real-time user feature data. In addition, the head posture change, hand movement, and body movement of the second conference participant are captured in real time by analyzing the human body posture of the continuous frame images, and these action information also constitutes part of the real-time user feature data.

[0061] Wherein, a pre-trained model or tool (such as the open source tool MediaPipe) can be used to extract the control point positions in the standard SMPL human body model (as shown in Fig. 3) and the control point positions in the standard FLAME human head model space (as shown in Fig. 4), and based on the regression method, the SMPL-based action information and the FLAME-based expression information are obtained for the image frames.

[0062] S12, obtaining a second rendering model of the second conference participant, wherein the second rendering model is constructed based on the head posture information, expression information, and action information in the multiple frames of human body images of the second conference participant, and at least one of the head posture and facial expression of any two frames of human body images is different.

[0063] The second rendering model is a 3D rendering model constructed according to the plurality of human body images of the second participant, and the 3D rendering model can dynamically present the head posture, facial expression and action change of the second participant.

[0064] The process of obtaining the second rendering model can include: collecting a plurality of human body images, and constructing the second rendering model in advance before the implementation method of the immersive conference is started, the second electronic device collects an image stream of the second participant, and selects a plurality of human body image frames at different time points from the image stream; for the plurality of human body image frames, computer vision technology can be applied to analyze the head posture information (such as head rotation, elevation angle, etc.), facial expression information and body action change information of the second participant in each frame, which will be extracted and recorded; wherein for the selected plurality of human body images, it is required that there is at least a difference in head posture or facial expression between any two frames, that is, frames that are completely the same are not selected, so as to ensure that there is a certain dynamic change; the plurality of human body images can include human body images at multiple different angles. Next, based on the plurality of human body images containing head posture, expression and action information, a 3D model reconstruction and rendering technology is used to construct the second rendering model of the second participant, and the second rendering model can dynamically present the head movement, facial expression change and body movement of the second participant.

[0065] S13, processing the real-time user feature data by using the second rendering model to obtain a stereoscopic image of the second participant.

[0066] The second rendering model is used to process the expression information and action information of the second participant to obtain a stereoscopic image of the second participant. Based on the extracted head posture, expression and action features, the second rendering model reconstructs a 3D stereoscopic image of the second participant in real time, and the 3D stereoscopic image dynamically presents the real-time state change of the second participant.

[0067] S14, presenting the stereoscopic image and the audio data.

[0068] The reconstructed 3D stereoscopic image and the associated audio data can be presented to the first participant of the first electronic device through a stereoscopic display device (such as a naked-eye 3D display screen, a VR head-mounted display, etc.), so that the first participant can feel the vivid and stereoscopic image of the second participant and observe the changes of the expression and action of the second participant.

[0069] Compared with directly transmitting high-definition video, the implementation method of the immersive conference of the embodiment of the application uses a second rendering model to construct a 3D stereoscopic image of the second conference participant, only needs to transmit relatively small feature data (such as expression and action data) instead of an original video stream, which effectively reduces the network bandwidth required for transmission, so that a good remote conference experience can be achieved even in a bandwidth-limited scenario. Secondly, each conference participant can obtain a 3D stereoscopic image of the second conference participant belonging to his own perspective, thereby supporting multi-person and multi-perspective. Finally, the second rendering model constructed in advance can reconstruct the 3D stereoscopic image of the second conference participant in real time according to the feature data, without complex real-time rendering calculation, which improves the overall rendering efficiency, so that even a terminal device with weak performance can smoothly present vivid 3D visual effects. Therefore, the method of the embodiment of the application constructs a stereoscopic image of a participant by using a second rendering model, not only greatly reduces the network bandwidth required for transmission, but also supports each participant to obtain an immersive experience belonging to his own perspective, and meanwhile improves the overall rendering efficiency.

[0070] In some embodiments, referring to FIG. 6, the above step S13 specifically includes:

[0071] S131, acquiring preset multi-view information.

[0072] The multi-view information includes pose parameters of a plurality of virtual cameras, and the pose parameters of the virtual cameras represent different perspectives and are used to generate a sequence of rendered images with parallax.

[0073] The preset multi-view information includes: determining a preset number of virtual camera poses according to a set field of view angle and an eye parallax angle, and the preset number of virtual camera poses constitute the multi-view information. For example, the field of view angle is set to be a horizontal perspective of plus or minus 30 degrees, the total field of view angle is 60 degrees, and the binocular parallax angle is set to be 0.5 degrees. According to the formula (30*2) / 0.5=120, 120 virtual camera poses are needed, which are set every 0.5 degrees from -30 degrees to +30 degrees. The pose matrix of the 120 virtual cameras includes the position and orientation information of each virtual camera, and the 120 virtual camera pose matrices constitute the multi-view information.

[0074] S132, inputting the multi-view information and real-time user feature data into a second rendering model for processing to obtain an image sequence.

[0075] The image sequence includes a plurality of rendered images, each rendered image being rendered from a preset virtual camera pose. There is a parallax between any two rendered images. The parallax between the rendered images is caused by the difference in camera position and orientation, and simulates the parallax of the human eyes. The parallax between adjacent two of the plurality of rendered images is equal, because the interval between adjacent virtual camera positions is constant (for example, 0.5 degrees), so the parallax between adjacent rendered images is also constant.

[0076] The above design can achieve the naked-eye 3D effect of multiple people viewing from multiple perspectives. Because the positions and orientations of each participant are different, the image sequences they see are also different. However, because the parallax between adjacent images is constant, each participant can feel stereoscopic, that is, even if multiple people are viewing at the same time, the naked-eye 3D effect can be achieved.

[0077] The image sequence rendered in this way can provide a more immersive and immersive experience for participants in the meeting. Each person can view the stereoscopic conference picture from their own perspective, achieving a truly multi-person participation naked-eye 3D effect.

[0078] The multi-viewpoint information includes a preset number of virtual camera poses (such as 120 virtual camera poses). The processing of the multi-viewpoint information and the real-time user feature data in the second rendering model to obtain the image sequence includes: obtaining a reference rendered image obtained by processing the real-time user feature data by the second rendering model; adjusting the pose of the second participant in the reference rendered image according to the pose of the virtual camera in each virtual camera pose, to generate a rendered image corresponding to the second participant for each of the preset number of virtual camera poses, wherein the second participant in each rendered image corresponds to a view angle associated with the pose of a virtual camera. The second rendering model generates a reference rendered image according to real-time user feature data (such as expressions, actions, etc.). For each preset virtual camera pose, the position and angle of the second participant in the reference rendered image can be adjusted according to the position and orientation of the virtual camera, so that a rendered image corresponding to each virtual camera view angle can be generated. The second participant in each rendered image can correspond to a view angle associated with the pose of a virtual camera. Through the above steps, 120 rendered images corresponding to 120 virtual camera poses can be obtained, which constitute the final image sequence.

[0079] The image sequence is generated in this way, which can make full use of the preset multi-viewpoint information to render the second participant picture with parallax at different view angles, so that the first participant corresponding to the first electronic device can observe the second participant from different view angles, producing a real naked-eye 3D effect.

[0080] The second rendering model comprises a second 3D head model and a second human body motion capture model, and the obtaining of the reference rendering image processed by the second rendering model on the real-time user feature data comprises:

[0081] inputting the expression information in the real-time user feature data into the second 3D head model to obtain a head rendering picture of the second participant;

[0082] inputting the motion information in the real-time user feature data into the second human body motion capture model to obtain a human body motion rendering picture of the second participant;

[0083] combining and rendering the head rendering picture and the human body motion rendering picture of the second participant to obtain the reference rendering image of the second participant.

[0084] The above method processes the expression information in the real-time user feature data by using the second 3D head model, processes the motion information in the real-time user feature data by using the second human body motion capture model, and then combines and renders the head rendering picture and the human body motion rendering picture, that is, superimposes the head and the human body motion rendering results of the second participant to obtain the final reference rendering image of the second participant.

[0085] The method models and renders the head expression and the human body motion respectively, can better capture the real-time features of the second participant, and finally generates a complete reference rendering image of the second participant through combined rendering. Thus, the technical advantages of 3D head reconstruction and human body motion capture can be fully utilized, and the real-time state and motion performance of the second participant in the conference can be more accurately reflected.

[0086] S133, generating a stereoscopic image of the second participant according to the image sequence.

[0087] The stereoscopic image refers to an image that can present a real three-dimensional space effect, has a parallax effect, can present a three-dimensional sense, and utilizes the characteristics of binocular parallax of human eyes. By inputting the preset multi-viewpoint information and the real-time user feature data into the second rendering model, the finally generated image sequence has such a stereoscopic effect. When the viewer watches these images, he or she can feel the positional relationship and dynamic changes of the second participant in the three-dimensional space and produce an immersive feeling.

[0088] Referring to FIG. 7, the generation of the stereoscopic image of the second participant according to the image sequence comprises:

[0089] S1331, obtaining a preset multi-viewpoint sub-pixel mapping table and a pixel layout of a display interface of the first electronic device.

[0090] wherein, first, a preset multi-view sub-pixel mapping table is obtained, which defines the corresponding position of each pixel position in the image sequence. It provides a mapping relationship between each pixel position in the stereoscopic image to be synthesized and the corresponding position in the image sequence. Through the multi-view sub-pixel mapping table, it can be determined which pixel values in the image sequence are extracted to synthesize the final stereoscopic image.

[0091] At the same time, the pixel layout of the display interface of the first electronic device is also obtained, that is, the pixel size and layout of the display interface are determined. The pixel layout of the display interface includes the pixel size and layout method. The pixel size refers to the number and resolution of the pixels of the display interface, such as the number of pixels in width and height. The layout method describes the arrangement method of the pixels on the display interface, such as the number of pixels in each row and each column, and the arrangement rule between them. By determining the pixel size and layout, it can be ensured that the generated stereoscopic image can be correctly displayed on the first electronic device, and the position and layout of other images or elements on the interface are consistent, so as to provide a better viewing experience, so that the generated stereoscopic image can be accurately presented to the second participant.

[0092] S1332, according to the multi-view sub-pixel mapping table and the pixel layout, determine the size and pixel layout of the stereoscopic image to be synthesized.

[0093] wherein, according to the pixel size information in the pixel layout, the width and height of the stereoscopic image to be synthesized can be determined, such as the size of the stereoscopic image matching or adapting to the size of the display interface of the first electronic device. According to the arrangement method information in the pixel layout, the arrangement method of the pixels in the stereoscopic image to be synthesized is determined, such as determining the number of pixels in each row and each column, and the arrangement rule between them. By ensuring that the size and pixel layout of the stereoscopic image to be synthesized match the multi-view sub-pixel mapping table, that is, ensuring that each pixel position has a corresponding position in the image sequence, so that the pixel values extracted from the image sequence can be correctly matched.

[0094] This step ensures that the stereoscopic image to be synthesized is compatible with the display interface of the first electronic device, and conforms to the preset multi-view sub-pixel mapping table, so as to ensure that the generated stereoscopic image can be correctly displayed on the display device, and the position and layout of other images or elements on the interface are consistent.

[0095] S1333, for each pixel position in the stereoscopic image to be synthesized, according to the multi-view sub-pixel mapping table, extract the pixel value corresponding to the pixel position from the image sequence.

[0096] S1334, fill the extracted pixel value to the corresponding pixel position in the stereoscopic image to be synthesized to generate the stereoscopic image corresponding to the second participant, wherein the filled position and pixel value match the multi-view sub-pixel mapping table and the pixel layout of the display interface.

[0097] Specifically, each pixel position of the stereoscopic image to be synthesized can be traversed, such as starting from the top left corner, traversing each pixel position in the stereoscopic image to be synthesized row by row and column by column. During the traversal process, for the current pixel position, the position in the image sequence corresponding to the pixel position is found according to the multi-view sub-pixel mapping table, wherein the multi-view sub-pixel mapping table can provide position information of each pixel position in the stereoscopic image to be synthesized in the image sequence. Next, according to the position obtained from the image sequence, the pixel value of the corresponding position in the image sequence is extracted, which can be a single image or multiple image pixel values, depending on the number of viewpoints defined in the multi-view sub-pixel mapping table. Then, the extracted pixel value is assigned to the corresponding pixel position in the stereoscopic image to be synthesized, so that each pixel position obtains the pixel value of the corresponding viewpoint.

[0098] In the process of generating the stereoscopic image, the embodiment determines the correct filling position according to the multi-view sub-pixel mapping table, and correctly arranges the corresponding pixel value according to the pixel layout of the display interface, to ensure that the generated stereoscopic image is consistent with the preset requirements. Thus, a high-quality stereoscopic image can be generated, with accuracy and consistency, and a multi-view experience can be provided, which can adapt to different display devices, thereby helping to improve the user's viewing experience and meet the demand for stereoscopic images.

[0099] In some embodiments, referring to FIG. 8, the method further comprises training the second 3D head model; the training the second 3D head model comprises:

[0100] S21, obtaining target video data, the target video data comprising a plurality of original head images of a target head in different postures, the second 3D head model comprising a Gaussian sphere learning model and a head rendering model;

[0101] S22, determining head control information corresponding to a target head image, the target head image being an original head image in a to-be-trained state;

[0102] S23, obtaining Gaussian sphere parameters output by the Gaussian sphere learning model when the target head image is used as a training input sample, the Gaussian sphere parameters being used to constrain a Gaussian sphere of the head rendering model;

[0103] S24, inputting the Gaussian sphere parameters and the head control information into the head rendering model to obtain a head rendering image;

[0104] S25, determine image difference information of the head rendering image and the target head image;

[0105] S26, train the second 3D head model according to the image difference information.

[0106] The Gaussian sphere parameters include Gaussian sphere rendering parameters and position offset parameters, the Gaussian sphere learning model includes a neural network model, the head rendering model includes a human head FLAME model and a 3D Gaussian rendering model, and the Gaussian sphere parameters output by the Gaussian sphere learning model when the target head image is used as a training input sample include: determining vertex features of vertices in the human head FLAME model that are associated with the posture of the target head when the target head image is used as a training input sample; inputting the vertex features of the vertices into the neural network model to obtain Gaussian sphere rendering parameters and position offset parameters, wherein the Gaussian sphere rendering parameters are used to represent rendering attributes of a Gaussian sphere of the 3D Gaussian rendering model, and the position offset parameters are used to indicate position offset of the Gaussian sphere of the 3D Gaussian rendering model.

[0107] The Gaussian sphere rendering parameters include spherical harmonics, opacity, a rotation matrix, and a scale vector, and the position offset parameters include an offset vector of a Gaussian sphere center point and a weight value of a vertex associated with the posture of the target head.

[0108] The Gaussian sphere parameters and the head control information are input into the head rendering model to obtain a head rendering image, including: determining target positions of target vertices after the target head changes according to the position offset parameters and the head control information, wherein the target vertices are vertices in the target head image; inputting the target positions of the target vertices and the Gaussian sphere rendering parameters into the human head FLAME model to obtain head rendering parameters corresponding to each target vertex; and inputting the head rendering parameters corresponding to each target vertex into the 3D Gaussian rendering model to obtain the head rendering image.

[0109] The target positions of the target vertices after the target head changes are determined according to the position offset parameters and the head control information, including: updating vertex positions of the target vertices according to the position offset parameters to obtain static positions of the target vertices before the target head changes; and moving the static positions according to the head control information to obtain the target positions of the target vertices after the target head changes.

[0110] The position offset parameter includes an offset vector of a Gaussian sphere center point, and the updating of the vertex position of the target vertex according to the position offset parameter to obtain the static position of the target vertex before the target head changes includes: adding the vertex position of the target vertex and the offset vector of the Gaussian sphere center point to obtain the static position of the target vertex before the target head changes.

[0111] The head control information includes a camera pose and a head key point of the target head image, the camera pose is used to indicate a pose of the camera shooting the target head, the position offset parameter includes a weight value of a vertex associated with the pose of the target head, and the moving of the static position according to the head control information to obtain a target position of the target vertex after the target head changes includes: performing position updating on the target vertex according to a preset linear blending skinning algorithm, using the static position, the camera pose, the control point and the vertex weight to obtain the target position of the target vertex after the target head changes.

[0112] The determining of the vertex feature of the vertex in the human head FLAME model for association with the pose of the target head includes: obtaining a vertex position of the vertex in the human head FLAME model for association with the pose of the target head; and performing a feature sampling operation on the vertex position of the vertex in a preset sampling space to obtain the vertex feature of the vertex.

[0113] Optionally, when the electronic device establishes a conference connection with a target device, real-time head control information of a target user sent by the target device can be received, the target user being a user participating in the conference through the target device; the real-time head control information is input into a 3D head model corresponding to the target user, so that the 3D head model outputs a real-time head rendering image, wherein the 3D head model is a second 3D head model obtained by training in the above embodiment; a 3D head image is generated according to the real-time head rendering image; and a 3D head image of the target user is presented.

[0114] It should be noted that the above-mentioned second 3D head model is a 3D head reconstruction model of a second conference participant corresponding to a second electronic device. Based on the same principle, a first 3D head model of a first conference participant corresponding to a first electronic device can also be trained, thereby obtaining a first 3D head reconstruction model of the first conference participant.

[0115] In the training method of the 3D head model, first, target video data is obtained, the target video data being a plurality of original head images of a target head in different postures. Second, head control information corresponding to the target head image is determined, the target head image being an original head image in a to-be-trained state. In this step, each original head image in the to-be-trained state can be sequentially selected as the target head image to be added to the training operation. Third, Gaussian sphere parameters output by a Gaussian sphere learning model are obtained, the Gaussian sphere parameters being used to constrain a Gaussian sphere of a 3D Gaussian rendering model. In this step, the Gaussian sphere parameters can be automatically selected to compensate for the deviation caused by the head control information to the image generation. Fourth, the Gaussian sphere parameters and the head control information are input into the head rendering model to obtain a head rendering image. In this step, the head rendering image obtained through training is generated to adjust the 3D head model in combination with the target head image. Finally, image difference information between the head rendering image and the target head image is determined, and the 3D head model is trained according to the image difference information. In this step, the image difference information is used as supervision information to feedback the training of the 3D head model to obtain an optimal 3D head model. Overall, the training method uses the Gaussian sphere parameters automatically learned and output by the Gaussian sphere learning model to compensate for the deviation caused by the head control information to the image generation. In this way, the performance of the 3D head model can be improved, and the 3D head model can output a head rendering image that is highly matched with the original head image.

[0116] In some embodiments, referring to FIG. 9, the method further includes training the second human motion capture model, and the training the second human motion capture model includes:

[0117] S31, obtaining target video data, the target video data including video data of a target human body in different postures, the video data including a plurality of original human motion images, and the second human motion capture model including a Gaussian sphere learning model and a human rendering model.

[0118] S32, obtaining human control information corresponding to a target human motion image, the target human motion image being an original human motion image in a training state.

[0119] S33, obtaining Gaussian sphere parameters output by the Gaussian sphere learning model when the target human image is used as a training input sample, the Gaussian sphere parameters being used to constrain a Gaussian sphere of the human rendering model.

[0120] S34, inputting the Gaussian sphere parameters and the human control information into the human rendering model to obtain a human motion rendering image.

[0121] S35, determining image difference information between the human motion rendering image and the target human motion image.

[0122] S36, training the second human action capture model according to the image difference information.

[0123] The Gaussian sphere parameters comprise Gaussian sphere rendering parameters and position offset parameters, the Gaussian sphere learning model comprises a neural network model, and the human body rendering model comprises an SMPL human body model and a 3D Gaussian rendering model. The obtaining of the Gaussian sphere parameters output by the Gaussian sphere learning model when the target human body image is used as a training input sample comprises: when the target human action picture is used as a training input sample, determining vertex features of vertices in the SMPL human body model that are used to be associated with a posture of the target human body; inputting the vertex features of the vertices into the neural network model to obtain Gaussian sphere rendering parameters and position offset parameters, the Gaussian sphere rendering parameters being used to represent rendering attributes of Gaussian spheres of the 3D Gaussian rendering model, and the position offset parameters being used to indicate position offsets of the Gaussian spheres of the 3D Gaussian rendering model.

[0124] The determining of the vertex features of the vertices in the SMPL human body model that are used to be associated with the posture of the target human body comprises: obtaining control points of the SMPL human body model, the control points being key points of specific joint positions in the SMPL human body model; performing posture transformation and shape adjustment on the key points to obtain vertex positions of the vertices in the SMPL human body model that are used to be associated with the posture of the target human body; and performing feature sampling operations on the vertex positions of the vertices in a preset sampling space to obtain the vertex features of the vertices.

[0125] The inputting of the vertex features of the vertices into the neural network model to obtain the Gaussian sphere rendering parameters and the position offset parameters comprises: inputting the vertex features of the vertices into the neural network model to output Gaussian sphere rendering parameters composed of spherical harmonics, opacity, a rotation matrix and a scale vector, and position offset parameters composed of an offset vector of a center point of a Gaussian sphere and a weight value of a vertex.

[0126] The inputting of the Gaussian sphere parameters and the human body control information into the human body rendering model to obtain the human action rendering image comprises: determining target positions of target vertices after the target human body changes according to the position offset parameters and the human body control information, the target vertices being vertices in the target human action image; inputting the target positions of the target vertices and the Gaussian sphere rendering parameters into the SMPL human body model to obtain rendering parameters of Gaussian spheres corresponding to each of the target vertices; and inputting the rendering parameters of the Gaussian spheres corresponding to each of the target vertices into the 3D Gaussian rendering model to obtain the human action rendering image.

[0127] The target position of the target vertex after the action of the target human body changes is determined according to the position offset parameter and the human body control information, including: updating the vertex position of the target vertex according to the position offset parameter to obtain a static position of the target vertex before the action of the target human body changes; and moving the static position according to the human body control information to obtain the target position of the target vertex after the action of the target human body changes.

[0128] The position offset parameter includes an offset vector of a Gaussian sphere center point, and the updating the vertex position of the target vertex according to the position offset parameter to obtain a static position of the target vertex before the action of the target human body changes includes: adding the coordinates corresponding to the vertex position of the target vertex and the offset vector of the Gaussian sphere center point to obtain the static position of the target vertex before the action of the target human body changes.

[0129] The human body control information includes a camera pose when the target human body action image in the target video data is acquired, and a key point associated with a specific joint position of a human body in the target human body action image; the position offset parameter includes a weight value of a vertex; and the moving the static position according to the human body control information to obtain the target position of the target vertex after the action of the target human body changes includes: updating the position of the target vertex of the static position according to the camera pose, the coordinates of the key point and the weight value of the vertex to obtain the target position of the target vertex after the action of the target human body changes.

[0130] The image difference information includes a loss value, and the training the second human body motion capture model according to the image difference information includes: optimizing neural network parameters in the neural network model and parameters of the SMPL human body model by using a gradient descent algorithm according to the loss value.

[0131] Optionally, when the electronic device establishes a conference connection with a target device, real-time human body control information of a target user sent by the target device is received, the target user being a user participating in the conference through the target device; the real-time human body control information is input into a human body motion capture model corresponding to the target user, so that the human body motion capture model generates a human body action rendering image of the target user; and a conference interface of the electronic device is controlled to present the human body action rendering image; wherein the human body motion capture model can be trained by the training method of the second human body motion capture model.

[0132] It should be noted that the second human motion capture model is a human motion capture model of the second participant corresponding to the second electronic device. Based on the same principle, the first human motion capture model of the first participant corresponding to the first electronic device can also be trained, so as to obtain the first human motion capture model of the first participant.

[0133] The training method can generate high-quality and accurate human motion rendering images by obtaining target video data, applying the human motion capture model and the human rendering model, and automatically learning the Gaussian sphere parameters output by the Gaussian sphere learning model. This automatic learning and parameter constraint method can compensate for the deviation of the human control information to the image generation, thereby improving the performance of the model and making the generated image more consistent with the real human motion. Moreover, the application of the Gaussian sphere parameters can reduce the computational complexity, speed up the image generation, and improve the real-time performance and interactivity. In addition, by obtaining video data of the target human in different postures and human control information corresponding to the target human motion image, the method can more accurately generate human motion rendering images. By using the human motion capture model and the human rendering model, the posture and motion of the human body can be better captured and presented, and the accuracy of image generation can be improved.

[0134] In some embodiments, before performing step S11 in FIG. 5, the method further includes: generating a first rendering model of a first participant participating in the conference by the first electronic device; and sending the first rendering model of the first participant to the second electronic device, so that the second electronic device saves the first rendering model of the first participant locally. Wherein, the real-time image or depth data of the first participant can be obtained by using the camera or depth sensor on the first electronic device. Then, the first rendering model of the first participant can be generated by processing and analyzing the obtained data through a corresponding algorithm, and the generated first rendering model of the first participant is sent to the second electronic device through the network or other communication mode. After receiving the first rendering model, the second electronic device saves it in the local storage. In this way, the second electronic device can use the first rendering model in the subsequent conference process without having to obtain it from the first electronic device every time. Thus, real-time rendering effect can be achieved, network burden can be reduced, and privacy and security can be improved. These features help to improve the conference experience and ensure the security of data.

[0135] The first rendering model of the first participant includes a first 3D head model and a first human motion capture model. Please refer to FIG. 10, the generation of the first rendering model of the first participant participating in the conference by the first electronic device includes:

[0136] S41, acquire first video data, the first video data is the video data of the first participant in different postures collected by the camera connected with the first electronic device, and the first video data includes a plurality of first full-body pictures;

[0137] S42, extract the expression information and the action information of the first participant according to the plurality of first full-body pictures.

[0138] S43, process the expression information according to a preset 3D head reconstruction algorithm to obtain a first 3D head model.

[0139] S44, process the action information according to a preset human action capture algorithm to obtain a first human action capture model.

[0140] Among them, the plurality of first full-body pictures provides perspective information in different angles and postures. By analyzing and processing the plurality of first full-body pictures, the expression information and the action information of the first participant can be extracted, which may involve face expression recognition, posture estimation or action analysis technology to obtain features related to human expression and action.

[0141] The first 3D head model can realize accurate presentation of head posture and expression through modeling and reconstruction of facial features. The method for obtaining the first 3D head model can refer to the above embodiments.

[0142] The first human action capture model can realize accurate presentation of human head posture and body posture through posture, motion tracking or joint skeleton analysis of the human body. The method for obtaining the first human action capture model can refer to the above embodiments.

[0143] The embodiment can realize more realistic rendering effect, natural action presentation, improve visual communication effect and enhance user's participation by generating the first rendering model of the first participant, including the 3D head model and the human action capture model.

[0144] Please refer to FIG. 11, which is a flowchart of an implementation method of an immersive conference according to another embodiment of the present application. The method is applied to a second electronic device and includes the following steps:

[0145] S51, when the second electronic device establishes a conference connection with the first electronic device, acquire an image stream of a second participant. The image stream of the second participant can be acquired by a camera or other image acquisition device.

[0146] S52, perform feature extraction on the image stream of the second participant to obtain real-time user feature data, wherein the real-time user feature data includes expression information and action information.

[0147] S53, resolving audio data from the image stream of the second participant.

[0148] S54, packaging the audio data and the real-time user feature data into video stream data and sending the video stream data to the first electronic device, so that the first electronic device inputs the real-time user feature data into a second rendering model corresponding to the second participant, generates a stereoscopic image, and causes the first electronic device to present the stereoscopic image of the second participant and the audio data.

[0149] The method of the embodiments of the present application can provide real-time visual and audio information and generate immersive presentation of stereoscopic images and audio data by steps such as image stream acquisition, feature extraction, audio resolution, and audio and video presentation, thereby enhancing the communication and interaction experience of the conference. In addition, the method only needs to transmit small feature data (expression, motion information) and audio data, instead of original high-definition video stream, which greatly reduces the bandwidth required for network transmission, so that good remote conference experience can be achieved even in a bandwidth-limited scenario; the second rendering model constructed in advance can efficiently reconstruct the 3D stereoscopic image of the second participant in real time according to the feature data, without complex real-time rendering calculation, which improves the overall rendering efficiency; the first electronic device can render the stereoscopic image of the second participant locally according to the received real-time feature data, and each participant can render the stereoscopic image of the second participant according to his own perspective, realizing multi-perspective immersive experience. Therefore, the implementation method of the immersive conference applied to the second electronic device not only effectively reduces the bandwidth requirement, but also supports multi-perspective interaction, and improves the rendering efficiency.

[0150] In some embodiments, before the image stream of the second participant is acquired, the method further includes: generating a second rendering model of the second participant; and sending the second rendering model of the second participant to the first electronic device, so that the first electronic device saves the second rendering model of the second participant locally. The second rendering model of the second participant includes a second 3D head model and a second human motion capture model, please refer to FIG. 12, and the generation of the second rendering model of the second participant includes:

[0151] S61, acquiring second video data, the second video data being video data of the second participant in different postures collected by a camera connected to the second electronic device, the second video data including a plurality of second full-body pictures;

[0152] S62, extracting expression information and motion information of the second participant according to the plurality of second full-body pictures;

[0153] S63, processing the expression information according to a preset 3D head reconstruction algorithm to obtain a second 3D head model;

[0154] S64, processing the motion information according to a preset human motion capture algorithm to obtain a second human motion capture model.

[0155] The second 3D head model can realize accurate presentation of head posture and expression through modeling and reconstruction of facial features. The method for obtaining the second 3D head model can refer to the above embodiments. The second human motion capture model can realize accurate presentation of human head posture and body posture through posture, motion tracking or joint skeleton analysis of the human body. The method for obtaining the second human motion capture model can refer to the above embodiments.

[0156] Please refer to FIG. 13, which is a structural schematic diagram of an implementation device of an immersive conference provided by an embodiment of the present application. The device is applied to a first electronic device, and the device 40 includes:

[0157] A video stream data receiving module 41 is configured to receive video stream data sent by a second electronic device participating in a conference together with the first electronic device. The video stream data includes audio data and real-time user feature data. The real-time user feature data is obtained by performing feature extraction on image streams of a second conference participant collected in real time by the second electronic device. The real-time user feature data includes expression information and motion information of the second conference participant.

[0158] A second rendering model obtaining module 42 is configured to obtain a second rendering model of the second conference participant. The second rendering model is constructed based on head posture information, expression information and motion information in multiple frames of human body images of the second conference participant. The head posture and facial expression of any two frames of the human body images are different at least in one aspect.

[0159] A stereoscopic image determining module 43 is configured to process the real-time user feature data by using the second rendering model to obtain a stereoscopic image of the second conference participant.

[0160] A 3D presentation module 44 is configured to present the stereoscopic image and the audio data.

[0161] The implementation device 40 of the immersive conference described above can be a software module. The software module includes a plurality of instructions stored in a memory. A processor can access the memory and call the instructions to perform the implementation method of the immersive conference described in the above embodiments.

[0162] In some embodiments, the implementation device 40 of the immersive conference can also be built by hardware devices, for example, the implementation device 40 of the immersive conference can be built by one or more chips, and each chip can work in coordination with each other to complete the implementation method of the immersive conference described in each embodiment. For another example, the implementation device 40 of the immersive conference can also be built by various logic devices, such as general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), single-chip microcomputers, ARM (Acorn RISC Machine), or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or any combination of these components.

[0163] It should be noted that the implementation device 40 of the immersive conference can execute the implementation method of the immersive conference applied to the first electronic device provided by the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method. Technical details not described in detail in the implementation device 40 of the immersive conference can be referred to the implementation method of the immersive conference applied to the first electronic device provided by the embodiments of the present application.

[0164] Please refer to FIG. 14, which is a structural schematic diagram of an implementation device of an immersive conference according to another embodiment of the present application. The device is applied to a second electronic device, and the device 50 includes:

[0165] An image stream acquisition module 51 is configured to acquire an image stream of a second conference participant when the second electronic device establishes a conference connection with a first electronic device.

[0166] An image stream feature extraction module 52 is configured to perform feature extraction on the image stream of the second conference participant to obtain real-time user feature data, wherein the real-time user feature data includes expression information and action information.

[0167] An audio data extraction module 53 is configured to parse audio data from the image stream of the second conference participant.

[0168] A video stream data sending module 54 is configured to package the audio data and the real-time user feature data into video stream data and send the video stream data to the first electronic device, so that the first electronic device inputs the real-time user feature data into a second rendering model corresponding to the second conference participant, generates a stereoscopic image, and makes the first electronic device present the stereoscopic image of the second conference participant and the audio data.

[0169] The implementation device 50 of the immersive conference can be a software module including a plurality of instructions stored in a memory accessible by a processor to invoke the instructions for execution to complete the implementation method of the immersive conference described in the embodiments.

[0170] In some embodiments, the implementation device 50 of the immersive conference can also be built by hardware devices, for example, the implementation device 50 of the immersive conference can be built by one or more chips, and each chip can work in coordination to complete the implementation method of the immersive conference described in the embodiments. For another example, the implementation device 50 of the immersive conference can also be built by various logic devices, such as a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine), or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or any combination of these components.

[0171] It should be noted that the implementation device 50 of the immersive conference can execute the implementation method of the immersive conference applied to the second electronic device provided in the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method. The technical details not described in detail in the embodiments of the implementation device 50 of the immersive conference can be referred to the implementation method of the immersive conference applied to the second electronic device provided in the embodiments of the present application.

[0172] Referring to FIG. 15, FIG. 15 is a structural schematic diagram of an electronic device provided in an embodiment of the present application. The electronic device includes one or more processors and a memory. The memory is connected to the one or more processors, for example, connected to the processor through a bus.

[0173] The processor is configured to support the computer device to perform the corresponding functions in the methods in the above method embodiments. The processor can be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The hardware chip can be an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0174] The memory is used to store program codes and the like. The memory can include a volatile memory (VM), such as a random access memory (RAM); the memory can also include a non-volatile memory (NVM), such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); and the memory can further include a combination of the above types of memories.

[0175] The memory can be used to store non-volatile software programs, non-volatile computer executable programs, and modules, such as program instructions / modules corresponding to the implementation method of the immersive conference in the embodiments of the present application. The processor performs various functional applications and data processing of the implementation method of the immersive conference and the implementation device of the immersive conference by running the non-volatile software programs, instructions, and modules stored in the memory, that is, implements the functions of each module or unit of the implementation method of the immersive conference and the implementation device of the immersive conference provided by the above method embodiments.

[0176] The memory can include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required by at least one function. The data storage area can store data created according to the use of the immersive conference implementation device, etc. In some embodiments, the memory can optionally include a memory remotely arranged with respect to the processor, which can be connected to the immersive conference implementation device through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0177] The one or more modules are stored in the memory and, when executed by the one or more processors, perform the immersive conference implementation method in any of the above method embodiments, for example, perform the method steps described in the above method embodiments, and implement the functions of the modules described in the above device embodiments.

[0178] The electronic device of the embodiments of the present application can specifically be the first electronic device or the second electronic device, which corresponds to a super-mobile personal computer device, a smart display or an all-in-one machine, a server or a server cluster, etc.

[0179] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, the computer program including program instructions, the program instructions causing a computer to perform the method described in the above embodiments when executed by the computer.

[0180] The embodiments of the present application also provide a computer program product, which includes a computer program stored on a computer readable storage medium, the computer program including program instructions, the program instructions causing an electronic device to perform the method described in the above embodiments when executed by the electronic device. The electronic device can be the first electronic device or the second electronic device.

[0181] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the program can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM) or a random access memory (RAM), etc.

[0182] The above disclosure is only the preferred embodiments of the present application, and of course cannot limit the scope of the rights of the present application, so equivalent changes made according to the claims of the present application are still within the scope of the present application.

Claims

1. An implementation method of an immersive conference, applied to a first electronic device, characterized in that, Comprising: receiving video stream data sent by the second electronic device participating in the conference together with the first electronic device, the video stream data including audio data and real-time user feature data, the real-time user feature data being obtained by the second electronic device performing feature extraction on a real-time collected image stream of a second conference participant, the real-time user feature data including expression information and action information of the second conference participant; obtaining a second rendering model of the second conference participant, wherein the second rendering model is constructed based on head posture information, expression information and action information in a plurality of frames of human body images of the second conference participant, and the head posture and facial expression of any two frames of the human body images are different at least in one; processing the real-time user feature data using the second rendering model to obtain a stereoscopic image of the second conference participant; presenting the stereoscopic image and the audio data.

2. The implementation method of claim 1, wherein, The processing of the real-time user feature data using the second rendering model to obtain the stereoscopic image of the second conference participant comprises: obtaining preset multi-view information; inputting the multi-view information and the real-time user feature data into the second rendering model for processing to obtain an image sequence; wherein the image sequence includes a plurality of rendering images, and there is a parallax between any two rendering images; generating the stereoscopic image of the second conference participant according to the image sequence.

3. The implementation method of claim 2, wherein, The obtaining of the preset multi-view information comprises: determining a preset number of virtual camera poses according to a set field of view angle and eye parallax angle, and the preset number of virtual camera poses constitute the multi-view information.

4. The implementation method of claim 2, wherein, The parallax between adjacent two rendering images in the plurality of rendering images is equal.

5. The implementation method of claim 2, wherein, The multi-view information includes a preset number of virtual camera poses. The inputting of the multi-view information and the real-time user feature data into the second rendering model for processing to obtain an image sequence comprises: obtaining a reference rendering image processed by the second rendering model on the real-time user feature data; adjusting the pose of the second conference participant in the reference rendering image according to the pose of the virtual camera in each virtual camera pose to generate a preset number of rendering images corresponding to the second conference participant, wherein the second conference participant in each rendering image corresponds to the perspective angle associated with the pose of a virtual camera.

6. The implementation method of claim 2, wherein, The generation of the stereoscopic image of the second conference participant according to the image sequence comprises: obtaining a preset multi-view sub-pixel mapping table and a pixel layout of a display interface of the first electronic device; determining the size and pixel layout of a to-be-combined stereoscopic image according to the multi-view sub-pixel mapping table and the pixel layout; for each pixel position in the to-be-combined stereoscopic image, extracting a pixel value corresponding to the pixel position from the image sequence according to the multi-view sub-pixel mapping table; Fill the extracted pixel values to corresponding pixel positions in the stereoscopic image to be synthesized to generate a stereoscopic image corresponding to the second conference participant, wherein the filled positions and pixel values match the multi-view sub-pixel mapping table and the pixel layout of the display interface.

7. The implementation method of claim 5, wherein, The second rendering model includes a second 3D head model and a second human body motion capture model, and the obtaining of the reference rendering image processed by the second rendering model on the real-time user feature data includes: inputting the expression information in the real-time user feature data into the second 3D head model to obtain a head rendering picture of the second conference participant; inputting the motion information in the real-time user feature data into the second human body motion capture model to obtain a human body motion rendering picture of the second conference participant; combining and rendering the head rendering picture and the human body motion rendering picture of the second conference participant to obtain a reference rendering image of the second conference participant.

8. The implementation method of claim 7, wherein, The method further includes training the second 3D head model. The training of the second 3D head model includes: obtaining target video data, the target video data including multiple original head images of a target head in different postures, and the second 3D head model including a Gaussian sphere learning model and a head rendering model; determining head control information corresponding to a target head image, the target head image being an original head image in a training state; obtaining Gaussian sphere parameters output by the Gaussian sphere learning model when the target head image is used as a training input sample, the Gaussian sphere parameters being used to constrain a Gaussian sphere of the head rendering model; inputting the Gaussian sphere parameters and the head control information into the head rendering model to obtain a head rendering image; determining image difference information between the head rendering image and the target head image; training the second 3D head model according to the image difference information.

9. The implementation method of claim 7, wherein, The method further includes training the second human body motion capture model. The training of the second human body motion capture model includes: obtaining target video data, the target video data including video data of a target human body in different postures, the video data including multiple original human body motion images, and the second human body motion capture model including a Gaussian sphere learning model and a human body rendering model; obtaining human body control information corresponding to a target human body motion image, the target human body motion image being an original human body motion image in a training state; obtaining Gaussian sphere parameters output by the Gaussian sphere learning model when the target human body image is used as a training input sample, the Gaussian sphere parameters being used to constrain a Gaussian sphere of the human body rendering model; inputting the Gaussian sphere parameters and the human body control information into the human body rendering model to obtain a human body motion rendering image; determining image difference information between the human body motion rendering image and the target human body motion image; training the second human body motion capture model according to the image difference information.

10. The implementation method according to any one of claims 1 to 9, characterized in that, Before receiving the video stream data sent by the second electronic device, the method further includes: generating a first rendering model of a first conference participant participating in a conference through the first electronic device; transmitting the first rendering model of the first participant to the second electronic device, so that the second electronic device saves the rendering model of the first participant locally.

11. The implementation method of claim 10, wherein, The first rendering model of the first participant includes a first 3D head model and a first human motion capture model, and the generating the first rendering model of the first participant participating in the conference through the first electronic device includes: obtaining first video data, the first video data being video data of the first participant in different postures collected by a camera connected to the first electronic device, the first video data including a plurality of first full-body pictures; extracting expression information and motion information of the first participant according to the plurality of first full-body pictures; processing the expression information according to a preset 3D head reconstruction algorithm to obtain a first 3D head model; processing the motion information according to a preset human motion capture algorithm to obtain a first human motion capture model.

12. An implementation method of an immersive meeting, applied to a second electronic device, comprising: including: when the second electronic device establishes a conference connection with the first electronic device, obtaining an image stream of a second participant; extracting features from the image stream of the second participant to obtain real-time user feature data, wherein the real-time user feature data includes expression information and motion information; parsing audio data from the image stream of the second participant; packaging the audio data and the real-time user feature data into video stream data and sending the video stream data to the first electronic device, so that the first electronic device inputs the real-time user feature data into a second rendering model corresponding to the second participant, generates a stereoscopic image, and causes the first electronic device to present the stereoscopic image and the audio data of the second participant.

13. The implementation method of claim 12, wherein, Before the image stream of the second participant is obtained, further including: generating a second rendering model of the second participant; transmitting the second rendering model of the second participant to the first electronic device, so that the first electronic device saves the second rendering model of the second participant locally.

14. The implementation method of claim 13, wherein, The second rendering model of the second participant includes a second 3D head model and a second human motion capture model, and the generating the second rendering model of the second participant includes: obtaining second video data, the second video data being video data of the second participant in different postures collected by a camera connected to the second electronic device, the second video data including a plurality of second full-body pictures; extracting expression information and motion information of the second participant according to the plurality of second full-body pictures; processing the expression information according to a preset 3D head reconstruction algorithm to obtain a second 3D head model; processing the motion information according to a preset human motion capture algorithm to obtain a second human motion capture model.

15. A first electronic device, comprising: including: a memory connected to the processor, the processor being configured to execute one or more computer programs stored in the memory, and the processor being configured to implement the method for implementing an immersive conference according to any one of claims 1-11 when executing the one or more computer programs.

16. A second electronic device, comprising: including: a memory connected to the processor, the memory for storing one or more computer programs to be executed by the processor, the processor, when executing the one or more computer programs, causing the first electronic device to implement the method for implementing an immersive conference according to any one of claims 12-14.

Citation Information

Patent Citations

  • Method, device and equipment for playing real-time video stream

    CN115883814A

  • Immersive virtual network conference method and device

    CN116016837A

  • Conference real-time interaction video generation method and device and conference system

    CN118138708A

  • Virtual 3D communications with participant viewpoint adjustment

    US20220286657A1

  • Supporting an immersive communication session between communication devices

    WO2023232267A1

Cited By

  • Display method and system, electronic equipment and storage medium

    CN122027777A