Immersive Video Conferencing System
By building virtual meeting spaces and generating immersive images based on viewpoints, the problem of lack of visual information in remote video conferencing is solved, and a more efficient immersive communication experience is achieved.
Patent Information
- Application Number
- CN202111522154.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-12-13
AI Technical Summary
During remote video conferencing, it is difficult for participants to feel visual information such as eye contact, resulting in low communication efficiency and difficult to achieve efficient communication effects in face-to-face conversations.
By determining the conference mode of video conference, a virtual conference space layout is constructed, and immersive conference images are generated based on viewpoint information, and an immersive video conference experience is provided using a display device and an image capture device.
It improves the flexibility and realism of video conferencing, allows participants to obtain a more realistic immersive communication experience, and enhances communication efficiency.
Smart Images

Figure CN114339120B_ABST
Abstract
Description
Background Art
[0001] In recent years, under the influence of various factors, remote video conferencing has gradually been applied to many aspects such as people's work or entertainment. Remote video conferencing can effectively help participants overcome limitations such as distance and achieve remote collaboration.
[0002] However, compared with face-to-face conversations, it is difficult for participants in a video conference to perceive visual information such as eye contact and conduct natural interactions (including turning the head, turning the head and shifting attention in a multi-person conference, having a private conversation, and sharing documents, etc.), which makes it difficult for video conferencing to provide efficient communication like face-to-face conversations. Summary of the Invention
[0003] According to an implementation of the present disclosure, a solution for immersive video conferencing is provided. In this solution, first, the conference mode of the video conference is determined, and this conference mode can indicate the layout of the virtual conference space of the video conference. Further, the viewpoint information associated with the second participant in the video conference can be determined based on the layout, and this viewpoint information is used to indicate the virtual viewpoint of the second participant watching the first participant in the video conference. Further, the first view of the first participant can be determined based on the viewpoint information, and the first view is sent to the conference device associated with the second participant for displaying the conference image generated based on this first view to the second participant. Thus, on the one hand, it can enable video conference participants to obtain a more realistic immersive video conference experience, and on the other hand, it can more flexibly obtain the desired virtual conference space layout according to needs.
[0004] The Summary of the Invention section is provided to introduce the identification of concepts in a simplified form, which will be further described in the detailed implementation below. The Summary of the Invention section is not intended to identify the key features or main features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Brief Description of the Drawings
[0005] Figure 1 A schematic diagram showing an example conference system arrangement according to some implementations of the present disclosure;
[0006] Figure 2A and Figure 2B A schematic diagram showing the conference mode according to some implementations of the present disclosure;
[0007] Figure 3A and Figure 3B A schematic diagram showing the conference mode according to some other implementations of the present disclosure;
[0008] Figure 4A and Figure 4B A schematic diagram showing the conference mode according to some further implementations of the present disclosure;
[0009] Figure 5 shows a schematic block diagram of an example conference system according to some implementations of the present disclosure;
[0010] Figure 6 shows a schematic diagram for determining viewpoint information according to some implementations of the present disclosure;
[0011] Figure 7 shows a schematic diagram of a view generation module according to some implementations of the present disclosure;
[0012] Figure 8 shows a schematic diagram of a depth prediction module according to some implementations of the present disclosure;
[0013] Figure 9 shows a schematic diagram of a view drawing module according to some implementations of the present disclosure;
[0014] Figure 10 shows a flowchart of an example method for video conferencing according to some implementations of the present disclosure;
[0015] Figure 11 shows a flowchart of an example method for generating a view according to some implementations of the present disclosure; and
[0016] Figure 12 shows a block diagram of an example computing device according to some implementations of the present disclosure.
[0017] In these drawings, like or similar reference numerals are used to denote like or similar elements. Detailed Description
[0018] The present disclosure will now be described with reference to several example implementations. It should be understood that these implementations are described only to enable those of ordinary skill in the art to better understand and thus implement the present disclosure, and do not imply any limitation on the scope of the subject matter.
[0019] As used herein, the term "comprising" and its variants are to be construed as open-ended terms meaning "including but not limited to". The term "based on" is to be construed as "at least partially based on". The terms "one implementation" and "an implementation" are to be construed as "at least one implementation". The term "another implementation" is to be construed as "at least one other implementation". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included hereinafter.
[0020] As discussed above, compared with face-to-face conversations, it is difficult for participants to perceive visual information such as eye contact in a video conference, which makes it difficult for video conferences to provide efficient communication like face-to-face conversations. People expect to obtain a more realistic and efficient communication experience in video conferences.
[0021] According to an implementation of the present disclosure, a solution for video conferencing is provided. In this solution, first, the conference mode of the video conference is determined, and this conference mode can indicate the layout of the virtual conference space of the video conference. Further, based on the layout, the viewpoint information associated with the second participant in the video conference can be determined, and this viewpoint information is used to indicate the virtual viewpoint of the second participant watching the first participant in the video conference. Further, based on the viewpoint information, the first view of the first participant can be determined, and the first view is sent to the conference device associated with the second participant for displaying the conference image generated based on this first view to the second participant.
[0022] By flexibly constructing the virtual conference space according to the conference mode, the embodiments of the present disclosure can improve the flexibility of the conference system. In addition, by generating a viewpoint-based view based on the viewpoint information, the embodiments of the present disclosure can also enable video conference participants to obtain a more realistic video conference experience.
[0023] The following describes the basic principles and several example implementations of the present disclosure with reference to the accompanying drawings.
[0024] Example Arrangement
[0025] Figure 1 An example conference system arrangement 100 according to an embodiment of the present disclosure is shown. As Figure 1 shown, the arrangement 100 (also referred to as a conference unit) may include, for example, a cubic physical conference space, and such a physical conference space may also be referred to as a Cubicle, for example. As will be described in detail below, such a physical conference space can be dynamically constructed into a virtual conference space for video conferencing according to the layout indicated by the conference mode, thereby improving the flexibility of the conference system.
[0026] As Figure 1 shown, the arrangement 100 may further include display devices 110-1, 110-2, and 110-3 (collectively or individually referred to as the display device 110). In Figure 1 the example arrangement 100, the display device 110 may include three separate display screens provided on three walls of the physical conference space, and it can be configured to provide an immersive conference image to the participant sitting on the chair 130. In some implementations, the display device 110 may also be provided on one wall or two walls of the physical conference space, for example.
[0027] In some implementations, the explicit device 110 may also include an integrally formed flexible screen (e.g., a circular screen). The flexible screen may, for example, have a 180-degree viewing angle to provide an immersive meeting image to the participants.
[0028] In some implementations, the explicit device 110 may also provide an immersive meeting image to the participants through other appropriate image presentation techniques. Exemplarily, the explicit device 110 may include a projection device for providing an immersive image to the participants. The projection device may, for example, project the meeting image on the wall of the physical meeting space.
[0029] As will be described in detail below, the immersive meeting image may include the views of other meeting participants in a video conference. In some implementations, the display device 110 may have an appropriate size, or the immersive image may have an appropriate size, such that the views of other meeting participants seen by the participants in the immersive image have a true scale, thereby enhancing the realism of the meeting system.
[0030] Additionally, the immersive meeting image may also include a virtual background to enhance the realism of the video conference. Additionally, the immersive meeting image may, for example, also include an operable image area, which may, for example, provide functions such as an electronic whiteboard to provide corresponding responses in response to the operations of appropriate participants in the video conference.
[0031] As Figure 1 shown, the arrangement 100 may also include a set of image capture devices 120. In some implementations, as Figure 1 shown, to improve the quality of the generated participant views, a set of image capture devices 120 may include multiple cameras that capture the participants from different directions. As Figure 1 shown, the set of image capture devices 120 may, for example, be arranged on a wall in the physical meeting space.
[0032] In some implementations, the image capture device 120 may, for example, include a depth camera to capture the image data and corresponding depth data of the participants. Alternatively, the image capture device 120 may also include ordinary RGB cameras and may determine the corresponding depth information through techniques such as binocular vision. In some implementations, all the cameras included in the image capture device 120 may be configured to be able to synchronously acquire images.
[0033] In some implementations, other corresponding components may also be provided in the arrangement 100 according to the needs of the meeting mode. For example, a semi-circular tabletop for a round-table meeting mode, an L-shaped corner tabletop for a side-by-side meeting mode, etc.
[0034] In this way, the participants in the video conference can obtain an immersive video conference experience through such a physical conference space. In addition, as will be described in detail below, such a modular physical conference space setup also helps to more flexibly construct the required virtual conference space.
[0035] In some implementations, the arrangement 100 may further include a control device 140 communicatively connected to the image capture device 120 and the display device 110. As will be described in detail below, the control device 140 can control processes such as participant image capture, video conference image generation, and display.
[0036] In some implementations, the explicit devices 110, image capture devices 120, and other components (semicircular desks, L-shaped corner desks, etc.) included in the arrangement 100 can also be pre-calibrated to determine the positions of all components in the arrangement 100.
[0037] Example Meeting Mode
[0038] Using the modular physical conference space as discussed above, embodiments of the present disclosure can virtualize multiple modular physical conference spaces into multiple sub-virtual spaces and accordingly construct virtual conference spaces with different layouts to support different types of meeting modes. Example meeting modes will be described below.
[0039] Example 1: Face-to-Face Meeting Mode
[0040] In some implementations, the meeting system of the present disclosure can support a face-to-face meeting mode. Figure 2A and Figure 2B shows a schematic diagram of a face-to-face meeting mode according to some implementations of the present disclosure. As Figure 2A shown, in the face-to-face meeting mode, the meeting system can construct a virtual conference space 200A by splicing face-to-face the sub-virtual spaces corresponding to the physical conference spaces where two participants 210 and 220 are located.
[0041] As Figure 2B shown, from the perspective of participant 210, the meeting system can use the front display device 110-1 in the physical conference space where participant 210 is located to provide a meeting image 225. As Figure 2B shown, the meeting image 225 may include a view of another participant 220. In some implementations, the meeting image 225 may also include, for example, a virtual background, such as a background wall and a semicircular desk.
[0042] In the face-to-face meeting mode, embodiments of the present disclosure enable two participants to obtain an experience as if having a face-to-face conversation at a table.
[0043] Example 2: Round Table Meeting Mode
[0044] In some implementations, the conferencing system of the present disclosure can support a round-table conferencing mode. Figure 3A and Figure 3B FIG. shows a schematic diagram of a round-table conferencing mode according to some implementations of the present disclosure. As Figure 3A shown, in the round-table conferencing mode, the conferencing system can combine sub-virtual spaces corresponding to the physical conferencing spaces where multiple participants (e.g., the participants 310, 320-1, and 320-2 shown in Figure 3A ) are located to construct a virtual conferencing space 300A. It can be seen that, different from the layout of the face-to-face conferencing mode, in the round-table conferencing mode, multiple participants can be arranged at a certain angle.
[0045] As Figure 3B shown, from the perspective of participant 310, the conferencing system can use the front display device 110-1 in the physical conferencing space where participant 310 is located to provide a conference image 325. As Figure 3B shown, the conference image 325 can include views of participants 320-1 and 320-2. In some implementations, the conference image 325 can also include, for example, a virtual background, such as a background wall, a semi-circular tabletop, or an electronic whiteboard area.
[0046] In some implementations, the electronic whiteboard area can be used to provide content related to video conferencing, such as documents, pictures, videos, slides, etc. Alternatively, the content of the electronic whiteboard area can change in response to instructions from appropriate participants. For example, the electronic whiteboard area can be used to play slides and can perform page-turning actions in response to gesture instructions, voice instructions, or other appropriate types of instructions from the slide presenter.
[0047] In the round-table conferencing mode, the embodiments of the present disclosure enable participants to obtain an experience of having a face-to-face conversation with multiple other participants as if they were sitting at the same table.
[0048] Example 3: Side-by-Side Meeting Mode
[0049] In some implementations, the conferencing system of the present disclosure can support a round-table conferencing mode. Figure 4A and Figure 4B FIG. shows a schematic diagram of a side-by-side conferencing mode according to some implementations of the present disclosure. As Figure 4AAs shown, in the side-by-side meeting mode, the meeting system can laterally combine the sub-virtual spaces corresponding to the physical meeting spaces where participants 410 and 420 are located to construct the virtual meeting space 400A. It can be seen that, different from the layout of the face-to-face meeting mode, in the side-by-side meeting mode, participant 420 will be presented on the side of participant 410, rather than in the front.
[0050] As Figure 4B shown, from the perspective of participant 410, the meeting system can use the display devices 110-1 and 110-2 in the physical meeting space where participant 410 is located to provide the meeting image 425.
[0051] As Figure 4B shown, the display device 110-1 on the side of participant 410 can be used to display the view of participant 420. In some implementations, the display device 110-1 can also display a virtual background associated with participant 420, such as a virtual desktop and a virtual display in front of participant 420. Thus, in the side-by-side meeting mode, participant 410 can obtain a visual experience as if participant 420 is located at an adjacent work station.
[0052] In some implementations, as Figure 4B shown, the display device 110-2 in front of participant 410 can also present, for example, an operable image area that supports interaction, such as the virtual screen area 430. In some implementations, the virtual screen can be, for example, the graphical interface of a cloud operating system, and participant 410 can interact with this graphical interface in an appropriate manner. For example, the participant can use the cloud operating system to perform online editing of documents through control devices such as a keyboard and a mouse.
[0053] In some implementations, the virtual screen area 430 can also be presented in real time through the display device in the physical meeting space where participant 420 is located, thereby realizing online remote interaction.
[0054] In an example scenario, participant 410 can, for example, use a keyboard to modify the code in the virtual screen area 430 in real time and, for example, solicit opinions from another participant 420 in real time through voice. Another participant 420 can view the modifications made by participant 410 in real time through the meeting image and can provide opinions through voice. Or, another participant 420 can, for example, also request control of the virtual screen area 430 and perform modifications through an appropriate control device (such as a mouse or a keyboard, etc.).
[0055] In another example scenario, the participants 410 and 420 may each have different virtual screen areas, similar to different work devices in a real work scenario. Further, such virtual screen areas can be implemented, for example, through a cloud operating system and can support the participants 410 or 420 to initiate real-time interaction between two different virtual screen areas. For example, in a drag-and-drop manner, a file can be dragged in real time from one virtual screen area to another virtual screen area, etc.
[0056] Thus, in the side-by-side meeting mode, the implementation of the present disclosure can utilize other areas of the display device to further provide work such as remote collaboration, thereby enriching the functions of the video conference.
[0057] In some implementations, the distance between the participants 410 and 420 in the virtual meeting space 400A can be dynamically adjusted according to input, for example, so that the two participants feel closer or farther.
[0058] Other Meeting Modes
[0059] Some example meeting modes are introduced above. It should be understood that other suitable meeting modes are also possible. Exemplarily, the meeting system of the present disclosure can also support, for example, a lecture meeting mode, in which one or more participants can be designated as speakers, and one or more other participants can be designated as listeners. Accordingly, the meeting system can construct a virtual meeting scene such that, for example, the speakers can be drawn on one side of the podium, and the listeners are drawn on the other side of the podium.
[0060] It should be understood that other suitable virtual meeting space layouts are also possible. Based on the modular physical meeting space discussed above, the meeting system of the present disclosure can flexibly construct different types of virtual meeting space layouts according to needs.
[0061] In some implementations, the meeting system can automatically determine the meeting mode according to the number of participants included in the video conference. For example, when it is determined that there are two participants, the system can automatically determine the face-to-face meeting mode.
[0062] In some implementations, the meeting system can automatically determine the meeting mode according to the number of meeting devices associated with the video conference. For example, when it is determined that the number of access terminals for the video conference is greater than two, the system can automatically determine the round-table meeting mode.
[0063] In some implementations, the meeting system can also determine the meeting mode according to the configuration information associated with the video conference. For example, the participants or organizers of the video conference can configure the meeting mode through input before initiating the video conference.
[0064] In some implementations, the conferencing system can also dynamically change the conferencing mode during a video conference based on the interactions of the participants in the video conference or in response to changes in the environment. For example, the conferencing system can default to recommending a two-person conferencing mode as a face-to-face mode and, upon receiving a participant instruction, dynamically adjust to a side-by-side conferencing mode. Or, the conferencing system initially only detects two participants and starts a face-to-face conferencing mode, and can automatically switch to a round-table conferencing mode after detecting that new participants have joined the video conference.
[0065] System Architecture
[0066] Figure 5 Further shown is an example architecture diagram of a conferencing system 500 implemented according to the present disclosure. As Figure 5 shown, the sender 550 represents a remote participant in the conferencing system 500, which can be, for example, Figure 2A participant 220 in Figure 3A participants 320-1 and 320-2 in Figure 4A or participant 420 in Figure 2A The receiver 560 represents a local participant in the conferencing system 500, such as Figure 3A participant 210 in Figure 4A participant 310 in
[0067] As Figure 5 shown, taking the sender 550 as an example, the conferencing system 500 can include an image acquisition module 510-1, which is configured to acquire an image of the sender 550 using the image capture device 120.
[0068] The conferencing system 500 further includes a viewpoint determination module 520-1, which is configured to determine the viewpoint information of the sender 550 based on the acquired image of the sender 550. This viewpoint information can be further provided to the view generation module 530-2 corresponding to the receiver 560.
[0069] The conferencing system 500 further includes a view generation module 530-1, which is configured to receive the viewpoint information of the receiver 560 determined by the viewpoint determination module 520-2 corresponding to the receiver 560 and generate a view of the sender 550 based on the image of the sender 550. This view can be further provided to the rendering module 540-2 corresponding to the receiver 560.
[0070] The conferencing system 500 further includes a rendering module 540-1 configured to generate a final conferencing image based on the received view of the recipient 560 and the background image for providing to the sender 550. In some implementations, the rendering module 540-1 can directly present the received view of the recipient 560. Alternatively, the rendering module 540-1 can also perform corresponding processing on the received view to obtain the final image of the recipient 560 for display.
[0071] The following will describe in detail the implementation of each module in conjunction with Figures 6 to 9 to describe the implementation of each module in detail.
[0072] Viewpoint Determination
[0073] As introduced above, the viewpoint determination module 520-2 is configured to determine the viewpoint information of the recipient 560 based on the captured image of the recipient 560. Figure 6 A schematic diagram of determining viewpoint information according to some implementations of the present disclosure is further shown.
[0074] As Figure 6 shown, the viewpoint determination module 520-1 or the viewpoint determination module 520-2 can determine the global coordinate system corresponding to the virtual conferencing space 630 based on the layout information indicated by the conferencing mode. Further, the viewpoint determination module 520 can further determine the coordinate transformation from the first physical conferencing space 620 of the sender 550 to the virtual conferencing space 630 and the coordinate transformation from the second physical conferencing space 610 of the recipient 560 to the virtual conferencing space so as to determine the coordinate transformation from the second physical conferencing space 610 to the first physical conferencing space 620
[0075] Further, the viewpoint determination module 520-1 or the viewpoint determination module 520-2 can determine the first viewpoint position of the recipient 560 in the second physical conferencing space 610. In some implementations, the viewpoint position can be determined by detecting the facial features of the recipient 560. Exemplarily. The viewpoint determination module 520 can detect the positions of the two eyes of the recipient 560 and determine the midpoint position of the two eyes as the first viewpoint position of the recipient 560. It should be understood that other appropriate feature points can also be used to determine the first viewpoint position of the recipient 560.
[0076] In some implementations, in order to determine the first viewpoint position, the system can be calibrated first to determine the relative position relationship between the display device 110 and the image capture device 120, and their positions relative to the ground.
[0077] Further, the image acquisition module 510-2 can acquire multiple images from the image capture device 120 for each frame, and the number thereof depends on the number of image capture devices 120. Face detection can be performed on each image. If a face can be detected, the pixel coordinates of the centers of the two eyes of the face are acquired, denoted as respectively, and the midpoint of these two pixels is denoted as the viewing point. If no face can be detected, or multiple faces are detected, then this image is skipped.
[0078] In some implementations, if two or more images can detect eyes, then the three-dimensional coordinates eye_pos of the viewing point of the current frame are calculated by triangulation. Then, filtering is performed on the three-dimensional coordinates eye_pos of the viewing point of the current frame. The filtering method is eye_pos’ = w * eye_pos + (1 - w) * eye_pos_prev. Where eye_pos_prev is the three-dimensional coordinates of the viewing point of the previous frame, and w is the weight coefficient of the current viewing point. The weight coefficient can, for example, be proportional to the distance L (in meters) between eye_pos and eye_pos_prev, and the time interval T (in seconds) between two frames. Exemplarily, w can be determined as (100 * L) * (5 * T), and finally its value is truncated between 0 and 1.
[0079] In some implementations, the viewing point determination module 520-1 or the viewing point determination module 520-2 can convert the first viewing point position into the second viewing point position (also referred to as the virtual viewing point) in the first physical conference space 620 according to the coordinate transformation from the second physical conference space 610 to the first physical conference space 620 for use in further determining the viewing point information of the view of the sender 550.
[0080] Exemplarily, the viewing point determination module 520-2 of the receiver 560 can determine the second viewing point position of the receiver 560 and send the second viewing point position to the sender 550. Alternatively, the viewing point determination module 520-2 of the receiver 560 can determine the first viewing point position of the receiver 560 and send the first viewing point position to the sender 550, so that the viewing point determination module 520-1 determines the second viewing point position of the receiver 560 in the first physical conference space 620 according to the first viewing point position.
[0081] By sending the viewing point position of the receiver 560 to the sender 550 for use in determining the view of the sender 550, the implementation of the present disclosure can save the transmission of the captured images of the sender 550, thereby reducing the network transmission overhead and reducing the transmission delay of the video conference.
[0082] View Generation
[0083] As introduced above, the view generation module 530-1 is configured to generate a view of the sender 550 based on the captured image of the sender 550 and the viewpoint information of the receiver 560. Figure 7 FIG. 700 is a schematic diagram further showing a view generation module according to some implementations of the present disclosure.
[0084] As Figure 7 shown, the view generation module 530-1 mainly includes a depth prediction module 740 and a view rendering module 760. The depth prediction module 740 is configured to determine a target depth map 750 based on a set of images 710 of the sender 550 captured by a set of image capture devices 120 and a corresponding set of depth maps 720. The view rendering module 760 is further configured to generate a view 770 of the sender 550 based on the target depth map 750, the set of images 710, and the set of depth maps 720.
[0085] In some implementations, the view generation module 540-1 may perform image segmentation on the set of images 710 to retain the image portion associated with the sender 550. It should be understood that any suitable image segmentation algorithm may be employed to process the set of images 710.
[0086] In some implementations, the set of images 710 used to determine the target depth map 750 and the view 770 may be selected from multiple image capture devices for capturing images of the sender 550 based on the viewpoint information. Exemplarily, taking the arrangement 100 Figure 1 shown as an example, the image capture devices may include, for example, six depth cameras installed at different positions.
[0087] In some implementations, the view generation module 530-1 may determine a set of image capture devices from multiple image capture devices based on the distance between the viewpoint position indicated by the viewpoint information and the installation positions of the multiple image capture devices for capturing images of the first participant, and obtain a set of images 710 and the corresponding depth maps 720 captured by the set of image capture devices. For example, the view generation module 530 may select the four depth cameras with the closest installation positions to the viewpoint position and obtain the images captured by the four depth cameras.
[0088] In some implementations, to improve the processing efficiency, the view generation module 530-1 may further include a downsampling module 730 to downsample the set of images 710 and the set of depth maps 720 to improve the operation efficiency.
[0089] Depth Prediction
[0090] The following will refer to Figure 8 to describe in detail the specific implementation of the depth prediction module 740. As Figure 8As shown, the depth prediction module 740 can first project a set of depth maps 720, denoted as {D i}, onto the virtual viewpoints indicated by the viewpoint information to obtain the projected depth maps {D′ i}. Further, the virtual viewpoint depth prediction module 740 can obtain the initial depth map 805 by averaging:
[0091]
[0092] where M′ i represents the visibility mask of {D′ i}.
[0093] Further, the depth prediction module 740 can also construct a set of candidate depth maps 810 based on the initial depth map 805. Specifically, the depth prediction module 740 can define a depth correction range [-Δd, Δd], uniformly sample N correction values {σ k} from this range, and add them to the initial depth map 805 to determine a set of candidate depth maps 810:
[0094]
[0095] Further, the depth prediction module 740 can determine the probability information associated with the set of candidate depth maps 810 by warping the set of images 720 to the virtual viewpoints using the set of candidate depth maps 810.
[0096] Specifically, as Figure 8 shown, the depth prediction module 740 can use a convolutional neural network CNN 815 to process a set of images 710, denoted as {I i}, to determine a set of image features 820, denoted as {F i}. Further, the depth prediction module 740 can include a warping module 825, which is configured to warp the set of image features 820 to the virtual viewpoints according to a set of virtual depth maps 710.
[0097] Further, the warping module 825 can further calculate the feature variance between multiple image features warped by different depth maps as the cost for the corresponding pixel points. Exemplarily, the cost matrix 830 can be represented as: H×W×N×C, where H represents the height of the image, W represents the width of the image, and C represents the number of feature channels.
[0098] Further, the depth prediction module 740 can use a convolutional neural network CNN 835 to process the cost matrix 830 to determine the probability information 840 associated with the set of candidate depth maps 810, denoted as P, whose size is H×W×N.
[0099] Furthermore, the depth prediction module 740 further includes a weighting module 845, which is configured to determine a target depth map 750 based on probability information and according to a set of candidate depth maps 710:
[0100]
[0101] Based on this approach, the implementation of the present disclosure can determine a more accurate depth map.
[0102] View Drawing
[0103] The following will refer to Figure 9 to describe in detail the specific implementation of the view rendering module 760. As Figure 9 shown, the view rendering module 760 may include a weight prediction module 920, which is configured to determine a set of blending weights based on the input features 910.
[0104] In some implementations, the weight prediction module 920 may be implemented as a machine learning model such as a convolutional neural network, for example. In some implementations, the input features 910 to the machine learning model may include the features of a set of projected images, which may be represented as: In some implementations, the set of projected images is determined by projecting a set of images 710 onto a virtual view point according to the target depth map 750.
[0105] In some implementations, the input features 910 may further include a visibility mask corresponding to the set of projected images
[0106] In some implementations, the input features 910 may further include depth difference information associated with a set of image capture view points, where the set of image capture view points indicates the view point positions of the set of image capture devices 120. Specifically, for each pixel p in the depth map D, the view rendering module 760 may determine depth information Specifically, the view rendering module 760 may project the depth map D onto the set of image capture view points to determine a set of projected depth maps. Further, the view rendering module 760 may further warp the set of depth maps back to the virtual view point to determine the depth information Furthermore, the view rendering module 760 may determine the difference between the two It should be understood that the warping operation is intended to represent corresponding the pixels in the projected depth map to the corresponding pixels in the depth map D without changing the depth values of the pixels in the projected depth map.
[0107] In some implementations, the input feature 910 may further include angular difference information, where the angular difference information indicates the difference between a first angle associated with a corresponding image capture view point and a second angle associated with a virtual view point. The first angle is determined based on the surface point corresponding to the pixel in the target depth map and the corresponding image capture view point, and the second angle is determined based on the surface point and the virtual view point.
[0108] Specifically, for the first capture view point in the set of image capture view points, the view rendering module 760 may determine the first angle from the surface point corresponding to the pixel in the depth map D to the first capture view point, denoted as Further, the view rendering module 760 may also determine the second angle from the surface point to the virtual view point, denoted as N. Further, the view rendering module 760 may determine the angular difference information based on the first angle and the second angle, which is denoted as
[0109] In some implementations, the input feature 910 may be represented as: It should be understood that the view rendering module 760 may also use only some of the above information as the input feature 910.
[0110] Further, the weight prediction module 920 may determine a set of blending weights based on the input feature 910. In some implementations, as Figure 9 shown, the view rendering module 760 may further include an upsampling module 930 to upsample the set of blending weights to obtain the weight information W that matches the original resolution i . Further, the weight prediction module 920 may also normalize the weight information, for example:
[0111]
[0112] Further, the view rendering module 760 may include a blending module 940 to blend a set of projected images based on the determined weight information to determine a blended image:
[0113]
[0114] In some implementations, the weight prediction module 920 may further include a post-processing module 950 to determine the first view 770 based on the blended image. In some embodiments, the post-processing module 950 may include a convolutional neural network for post-processing operations on the blended image, examples of which may include but are not limited to: optimizing contour boundaries, filling holes, or optimizing facial regions, etc.
[0115] Based on the view drawing module introduced above, by considering depth differences and angle differences during the determination of the blending weights, the implementation of the present disclosure can increase the weights of images with smaller depth differences and / or smaller angle differences during the blending process, thereby further improving the quality of the generated views.
[0116] Model Training
[0117] As introduced in the reference Figures 7 to 9 The view generation module 530-1 may include multiple machine learning models. In some implementations, the multiple machine learning models can be trained collaboratively through end-to-end training.
[0118] In some implementations, the loss function for training may include the difference between the blended image I a and the warped images {I′ i} obtained by warping a set of images 710:
[0119]
[0120] where x represents an image pixel, M = ∪ i M′ i represents the valid pixel mask of I a and ||·||1 represents the l1 norm operation.
[0121] In some implementations, the loss function for training may include the difference between the blended image I a and the ground-truth image I * :
[0122]
[0123] where the ground-truth image can be obtained, for example, using an additional image capture device.
[0124] In some implementations, the loss function for training may include the smoothness loss of the depth map:
[0125]
[0126] Wherein represents the Laplacian operator.
[0127] In some implementations, the loss function for training may include the difference between the blended image output by the blending module 940 and the ground-truth image I * :
[0128]
[0129] In some implementations, the loss function for training may include the rgba difference between the view output by the post-processing module 950 and the ground-truth image I * :
[0130]
[0131] In some implementations, the loss function for training may include the color difference between the view output by the post-processing module 950 and the ground-truth image I * :
[0132]
[0133] In some implementations, the loss function for training may include the α-map loss:
[0134]
[0135] In some implementations, the loss function for training may include the perceptual loss associated with the human face:
[0136]
[0137] where crop(·) represents the face detection box excision operation, and φ l (·) represents the feature extraction operation of the trained network.
[0138] In some implementations, the loss function for training may include the GAN loss:
[0139]
[0140] where D represents the discriminator network.
[0141] In some implementations, the loss function for training may include the adversarial loss:
[0142]
[0143] It should be understood that a combination of one or more of the above loss functions can be used as the objective function for training the view generation module 530-1.
[0144] Example Process
[0145] Figure 10 FIG. shows a flowchart of an example process 1000 for video conferencing according to some implementations of the present disclosure. The process 1000 may be performed, for example, by Figure 1 the control device 140 in or other suitable devices (e.g., as will be described in connection with Figure 11implemented by the device under discussion
[0146] As Figure 10 shown, at block 1002, the control device 140 determines the meeting mode of a video conference, which includes at least a first participant and a second participant, and the meeting mode indicates the layout of the virtual meeting space of the video conference.
[0147] At block 1004, based on the layout, the control device 140 determines the viewpoint information associated with the second participant, and the viewpoint information indicates the virtual viewpoint from which the second participant views the first participant in the video conference.
[0148] At block 1006, based on the viewpoint information, the control device 140 determines the first view of the first participant.
[0149] At block 1008, the control device 140 sends the first view to the meeting device associated with the second participant for displaying a meeting image to the second participant, and the meeting image is generated based on the first view.
[0150] In some implementations, the virtual meeting space includes a first sub-virtual space and a second sub-virtual space. The first sub-virtual space is determined by virtualizing the first physical meeting space where the first participant is located, the layout indicates the distribution of the first sub-virtual space and the second sub-virtual space in the virtual meeting space, and the second sub-virtual space is determined by virtualizing the second physical meeting space where the second participant is located.
[0151] In some implementations, determining the viewpoint information associated with the second participant based on the layout includes: based on the layout, determining a first coordinate transformation between the first physical meeting space and the virtual meeting space and a second coordinate transformation between the second physical meeting space and the virtual meeting space; based on the first coordinate transformation and the second coordinate transformation, transforming the first viewpoint position of the second participant in the second physical meeting space into a second viewpoint position in the first physical meeting space; and based on the second viewpoint position, determining the viewpoint information.
[0152] In some implementations, the first viewpoint position is determined by detecting the facial feature points of the second participant.
[0153] In some implementations, generating the first view of the first participant based on the viewpoint information includes: obtaining a set of images of the first participant captured by a set of image capture devices, where the set of images corresponds to a set of depth maps; based on the set of images and the set of depth maps, determining a target depth map corresponding to the viewpoint information; and based on the target depth map and the set of images, determining the first view of the first participant corresponding to the viewpoint information.
[0154] In some implementations, the method further includes: determining a set of image capture devices from the multiple image capture devices based on the distance between the viewpoint position indicated by the viewpoint information and the installation positions of the multiple image capture devices for capturing the image of the first participant.
[0155] In some implementations, the video conference further includes a third participant, and the generation of the conference image is further based on the second view of the third participant.
[0156] In some implementations, the conference image further includes an operable image area, and the graphic elements in the operable image area change in response to the interaction actions of the first participant or the second participant.
[0157] In some implementations, the conference mode includes at least one of the following: face-to-face conference mode, multi-person round-table conference mode, side-by-side conference mode, or lecture conference mode.
[0158] In some implementations, determining the conference mode of the video conference includes: determining the conference mode based on at least one of the following: the number of participants included in the video conference, the number of conference devices associated with the video conference, or the configuration information associated with the video conference.
[0159] Figure 11 A flowchart of an example process 1100 for determining a view according to some implementations of the present disclosure is shown. Process 1100 may be implemented, for example, by Figure 1 the control device 140 in or other suitable devices (such as the devices to be discussed in conjunction with Figure 11 .
[0160] As Figure 11 shown, at block 1102, the control device 140 determines a target depth map associated with the virtual viewpoint based on a set of images and a set of depth maps corresponding to the set of images, the set of images being captured by a set of image devices associated with a set of image capture viewpoints.
[0161] At block 1104, the control device 140 determines depth difference information or angle difference information associated with the set of image capture viewpoints; wherein, the depth difference information indicates the difference between the depth of a pixel in the projected depth map corresponding to the respective image capture viewpoint and the depth of the corresponding pixel in the target depth map, the projected depth map being determined by projecting the target depth map onto the respective image capture viewpoint, and the angle difference information indicates the difference between the first angle associated with the respective image capture viewpoint and the second angle associated with the virtual viewpoint, the first angle being determined based on the surface point corresponding to the pixel in the target depth map and the respective image capture viewpoint, and the second angle being determined based on the surface point and the virtual viewpoint.
[0162] At block 1106, the control device 140 determines a set of blending weights associated with a set of image capture viewpoints based on depth difference information or angle difference information.
[0163] At block 1108, the control device 140 blends a set of projected images based on a set of blending weights to determine a target view corresponding to a virtual viewpoint. The set of projected images is generated by projecting a set of images onto the virtual viewpoint. In some implementations, determining a target depth map associated with the virtual viewpoint includes: downsampling a set of images and a set of depth maps; and using the downsampled set of images and a set of depth maps to determine a target depth map corresponding to the viewpoint information.
[0164] In some implementations, blending a set of projected images based on a set of blending weights includes: upsampling the set of blending weights to determine weight information; and based on the weight information, blending the set of projected images to determine a target view corresponding to the virtual viewpoint.
[0165] In some implementations, determining a target depth map associated with the virtual viewpoint includes: based on a set of depth maps, determining an initial depth map corresponding to the virtual viewpoint; based on the initial depth map, constructing a set of candidate depth maps; determining probability information associated with the set of candidate depth maps by warping the set of images to the virtual viewpoint using the set of candidate depth maps; and based on the probability information, determining the target depth map from the set of candidate depth maps.
[0166] In some implementations, blending a set of projected images based on a set of blending weights includes: blending the set of projected images based on the set of blending weights to determine a blended image; and the method further includes: post-processing the blended image using a neural network to determine the target view.
[0167] Example Device
[0168] Figure 12 A schematic block diagram of an example device 1200 that can be used to implement embodiments of the present disclosure is shown. It should be understood that Figure 12 The device 1200 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described in the present disclosure. As Figure 12 shown, the components of the device 1200 may include, but are not limited to, one or more processors or processing units 1210, a memory 1220, a storage device 1230, one or more communication units 1240, one or more input devices 1250, and one or more output devices 1260.
[0169] In some implementations, device 1200 may be implemented as various user terminals or service terminals. The service terminal may be a server, a large computing device, etc. provided by various service providers. The user terminal is, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, multimedia computers, multimedia tablets, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is also foreseeable that device 1200 can support any type of user interface (such as a "wearable" circuit, etc.).
[0170] The processing unit 1210 may be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 1220. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of device 1200. The processing unit 1210 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0171] Device 1200 generally includes multiple computer storage media. Such media may be any accessible media available to device 1200, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 1220 may be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The memory 1220 may include one or more conference modules 1225, and these program modules are configured to perform the video conferencing functions of various implementations described herein. The conference module 1225 may be accessed and run by the processing unit 1210 to implement the corresponding functions. The storage device 1230 may be removable or non-removable media and may include machine-readable media that can be used to store information and / or data and can be accessed within device 1200.
[0172] The functions of the components of device 1200 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, device 1200 can operate in a networked environment using a logical connection to one or more other servers, personal computers (PCs), or another general network node. Device 1200 can also communicate with one or more external devices (not shown) as needed via communication unit 1240, such as databases, other storage devices, servers, display devices, etc., communicate with one or more devices that enable a user to interact with device 1200, or communicate with any device that enables device 1200 to communicate with one or more other computing devices (e.g., network cards, modems, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0173] Input device 1250 can be one or more various input devices, such as a mouse, keyboard, trackball, voice input device, camera, etc. Output device 1260 can be one or more output devices, such as a display, speaker, printer, etc.
[0174] Example Implementation
[0175] Some example implementations of the present disclosure are listed below.
[0176] In a first aspect of the present disclosure, a method for video conferencing is provided. The method includes: determining a conference mode of the video conference, the video conference including at least a first participant and a second participant, the conference mode indicating a layout of a virtual conference space of the video conference; based on the layout, determining viewpoint information associated with the second participant, the viewpoint information indicating a virtual viewpoint of the second participant viewing the first participant in the video conference; based on the viewpoint information, determining a first view of the first participant; and sending the first view to a conference device associated with the second participant for displaying a conference image to the second participant, the conference image being generated based on the first view.
[0177] In some implementations, the virtual conference space includes a first sub-virtual space and a second sub-virtual space, the layout indicating the distribution of the first sub-virtual space and the second sub-virtual space in the virtual conference space, the first sub-virtual space being determined by virtualizing a first physical conference space where the first participant is located, and the second sub-virtual space being determined by virtualizing a second physical conference space where the second participant is located.
[0178] In some implementations, determining the viewpoint information associated with a second participant based on the layout includes: determining, based on the layout, a first coordinate transformation between a first physical meeting space and a virtual meeting space and a second coordinate transformation between a second physical meeting space and the virtual meeting space; transforming, based on the first coordinate transformation and the second coordinate transformation, a first viewpoint position of the second participant in the second physical meeting space into a second viewpoint position in the first physical meeting space; and determining the viewpoint information based on the second viewpoint position.
[0179] In some implementations, the first viewpoint position is determined by detecting facial feature points of the second participant.
[0180] In some implementations, generating a first view of a first participant based on the viewpoint information includes: obtaining a set of images of the first participant captured by a set of image capture devices, the set of images corresponding to a set of depth maps; determining, based on the set of images and the set of depth maps, a target depth map corresponding to the viewpoint information; and determining, based on the target depth map and the set of images, a first view of the first participant corresponding to the viewpoint information.
[0181] In some implementations, the method further includes: determining a set of image capture devices from the plurality of image capture devices based on the distance between the viewpoint position indicated by the viewpoint information and the installation positions of the plurality of image capture devices for capturing images of the first participant.
[0182] In some implementations, the video conference further includes a third participant, and the generation of the conference image is further based on a second view of the third participant.
[0183] In some implementations, the conference image further includes an operable image area, and the graphic elements in the operable image area change in response to interaction actions of the first participant or the second participant.
[0184] In some implementations, the conference mode includes at least one of the following: face-to-face conference mode, multi-person round table conference mode, side-by-side conference mode, or lecture conference mode.
[0185] In some implementations, determining the conference mode of the video conference includes: determining the conference mode based on at least one of the following: the number of participants included in the video conference, the number of conference devices associated with the video conference, or the configuration information associated with the video conference.
[0186] In a second aspect of the present disclosure, an electronic device is provided. The device includes: a processing unit; and a memory coupled to the processing unit and containing instructions stored thereon, which when executed by the processing unit cause the device to perform the following actions: determining a meeting mode of a video conference, the video conference including at least a first participant and a second participant, the meeting mode indicating a layout of a virtual meeting space of the video conference; based on the layout, determining viewpoint information associated with the second participant, the viewpoint information indicating a virtual viewpoint of the second participant watching the first participant in the video conference; based on the viewpoint information, determining a first view of the first participant; and sending the first view to a meeting device associated with the second participant for displaying a meeting image to the second participant, the meeting image being generated based on the first view.
[0187] In some implementations, the virtual meeting space includes a first sub-virtual space and a second sub-virtual space, the layout indicating the distribution of the first sub-virtual space and the second sub-virtual space in the virtual meeting space, the first sub-virtual space being determined by virtualizing a first physical meeting space where the first participant is located, and the second sub-virtual space being determined by virtualizing a second physical meeting space where the second participant is located.
[0188] In some implementations, determining the viewpoint information associated with the second participant based on the layout includes: based on the layout, determining a first coordinate transformation between the first physical meeting space and the virtual meeting space and a second coordinate transformation between the second physical meeting space and the virtual meeting space; based on the first coordinate transformation and the second coordinate transformation, transforming a first viewpoint position of the second participant in the second physical meeting space into a second viewpoint position in the first physical meeting space; and based on the second viewpoint position, determining the viewpoint information.
[0189] In some implementations, the first viewpoint position is determined by detecting facial feature points of the second participant.
[0190] In some implementations, generating the first view of the first participant based on the viewpoint information includes: obtaining a set of images of the first participant captured by a set of image capture devices, the set of images corresponding to a set of depth maps; based on the set of images and the set of depth maps, determining a target depth map corresponding to the viewpoint information; and based on the target depth map and the set of images, determining the first view of the first participant corresponding to the viewpoint information.
[0191] In some implementations, the method further includes: determining a set of image capture devices from the plurality of image capture devices based on the distance between the viewpoint position indicated by the viewpoint information and the installation positions of the plurality of image capture devices for capturing images of the first participant.
[0192] In some implementations, the video conference further includes a third participant, and the generation of the meeting image is further based on a second view of the third participant.
[0193] In some implementations, the conference image further includes an operable image area, and the graphic elements in the operable image area change in response to the interaction actions of the first participant or the second participant.
[0194] In some implementations, the conference mode includes at least one of the following: face-to-face conference mode, multi-person round-table conference mode, side-by-side conference mode, or lecture conference mode.
[0195] In some implementations, determining the conference mode of a video conference includes: determining the conference mode based on at least one of the following: the number of participants included in the video conference, the number of conference devices associated with the video conference, or the configuration information associated with the video conference.
[0196] In a third aspect of the present disclosure, a computer program product is provided. The computer program product is tangibly stored in a non-transitory computer storage medium and includes machine-executable instructions that, when executed by a device, cause the device to perform the following actions: determining the conference mode of a video conference, the video conference including at least a first participant and a second participant, the conference mode indicating the layout of the virtual conference space of the video conference; based on the layout, determining viewpoint information associated with the second participant, the viewpoint information indicating the virtual viewing point of the second participant viewing the first participant in the video conference; based on the viewpoint information, determining a first view of the first participant; and sending the first view to a conference device associated with the second participant for displaying a conference image to the second participant, the conference image being generated based on the first view.
[0197] In some implementations, the virtual conference space includes a first sub-virtual space and a second sub-virtual space, the layout indicating the distribution of the first sub-virtual space and the second sub-virtual space in the virtual conference space, the first sub-virtual space being determined by virtualizing a first physical conference space where the first participant is located, and the second sub-virtual space being determined by virtualizing a second physical conference space where the second participant is located.
[0198] In some implementations, determining the viewpoint information associated with the second participant based on the layout includes: based on the layout, determining a first coordinate transformation between the first physical conference space and the virtual conference space and a second coordinate transformation between the second physical conference space and the virtual conference space; based on the first coordinate transformation and the second coordinate transformation, transforming a first viewpoint position of the second participant in the second physical conference space into a second viewpoint position in the first physical conference space; and based on the second viewpoint position, determining the viewpoint information.
[0199] In some implementations, the first viewpoint position is determined by detecting facial feature points of the second participant.
[0200] In some implementations, generating a first view of a first participant based on viewpoint information includes: obtaining a set of images of the first participant captured by a set of image capture devices, where the set of images corresponds to a set of depth maps; determining a target depth map corresponding to the viewpoint information based on the set of images and the set of depth maps; and determining a first view of the first participant corresponding to the viewpoint information based on the target depth map and the set of images.
[0201] In some implementations, the method further includes: determining a set of image capture devices from the plurality of image capture devices based on the distance between the viewpoint position indicated by the viewpoint information and the installation positions of the plurality of image capture devices for capturing images of the first participant.
[0202] In some implementations, the video conference further includes a third participant, and generating the conference image is further based on a second view of the third participant.
[0203] In some implementations, the conference image further includes an operable image area, and graphic elements in the operable image area change in response to interaction actions of the first participant or the second participant.
[0204] In some implementations, the conference mode includes at least one of the following: face-to-face conference mode, multi-person round table conference mode, side-by-side conference mode, or lecture conference mode.
[0205] In some implementations, determining the conference mode of the video conference includes: determining the conference mode based on at least one of the following: the number of participants included in the video conference, the number of conference devices associated with the video conference, or the configuration information associated with the video conference.
[0206] In a fourth aspect of the present disclosure, a method for video conferencing is provided. The method includes: determining a target depth map associated with a virtual view point based on a set of images and a set of depth maps corresponding to the set of images, the set of images being captured by a set of image devices associated with a set of image capture viewpoints; determining depth difference information or angle difference information associated with the set of image capture viewpoints; wherein the depth difference information indicates a difference between the depth of a pixel in a projected depth map corresponding to a respective image capture viewpoint and the depth of the corresponding pixel in the target depth map, the projected depth map being determined by projecting the target depth map onto the respective image capture viewpoint, and the angle difference information indicates a difference between a first angle associated with the respective image capture viewpoint and a second angle associated with the virtual view point, the first angle being determined based on a surface point corresponding to the pixel in the target depth map and the respective image capture viewpoint, and the second angle being determined based on the surface point and the virtual view point; determining a set of blending weights associated with the set of image capture viewpoints based on the depth difference information or the angle difference information; and blending a set of projected images based on the set of blending weights to determine a target view corresponding to the virtual view point, the set of projected images being generated by projecting the set of images onto the virtual view point.
[0207] In some implementations, determining the target depth map associated with the virtual view point includes: downsampling the set of images and the set of depth maps; and using the downsampled set of images and the set of depth maps to determine the target depth map corresponding to the viewpoint information.
[0208] In some implementations, blending the set of projected images based on the set of blending weights includes: upsampling the set of blending weights to determine weight information; and blending the set of projected images based on the weight information to determine the target view corresponding to the virtual view point.
[0209] In some implementations, determining the target depth map associated with the virtual view point includes: determining an initial depth map corresponding to the virtual view point based on the set of depth maps; constructing a set of candidate depth maps based on the initial depth map; determining probability information associated with the set of candidate depth maps by warping the set of images to the virtual view point using the set of candidate depth maps; and determining the target depth map according to the set of candidate depth maps based on the probability information.
[0210] In some implementations, blending the set of projected images based on the set of blending weights includes: blending the set of projected images based on the set of blending weights to determine a blended image; and the method further includes: post-processing the blended image using a neural network to determine the target view.
[0211] In a fifth aspect of the present disclosure, an electronic device is provided. The device includes: a processing unit; and a memory coupled to the processing unit and containing instructions stored thereon, which when executed by the processing unit cause the device to perform the following actions: determining a target depth map associated with a virtual view point based on a set of images and a set of depth maps corresponding to the set of images, the set of images being captured by a set of image devices associated with a set of image capture viewpoints; determining depth difference information or angle difference information associated with the set of image capture viewpoints; wherein the depth difference information indicates the difference between the depth of a pixel in a projected depth map corresponding to a respective image capture viewpoint and the depth of the corresponding pixel in the target depth map, the projected depth map being determined by projecting the target depth map onto the respective image capture viewpoint, and the angle difference information indicates the difference between a first angle associated with the respective image capture viewpoint and a second angle associated with the virtual view point, the first angle being determined based on a surface point corresponding to a pixel in the target depth map and the respective image capture viewpoint, and the second angle being determined based on the surface point and the virtual view point; determining a set of blending weights associated with the set of image capture viewpoints based on the depth difference information or the angle difference information; and blending a set of projected images based on the set of blending weights to determine a target view corresponding to the virtual view point, the set of projected images being generated by projecting the set of images onto the virtual view point.
[0212] In some implementations, determining the target depth map associated with the virtual view point includes: downsampling the set of images and the set of depth maps; and using the downsampled set of images and the set of depth maps to determine the target depth map corresponding to the viewpoint information.
[0213] In some implementations, blending the set of projected images based on the set of blending weights includes: upsampling the set of blending weights to determine weight information; and blending the set of projected images based on the weight information to determine the target view corresponding to the virtual view point.
[0214] In some implementations, determining the target depth map associated with the virtual view point includes: determining an initial depth map corresponding to the virtual view point based on the set of depth maps; constructing a set of candidate depth maps based on the initial depth map; determining probability information associated with the set of candidate depth maps by warping the set of images to the virtual view point using the set of candidate depth maps; and determining the target depth map according to the set of candidate depth maps based on the probability information.
[0215] In some implementations, blending the set of projected images based on the set of blending weights includes: blending the set of projected images based on the set of blending weights to determine a blended image; and the method further includes: post-processing the blended image using a neural network to determine the target view.
[0216] In a sixth aspect of the present disclosure, there is provided a computer program product. The computer program product is tangibly stored in a non-transitory computer storage medium and includes machine-executable instructions that, when executed by a device, cause the device to perform the following actions: determining a target depth map associated with a virtual viewpoint based on a set of images and a set of depth maps corresponding to the set of images, the set of images being captured by a set of image devices associated with a set of image capture viewpoints; determining depth difference information or angle difference information associated with the set of image capture viewpoints; wherein the depth difference information indicates a difference between the depth of a pixel in a projected depth map corresponding to a respective image capture viewpoint and the depth of the corresponding pixel in the target depth map, the projected depth map being determined by projecting the target depth map onto the respective image capture viewpoint, and the angle difference information indicates a difference between a first angle associated with the respective image capture viewpoint and a second angle associated with the virtual viewpoint, the first angle being determined based on a surface point corresponding to a pixel in the target depth map and the respective image capture viewpoint, and the second angle being determined based on the surface point and the virtual viewpoint; determining a set of blending weights associated with the set of image capture viewpoints based on the depth difference information or the angle difference information; and blending a set of projected images based on the set of blending weights to determine a target view corresponding to the virtual viewpoint, the set of projected images being generated by projecting the set of images onto the virtual viewpoint.
[0217] In some implementations, determining the target depth map associated with the virtual viewpoint includes: downsampling the set of images and the set of depth maps; and using the downsampled set of images and the set of depth maps to determine the target depth map corresponding to the viewpoint information.
[0218] In some implementations, blending the set of projected images based on the set of blending weights includes: upsampling the set of blending weights to determine weight information; and blending the set of projected images based on the weight information to determine the target view corresponding to the virtual viewpoint.
[0219] In some implementations, determining the target depth map associated with the virtual viewpoint includes: determining an initial depth map corresponding to the virtual viewpoint based on the set of depth maps; constructing a set of candidate depth maps based on the initial depth map; determining probability information associated with the set of candidate depth maps by warping the set of images to the virtual viewpoint using the set of candidate depth maps; and determining the target depth map based on the probability information according to the set of candidate depth maps.
[0220] In some implementations, blending the set of projected images based on the set of blending weights includes: blending the set of projected images based on the set of blending weights to determine a blended image; and the method further includes: post-processing the blended image using a neural network to determine the target view.
[0221] In a seventh aspect of the present disclosure, a video conferencing system is provided. The system includes: at least two conference units, each of the at least two conference units including: a set of image capture devices configured to capture images of participants in a video conference, the participants being in a physical conference space; and a display device disposed in the physical conference space for providing an immersive conference image to the participants, the immersive conference image including views of at least one other participant in the video conference; wherein at least two physical conference spaces of the at least two conference units are virtualized into at least two sub-virtual spaces, and the at least two sub-virtual spaces are organized into a virtual conference space for the video conference according to a layout indicated by a conference mode of the video conference.
[0222] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, the types of hardware logic components that may be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and the like.
[0223] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0224] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash memory), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0225] In addition, although the operations are depicted in a particular order, this should be understood as requiring that the operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features that are described in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, the various features that are described in the context of a single implementation may also be implemented separately or in any suitable subcombination in multiple implementations.
[0226] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for video conferencing, comprising: Determining a conference mode of the video conference, where the video conference includes at least a first participant and a second participant, and the conference mode indicates the layout of the virtual conference space of the video conference; Based on the layout, determining viewpoint information associated with the second participant, where the viewpoint information indicates the virtual viewpoint of the second participant watching the first participant in the video conference; Based on the viewpoint information, generating a first view of the first participant; And Sending the first view to a conference device associated with the second participant for displaying a conference image to the second participant, where the conference image is generated based on the first view; Where generating the first view of the first participant based on the viewpoint information includes: Based on a set of images of the first participant and a set of depth maps corresponding to the set of images, determining a target depth map associated with the viewpoint information, where the set of images of the first participant is captured by a set of image devices associated with a set of image capture viewpoints; Determining depth difference information or angle difference information associated with the set of image capture viewpoints; Based on the depth difference information or the angle difference information, determining a set of blending weights associated with the set of image capture viewpoints; And Based on the set of blending weights, blending a set of projected images to determine the first view of the first participant corresponding to the viewpoint information, where the set of projected images is generated by projecting the set of images onto the viewpoint information.
2. The method according to claim 1, where the virtual conference space includes a first sub-virtual space and a second sub-virtual space, the layout indicates the distribution of the first sub-virtual space and the second sub-virtual space in the virtual conference space, the first sub-virtual space is determined by virtualizing a first physical conference space where the first participant is located, and the second sub-virtual space is determined by virtualizing a second physical conference space where the second participant is located.
3. The method according to claim 2, where determining viewpoint information associated with the second participant based on the layout includes: Based on the layout, determining a first coordinate transformation between the first physical conference space and the virtual conference space and a second coordinate transformation between the second physical conference space and the virtual conference space; Based on the first coordinate transformation and the second coordinate transformation, transforming a first viewpoint position of the second participant in the second physical conference space into a second viewpoint position in the first physical conference space; and Based on the second viewpoint position, determining the viewpoint information.
4. The method according to claim 3, where the first viewpoint position is determined by detecting facial feature points of the second participant.
5. The method according to claim 1, further comprising: Determine the set of image capture devices from the plurality of image capture devices based on the distance between the viewpoint position indicated by the viewpoint information and the installation positions of the plurality of image capture devices for capturing an image of the first participant.
6. The method according to claim 1, wherein the video conference further includes a third participant, and the generation of the conference image is further based on a second view of the third participant.
7. The method according to claim 1, wherein the conference image further includes an operable image area, and graphic elements in the operable image area can change in response to interaction actions of the first participant or the second participant.
8. The method according to claim 1, wherein the conference mode includes at least one of the following: face-to-face conference mode, multi-person round table conference mode, side-by-side conference mode, lecture conference mode.
9. The method according to claim 1, wherein determining the conference mode of the video conference includes: Determining the conference mode based on at least one of the following: the number of participants included in the video conference, the number of conference devices associated with the video conference, and the configuration information associated with the video conference.
10. A method for generating a view, comprising: Determining a target depth map associated with a virtual viewpoint based on a set of images and a set of depth maps corresponding to the set of images, the set of images being captured by a set of image devices associated with a set of image capture viewpoints; Determining depth difference information or angle difference information associated with the set of image capture viewpoints; wherein the depth difference information indicates the difference between the depth of a pixel in a projected depth map corresponding to a respective image capture viewpoint and the depth of the corresponding pixel in the target depth map, the projected depth map being determined by projecting the target depth map onto the respective image capture viewpoint, and wherein the angle difference information indicates the difference between a first angle associated with the respective image capture viewpoint and a second angle associated with the virtual viewpoint, the first angle being determined based on a surface point corresponding to a pixel in the target depth map and the respective image capture viewpoint, and the second angle being determined based on the surface point and the virtual viewpoint; Determining a set of blending weights associated with the set of image capture viewpoints based on the depth difference information or the angle difference information; and Blending a set of projected images based on the set of blending weights to determine a target view corresponding to the virtual viewpoint, the set of projected images being generated by projecting the set of images onto the virtual viewpoint.
11. The method according to claim 10, wherein determining the target depth map associated with the virtual viewpoint includes: Downsampling the set of images and the set of depth maps; and Using the downsampled set of images and the set of depth maps to determine the target depth map corresponding to the virtual viewpoint.
12. The method according to claim 11, wherein blending the set of projected images based on the set of blending weights includes: Upsample the set of mixed weights to determine weight information; and Based on the weight information, mix the set of projected images to determine a target view corresponding to the virtual viewpoint.
13. The method according to claim 10, wherein determining a target depth map associated with a virtual viewpoint includes: Based on the set of depth maps, determine an initial depth map corresponding to the virtual viewpoint; Based on the initial depth map, construct a set of candidate depth maps; By using the set of candidate depth maps to warp the set of images to the virtual viewpoint, determine probability information associated with the set of candidate depth maps; and Based on the probability information, determine the target depth map according to the set of candidate depth maps.
14. The method according to claim 10, wherein mixing the set of projection images based on the set of hybrid weights comprises: Based on the set of mixed weights, mix the set of projected images to determine a mixed image; and The method further includes: Post-process the mixed image by using a neural network to determine the target view.
15. An electronic device, comprising: A processing unit; and A memory, coupled to the processing unit and containing instructions stored thereon, the instructions, when executed by the processing unit, cause the electronic device to perform the method according to any one of claims 1 to 9.
16. An electronic device, comprising: A processing unit; and A memory, coupled to the processing unit and containing instructions stored thereon, the instructions, when executed by the processing unit, cause the electronic device to perform the method according to any one of claims 10 to 14.
17. A computer program product, the computer program product being tangibly stored in a non-transitory computer storage medium and including machine-executable instructions, the machine-executable instructions, when executed by a device, cause the device to perform the method according to any one of claims 1 to 9.
18. A computer program product, the computer program product being tangibly stored in a non-transitory computer storage medium and including machine-executable instructions, the machine-executable instructions, when executed by a device, cause the device to perform the method according to any one of claims 10 to 14.
19. A video conferencing system, comprising: At least two conference units, each of the at least two conference units including: A set of image capture devices, configured to capture a set of images of participants in a video conference, the participants being in a physical conference space, wherein the set of image capture devices is associated with a set of image capture viewpoints; and A display device, disposed in the physical conference space for providing immersive conference images to the participants, the immersive conference images including views of at least one other participant in the video conference; Wherein at least two physical conference spaces of the at least two conference units are virtualized into at least two sub-virtual spaces, and the at least two sub-virtual spaces are organized into a virtual conference space for the video conference according to a layout indicated by a conference mode of the video conference; Wherein each of the at least two conference units further includes: A view generation module, configured to: Determine a target depth map associated with the viewpoint information of the at least one other participant based on the set of images of the participant and the set of depth maps corresponding to the set of images; Determine depth difference information or angle difference information associated with the set of image capture viewpoints; Determine a set of blending weights associated with the set of image capture viewpoints based on the depth difference information or the angle difference information; and Blend a set of projected images based on the set of blending weights to determine the view of the participant corresponding to the viewpoint information of the at least one other participant, the set of projected images being generated by projecting the set of images onto the viewpoint information of the at least one other participant.
Citation Information
Patent Citations
Controlled three-dimensional communication endpoint
US20140098183A1
Light-field viewpoint and pixel culling for a head mounted display device
US20190088023A1