Video telephone conference method and device
By capturing and displaying a real view of end users in a video conference device, the problem of face occlusion in existing systems is solved, improving user experience and supporting interoperability of multiple devices.
Patent Information
- Application Number
- CN202380090527.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-14
- Filing Date
- 2023-08-18
- Publication Date
- 2025-08-08
AI Technical Summary
The existing video conference call system cannot effectively capture and display the real view of the end user, resulting in poor user experience, especially when using HMD and AR glasses, which affects the immersive experience.
Using a video teleconferencing device, the end user's own view on the mirror reflecting surface is captured through the camera, and the pixel area is extracted from the video and sent to a remote device for display. The mirror acts as both a capture area and a display area to realize the real view display of the end user.
Improve the video call experience, allowing both parties to see real people instead of avatars, reduce computing resource consumption, support the interoperability of devices such as HMD, phones, PCs, AR glasses, etc., and eliminate the uncanny valley effect and occlusion problems.
Smart Images

Figure CN120457671A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is based upon and claims the benefit of priority from European patent application number “23305944.3” filed on June 14, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure relates to a teleconferencing system comprising at least one video teleconferencing device including a camera whose field of view at least partially covers the viewpoint of an end user, such as, for example, a camera of an optical or video see-through head-mounted display device. Background Art
[0004] This section is intended to introduce the reader to various aspects of the art that may be related to various aspects of the present principles described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present principles. Therefore, it should be understood that these statements are to be read in this light, and not as admissions of prior art.
[0005] In recent years, the emergence of large field of view content (up to 360°) has become increasingly common. For users who view content using immersive display devices (such as head-mounted displays (HMDs) or smart glasses, PC screens, tablets, smartphones, etc.), such content may not be fully displayed. This means that at a given moment, the user may only be able to see a portion of the content. Although the user can navigate in the content through various means such as head movement, mouse movement, touch screen, voice, etc., large field of view content can specifically include three-dimensional computer graphics scenes (3D CGI scenes), point clouds, or immersive videos.
[0006] Many terms can be used to design this type of immersive video: virtual reality (VR), 360, panoramic, 4π steradian, omnidirectional, large field of view, etc.
[0007] To encode and decode omnidirectional video into a bitstream, for example for transmission over a data network, a conventional video codec / decoder such as HEVC (ISO / IEC 23008-2 High Efficiency Video Coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en) or VVC (ISO / IEC 23090-3 Versatile Video Coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / T-REC-H.266-202008-I / en) can be used. Encoding immersive video involves projecting each frame of the immersive video onto one or more 2D frames (e.g., one or more rectangular frames) using a suitable projection function. In practice, the frames of the immersive video are represented as 3D surfaces. To facilitate projection, a convex and simple surface is typically used for projection, such as a sphere, a cube, or a pyramid. The projected 2D pictures representing the pictures of the immersive video are then encoded / decoded using conventional video codecs / decoders such as HEVC, VVC, etc. Rendering the immersive video includes decoding the 2D pictures and deprojecting the decoded 2D pictures.
[0008] One of the most promising and challenging future uses for HMDs is in applications where virtual environments augment, rather than replace, the real world. To gain an augmented view of their real environment, users can wear a see-through HMD to see 3D computer-generated objects superimposed on their real-world view. This see-through capability can be achieved using either optical or video see-through HMDs.
[0009] With an optical see-through HMD, the general principle is that the user simultaneously perceives the real world and computer-generated imagery (CGI) signals via an optical system. The real world is viewed through a translucent display panel, which is typically equipped with mirrors placed in front of the user's eyes. These mirrors are also used to reflect the CGI images into the user's eyes, thereby combining the real-world and virtual-world views.
[0010] With a video see-through HMD, the real-world view is captured using multiple miniature video cameras mounted on the head-mounted device, and the CGI signal is electronically combined with a video representation of the real world and then displayed to the end user.
[0011] Video conferencing relies on two or more parties exchanging video and audio representing each party in real time. This is typically achieved using a camera (recording the user), a display (showing the other users), and microphones / speakers. However, in extended reality (XR), although the corresponding devices (HMDs, AR glasses, etc.) are often equipped with cameras, they face outward and cannot record the user wearing the device, so they cannot provide a self-view for video calls. In order to achieve the capture of the self-view, an external camera is required, thus limiting the independence and usability of dedicated XR devices.
[0012] Current approaches to teleconferencing in XR do not attempt to mimic existing video call experiences that users are familiar with, but rather to create new, rich video call experiences. This is done either by introducing avatars that can be adjusted to fit the scene more appropriately (Taeheon Kim, Ashwin Kachhara, and Blair Mac Intyre. 2016. Redirected head gaze) or by capturing and transmitting a 3D representation of the caller (Alexander Gerd Reis and Didier Stricker. 2022. A Survey on Synchronous Augmented, Virtual, and Mixed Reality Remote Collaboration Systems. Comput. Surveys 55, 6 (2022), 1-27), or by rendering video capture of the user in space (Mark Billinghurst and Hirokazu Kato. 1999. Real world teleconferencing. In CHI'99 extended abstracts on Human factors in computing systems. 194-195).
[0013] While these features can enhance immersion, the video call experience can be further enhanced by introducing a true view of the end user using an HMD. Currently, this is impossible because the HMD obscures the end user's face. New optical see-through AR glasses are developing towards more compact and slimmer forms, more similar to regular glasses, minimizing obstruction of the wearer's facial features.
[0014] However, a major obstacle to using face-obstructing devices in video teleconferencing systems is the lack of means to capture and display the end-user's self-view. Summary of the Invention
[0015] The following section provides a brief overview of at least one exemplary embodiment to provide a basic understanding of some aspects of the present disclosure. This overview is not an exhaustive overview of the exemplary embodiments. Its purpose is not to identify key or core elements of the exemplary embodiments. The following overview merely presents some aspects of at least one exemplary embodiment in a simplified form as a prelude to the more detailed descriptions provided elsewhere in this document.
[0016] According to a first aspect of the present disclosure, a video teleconference method is provided. The video teleconference method is performed between at least two remote video teleconference devices communicating via a communication network, wherein at least one video teleconference device, referred to as a first video teleconference device, includes: a camera, referred to as a first camera, and a display, referred to as a first display, wherein a field of view of the first camera at least partially covers a viewpoint of an end user of the first video teleconference device, wherein the video teleconference method includes the following steps:
[0017] - capturing a first video by a first camera, the first video representing a view of an end user of the first video teleconferencing device by themselves on a reflective surface of a mirror, the mirror being located in front of the end user and within the field of view of the first camera;
[0018] - extracting at least one pixel region from the first video, the at least one pixel region surrounding the end user's self-view; and
[0019] - sending first video data to at least one remote video teleconferencing device, the first video data representing the at least one extracted pixel region.
[0020] In one variation, the video teleconferencing method further comprises:
[0021] - receiving second video data from a remote video teleconferencing device; and
[0022] - Rendering a second video on the first display based on the received second video data so as to cover at least part of the reflective surface on the first display.
[0023] In an exemplary embodiment, wherein the first video teleconferencing device is a video see-through device, and wherein rendering the second video includes replacing pixels of the first display belonging to the extracted at least one pixel region with pixels of the second video.
[0024] In one exemplary embodiment, wherein the first video teleconferencing device is an optical see-through device, and wherein the second video is rendered based on the received second video data on an area of the first display so as to overlay a projection of the reflective surface on the first display.
[0025] In an exemplary embodiment, pixels of the second video on the first display correspond to a projection of the reflective surface, the pixels of the second video have an opacity value adapted to hide the end user's view of themselves on the reflective surface, and the projection of the reflective surface on the first display is adapted to the end user's viewpoint.
[0026] In one variation, the video teleconferencing method further includes receiving third video data representing a third video; and rendering the third video based on the received third video data on the display over an area extending beyond the projection of the reflective surface so as to overlay the second video.
[0027] In one variation, the video teleconferencing method further includes receiving third video data representing a third video and rendering the third video based on the received third video data over an area surrounding the area corresponding to the projection of the reflective surface on the first display.
[0028] In an exemplary embodiment, based on the detection of the face, at least one pixel region is extracted and detected in the first video.
[0029] In an exemplary embodiment, at least one extracted pixel region is detected in the first video based on detecting a view of the first video teleconferencing device in the first video.
[0030] In one exemplary embodiment, the extracted at least one pixel region is detected based at least in part on a correlation of the motion detected in the first video with sensor data captured by a sensor of the first video teleconferencing device.
[0031] In one variation, the video teleconferencing method further includes removing or replacing a background of the extracted at least one pixel region prior to transmission.
[0032] In one variation, the video teleconferencing method further includes formatting the extracted at least one pixel region into a video format or a volume format, or encapsulating the extracted at least one pixel region into a real-time transport protocol packet suitable for real-time communication.
[0033] According to a second aspect of the present disclosure, a computer program product comprising instructions is provided. When the program is executed by one or more processors, the one or more processors are caused to perform the video teleconferencing method according to the first aspect of the present disclosure.
[0034] According to a third aspect of the present disclosure, a video teleconference apparatus is provided, which includes components for executing the steps of the video teleconference method according to the first aspect of the present disclosure.
[0035] According to a fourth aspect of the present disclosure, a video teleconference system is provided, which includes at least one video teleconference device according to the third aspect of the present disclosure.
[0036] The specific nature of at least one exemplary embodiment and other objects, advantages, features and uses of the at least one exemplary embodiment will become apparent from the following description of the examples in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 An illustrative example of a video teleconferencing system according to an exemplary embodiment of the present disclosure is shown;
[0038] Figure 2 A schematic block diagram illustrating an example of a video teleconferencing device according to an exemplary embodiment of the present disclosure is illustrated;
[0039] Figure 3 A schematic block diagram showing steps of a video teleconference method according to an exemplary embodiment of the present disclosure;
[0040] Figure 4 Shown Figure 1 Illustrative examples of implementation of the video teleconferencing method;
[0041] Figure 5 Shown Figure 1 Illustrative examples of implementation of the video teleconferencing method;
[0042] Figure 6 Shown Figure 1 An illustrative example of an implementation of a video teleconferencing method.
[0043] Similar or identical elements are referenced with the same reference numerals. DETAILED DESCRIPTION
[0044] At least one of the exemplary embodiments will be described more fully below with reference to the accompanying drawings, which depict examples of at least one of the exemplary embodiments. However, the exemplary embodiments may be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Therefore, it should be understood that the present invention is not intended to limit the exemplary embodiments to the specific forms disclosed.
[0045] The video includes consecutive video frames, and there is a temporal relationship between the consecutive video frames.
[0046] A video picture, also called a frame or picture frame, comprises at least one component (also called picture component or channel) determined by a specific picture / video format, which picture / video format specifies all information related to pixel values and all information that can be used by a display unit and / or any other device to display and / or decode video picture data related to the video picture.
[0047] A video picture comprises at least one component which is usually represented in the form of an array of samples.
[0048] A monochrome video picture includes a single component, while a color video picture may include three components.
[0049] For example, when the picture / video format is the well-known (Y, Cb, Cr) format, a color video picture may include one luminance (or luminance) component and two chrominance components, and when the picture / video format is the well-known (R, G, B) format, a color video picture may include three color components (one for red, one for green, and one for blue).
[0050] Each component of a video picture may comprise a number of samples relative to the number of pixels of the screen on which the video picture is to be displayed. In a variant, the number of samples comprised in a component may be a multiple (or fraction) of the number of samples comprised in another component of the same video picture.
[0051] For example, in case the video format includes one luma component and two chroma components (such as the (Y, Cb, Cr) format), the chroma components may contain half the number of samples in width and / or height relative to the luma component, depending on the color format considered.
[0052] A sample is the smallest unit of visual information that makes up a component of a video picture. A sample value can be, for example, a luminance or chrominance value, or a color value in (R, G, B) format. A luminance sample is a brightness value, and a chrominance sample is a chrominance or color value.
[0053] A pixel value is the value of a pixel on the screen. For monochrome video, a pixel value can be represented by a single sample, while for color video, a pixel value can be represented by multiple co-located samples. The co-located samples associated with a pixel are samples that correspond to the position of the pixel on the screen.
[0054] A video frame is usually viewed as a set of pixel values, with each pixel represented by at least one sample.
[0055] At least one exemplary embodiment is not limited to a specific picture / video format.
[0056] At least one aspect generally relates to a video teleconferencing system comprising, for example, at least two remote video teleconferencing devices communicating via a communications network, such as the Internet. At least one video teleconferencing device, referred to as a first video teleconferencing device, comprises a camera, referred to as a first camera, a display, referred to as a first display, and a mirror, wherein the field of view of the first camera at least partially covers the viewpoint of an end user of the first video teleconferencing device, and the mirror is located in front of the user and within the field of view of the first camera. A first video is captured by the first camera, the first video representing the end user's self-view of the first video teleconferencing device on the reflective surface of the mirror, and at least one pixel region is extracted from the first video, the at least one pixel region corresponding to the end user's self-view. The first video data is then sent to the at least one remote video teleconferencing device. The first video data represents the extracted at least one pixel region.
[0057] A first video teleconferencing device is used by a terminal user. This first video teleconferencing device uses a camera to detect the terminal user's view of themselves on a reflective surface of a mirror. The captured video of the terminal user is then transmitted to a remote video teleconferencing device during a video call. The remote video teleconferencing device can display the terminal user's actual view, rather than an avatar as in the prior art.
[0058] Compared to conventional video teleconferencing systems using a see-through HMD and / or AR glasses, using the first video teleconferencing device in a video call improves the end-user experience of the video call because both parties involved in the video call see real people without the need for avatars, which are often considered to have serious problems with the quality of experience (an effect known as the uncanny valley) or are not suitable for usage scenarios (e.g., professional and work environments).
[0059] Using the first video teleconferencing device in a video call further provides an illusion of proximity between people, i.e., the image of the other party is not floating around but anchored to physical objects in the local real scene.
[0060] Compared to avatar solutions that require tracking of the entire user's body and a subsequent avatar generation process (usually based on a series of artificial intelligence algorithms), using the first video teleconferencing device in a video call consumes fewer resources (computing power, memory, etc.).
[0061] Using the first video teleconferencing device in a video call is also advantageous because no external phone or camera is required to capture the user to provide a true view of the user.
[0062] In one exemplary embodiment, the first video teleconferencing device can receive the other party's video and then render the other party's video on top of a display area corresponding to the display area occupied by the mirror, thereby hiding the user's own view from the user themselves.
[0063] Thus, the mirror of the first video teleconferencing device acts both as a capture area (for the camera) to record the end user's own view and as an area for displaying the remote user's video.
[0064] Therefore, the present disclosure solves the problem of blocking the end user's own view and requires neither an additional display nor an external camera, and allows interoperability between types of video teleconferencing devices such as HMDs, phones, PCs, AR glasses that can be used in combination.
[0065] The present disclosure is also advantageous because capturing the end user's view on the reflective surface by a first camera and rendering the video received from another party can be processed in parallel because they are independent processes.
[0066] Figure 1 An example of the video teleconference system 1 according to an exemplary embodiment of the present disclosure is shown.
[0067] The video teleconference system 1 includes a first video teleconference apparatus 100 located at a site 10 and a second video teleconference apparatus 200 located at a remote site 20 .
[0068] The second video teleconference device is not limited to a specific type, but can be extended to any device including components for joining a video teleconference system. In particular, each video teleconference device should include a camera and a video display.
[0069] The end user 101 uses the first video teleconferencing device 100 during a video call, and the end user 201 uses the second video teleconferencing device 200 during the video call.
[0070] The video teleconference system is not limited to two video teleconference devices, but can be extended to an unlimited number of first video teleconference devices and / or an unlimited number of second video teleconference devices.
[0071] Figure 2 Shown is a schematic block diagram illustrating an example of a first video teleconferencing apparatus 100 in which various aspects and exemplary embodiments of the present disclosure are implemented.
[0072] The first video teleconferencing apparatus 100 may be implemented as one or more devices including the components described below. In various exemplary embodiments, the first video teleconferencing apparatus 100 may be configured to implement one or more of the aspects described in the present disclosure.
[0073] Example embodiments of equipment that may form part of the first video teleconferencing apparatus 100 include, for example, an optical or video HMD device or AR glasses.
[0074] According to the present disclosure, the first video teleconference device 100 further includes a mirror 105 placed in front of the end user 101 , ie, the reflective surface of the mirror 105 is located in front of the user's (end user 101 ) eyes so as to reflect the end user's 101 own view.
[0075] When the first video teleconferencing device includes a video see-through device (e.g., an HMD), the video to be displayed is calculated and displayed on the display 104 of the video see-through device. The displayed video may be a combination of two or more videos. The pixels of the display 104 corresponding to the pixels of the reflective surface of the mirror 105 are replaced with the pixel values of the displayed video. Thus, the end user's own view is hidden.
[0076] When the first video teleconferencing device includes an optical see-through device, the display 104 of the optical see-through device is a semi-transparent display, and the end user 101 wearing the optical see-through device sees themselves on the reflective surface of the mirror 105 before starting the teleconferencing session. When the session begins, the semi-transparent display 104 can display an image in such a way that the end user of the optical see-through device cannot see their own view on the reflective surface. For example, a CGI image can be displayed on top of the surface 105 to hide the end user's self-view.
[0077] According to the present disclosure, pixels of the video to be displayed on the semi-transparent display 104 correspond to pixels of the projection of the reflective surface of the mirror 105, and the pixels of the video to be displayed on the semi-transparent display have opacity values adapted to hide the end user's self-view on the reflective surface of the mirror 105. The video to be displayed on the semi-transparent display 104 then covers the reflective surface of the mirror 105, and the end user of the optical see-through device cannot see his or her self-view on the reflective surface of the mirror 105.
[0078] In an exemplary embodiment, the projection of the reflective surface of the mirror 105 is adapted to the end user viewpoint on the translucent display 104 by using well-known alignment methods (ie, methods of aligning virtual objects into the real scene according to the end user viewpoint).
[0079] The various elements of the first video teleconferencing device 100 may be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one exemplary embodiment, the processing and encoder / decoder elements of the first video teleconferencing device 100 may be distributed across multiple ICs and / or discrete components.
[0080] The first video teleconferencing device 100 may include at least one processor 1010 configured to execute instructions loaded therein to implement, for example, the various aspects described in the present disclosure. The processor 1010 may include embedded memory, input / output interfaces, and various other circuit systems known in the art. The first video teleconferencing device 100 may include at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). The first video teleconferencing device 100 may include a storage device 1040, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, a magnetic disk drive, and / or an optical disk drive. As non-limiting examples, the storage device 1040 may include an internal storage device, an attached storage device, and / or a network-accessible storage device.
[0081] The first video teleconferencing device 100 may include, for example, an encoder / decoder module 1030 configured to process data to provide encoded / decoded video data, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 may represent (one or more) modules that may be included in the device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. In addition, the encoder / decoder module 1030 may be implemented as a separate element of the first video teleconferencing device 100, or may be integrated into the processor 1010 as a combination of hardware and software known to those skilled in the art.
[0082] Program code to be loaded onto the processor 1010 or the encoder / decoder 1030 to perform various aspects described in the present disclosure may be stored in the storage device 1040 and subsequently loaded onto the memory 1020 for execution by the processor 1010. According to various exemplary embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of various items during the execution of the processes described in the present disclosure. Such stored items may include, but are not limited to, video data, information data for encoding / decoding video data, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0083] In several exemplary embodiments, memory within the processor 1010 and / or encoder / decoder module 1030 may be used to store instructions and provide working memory for processes that may be performed during encoding or decoding.
[0084] However, in other exemplary embodiments, memory external to a processing device (e.g., the processing device may be either processor 1010 or encoder / decoder module 1030) may be used for one or more of these functions. The external memory may be memory 1020 and / or storage device 1040, such as dynamic volatile memory and / or non-volatile flash memory. In at least one exemplary embodiment, a fast external dynamic volatile memory (such as RAM) may be used as working memory for video encoding and decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 Video), (ISO / IEC 14496-10 Advanced Video Coding for generic audio-visual services, ITU-T Recommendation H.264, https: / / www.itu.int / rec / T-REC-H.264-202108-P / en), EVC (ISO / IEC 23094-1 Essential video coding), HEVC, VVC, but may be applied to other standards and recommendations such as AV1 (AOMedia Video 1, http: / / aomedia.org / av1 / specification / ), etc.
[0085] In various exemplary embodiments, the first video teleconferencing device 100 may be communicatively coupled to other similar devices or other electronic devices via, for example, a wired or wireless communication network, a communication bus, or through dedicated input and / or output ports.
[0086] In an exemplary embodiment, input to elements of the first video teleconferencing apparatus 100 may be provided through various input devices as indicated in block 1090 .
[0087] According to this disclosure, Figure 1 As illustrated above, one of the input devices is a camera 102 , the field of view of which at least partially covers the viewpoint of the end user 101 .
[0088] As indicated in box 1090, the input device may include, but is not limited to: (i) an RF part that can receive an RF signal transmitted wirelessly, for example, by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, (iv) an HDMI input terminal, (v) a bus when the present disclosure is implemented in the automotive field, such as a CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-rate), FlexRay (ISO 17458) or Ethernet (ISO / IEC 802-3) bus.
[0089] In various exemplary embodiments, as indicated in block 1090, the input device may have associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements necessary for: (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a certain frequency band); (ii) down-converting the selected signal; (iii) band-limiting again to a narrower frequency band to select a signal band that may be referred to as a channel in some exemplary embodiments; (iv) demodulating the down-converted and band-limited signal; (v) performing error correction; and (vi) demultiplexing to select a desired data packet stream. The RF section in various exemplary embodiments may include one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband.
[0090] Various exemplary embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.
[0091] Adding an element may include inserting an element between existing elements, such as, for example, inserting an amplifier and an analog-to-digital converter.In various exemplary embodiments, the RF portion may include an antenna.
[0092] In addition, the USB and / or HDMI terminals may include corresponding interface processors for connecting the first video teleconferencing device 100 to other electronic devices via USB and / or HDMI connections. It will be understood that various aspects of input processing (e.g., Reed-Solomon error correction) can be implemented, for example, within a separate input processing IC or processor 1010, if necessary. Similarly, various aspects of USB or HDMI interface processing can be implemented within a separate interface IC or processor 1010, if necessary. The demodulated, error-corrected, and demultiplexed streams can be provided to various processing elements, including, for example, the processor 1010 and the encoder / decoder 1030, which operate in combination with memory and storage elements to process the data streams, if necessary, for presentation on an output device.
[0093] The various components of the first video teleconferencing device 100 can be provided in an integrated housing. Within the integrated housing, the various components can be interconnected and transmit data between them using a suitable connection arrangement 1090 (e.g., an internal bus known in the art, including an I2C bus, wires, and printed circuit boards).
[0094] The first video teleconferencing apparatus 100 may further include a communication interface 1050 that enables communication with other devices (such as the second video teleconferencing apparatus 200) via a communication channel 1051. The communication interface 1050 may include, but is not limited to, a transceiver configured to send and receive data through the communication channel 1051. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1051 may be implemented, for example, within a wired and / or wireless medium.
[0095] Video data can be streamed to or from the first video teleconferencing device 100 using a Wi-Fi network (such as IEEE 802.11). Wi-Fi signals of these exemplary embodiments can be received / sent via a communication channel 1051 suitable for Wi-Fi communication and a communication interface 1050. The communication channel 1051 of these exemplary embodiments can typically be connected to an access point or router that provides access to external networks (including the Internet) to allow streaming applications and other over-the-top communications.
[0096] Other exemplary embodiments may provide streaming video data to or receive streaming video data from an external device (such as the second video teleconferencing apparatus 200 ) via a connection (eg, an HDMI or RF connection) of input block 1090 .
[0097] The streamed video data may be used in the form of signaling information for use by the first video teleconferencing device 100 or the second video teleconferencing device 200. The signaling information may include a bitstream and / or information such as the number of pixels of a video screen and / or any codec / decoding setting parameters related to encoding / decoding and / or display and / or rendering of the streamed video data.
[0098] It will be appreciated that signaling can be implemented in various ways. For example, in various exemplary embodiments, one or more syntax elements, flags, etc. can be used to signal information.
[0099] The first video teleconferencing apparatus 100 may provide a video signal to a display 104 and possibly various other output devices, including speakers 1071 and other peripheral devices 1081. In various examples of the exemplary embodiments, the other peripheral devices 1081 may include one or more of a disk player, a stereo system, a lighting system, and other devices that provide functionality based on the output of the first video teleconferencing apparatus 100.
[0100] In various exemplary embodiments, control signals may be communicated between the first video teleconferencing apparatus 100 and an output device such as a display 104, a speaker 1071 or other peripheral device 1081 using signaling such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that enable device-to-device control with or without user intervention.
[0101] The output devices may be communicatively coupled to the first video teleconferencing apparatus 100 via dedicated connections through respective interfaces 1060 , 1070 , and 1080 .
[0102] Alternatively, the output device may be connected to the first video teleconferencing apparatus 100 via the communication interface 1050 using the communication channel 1051. The display 104 and the speaker 1071 may be integrated with other components of the first video teleconferencing apparatus 100 in a single unit in an electronic device such as, for example, an optical or video see-through HMD.
[0103] In various exemplary embodiments, the display interface 1060 may include a display driver, such as, for example, a timing controller (TCon) chip.
[0104] The display 104 and the speaker 1071 may alternatively be separate from one or more of the other components. In various exemplary embodiments where the display 104 and the speaker 1071 may be external components, the output signal may be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0105] Figure 1 The second video teleconferencing device 200 used by the end user 201 includes a camera 202 and a display 203.
[0106] Examples of equipment that may form all or part of the second video teleconferencing apparatus 200 include a personal computer, a laptop, a smartphone, a tablet, a digital television receiver, a personal video recording system, a connected home appliance, a connected vehicle and its associated processing systems, an optical or video HMD, AR glasses, a projector (beamer), a "cave" (a system comprising multiple displays), or any other communication device.
[0107] The camera 202 is used to capture the video of the terminal user 201 during the video call. The captured video of the terminal user 201 can be displayed on the display 203, such as Figure 1 As shown in .
[0108] Figure 3 A schematic block diagram showing steps of a video teleconference method according to an exemplary embodiment of the present disclosure is shown.
[0109] In step 100, if Figure 1 As illustrated in , a first video is captured by camera 102 (also referred to as a first camera), the first video representing the end user's 101 view of himself on the reflective surface of mirror 105.
[0110] In step 110 , at least one pixel region is extracted from the first video. The at least one pixel region corresponds to a spatial region surrounding the end user's 101 self-view in the reflective surface of the mirror 105 .
[0111] The pixel region of the first video may be defined as a set of pixel positions of the first video and video data corresponding to the set of pixel positions. The video data corresponding to the set of pixel positions changes during the video call.
[0112] exist Figure 1 In one exemplary embodiment illustrated above, a single pixel region 106 is extracted from the first video.
[0113] In a variant, the area of the extracted pixel areas is the area of the reflective surface of the mirror 105 .
[0114] In an exemplary embodiment, at least one pixel region is extracted from the first video based on the detection of the face.
[0115] In one exemplary embodiment, at least one pixel region is extracted from the first video based on detecting a view of the first video teleconferencing device 100 in the first video.
[0116] In an exemplary embodiment, at least one pixel region is extracted from the first video based at least in part on correlation of motion detected in the first video with sensor data captured by a sensor of the first video teleconferencing device 100 .
[0117] For example, the sensor data may be movements derived by an object tracking component implemented by the first video teleconferencing device 100 that tracks over time objects detected in the first video. Well-known object tracking methods may be used.
[0118] exist Figure 4 In step 120 shown in the upper diagram, the first video data is sent to at least one remote video teleconference device (such as Figure 1 In the second remote video teleconferencing device 200), the first video data represents video data corresponding to the extracted at least one pixel area.
[0119] Then, at least one remote video teleconference device may calculate rendering of the first video on a display based on the received first video data.
[0120] In one variation, the video conferencing method further includes steps 130 and 140 .
[0121] In step 130 , the first video teleconferencing device 100 receives second video data from a remote video teleconferencing device (eg, from the second video teleconferencing device 200 ).
[0122] In step 140 , the second video is then rendered on a display of the video teleconferencing device (referred to as the first display) based on the received second video data so as to overlay the reflective surface of the mirror 105 on the first display.
[0123] In an exemplary embodiment of step 140 , the first video teleconferencing device is a video see-through device, and rendering the second video includes replacing pixels of the first display belonging to the extracted at least one pixel region with pixels of the second video.
[0124] exist Figure 5In an exemplary embodiment of step 140 illustrated above, when the first video teleconferencing device 100 is an optical see-through device (such as an HMD), the second video is rendered based on the received second video data on an area of the display 104 so as to overlay a projection of the reflective surface of the mirror 105 on the display 104.
[0125] In an exemplary embodiment, pixels of the second video corresponding to the projection of the reflective surface of the mirror 105 on the display 104 have opacity values adapted to hide the end user 101's own view on the reflective surface of the mirror 105, and the projection of the reflective surface of the mirror 105 on the display 104 is adapted to (aligned to) the viewpoint of the end user 101.
[0126] exist Figure 6 In one exemplary embodiment illustrated above, the second video is rendered on the entire display 104 .
[0127] When the first video teleconferencing device 100 is an optical see-through device, the pixels of the second video corresponding to the projection of the reflective surface of the mirror 105 on the display 104 have an opacity value adapted to hide the reflective surface of the mirror 105, and the projection of the reflective surface of the mirror 105 on the display 104 is adapted to (aligned to) the viewpoint of the end user 101.
[0128] In an exemplary embodiment, when the first video teleconferencing device 100 is a video see-through device, the first video is at least partially displayed on the display 104, and rendering the second video (step 140) includes replacing pixel values of the first video with pixel values of the second video.
[0129] In one variation, the video conferencing method further includes steps 150 and 160 .
[0130] In step 150 , third video data representing a third video is received.
[0131] The third video data may represent, for example, a shared view of an application, such as an office, a presentation, a desktop, or the like.
[0132] In step 160 , a third video is rendered on the display 104 based on the received third video data over an area of the display 104 that extends beyond the projection of the reflective surface of the mirror 105 so as to overlay the second video.
[0133] For example, the third video is rendered on the entire display 104 .
[0134] When the first video teleconferencing device 100 is an optical see-through device (e.g., an HMD), the third video is rendered based on the received third video data on an area of the display 104 so as to overlay the projection of the reflective surface of the mirror 105 on the display 104. Pixels of the third video corresponding to the projection of the reflective surface of the mirror 105 on the display 104 have opacity values adapted to hide the end user 101's view of themselves on the reflective surface of the mirror 105, and the projection of the reflective surface of the mirror 105 on the display 104 is adapted to (aligned with) the viewpoint of the end user 101.
[0135] In a variant of step 160 , a third video is rendered on the display 104 over an area surrounding the area corresponding to the projection of the reflective surface of the mirror 105 , according to the received third video data.
[0136] This variant is advantageous because it allows for simultaneous rendering of a second video (eg, the end user's 201 own view) on the area corresponding to the projection on the reflective surface of the mirror 105 and a third video on the display 104 .
[0137] In one variant, the video conferencing method further comprises a step 170 during which the background of the at least one extracted pixel region is removed or replaced before being transmitted (step 120).
[0138] In one variant, the video conferencing method further comprises a step 180 during which the at least one extracted pixel region is formatted into a video format or a volume format, such as, for example, based on ISOBMFF (ISO / IEC 14496-12) or one of its derivatives, or based on a real-time transport protocol (RTP) suitable for real-time communication.
[0139] This embodiment may be advantageous when a stereo camera or a depth perception camera (depth image based on time of flight, infrared, etc.) is used as the camera 102 of the first video teleconferencing device 100 .
[0140] exist Figure 1-6 In the present invention, various methods are described herein, and each method includes one or more steps or actions to achieve the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions can be modified or combined.
[0141] Some examples are described with respect to block diagrams and / or operational flow charts. Each block represents a portion of a circuit element, module, or code that includes one or more executable instructions for implementing (one or more) specified logical functions. It should also be noted that, in other embodiments, the (one or more) functions marked in the blocks may not occur in the order indicated. For example, depending on the functions involved, two blocks shown in succession may actually be executed substantially concurrently, or may sometimes be executed in the reverse order.
[0142] The various embodiments and aspects described herein may be implemented in, for example, a method or process, an apparatus, a computer program, a data stream, a bit stream, or a signal. Even if only discussed in the context of a single form of embodiment (e.g., discussed only as a method), the embodiments of the features discussed may also be implemented in other forms (e.g., an apparatus or a computer program).
[0143] The method may be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. A processor also includes a communication device.
[0144] In addition, the method can be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the implementation) can be stored on a computer-readable storage medium. The computer-readable storage medium can take the form of a computer-readable program product implemented in one or more computer-readable media and having a computer-readable program code implemented thereon that is executable by a computer. Considering the inherent ability to store information therein and the inherent ability to provide retrieval of information therefrom, the computer-readable storage medium used herein can be considered to be a non-transitory storage medium. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or apparatus, or any suitable combination of the foregoing. It should be appreciated that although more specific examples of computer-readable storage media to which the present exemplary embodiment can be applied are provided below, as will be readily appreciated by those skilled in the art, this is merely an illustrative and non-exhaustive list: a portable computer floppy disk; a hard disk; a read-only memory (ROM); an erasable programmable read-only memory (EPROM or flash memory); a portable compact disc read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination of the foregoing.
[0145] The instructions may form an application program tangibly embodied on a processor-readable medium.
[0146] For example, instructions may be in hardware, firmware, software, or a combination thereof. For example, instructions may be found in an operating system, a separate application, or a combination of the two. Thus, a processor may be characterized as both a device configured to perform a process and a device that includes a processor-readable medium (such as a storage device) having instructions for performing the process. Additionally, in addition to or in lieu of instructions, a processor-readable medium may store data values generated by an embodiment.
[0147] The computer software may be implemented by the processor 1010 or by hardware, or by a combination of hardware and software. As a non-limiting example, the exemplary embodiments may also be implemented by one or more integrated circuits. The memory 1020 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology (such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples). The processor 910 may be of any type suitable for the technical environment and may include one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture, as non-limiting examples.
[0148] As will be apparent to one of ordinary skill in the art, embodiments may generate various signals formatted to carry information that can be stored or transmitted, for example. The information may include, for example, instructions for performing a method or data generated by one of the described embodiments. For example, a signal may be formatted to carry video data of the described exemplary embodiments. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. The formatting may include, for example, encoding a video data stream and modulating a carrier with the encoded video data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.
[0149] The terms used herein are only used to describe the purpose of specific exemplary embodiments and are not intended to be limiting. As used herein, the singular forms "a", "a kind of" and "the" may also be intended to include plural forms, unless the context clearly indicates otherwise. It will be further understood that, when used in this specification, the terms "include / comprise" and / or "including / comprising" may specify the existence of stated, for example, features, integers, steps, operations, elements and / or components, but do not exclude the existence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. Moreover, when an element is referred to as "responsive to" or "connected to" another element or "associated with another element", it may directly respond to or be connected to another element, or there may be an intermediate element. In contrast, when an element is referred to as "directly responsive to" or "directly connected to" another element or "directly associated with another element", there is no intermediate element.
[0150] It should be appreciated that use of any of the symbols / terms " / ," "and / or," and "at least one of...", for example, in the context of "A / B," "A and / or B," and "at least one of A and B," may be intended to encompass selection of only the first listed option (A), or only the second listed option (B), or both options (A and B). As a further example, in the context of "A, B, and / or C" and "at least one of A, B, and C," such wording is intended to encompass selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). This may be extended to as many items as listed, as will be apparent to one of ordinary skill in this and related arts.
[0151] Various numerical values may be used in the present disclosure. Specific values may be used for illustrative purposes and the described aspects are not limited to these specific values.
[0152] It will be understood that although the terms first, second, etc. can be used to describe various elements in this article, these elements are not limited by these terms. These terms are only used to distinguish one element from another element. For example, without departing from the teachings of the present disclosure, the first element can be referred to as the second element, and similarly, the second element can be referred to as the first element. There is no suggestion of sorting between the first element and the second element.
[0153] References to "one exemplary embodiment" or "an exemplary embodiment" or "one implementation" or "an implementation" and other variations thereof are frequently used to convey that a particular feature, structure, characteristic, etc. (described in connection with the exemplary embodiment / implementation) is included in at least one exemplary embodiment / implementation. Thus, the appearances of the phrases "in one exemplary embodiment" or "in an exemplary embodiment" or "in one implementation" or "in an implementation" and any other variations appearing in various places throughout this disclosure are not necessarily all referring to the same exemplary embodiment.
[0154] Similarly, references herein to "according to an exemplary embodiment" or "in an exemplary embodiment" and other variations thereof are frequently used to convey that a particular feature, structure, or characteristic (described in conjunction with an exemplary embodiment) may be included in at least one exemplary embodiment. Thus, the phrases "according to an exemplary embodiment" or "in an exemplary embodiment" appearing in various places throughout this disclosure are not necessarily all referring to the same exemplary embodiment, nor are separate or alternative exemplary embodiments necessarily mutually exclusive of other exemplary embodiments.
[0155] Reference numerals appearing in the claims are for illustration purposes only and have no limiting effect on the scope of the claims.The present exemplary embodiments / examples and variations may be employed in any combination or subcombination although not explicitly described.
[0156] When a figure is presented as a flow chart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow chart of the corresponding method / process.
[0157] While some diagrams include arrows on communication paths to illustrate a primary direction of communication, it should be understood that communication can occur in the opposite direction to the depicted arrows.
[0158] Various embodiments relate to decoding. As used in this disclosure, "decoding" can encompass, for example, all or part of a process performed on a received video picture (which may include a received bitstream encoding one or more video pictures) to produce a final output suitable for display or suitable for further processing in a reconstructed video domain. In various exemplary embodiments, such processes include one or more of the processes typically performed by a decoder. In various exemplary embodiments, for example, such processes also include or alternatively include processes performed by a decoder of the various embodiments described in this disclosure.
[0159] As a further example, in one exemplary embodiment, "decoding" may refer only to dequantization, in one exemplary embodiment, "decoding" may refer to entropy decoding, in another exemplary embodiment, "decoding" may refer only to differential decoding, and in another exemplary embodiment, "decoding" may refer to a combination of dequantization, entropy decoding, and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations, or generally to a broader decoding process, will be clear based on the context of the particular description and is believed to be well understood by those skilled in the art.
[0160] Various embodiments relate to encoding. In a manner similar to the above discussion of "decoding," "encoding" as used in this disclosure may encompass, for example, all or part of a process performed on an input video frame to produce an output bitstream. In various exemplary embodiments, such a process includes one or more of the processes typically performed by an encoder. In various exemplary embodiments, such a process also includes or alternatively includes a process performed by an encoder of the various embodiments described in this disclosure.
[0161] As a further example, in one exemplary embodiment, "encoding" may refer only to quantization, in one exemplary embodiment, "encoding" may refer only to entropy coding, in another exemplary embodiment, "encoding" may refer only to differential coding, and in another exemplary embodiment, "encoding" may refer to a combination of quantization, differential coding, and entropy coding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or generally to a broader encoding process will be clear, and is believed to be well understood by those skilled in the art, based on the context of the particular description.
[0162] Furthermore, this disclosure may refer to "receiving" various pieces of information. Receiving information may include, for example, one or more of: accessing information or receiving information from a communication network.
[0163] Furthermore, as used herein, the term "signal" specifically refers to indicating something to a corresponding decoder. For example, in certain exemplary embodiments, an encoder signals specific information, such as encoded video data. In this way, in exemplary embodiments, the same parameters can be used on both the encoder and decoder sides. Thus, for example, the encoder can send specific parameters to the decoder (explicit signaling) so that the decoder can use the same specific parameters. Conversely, if the decoder already has the specific parameters along with other parameters, signaling can be used without sending them (implicit signaling) to simply allow the decoder to know and select the specific parameters. By avoiding the transmission of any actual functionality, bit savings are achieved in various exemplary embodiments. It should be appreciated that signaling can be accomplished in a variety of ways. For example, in various exemplary embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. Although the verb form of the term "signal" has been previously described, the term "signal" can also be used herein as a noun.
[0164] Furthermore, as used herein, the term "rendering" refers to, among other things, instructing a corresponding display to display a video. For example, rendering may include all operations that produce data for an immersive video ready for display on display 104 or on the reflective surface of mirror 105. Rendering operations may include, for example, parsing, projecting / inverting projection data, combining multiple video data to obtain a combined video, and the like.
[0165] A number of embodiments have been described. However, it should be understood that various modifications may be made. For example, elements of different embodiments may be combined, supplemented, modified, or removed to produce other embodiments. Furthermore, it will be understood by those of ordinary skill that other structures and processes may replace the disclosed structures and processes, and that the resulting embodiments will perform at least substantially the same function(s) in at least substantially the same manner(s) to achieve at least substantially the same result(s) as the disclosed embodiments. Thus, the present disclosure contemplates these and other embodiments.
Claims
1. A video teleconference method, the video teleconference method being performed between at least two remote video teleconference devices communicating via a communication network, wherein at least one video teleconference device, referred to as a first video teleconference device, comprises: a camera, referred to as a first camera, and a display, referred to as a first display, the field of view of the first camera at least partially covering the viewpoint of an end user of the first video teleconferencing device, wherein the method comprises the following steps: - capturing (100) a first video by the first camera, the first video representing the end user's view of the first video teleconferencing device on the reflective surface of a mirror, the mirror being located in front of the end user and within the field of view of the first camera; - extracting at least one pixel region from the first video, the at least one pixel region surrounding the self-view of the end user; and - sending (120) first video data to at least one remote video teleconferencing device, the first video data representing the at least one extracted pixel region.
2. The video teleconference method according to claim 1, wherein the method further comprises: - receiving (130) second video data from a remote video teleconference device; and - rendering (140) a second video on the first display based on the received second video data so as to cover at least part of the reflective surface on the first display.
3. A video teleconferencing method as described in claim 2, wherein the first video teleconferencing device is a video see-through device, and wherein rendering (140) the second video includes replacing pixels of the first display belonging to the at least one extracted pixel area with pixels of the second video.
4. A video teleconferencing method as claimed in claim 2, wherein the first video teleconferencing device is an optical see-through device, and wherein the second video is rendered (140) on an area of the first display based on the received second video data so as to overlay a projection of the reflective surface on the first display.
5. A video teleconferencing method as described in claim 4, wherein the pixels of the second video on the first display correspond to the projection of the reflective surface, the pixels of the second video have an opacity value adapted to hide the terminal user's own view on the reflective surface, and the projection of the reflective surface on the first display is adapted to the viewpoint of the terminal user.
6. The video teleconference method according to any one of claims 2 to 5, wherein the method further comprises: - receiving (150) third video data representing a third video; and - rendering (160) a third video on the first display over an area of the projection extending beyond the reflective surface based on the received third video data so as to overlay the second video.
7. The video teleconference method according to any one of claims 2 to 5, wherein the method further comprises: - receiving (150) third video data representing a third video; and - rendering a third video according to the received third video data over an area surrounding the area corresponding to the projection of the reflective surface on the first display. 8 . The video teleconference method according to claim 1 , wherein the at least one pixel region extracted is detected in the first video based on face detection.
9. The video teleconferencing method according to any one of claims 1 to 7, wherein the extracted at least one pixel region is detected in the first video based on detection of a view of the first video teleconferencing device in the first video.
10. A video teleconferencing method as described in any one of claims 1 to 7, wherein the extracted at least one pixel region is detected at least in part based on a correlation between movement detected in the first video and sensor data captured by a sensor of the first video teleconferencing device.
11. The video teleconferencing method according to any one of claims 1 to 10, wherein the method further comprises removing or replacing (170) the background of the at least one extracted pixel region before transmitting.
12. The video teleconferencing method according to any one of claims 1 to 11, wherein the method further comprises formatting (180) the extracted at least one pixel region into a video format or a volume format, or encapsulating the extracted at least one pixel region into a real-time transport protocol packet suitable for real-time communication.
13. A computer program product comprising instructions, which, when executed by one or more processors, cause the one or more processors to perform the video teleconferencing method according to any one of claims 1 to 12.
14. A video teleconference device, comprising components for executing the steps of the video teleconference method according to any one of claims 1 to 12.
15. A video teleconference system comprising at least one video teleconference device according to claim 14.