Video Teleconferencing

The described solution enhances video teleconferencing by identifying and displaying user devices with spatial audio capture capabilities in enlarged windows, improving the visualization of user movements and spatial audio rendering, thus addressing the limitations of existing systems.

JP7679432B2Active Publication Date: 2025-05-19NOKIA TECHNOLOGIES OY
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023142001
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-09-05
Filing Date
2023-09-01
Publication Date
2025-05-19
Estimated Expiration
2043-09-01

AI Technical Summary

Technical Problem

Existing video teleconferencing systems do not effectively differentiate and display the audio and video data from user devices with spatial audio capture capabilities, limiting the immersive experience and accurate audio localization for users.

Method used

An apparatus and method that receive audio and video data from multiple user devices, identify those with spatial audio capture capabilities, and display video data in enlarged windows with wider background areas, based on the tracking capabilities (3DoF or 6DoF) of the user devices, to enhance the visualization of user movements and spatial audio rendering.

Benefits of technology

This solution provides a more intuitive and immersive video conferencing experience by visually indicating the spatial audio capabilities of user devices and enhancing the rendering of spatial audio based on user movements, thereby improving user engagement and understanding of audio sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007679432000001
    Figure 0007679432000001
  • Figure 0007679432000002
    Figure 0007679432000002
  • Figure 0007679432000003
    Figure 0007679432000003
Patent Text Reader

Abstract

To disclose an apparatus and a method.SOLUTION: A method may comprise receiving audio data and video data from a plurality of user devices as part of a conference call, the video data representing a user of the respective user device and identifying one or more of the user devices as having a spatial audio capture capability. Another operation may comprise displaying, or causing display of, the video data from the user devices in different respective windows of a user interface. The respective windows for the identified one or more user devices may be displayed in an enlarged format so as to have a wider background region than for windows for the user devices without a spatial capture capability.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments relate to apparatuses, methods, and computer programs related to video teleconferencing.

Background Art

[0002] A teleconference is a communication session among a plurality of user devices, and thus a plurality of users or parties associated with each user device. At the time when a teleconference is established, one or more communication channels can be established between the user devices, possibly using a conference server.

[0003] A video teleconference can include a user device that transmits captured video and audio data to other user devices as part of the communication session. The video and audio data can be streamed, for example, via a conference server. The video data captured by a particular user device can be displayed on the display device of one or more other user devices, for example, within a window of a user interface. Each window can be associated with a different user device that transmits those video data as part of the communication session. The audio data captured by a particular user device can be output via one or more loudspeakers, earbuds, headphones, or headsets of one or more other user devices.

Summary of the Invention

[0004] The scope of protection sought for the various embodiments of the present invention is set forth by the independent claims. Embodiments and features described herein that do not fall within the scope of the independent claims are construed as useful examples for understanding the various embodiments of the present invention.

[0005] According to a first aspect, the present document describes an apparatus, the apparatus comprising means for receiving audio data and video data from a plurality of user devices as part of a conference call, the video data indicating a user of each respective user device, means for identifying one or more of the user devices as having spatial audio capture capabilities, and means for displaying or causing the display of video data from the user devices within respective different windows of a user interface, wherein each window for an identified one or more of the user devices is displayed in an enlarged format so as to have a wider background area compared to the case of a window for a user device without spatial capture capabilities.

[0006] The audio data can at least partially represent the speech of the user of each respective user device captured using the audio capture capabilities of the user device.

[0007] The means for identifying can be configured to identify the user device as having spatial audio capture capabilities, and the user device can track the position of the user over time.

[0008] The means for displaying can be configured to display one or more enlarged format windows such that the width and / or format of the background area is based on whether the tracking uses three degrees of freedom (3DoF) or six degrees of freedom (6DoF).

[0009] The width of the background area may be wider when the tracking uses six degrees of freedom (6DoF) compared to when the tracking uses three degrees of freedom (3DoF).

[0010] The enlarged format of the background area can use a wide-angle image format when the tracking uses six degrees of freedom (6DoF) compared to when the tracking uses three degrees of freedom (3DoF).

[0011] The means for displaying may be configured to display one or more enlarged format windows in response to the tracked movement amount of the user exceeding a predetermined threshold value.

[0012] The width of the background area can increase with an increase in the tracked movement amount of the user.

[0013] The enlarged format of the background area may be set to a wide-angle image format in response to the tracked movement amount of the user exceeding a predetermined threshold value.

[0014] The audio data for one or more identified user devices may be rendered to be recognized as coming from a direction based on the tracked position of the user.

[0015] The apparatus may further comprise means for receiving a selection of one or more specific user devices based on one or more selections of the enlarged format window, and the means for displaying is configured to further increase the size of the enlarged format window based on the selection.

[0016] The audio data from the selected one or more user devices may be rendered to come from positions within a wider range of positions based on the tracked position of each user as compared to the case of audio data from non-selected user devices.

[0017] The selection means may be configured to receive the selection via one or both of an input received using a user interface corresponding to a specific enlarged window and the current speaker identified using audio data from a user device associated with the specific enlarged window.

[0018] The background area can include video data showing captured video around or outside at least a part of the user as part of a conference call, or a predetermined image or video clip.

[0019] According to a second aspect, the present document describes a method, which includes receiving audio data and video data from a plurality of user devices as part of a conference call, where the video data represents the users of the respective user devices, receiving the data, identifying one or more of the user devices as having spatial audio capture capabilities, and displaying or causing the display of video data from the user devices within respective different windows of a user interface, and each window for the identified one or more user devices is displayed in an enlarged format so as to have a wider background area compared to the case of a window for a user device having no spatial capture capabilities.

[0020] The audio data can at least partially represent the speech of the users of the respective user devices captured using the audio capture capabilities of the user devices.

[0021] Identifying can identify the user device as having spatial audio capture capabilities, and the user device can track the position of the user over time.

[0022] Displaying can display one or more enlarged format windows such that the width and / or format of the background area is based on whether the tracking uses three degrees of freedom (3DoF) or six degrees of freedom (6DoF).

[0023] The width of the background area may be wider when the tracking uses six degrees of freedom (6DoF) compared to when the tracking uses three degrees of freedom (3DoF).

[0024] The enlarged format of the background area can use a wide-angle image format when the tracking uses 6 degrees of freedom (6DoF) compared to when the tracking uses 3 degrees of freedom (3DoF).

[0025] Displaying can display one or more enlarged format windows in response to the tracked movement amount of the user exceeding a predetermined threshold value.

[0026] The width of the background area can increase with an increase in the tracked movement amount of the user.

[0027] The enlarged format of the background area can be set to a wide-angle image format in response to the tracked movement amount of the user exceeding a predetermined threshold value.

[0028] The audio data for one or more identified user devices can be rendered to be recognized as coming from a direction based on the tracked position of the user.

[0029] The apparatus can further include receiving a selection of one or more specific user devices based on one or more selections of the enlarged format window, and displaying can further increase the size of the enlarged format window based on the selection.

[0030] The audio data from the selected one or more user devices can be rendered to come from positions within a wider range of positions based on the tracked position of each user compared to the case of audio data from non-selected user devices.

[0031] The selection can be received via one or both of an input received using a user interface corresponding to a specific enlarged window and the current speaker identified using audio data from a user device associated with the specific enlarged window.

[0032] The background area can include video data showing captured video around or outside at least a portion of the user, or a predetermined image or video clip, as part of a conference call.

[0033] According to a third aspect, the present specification describes a computer program, the computer program including instructions for causing at least the following to be performed by an apparatus, where at least the following is: receiving audio data and video data from a plurality of user devices as part of a conference call, where the video data shows the user of each user device, receiving, identifying one or more of the user devices as having spatial audio capture capabilities, displaying or causing the display of video data from the user devices within respective different windows of a user interface, where each window for the identified one or more user devices is displayed in an enlarged format so as to have a wider background area compared to the windows for user devices having no spatial capture capabilities.

[0034] The third aspect can also include any feature of the second aspect.

[0035] According to a fourth aspect, the present specification describes a computer-readable medium (such as a non-transitory computer-readable medium) storing program instructions for performing at least the following, where at least the following is: receiving audio data and video data from a plurality of user devices as part of a conference call, where the video data shows the user of each user device, receiving, identifying one or more of the user devices as having spatial audio capture capabilities, displaying or causing the display of video data from the user devices within respective different windows of a user interface, where each window for the identified one or more user devices is displayed in an enlarged format so as to have a wider background area compared to the windows for user devices having no spatial capture capabilities.

[0036] The fourth aspect can also include any feature of the second aspect.

[0037] According to a fifth aspect, the present specification describes an apparatus comprising at least one processor and at least one memory including computer program code, which, when executed by the at least one processor, causes the apparatus to receive audio data and video data from a plurality of user devices as part of a conference call, wherein the video data shows the user of each respective user device, receive, identify one or more of the user devices as having spatial audio capture capabilities, and display or cause to be displayed the video data from the user devices within different respective windows of a user interface, wherein each window for the identified one or more user devices is displayed in an enlarged format so as to have a wider background area compared to the case of a window for a user device having no spatial capture capabilities.

[0038] The fifth aspect can also include any feature of the second aspect.

[0039] Exemplary embodiments are now described by way of non-limiting examples with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0040]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8A

Figure 8B

Figure 9A

Figure 9B

Figure 10

Figure 11

[0041] Exemplary embodiments relate to apparatuses, methods, and computer programs related to video teleconferencing.

[0042] A video telephone conference or simply a video conference is a communication session between multiple user devices and thus between multiple users or parties associated with each user device. The communication session can include the transmission of video and audio data captured by one user device, which is part of the call or communication session, to one or more other user devices.

[0043] The video data can be output to the display device of one or more other user devices, for example, within each window of the user interface on one or more other user devices.

[0044] The audio data can be output via one or more loudspeakers of one or more other user devices or associated therewith. For example, the audio data can be output to earbuds, earphones, headphones, or a head-mounted device (HMD) that is in wired or wireless communication with one or more user devices.

[0045] A video conference can include one or more communication channels established between two or more user devices through a communication network and thus between the users or parties associated with each device. The conference session can include, for example, one or more channels established between two or more devices that are participants in the conference session. One or more communication channels can be established at the time of establishing the conference session and can typically provide a multicast data feed from a given user device to each of the other user devices in real time or near real time. One or more communication channels can be two-way communication channels.

[0046] The user device can include any device that can be operated by one or more users and is configured to transmit and receive data through a communication network.

[0047] The user device can include a processing function for running one or more applications, such as a video conferencing application. The video conferencing application can include a subordinate function of another application, such as a social media application.

[0048] The user device can also include one or more input modules and one or more output modules.

[0049] For example, the user device can include one or more input transducers and one or more output transducers.

[0050] For example, one or more input transducers can include one or more microphones for converting sound waves into electrical signals that can be stored, processed, and transmitted as audio data.

[0051] When two or more microphones are provided, the user device can be capable of generating a spatial audio signal, which includes a spatial percept that enables a listening user of the receiving-side user device to recognize where one or more sounds, such as speech from a user of the transmitting-side user device, are coming from.

[0052] For example, a user device can also be made capable of tracking the position of a user with respect to that user device, e.g., by being able to determine one or both of the orientation and the translational position over time. Tracking only the orientation of the user can be referred to as 3 degrees of freedom (3DoF) tracking. Tracking the translational movement of the user as well can be referred to as 6 degrees of freedom (6DoF) tracking. Tracking the position of the user can, as is known, use various video and / or audio techniques. Tracking the position of the user can alternatively or additionally, but not limited to, use data from one or more of an accelerometer, a global positioning system (GPS) receiver or the like, a compass, radar technology, one or more optical sensors, one or more cameras, etc., which can be held by the user. Such sensors can be used in known technical fields such as head tracking, trajectory estimation, and location estimation.

[0053] For example, one or more output transducers can comprise one or more loudspeakers that convert an electrical signal into a sound wave. When two or more loudspeakers are provided, the user device, or a related device such as earbuds, earphones, headphones, or an HMD, can be made capable of stereo and even spatial audio rendering that takes into account the 3DoF or 6DoF tracking described above.

[0054] For example, audio data received as part of a video teleconference can be spatially rendered so that a first listening user can recognize the direction of the audio received from a second other user as coming from a particular part of the audio field around that listening user. As the second user moves within the window, the recognized direction of the audio data can be similarly changed, e.g., as the user device tracks the movement of the user from the left side to the right side of the audio field.

[0055] For example, the user device can also include one or more cameras that can capture video images that can be stored, processed, and transmitted as video data.

[0056] For example, the user device can include one or more displays that can be in any form, whether or not it is a touch-sensitive display. In the case of a touch-sensitive display, the display can also be provided with a certain form of input module for receiving and invoking selection commands based on detecting touch inputs corresponding to specific user interface elements displayed by the touch-sensitive display.

[0057] The user device can also include one or more other input modules, such as one or more accelerometers or gyroscopes that can generate motion data from which the motion characteristics of the user device can be determined. The user device can also include one or more positioning receivers, such as GNSS (Global Navigation Satellite System), that can determine the geographical location of the user device.

[0058] The user device can include, but is not limited to, smartphones, digital assistants, digital music players, personal computers, laptops, tablet computers, or wearable devices such as smartwatches. The user device can be enabled to establish a communication session with one or more other user devices via a communication network, for example, a conference session.

[0059] The user device may be configured to transmit and receive data using protocols for 3G, 4G, LTE, 5G, or any future generation communication protocol. The user device can be equipped with means for short-range communication using, for example, Bluetooth, Zigbee, or WiFi. The user device can be equipped with one or more antennas for communicating with external devices.

[0060] Referring to FIG. 1, a first user device of the example is shown in the form of a smartphone 100.

[0061] The smartphone 100 can be equipped with a touch-sensitive display (hereinafter “display”) 101, a microphone 102, a loudspeaker 103, and a front camera 104 on the front side. The smartphone 100 can further be equipped with a rear camera (not shown) on the rear side of the smartphone. The front camera 104 can be used, for example, during the enablement of a video conferencing application where video data captured by the front camera can be transmitted through an established conference session.

[0062] The smartphone 100 can also be equipped with additional microphones 106A - 106D at different positions on the smartphone.

[0063] As shown, the fourth additional microphones 106A to 106D are provided on the front side of the smartphone 100. Other microphones can be additionally or alternatively provided on the rear side of the smartphone 100 and / or on one or more sides of the smartphone. By providing two or more microphones 102, 106A to 106D on the smartphone 100, the smartphone can be configured to generate a spatial audio signal including the captured sound using spatial recognition. Spatial recognition can enable a user listening to the spatial audio signal to recognize from where one or more sounds, such as speech from the user of the smartphone 100, are coming. The spatial audio signal can be generated by the smartphone 100 using conventional processing techniques; the format of the spatial audio signal can be, for example, any of a channel-based, object-based, and / or Ambisonics-based spatial audio format, although the exemplary embodiments are not limited to such examples. The smartphone 100 can comprise hardware and / or software functionality to generate a spatial audio signal that can be transmitted in data form as part of a video conference.

[0064] The smartphone 100 can also comprise further first and second loudspeakers 108A, 108B at different positions on the smartphone. One or both of the first and second loudspeakers 108A, 108B can be configured to render received audio data in a monaural, stereo, or spatial audio format that can be received as part of a video conference.

[0065] For example, metadata associated with the received audio data can enable the smartphone 100 to determine the appropriate audio format and render the audio data appropriately. Additionally or alternatively, the number of audio channels comprised by the received audio data can indicate the appropriate audio format.

[0066] For example, spatial audio data can be decoded and output so that a user of the smartphone 100 can recognize one or more sounds from other user devices on the transmission side as coming from a specific direction within an audio field that at least partially surrounds the other user devices.

[0067] Additionally or alternatively, the received audio data can be rendered and output to a related user device such as earbuds, earphones, headphones, or a head-mounted device (HMD).

[0068] Referring to FIG. 2, a video conferencing system 200 is shown.

[0069] The video conferencing system 200 can include a first user device 100, a second user device 202, a third user device 203, and a conference server 204. It can be assumed that the smartphone 100 described in relation to FIG. 1 includes the first user device 100.

[0070] For illustration purposes, the video conferencing system 200 shown in FIG. 2 includes only two remote devices, namely the second user device 202 and the third user device 203, but the video conferencing system can include any number of user devices participating in a video conferencing session. The later example described with reference to FIG. 3 includes additional user devices.

[0071] A first user 210 can use the first user device 100, a second user 211 can use the second user device 202, and a third user 212 can use the third user device 203. The user devices 100, 202, 203 can typically be in different remote locations.

[0072] The second and third user devices 202, 203 can include, for example, any of a smartphone, digital assistant, digital music player, personal computer, laptop, tablet computer, or wearable device such as a smartwatch. The second and third user devices 202, 203 can have the same or similar functions as the first user device 100 and can each include, for example, a display screen, one or more microphones, one or more loudspeakers, and one or more front-facing cameras. As with the first user device 100, it may be the case that at least one of the second and third user devices 202, 203 is capable of spatial audio capture.

[0073] Each of the first, second, and third user devices 100, 202, 203 can communicate a stream of captured video and audio data with other user devices via the conference server 204 as part of a conference session, in this example a video conference.

[0074] For example, the first user device 100 can communicate the video stream and the accompanying audio stream of the first user 210 who is speaking, for example, when the first user is facing the front camera 104. The video and audio streams can be transmitted through a first channel 220 established between the first user device 100 and the conference server 204. The video and audio streams can then be transmitted by the conference server 204 to the second and third user devices 202, 203 through the respective second and third channels 221, 222 using or in a manner of the multicast transmission protocol established between the conference server and the second and third user devices. The first, second, and third channels 220, 221, 222 are shown by a single line indicating a bidirectional channel, but there may be separate channels, a transmission channel and a reception channel. The same operating principle applies to the second and third user devices 202, 203 when communicating video and audio streams as part of a conference session.

[0075] The video and audio streams can include video packets and related audio packets. The video packets and audio packets can conform to any suitable conference standard such as the Real Time Protocol (RTP). The video packets and audio packets can include, for example, a packet header containing control information and a packet body containing video or audio data content. The packet header can include, for example, a sequence number indicating the sequential position of the packet within the stream of transmitted packets. The packet header can also include a timestamp indicating the timing of packet transmission. The packet body can include encoded video or video audio captured during the time slot before transmitting the packet. For example, the video data of the packet can include a sequence of images indicating encoded pixels and spatial coordinates.

[0076] One or more of the first, second, and third user devices 100, 202, 203 and the conference server 204 can include a device such as the device shown and described below with reference to FIG. 10. One or more of the first, second, and third user devices 100, 202, 203 and the conference server 204 can be configured by hardware, software, firmware, or a combination thereof to perform the operations described below, for example, with reference to FIG. 4.

[0077] FIG. 3 shows another video conferencing system 300 that provides video conferencing.

[0078] The video conferencing system 300 is similar to the video conferencing system shown in FIG. 2 in that it includes a first user device 100, a second user device 202, a third user device 203, and a conference server 204.

[0079] The video conferencing system 300 further includes fourth and fifth user devices 204, 205. The fourth user device 204 is associated with a fourth user 213, and the fifth user device 205 is associated with a fifth user 214.

[0080] Again, it can be assumed that the smartphone 100 described in connection with FIG. 1 includes the first user device 100.

[0081] Example audio capture capability information for each of the first to fifth user devices 100, 202, 203, 204, 205 is indicated by a dashed box.

[0082] For example, the first user device 100 has the spatial audio capture capability described above and utilizes the first to fourth additional microphones 106A to 106D. The spatial audio capture capability information is labeled "6DoF" for the reasons shown below.

[0083] The location tracking data can be accompanied by spatial audio data generated by the first user device 100. That is, the spatial audio data and the location tracking data can be transmitted to the second to fifth user devices 202, 203, 204, 205 as part of a conference session. In this regard, the first user device 100 can track the position of the first user 210 relative to the first user device during video and audio capture. The location tracking data can be based on one or both of audio analysis and video analysis. For example, the location tracking data can be determined based on angle-of-arrival audio measurements performed using the first to fourth additional microphones 106A - 106D. Additionally or alternatively, the location tracking data can be determined based on image analysis during video capture, for example using head tracking.

[0084] The location tracking data can be related to the tracked orientation of the first user 210 (which can mean the first user, for example, at least a part of the user's head), and similarly, the translational movement of the user in Euclidean space. Thus, the first user 210 can be tracked using 6DoF, and thus this is indicated within the spatial audio capture capability information.

[0085] If the location tracking data is only related to the tracked orientation of the user, the spatial audio capture capability information is labeled as 3DoF.

[0086] From FIG. 3, it can be seen that the second user device 202 has 3DoF spatial audio capture capability, the third user device 203 has monaural ("mono") audio capture capability, the fourth user device 204 has 6DoF spatial audio capture capability, and the fifth user device 205 has monaural audio capture capability.

[0087] Exemplary embodiments can enable one or more users, e.g., one or more of the first to fifth users 210 - 214, to understand the audio capture capabilities of other user devices that are parties to a video teleconference.

[0088] FIG. 4 is a flow diagram showing processing operations that may be performed by, for example, a first user device 100, according to one or more exemplary embodiments. The processing operations may similarly or alternatively be performed by one or more of the second to fifth user devices 202, 203, 204, 205 and / or further by a conference server 204. The processing operations may be performed by hardware, software, firmware, or a combination thereof.

[0089] A first operation 401 can include receiving audio data and video data from a plurality of user devices as part of a teleconference, the video data showing a user or at least a portion of the users of each respective user device.

[0090] At least a portion of the audio data from a given user device can indicate the speech of the associated user of the given user device. There may be other background sounds that are also indicated in the audio data.

[0091] A second operation 402 can include identifying one or more of the user devices as having spatial audio capture capabilities.

[0092] For example, the identifying can be based on metadata or similar data that may accompany the audio data and video data.

[0093] The third operation 403 can include displaying or causing the display of video data from the user device within each of the different windows of the user interface, and each window for the identified one or more user devices is displayed in an enlarged format so as to have a wider background area compared to the window for a user device having no spatial capture ability.

[0094] The window in this regard can include the display area of the user interface and is not necessarily limited to a display area having a square or rectangular shape.

[0095] The background area can refer to one or more portions of an image surrounding the foreground object.

[0096] The foreground object can refer to an object indicated, for example, by a blob of pixels closest to the camera, and that object is at least initially positioned within a central region of the image being captured and corresponding to a particular object, such as the user. The foreground object can be the user or at least a part of the user facing the camera of the user device from which the video data is received. The foreground object can include, for example, the face of the user.

[0097] The background area can include live video data showing the captured video around or outside the foreground object as part of a video conference, or the background area can include a predetermined background image, animation, or video clip. Conventional image segmentation methods can distinguish the foreground area and the background area based on, for example, position and / or movement. The foreground area can be cropped and / or scaled, whether as the captured video or as a predetermined image, animation, or video clip, and applied on top of the background area. The foreground area can remain substantially the same size regardless of the expansion of the background area, but there may be reasons for the foreground area to change size, for example, when the user approaches or moves away from the camera of the user device the user is using.

[0098] In this way, a user, for example, the first user 210, can know the audio capture capabilities of one or more other user devices that are parties to the video conference based at least on the expanded format of the window.

[0099] For example, a standard size window can indicate that the associated user device has monaural audio capture capabilities, while an expanded format window can indicate that the associated user device has spatial audio capture capabilities. The standard size window can be square or, for example, have a 4:3 aspect ratio. The expanded format window can have, for example, a 16:9 or 16:10 aspect ratio or the like.

[0100] This can prompt the receiving user or user device to enable the spatial audio rendering function, for example, to experience spatial audio rendering for at least a part of the audio data.

[0101] FIG. 5 is a front view of the first user device 100 according to an exemplary embodiment. The display 101 of the first user device 100 shows a user interface 500 related to a video telephone conference established using the video conferencing system 300 of FIG. 3 according to the exemplary embodiment.

[0102] The user interface 500 displays first to fourth windows 502, 503, 504, 505. The first and second windows 502, 503 display video data received from the third and fifth user devices 203, 205, respectively. The third and fifth users 212, 214 are shown as foreground objects. The third window 504 displays video data received from the second user device 202. The second user 211 is shown as a foreground object. The fourth window 505 displays video data received from the fourth user device 204. The fourth user 213 is shown as a foreground object.

[0103] The sizes of the first and second windows 502, 503, particularly the widths, may be considered standard or default sizes based on the fact that the audio capture capabilities of the third and fifth user devices 203, 205 are limited to monaural. The size of the third window 504 is in an enlarged format, whereby the width of the background area 520 is wider than the widths of the first and second windows 502, 503. The wide display of the third window 504 is based on metadata from the second user device 202 indicating spatial audio capture capabilities. The size of the fourth window 505 is also in an enlarged format, whereby the width of the background area 522 is wider than the widths of the first and second windows 502, 503. The wide display of the fourth window 505 is based on metadata from the fourth user device 204 indicating spatial audio capture capabilities.

[0104] In this case, it can be seen that the fourth window 505 is wider with respect to the background area than the background area of the third window 504. This indicates that the fourth user device 204 is capable of 6DoF tracking, while the second user device 202 is only capable of 3DoF tracking.

[0105] In addition to showing the first user 210 the audio capture capabilities of the second to fifth user devices 202, 203, 204, 205, the enlarged third and fourth windows 504, 505 allow for a greater degree of movement than can be shown over time by the second and fourth users 211, 213, and the greater degree of movement can track the direction of the spatial audio data from the above users.

[0106] When the first user device 100 is capable of spatial audio rendering, the direction in which the audio from the second user 211 is recognized by the first user 210 can be centered on the user interface 500 based on the position of the first user 210 relative to the second user device 202. There may be a certain degree of spatial change as the orientation of the second user 211 changes.

[0107] Similarly, the direction in which the audio from the fourth user 213 is recognized by the first user 210 can be associated with a large translational change amount that tracks the user's movement. For example, the fourth user 213 is shown on the right side of the fourth window 505, and thus the audio can be recognized by the first user 210 from this direction.

[0108] It can thus be seen that the enlarged formats of the third and fourth windows 504, 505 enable a more realistic and intuitive tracking of the positions of the respective second and fourth users, which is further reflected by the directional characteristics of the respective audio renderings.

[0109] FIG. 6 is a plan view of the first user device 100 according to another exemplary embodiment. FIG. 6 is similar to FIG. 5. In this case, a different user interface 600 is shown.

[0110] The user interface 600 includes the same first and second windows 502, 503. However, in this case, the third window 604 has the same or a similar width as the fourth window 605, and the difference in the audio capture capabilities of the second and fourth user devices 202, 204 is reflected by the format of the respective background areas of those user devices.

[0111] In particular, the 3DoF tracking ability of the second user device 202 is indicated by the background area 620 of the third window 604 using a standard image format. The 6DoF tracking ability of the fourth user device 205 is indicated by the background area 622 of the fourth window 605 using a wide-angle image format, for example, in a manner of a wide-angle lens having a longer focal length. For example, if the background area is a predetermined background image, the focal length information in the image metadata can be used to select the wide-angle version.

[0112] If two or more of the second to fifth user devices 202, 203, 204, 205 are capable of 6DoF tracking, discrimination can be made based on the amount of tracked movement of the respective users 211, 212, 213, 214.

[0113] For example, if the second user device 202 is capable of 6DoF tracking instead of only 3DoF tracking, the user interface examples shown in FIGS. 5 and 6 may still be applicable. For example, this may occur when the translational movement of the fourth user 213 is greater than the translational movement of the second user 211 over at least a predetermined period.

[0114] According to some exemplary embodiments, the display of the enlarged format windows, e.g., the third and fourth windows 504, 505, 604, 605 in the examples of FIGS. 5 and 6, may occur in response to the tracked movement amount exceeding a predetermined threshold value.

[0115] That is, first, all windows may be displayed in a default relatively small size. For example, all windows have the size of the first and second windows 502, 503. Thereafter, in response to one or more user devices having spatial audio capture capabilities (3DoF or 6DoF) providing position data indicating significant movement, one or more appropriate windows can be enlarged in the manner described in any of the examples shown above.

[0116] In some exemplary embodiments, an increase in the tracked movement amount, particularly the translational movement amount, can result in an increase in the appropriate window and thus a widening of the background region. A subsequent decrease in the tracked movement amount can result in a narrowing of the appropriate window and thus the background region.

[0117] In some exemplary embodiments, if a particular user, e.g., the fourth user 213, is determined to move outside the current boundary of the fourth window 505, 605, the background region can scroll to "catch-up" with the tracked movement.

[0118] In some exemplary embodiments, when it is determined that a particular user, e.g., a fourth user 213, approaches (advances) or moves away from (recedes from) the camera of the fourth user device 204, the background regions within the fourth windows 505, 605 can be appropriately scaled, e.g., enlarged or shrunk. In some exemplary embodiments, the spatial audio rendered for a user such as the fourth user 213 who is advancing or receding can be modified in coordination with such tracked motion, e.g., to adapt to the reverberation as the user approaches or moves away from the camera of the fourth user device 204. When in use, there may be set limits for distance-based rendering in this case due to the commonly used distance gain attenuation.

[0119] In some exemplary embodiments, the selection can be received with respect to one or more of the identified user devices, e.g., the second or fourth user devices 202, 204, using an enlarged window format.

[0120] The selection can be received by an input corresponding to one or more of the windows, received using the user interface 101. In response to the selection, the user interface 101 can be modified such that one or more corresponding enlarged format windows are further increased in size, which can be either a horizontal increase in size and / or a vertical increase in size. This provides a larger space for visual tracking of associated foreground objects, e.g., associated users, which can be appropriate when the user device has 6DoF tracking.

[0121] For ease of explanation, it will be hereafter assumed that the second user device 202 is also capable of 6DoF tracking and that the second user 211 moves with some translational motion.

[0122] Referring to FIG. 7 (FIG. 7A), the selection of the second user device 202 is performed by selecting a third window 604 on the user interface 600. In FIG. 7 (FIG. 7A), reference numeral 702 indicates a touch input performed by the first user 210. The selection can be received by cursor or mouse-based input, gesture input, or voice command. Alternatively or additionally, the selection can be determined by comparison with audio data from various other user devices, such that it can be based on which user is currently speaking or at least speaking with the greatest amplitude.

[0123] As shown in FIG. 7 (FIG. 7B), the result is an increase in the vertical extent of the third window 604, with a larger portion of its background area 620 being visible. In this case, no resizing of the first, second, and fourth windows 502, 503, 605 is necessary. However, if there is insufficient space available on the user interface 600, one or more of the first, second, and fourth windows 502, 503, 605 can be resized or repositioned to allow for further increase in size. Scrolling of the background area 620 can similarly be enabled if the fourth user approaches or crosses the boundary of the third window 604.

[0124] FIG. 8A is a front view of a first user device 100 according to another exemplary embodiment.

[0125] In the third window 604, it will be seen that the second user 211 has moved to the left by a translational movement. Otherwise, the respective positions of the third, fourth, and fifth users 212, 213, 214 have not changed.

[0126] FIG. 8B shows a top plan view of the spatial audio field 802 experienced by the first user 210 based on the position of FIG. 8A.

[0127] The audio data indicating the third and fifth users 212, 214 is in monaural format and thus has no directivity. The rendered audio can be recognized centrally by the first user 210.

[0128] The audio data indicating the second user 211 is in this case spatial audio data using 6DoF tracking, and thus is recognized from the direction in front of and to the left of the first user 210. Arrow 804 indicates the range of positions within the spatial audio field 802 where the rendered audio can be recognized by the translational movement of the second user 211.

[0129] The audio data indicating the fourth user 213 is similarly spatial audio data using 6DoF tracking, and thus is recognized from the direction in front of and to the right of the first user 210. Arrow 806 indicates the range of positions within the audio field 802 where the rendered audio can be recognized by the translational movement of the fourth user 213.

[0130] FIG. 9A is a front view of the first user device 100 following the reception of the selection of one of the identified user devices, in this case the second user device 202, via the third window 604.

[0131] This selection of the second user device 202 can result in a wider range of positions within the audio field where the rendered audio can be recognized by the translational movement of the second user 211. Arrow 904 indicates the range of positions within the audio field 802 where the rendered audio can be recognized. Audio data from non-selected user devices such as the fourth user device 204 can be rendered in a limited audio format such as mono or stereo, in a spatial format using only 3DoF tracking, or in a spatial format using 6DoF tracking, over a narrower range of positions. FIG. 9B shows, for example, that the fourth user 213 is heard here from behind the first user 210, and arrow 906 indicates a narrower range of positions within the audio field 802 where the audio data can be recognized.

[0132] The exemplary embodiments thus provide an intuitive and more meaningful way for users involved in a video conference to understand the audio capture capabilities of other user devices. This can result in the enabling of spatial audio rendering capabilities on their own devices. Further, widening the background area for users with spatial audio capture capabilities allows visualization of possible or actual user movements, which helps for a better user experience. Because the spatial nature of the audio at least partially tracks such visible movements.

[0133] Example apparatus FIG. 10 shows an apparatus according to some exemplary embodiments. The apparatus may be configured to perform operations described herein, such as operations described with reference to any of the disclosed processes. The apparatus includes at least one processor 1000 and at least one memory 1001 that is directly or closely connected to the processor. The memory 1001 includes at least one random access memory (RAM) 1001a and at least one read-only memory (ROM) 1001b. Computer program code (software) 1005 is stored in the ROM 1001b. The apparatus may be connected to a transmitter (TX) and a receiver (RX). The apparatus may optionally be connected to a user interface (UI) for instructing the apparatus and / or for outputting data. At least one processor 1000 having at least one memory 1001 and computer program code 1005 is arranged to cause the apparatus to perform at least the method according to at least any of the preceding processes disclosed, for example, with respect to the flowchart of FIG. 4 and its related features.

[0134] FIG. 11 shows a non-transitory medium 1100 according to some embodiments. The non-transitory medium 1100 is a computer-readable storage medium. The non-transitory medium 1100 can be, for example, a CD, DVD, USB stick, Blu-ray disk, etc. The non-transitory medium 1100 stores computer program instructions and causes an apparatus to perform the method of any of the preceding processes disclosed, for example, with respect to the flowchart of FIG. 4 and its related features.

[0135] The names of network elements, protocols, and methods are based on current standards. In other versions or other technologies, these network elements and / or protocols and / or method names can be different as long as they provide the corresponding functions. For example, embodiments can be deployed in 2G / 3G / 4G / 5G networks and further generations of 3GPP, but can also be deployed in non-3GPP wireless networks such as WiFi.

[0136] The memory can be volatile or non - volatile. The memory can be, for example, RAM, SRAM, flash memory, FPGA block RAM, DCD, CD, USB stick, and Blu - ray disk.

[0137] Unless otherwise stated or otherwise apparent from the context, a statement that two entities are different means that the two entities perform different functions. It does not necessarily mean that the two entities are based on different hardware. That is, each of the entities described in this description can be based on different hardware, or some or all of the entity can be based on the same hardware. It does not necessarily mean that the two entities are based on different software. That is, each of the entities described in this description can be based on different software, or some or all of the entity can be based on the same software. Each of the entities described in this description can be implemented in the cloud.

[0138] Embodiments of any of the blocks, devices, systems, techniques, or methods described above include, by way of non - limiting example, embodiments as hardware, software, firmware, dedicated circuits or logic, general - purpose hardware or controllers, or other computing devices, or any combination thereof. Some embodiments can be implemented in the cloud.

[0139] It is understood that what is described above is currently considered to be the preferred embodiments. However, it should be noted that the description of the preferred embodiments is given merely as an example, and various modifications can be made without departing from the scope defined by the appended claims.

Description of Reference Numerals

[0140] 100 Smartphone (First User Device) 101 Touch-Sensitive Display 102 Microphone 103 Loudspeaker 104 Front-Facing Camera 106A, 106B, 106C, 106D Microphones 108A, 108B Loudspeakers 200 Video Conference System 202 Second User Device 203 Third User Device 204 Conference Server or Fourth User Device 205 Fifth User Device 210 First User 211 Second User 212 Third User 213 Fourth User 214 Fifth User 220 First Channel 221 Second Channel 222 Third Channel 500, 600 User Interface 502 First Window 503 Second Window 504, 604 Third Window 505, 605 Fourth Window 520, 522, 620, 622 Background Area 702 Touch Input 802 Spatial Audio Field 804, 806 Range of Positions within the Audio Field 802 where Rendered Audio can be Recognized by the User's Translational Movement 904, 906 Range of Positions within the Audio Field 802 where Rendered Audio can be Recognized 1000 Processor 1001 Memory 1001a RAM 1001b ROM 1005 Software 1100 Non-volatile medium

Claims

1. means for receiving audio and video data from a plurality of user devices as part of a conference call, the video data being indicative of a user of each of the user devices; means for identifying one or more of said user devices as having spatial audio capture capability; means for displaying or causing a display of the video data from the user device in different respective windows of a user interface; the respective windows for the identified one or more user devices are displayed in an enlarged format having a wider background area than would a window for the user device that does not have spatial capture capabilities; the means for identifying is configured to identify a user device as having spatial audio capture capability, the user device being capable of tracking a location of the user over time; the displaying means is configured to display the one or more enlarged format windows such that a width and / or format of the background region is based on whether the tracking uses three degrees of freedom (3DoF) or six degrees of freedom (6DoF); An apparatus, wherein the width of the background region is wider when the tracking uses six degrees of freedom (6DoF) compared to when the tracking uses three degrees of freedom (3DoF).

2. The apparatus of claim 1 , wherein the audio data at least partially represents speech of the user of the respective user device captured using an audio capture capability of the user device.

3. 2. The apparatus of claim 1, wherein the enlarged format of the background region uses a wide-angle image format when the tracking uses six degrees of freedom (6DoF) compared to when the tracking uses three degrees of freedom (3DoF).

4. The apparatus of claim 1 , wherein the means for displaying is configured to display the one or more enlarged format windows in response to the tracked amount of movement of the user exceeding a predetermined threshold.

5. The apparatus of claim 3 , wherein the width of the background region increases with increasing tracked movement of the user.

6. The apparatus of claim 1 , wherein the enlarged format of the background region is set to a wide-angle image format in response to an amount of tracked movement of the user exceeding a predetermined threshold.

7. The apparatus of claim 1 , wherein the audio data for the identified one or more user devices is rendered to be perceived as coming from a direction based on a tracked location of the user.

8. 8. The apparatus of claim 7, further comprising: means for receiving a selection of one or more particular user devices based on one or more selections of the enlarged format window, the means for displaying configured to further increase a size of the enlarged format window based on the selection.

9. 10. The apparatus of claim 8, wherein the audio data from a selected one or more user devices is rendered to come from locations within a wider range of locations based on the tracked locations of respective users of the one or more particular user devices than is the case for the audio data from non-selected user devices.

10. The means for receiving a selection of the one or more particular user devices comprises: an input received using the user interface corresponding to a particular enlargement window of the enlarged format window; and a current speaker identified using the audio data from the user device associated with the particular magnification window; The apparatus of claim 8 , configured to receive the selection via one or both of:

11. The background region is video data indicative of a captured video around or outside at least a portion of the user as part of the conference call; or Prescribed image or video clip The apparatus according to any one of claims 1 to 10, comprising:

12. receiving audio and video data from a plurality of user devices as part of a conference call, the video data indicative of a user of each of the user devices; identifying one or more of the user devices as having spatial audio capture capability; displaying or causing a display of the video data from the user device in different respective windows of a user interface; the respective windows for the identified one or more user devices are displayed in an enlarged format having a wider background area than would a window for the user device that does not have spatial capture capabilities; the identifying is configured to identify a user device as having spatial audio capture capability, the user device being capable of tracking a location of the user over time; the displaying is configured to display the one or more enlarged format windows such that a width and / or format of the background region is based on whether the tracking uses three degrees of freedom (3DoF) or six degrees of freedom (6DoF); A method, wherein the width of the background region is wider when the tracking uses six degrees of freedom (6DoF) compared to when the tracking uses three degrees of freedom (3DoF).

Citation Information

Patent Citations

  • Video conference system and terminal equipment used for the same

    JP1997284404A

  • Video conference terminal equipment

    JP2017028608A

  • Viewing the presenter during a video conference

    JP2017512427A

  • Data processing apparatus, data processing method and program

    JP2019220848A

  • Information processing device and program

    JP2022109048A