Proximity-based audio meeting

By using a processor in an audio conferencing facilitation device to predict user proximity and location, and dynamically adjusting audio reproduction strategies, the problem of echo and lip-syncing in multi-user augmented reality applications is solved, achieving echo-free synchronous audio communication.

CN120814218APending Publication Date: 2025-10-17KONINK KPN NV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480019674.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-03
Filing Date
2024-01-09
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In multi-user augmented reality applications, when users are physically close to but not close to the meeting speaker system, the audio meeting system is prone to echo and lip-sync problems, which are difficult to solve effectively with existing technologies.

Method used

The audio conferencing system facilitates the use of a processor to predict the proximity and location of users, dynamically adjusting audio reproduction strategies to prevent users from directly hearing each other's audio or muting each other's audio, thus avoiding echoes and ensuring lip-sync.

Benefits of technology

It enables echo-free and synchronized audio communication even when users are physically close, enhancing the user experience and making it suitable for various user devices and environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120814218A_ABST
    Figure CN120814218A_ABST
Patent Text Reader

Abstract

A method of facilitating an audio meeting includes predicting or determining whether a first user (76) of a first user device (71) is able to directly hear a second user (77) of a second user device (72), and if the first user is predicted or determined not to directly hear the second user, determining that the first user is not able to directly hear the second user. If it is predicted or determined that the first user can directly hear the second user, the first user device is enabled to reproduce the audio captured by the second user device when the first user device reproduces the audio captured by the third user device (73), or if it is predicted or determined that the first user can directly hear the second user. If so, the first user device is prevented from reproducing the audio captured by the second user device while the first user device reproduces the audio captured by the third user device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an audio conference facilitation device and an audio conference system comprising such an audio conference facilitation device.

[0002] The present invention also relates to a method of facilitating audio conferencing.

[0003] The invention also relates to a computer program product enabling a computer system to perform such a method. Background Art

[0004] Conventional multi-user audio conferencing (optionally with video) remains very popular in business environments. However, multi-user augmented reality (AR) applications typically do not offer audio conferencing. Most such AR applications focus on multiple users in close physical proximity, viewing the same virtual object. These users are thus able to communicate directly (in the physical environment) without the need for any audio conferencing facilitation device.

[0005] For example, a group of museum visitors might use such a multi-user augmented reality application. Such groups (groups of at least two people) are usually together, but may also be spread out, with people going their own ways. It is difficult for these people to stay in touch with each other in such situations, and it is common to say "let's meet there in an hour" or call or text each other to find each other again. Similar group situations may occur at work, outdoor events, school, etc. As many users (especially the younger generation) often wear earbuds all day, such use cases may become more and more common. An audio conferencing system may be helpful here, allowing users to communicate with each other even when they are physically separated. While such a system is great when they are physically separated, it is unnecessary at best when they are physically together.

[0006] Ideally, users can talk to each other naturally without having to consider the current location of other users in the group, and without having to manually mute / unmute, share microphones, call each other while putting their smartphones in speakerphone mode to share the call with others, start and end calls, use push-to-talk, etc. While such an approach can be quite effective in enabling communication between group users, it places the burden on group members to determine how to communicate effectively given the current location of each group member. Such group communication should also function even in the worst-case scenario, where some users in the group are physically within talking distance, while others are in unknown locations.

[0007] When an audio conference operates with multiple locally physically present users, the problem that arises in this case is that these users can hear each other twice: once directly through the air, and once through the audio conference system. This gives a very uncomfortable echo, because the audio of the audio conference system will likely be delayed by at least 100-150 milliseconds, while the direct audio only suffers a negligible delay through the air. Even when a noise cancelling headset is used to cancel the direct audio and this is done perfectly, this will still not lead to a good user experience, because when people see each other directly, the audio will be delayed through the system, and then lip synchronization will not be achieved.

[0008] This problem also exists in a business environment when using a conference speaker system that includes one or more loudspeakers and one or more microphones. Certain conference systems, including applications such as Zoom Rooms and Microsoft Teams Rooms, provide the possibility to use proximity detection to mute the microphones and loudspeakers of the participants when they are in or enter a room with a room conference speaker system in the same meeting. This prevents echo / crosstalk. However, this solution is not applicable when multiple users are close to each other but not close to such a conference speaker system (e.g. if the conference speaker system is not used). SUMMARY

[0009] It is a first object of the present invention to provide a system that enables to facilitate an audio conference without annoying echoes and with good lip synchronization, even when multiple users are close to each other but not close to a conference speaker system.

[0010] It is a second object of the present invention to provide a method that can be used to facilitate an audio conference without annoying echoes and with good lip synchronization, even when multiple users are close to each other but not close to a conference speaker system.

[0011] In a first aspect of the invention, an audio conference facilitation apparatus comprises at least one processor that can be configured to predict or determine whether a first user of a first user apparatus can directly hear a second user of a second user apparatus, and if it is predicted or determined that the first user cannot directly hear the second user, can enable the first user apparatus to reproduce audio captured by a second user apparatus when the first user apparatus reproduces audio captured by a third user apparatus, or if it is predicted or determined that the first user can directly hear the second user, can prevent the first user apparatus from reproducing audio captured by the second user apparatus when the first user apparatus reproduces audio captured by the third user apparatus.

[0012] The first and second user devices form an audio conferencing system. The audio conferencing system can further comprise a central audio system, e.g. a conferencing server. The audio conferencing facilitating device can be e.g. such a user device or such a central audio system. The audio conferencing facilitating device need not be an audio conferencing device. For example, the audio conferencing facilitating device can only perform group management and control audio conferencing devices, e.g. as part of an AR service, without itself receiving any audio packets. The audio conferencing can comprise any form of real-time audio communication, e.g. including regular audio conferencing, multi-party telephony, real-time push-to-talk services, group audio communication as part of a multi-player game. The first and second user devices are preferably single-user devices.

[0013] If it is predicted that the first user of the first user device is able to hear the second user of the second user device, then by selectively reproducing on the first user device the audio captured by the second user device, i.e. by not reproducing on the first user device the audio captured by the second user device (but still reproducing on the first user device the audio captured by one or more other user devices), the first user is prevented from hearing the second user twice, i.e. once directly through the air and once through the audio conferencing system. In this case, the second user device preferably also does not reproduce the audio captured by the first user device. The use of such an audio conferencing facilitating device brings an audio conferencing system that provides a good user experience (i.e. no annoying echo and good lip-sync, even when multiple users are close to each other but not close to a single conferencing loudspeaker system).

[0014] For example, a room conferencing loudspeaker system is not practical in a museum environment, e.g. because there needs to be a sufficient number of room conferencing loudspeaker systems to cater for each group of visitors, the use of such room conferencing loudspeaker systems can disturb other visitors, and such room conferencing loudspeaker systems are typically placed and installed in fixed locations. The best results for such a museum environment can be obtained if the first and second users wear headphones / earphones that allow ambient audio to pass through. The audio conferencing facilitating device and the audio conferencing system can also support video conferencing and / or augmented reality.

[0015] A meeting facilitation device can be any kind of device that facilitates a meeting. These include user devices such as devices for capturing user audio or video, devices for rendering user audio or video, user devices for session establishment and control of audio and video management (i.e. capture, processing, transmission, rendering of audio and video), user devices for group management or floor control, etc. These are typically, but not limited to, smartphones, laptops, computers, AR and VR headsets, headsets, Bluetooth or other types of wireless or wired meeting headsets or speakers, room meeting systems, game consoles, handheld gaming devices. These also include central components or devices that play a certain role in the meeting, including but not limited to meeting servers, stream forwarding units, multipoint control units, network-based stream processors, rendezvous servers, echo cancellers, session border controllers, signaling servers, registration servers, etc. As is known, many of such servers and central components are implemented as software and run on general-purpose hardware, typically on a cloud platform or other distributed computing platform. Such software running on more general-purpose hardware is also considered a “meeting facilitation device” as it effectively implements the functionality or software that plays a role in any part of the meeting.

[0016] For example, the audio meeting facilitation device can be the first user device. For example, the at least one processor can be configured to prevent the first user device from rendering audio captured by the second user device by adapting its processing of the received audio packets such that audio captured by the second user device is not rendered by the first user device. For example, if it is predicted that the first user of the first user device is able to hear the second user of the second user device directly, the at least one processor of the first user device can simply skip the audio rendering step of rendering the audio captured by the second user device, or set the volume of that particular audio to zero, or replace the incoming audio packets received from the second user device with silent packets.

[0017] Alternatively, for example, the audio meeting facilitation device can be the second user device or a central audio system. For example, the at least one processor can be configured to prevent the first user device from rendering audio captured by the second user device by sending an instruction to the first user device over the network instructing the first user device not to render audio captured by the second user device. In this case, the audio stream with audio captured by the second user device will still be transmitted by the second user device and still be receivable by the first user device peer or forwarded by the central audio system. This is beneficial if the first user device would not work properly without receiving any audio stream belonging to the second user (device). This instruction can be provided in a signal or in metadata associated with the transmitted audio stream.

[0018] Alternatively, the at least one processor can be configured to prevent the first user device from reproducing the audio captured by the second user device by not transmitting the audio captured by the second user device to the first user device. In this case, the central audio device can be a stream forwarding unit. Either the stream of audio captured by the second user device is not transmitted, or the audio captured by the second user device is replaced by silence in the transmitted stream of audio. The former has the benefit of reducing the consumed bandwidth to the maximum. The latter has the benefit that the first user device can be a legacy user device.

[0019] The audio conferencing facilitation device can be a central audio system, and the at least one processor can be configured to mix the audio captured by the plurality of devices in a single stream tailored for the first user device, and to transmit the single stream to the first user device. In other words, the central audio device can be a multipoint control unit. The use of such a multipoint control unit reduces the bandwidth required on the link to the first user device. The at least one processor can be configured to prevent the first user device from reproducing the audio captured by the second user device by omitting the audio captured by the second user device from the single stream transmitted to the first user device.

[0020] If the audio conferencing facilitation device is the first user device, the at least one processor can be configured to prevent the first user device from reproducing the audio captured by the second user device by instructing the central audio system or the second user device to prevent the first user device from reproducing the audio captured by the second user device. Thus, it is the first user device that predicts whether the first user can directly hear the second user, e.g. based on proximity and / or location data, but the behavior of the first user device does not directly depend on this prediction. The central audio system can be a stream forwarding unit or a multipoint control unit as described above.

[0021] If the audio conferencing facilitation device is the second user device, the at least one processor can be configured to prevent the first user device from reproducing the audio captured by the second user device by instructing the central audio system to prevent the first user device from reproducing the audio captured by the second user device.

[0022] The at least one processor can be configured to receive audio captured by a third user device, determine whether the audio captured by the third user device comprises first audio information that is as much as the second audio information comprised in the audio captured by the second user device originated from the same source, and if the first audio information and the second audio information are determined to originate from the same source, remove the second audio information from the audio captured by the second user device or remove the first audio information from the audio captured by the third user device.

[0023] This solves a problem that can occur if the same (sound producing) source (e.g. a user) is simultaneously captured by multiple microphones, which is also referred to as the double capture problem in this specification. In this case, the remote user can hear each of the multiple nearby users twice. This can cause problems if the delay through the apparatus of one user is different from the delay through the apparatus of another user. If the delays are (approximately) the same, the double captured audio overlaps during playback and no strange effects occur. Removing the audio information can comprise audio processing such as cancelling or subtracting the audio information from the captured audio. This cancelling / subtracting typically uses techniques comparable to (regular) echo cancellation.

[0024] If the audio conferencing facilitating apparatus is a central audio system, determining whether the first audio information and the second audio information originate from the same source can also be used in the audio conferencing facilitating apparatus to predict whether the first user of the first user apparatus can directly hear the second user of the second user apparatus, in which case the user apparatus do not have to be modified.

[0025] Alternatively, if it is determined or expected that the audio captured by the first user apparatus and the audio captured by the second user apparatus comprises audio information from the same source, the at least one processor can be configured to remove the activation of capturing and / or transmitting audio by the first user apparatus or capturing and / or transmitting audio by the second user apparatus. This is another solution to the problem that can occur if the same (sound producing) source (e.g. a user) is simultaneously captured by multiple microphones, i.e. another solution to the double capture problem. By removing the activation of transmitting only and not removing the activation of capturing, the captured audio can still be analysed. For example, if it is no longer determined or expected that the audio captured by the first user apparatus and the audio captured by the second user apparatus comprises audio information from the same source, the audio transmission by the first user apparatus or the second user apparatus can be activated again.

[0026] For example, if it is predicted that the first user of the first user apparatus can directly hear the second user of the second user apparatus, it can be expected that the audio captured by the first user apparatus and the audio captured by the second user apparatus comprises audio information from the same source. As an extension to this, the audio captured in this way can be normalised, i.e. by for example increasing the speech volume to the highest detected speaking volume, ensuring that the volume of two or all speaking users is approximately the same. This can be used to ensure that the capturing and / or transmitting of audio by a user of an apparatus that is de-activated can hear as well as a user of a (nearby) apparatus whose capturing and / or transmitting of audio is not de-activated.

[0027] The process of analyzing the audio captured by the first user device and the audio captured by the second user device to determine or predict whether they comprise audio information from the same source (i.e. audio pattern detection) can be continuously performed, or can be started upon detecting a certain event. For example, this event can be detecting based on speech pattern recognition that a certain user is recognized to be speaking in both audio streams, or predicting that the first user is able to directly hear the second user.

[0028] This solution to the double capture problem can also be used in an audio conferencing facilitation device where the reproduction of audio captured by the second user device by the first user device is not prevented if it is predicted that the first user is able to directly hear the second user, and even in an audio conferencing facilitation device where it is not predicted whether the first user of the first user device is able to directly hear the second user of the second user device, but whether the audio captured by the first user device and the audio captured by the second user device comprise audio information from the same source is determined or predicted in a different way.

[0029] The at least one processor can be configured to obtain proximity and / or position data indicative of a proximity of the first user device to at least the second user device and / or indicative of a position of at least the first user device and a position of the second user device, and to predict whether the first user of the first user device is able to directly hear the second user of the second user device based on the proximity and / or position data.

[0030] As a first example, the at least one processor can be configured to obtain the proximity and / or position data by receiving at least part of the proximity and / or position data from one or more further devices. As a second example, the audio conferencing facilitation device can further comprise at least one sensor, and the at least one processor can be configured to obtain sensor information via the at least one sensor, and to obtain the proximity and / or position data by determining at least part of the proximity and / or position data based on the sensor information.

[0031] The proximity can be determined in different ways. The proximity can be determined using wireless probe signals, e.g. the second user device transmits an RF (e.g. Bluetooth or Wi-Fi), (ultra)sound or infrared signal, and the first user device receives this signal in the same time interval, i.e. listens to the signal and determines whether it is received and whether it is received in proximity with sufficient strength.

[0032] Predicting whether the first user can directly hear the second user can involve applying an audio volume threshold, which can depend on whether the users are wearing headphones / earbuds and, if so, whether they are allowing ambient audio to pass through, whether hear through functionality is enabled or whether ambient noise suppression functionality is activated, etc. The threshold can additionally or alternatively depend on the environment, e.g. lower for environments such as a conference room and higher for environments where there is more background noise, other people talking, music playing, etc.

[0033] In a second aspect of the application, an audio conferencing system comprises an audio conferencing facilitating device. If the audio conferencing facilitating device is a central audio system, the audio conferencing system can further comprise a first user device, a second user device and a third user device. If the audio conferencing facilitating device is one of the first user device and the second user device, the audio conferencing system can further comprise the other of the first user device and the second user device, and further comprise a third user device.

[0034] The at least one processor of the audio conferencing device or at least one further processor of the audio conferencing system can be configured to predict or determine whether a first user of the first user device can directly hear a third user of the third user device, and if it is predicted or determined that the first user cannot directly hear the third user, the first user device can be enabled to reproduce audio captured by the second user device and audio captured by the third user device if it is predicted or determined that the first user cannot directly hear the second user, or the first user device can be prevented from reproducing audio captured by the second user device while the first user device reproduces audio captured by the third user device if it is predicted or determined that the first user can directly hear the second user.

[0035] In a third aspect of the application, a method of facilitating an audio conference can comprise predicting or determining whether a first user of a first user device can directly hear a second user of a second user device, and enabling the first user device to reproduce audio captured by the second user device while the first user device reproduces audio captured by a third user device if it is predicted or determined that the first user cannot directly hear the second user, or preventing the first user device from reproducing audio captured by the second user device while the first user device reproduces audio captured by the third user device if it is predicted or determined that the first user can directly hear the second user. The method can be performed by software running on a programmable device. This software can be provided as a computer program product.

[0036] In a fourth aspect of the application, an audio conferencing facilitation apparatus comprises at least one processor which can be configured to predict or determine whether a first user of a first user apparatus can directly hear a second user of a second user apparatus at a first time, predict or determine whether the first user can directly hear the second user at a second time, cause the first user apparatus to stop reproducing audio captured by the second user apparatus if it is predicted or determined that the first user cannot directly hear the second user at the first time and it is predicted or determined that the first user can directly hear the second user at the second time, and cause the first user apparatus to start reproducing audio captured by the second user apparatus if it is predicted or determined that the first user can directly hear the second user at the first time and it is predicted or determined that the first user cannot directly hear the second user at the second time.

[0037] In a fifth aspect of the application, an audio conferencing facilitation apparatus comprises at least one processor which can be configured to receive audio captured by a user apparatus, determine whether the audio captured by the user apparatus comprises first audio information which is as derived from a same source as second audio information comprised in audio captured by a further user apparatus, and remove the second audio information from the audio captured by the further user apparatus or remove the first audio information from the audio captured by the user apparatus if it is determined that the first and second audio information are derived from the same source.

[0038] Further, a computer program for performing the methods described herein, and a non-transitory computer readable medium storing the computer program are provided. For example, the computer program can be downloaded or uploaded to an existing apparatus, or stored in the manufacturing of the apparatus.

[0039] A non-transitory computer readable storage medium stores at least software code portions which, when executed or processed by a computer, are configured to perform executable operations for facilitating an audio conference.

[0040] The executable operations can comprise predicting or determining whether a first user of a first user apparatus can directly hear a second user of a second user apparatus, and enabling the first user apparatus to reproduce audio captured by the second user apparatus while the first user apparatus reproduces audio captured by a third user apparatus if it is predicted or determined that the first user cannot directly hear the second user, or preventing the first user apparatus from reproducing audio captured by the second user apparatus while the first user apparatus reproduces audio captured by the third user apparatus if it is predicted or determined that the first user can directly hear the second user.

[0041] As will be appreciated by those skilled in the art, aspects of the present application can be embodied as a device, a method or a computer program product. Accordingly, aspects of the present application can take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, etc.) or an embodiment combining software and hardware aspects that can all generally be referred to herein as a "circuit," "module" or "system." The functionality described in the present disclosure can be implemented as an algorithm executed by a processor / microprocessor of a computer. Furthermore, aspects of the present application can take the form of a computer program product on one or more computer readable media (media) having computer readable program code embodied in the medium.

[0042] Any combination of one or more computer readable medium can be utilized. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a computer readable storage medium can include, but are not limited to, the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of the present application, a computer readable storage medium can be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0043] A computer readable signal medium can include a propagated data signal with computer readable program code embodied in the propagated data signal, for example, in baseband or as part of a carrier wave. Such a propagated signal can take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium can be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0044] Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, cable, RF, etc., or any suitable combination of the foregoing. Computer program code for carrying out operations for aspects of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java(TM), Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0045] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0046] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0047] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0048] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatuses, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing the specified logical function(s).

[0049] It should also be noted that in some alternative implementations, the functions noted in the blocks may not occur in the order noted in the figures. For example, two blocks shown in succession may actually be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It will also be noted that each block in the block diagrams and / or flow charts, and combinations of blocks in the block diagrams and / or flow charts, may be implemented by a dedicated hardware-based system or a combination of dedicated hardware and computer instructions that performs the specified functions or actions. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] These and other aspects of the invention will be apparent from and further elucidated by way of example with reference to the accompanying drawings, in which: Figure 1 is a flow chart of a first embodiment of a method for facilitating audio conferencing; Figure 2 Shown is a diagram Figure 1 Examples of methods; Figure 3 is a flow chart of a second embodiment of a method for facilitating audio conferencing; Figure 4 Three types of audio conferencing systems are shown; Figure 5 is a block diagram of a first embodiment of an audio conferencing system; Figure 6 is a block diagram of a second embodiment of an audio conferencing system; Figure 7 is a block diagram of a third embodiment of an audio conferencing system; Figure 8 is a block diagram of a fourth embodiment of an audio conferencing system; Figure 9 is a block diagram of a fifth embodiment of an audio conferencing system; Figure 10 is a flow chart of a method for solving the double capture problem; Figure 11 is a flow chart of a third embodiment of a method for facilitating audio conferencing; Figure 12 is a flow chart of a fourth embodiment of a method for facilitating audio conferencing; and Figure 13 is a block diagram of an exemplary data processing system for executing the method of the present invention.

[0051] Corresponding elements in the drawings are denoted by the same reference numerals. DETAILED DESCRIPTION

[0052] Figure 1 A first embodiment of a method of facilitating an audio conference is shown in Fig. 1. Step 101 comprises predicting or determining whether a first user of a first user device (UD1) is able to hear directly a second user of a second user device (UD2). This can be predicted based on proximity and / or location data, for example, or determined based on user input from the first user, for example. Step 102 comprises checking whether it was determined in step 101 that the first user is able to hear directly the second user. If so, step 105 is performed. If not, step 103 is performed.

[0053] Step 103 comprises enabling the first user device to reproduce audio captured by the second user device when the first user device reproduces audio captured by a third user device. Step 105 comprises preventing the first user device from reproducing audio captured by the second user device when the first user device reproduces audio captured by a third user device if it is predicted or determined that the first user is able to hear directly the second user.

[0054] Predicting whether the first user is able to hear directly the second user can involve applying an audio volume threshold, for example, which can depend on whether the user is wearing headphones / earbuds, and if so, whether they are allowing ambient audio to pass through, whether a hear-through functionality is enabled or whether an ambient noise suppression functionality is activated. The threshold can additionally or alternatively depend on the environment, for example, with a lower threshold for an environment such as a conference room, and a higher threshold for an environment where there is more background noise, for example, other people talking, music playing.

[0055] After step 103 or step 105 is performed, step 101 is repeated, and then the method proceeds as shown in Fig. 1. Thus, the method comprises predicting or determining whether the first user is able to hear the second user at least at a first time and at a second time. The method comprises causing the first user device to stop reproducing audio captured by the second user device in step 105 if it is predicted or determined in a first iteration of step 101 that the first user is not able to hear directly the second user at the first time, and it is predicted or determined in a second iteration of step 101 that the first user is able to hear directly the second user at the second time. Figure 1

[0056] The method further comprises causing the first user device to start reproducing audio captured by the second user device in step 103 if it is predicted or determined in a first iteration of step 101 that the first user is able to hear directly the second user at the first time, and it is predicted or determined in a second iteration of step 101 that the first user is not able to hear directly the second user at the second time.​

[0057] If using headphones / earbuds that allow ambient audio to pass through, Figure 1 The approach works best. Then, audio is added only for people who are not in the user's immediate vicinity (e.g., for people in the same group of users), while keeping the rest (i.e., the ambient sound) the same. These could be so-called "open-back" earbuds or headphones that physically let local sound through (also common in AR headsets), or active noise-canceling earbuds or headphones with an "ambient mode" that actively plays back external sounds picked up by the microphone.

[0058] Figure 2 Shown in the figure Figure 1 The illustrated audio conferencing system includes user devices 71-74 and an optional central audio device. Users 76 to 79 wear user devices 71-74, respectively. User devices 71-74 can be standalone devices (such as specialized AR audio devices) or devices connected to another device (e.g., a smartphone). User devices 71-74 (e.g., headphones) can include a microphone, or can use a microphone or other component to capture the user's voice, which is integrated into another device, such as a separate device or another user device used by the same user (e.g., a smartphone). Figure 2 Which user device reproduces the audio captured by which other user device is shown with dashed arrows.

[0059] Perform at least once for each pair of user devices Figure 1 Steps 101-105 are performed. In a first variant, step 101 is performed once per pair of user devices, and if the first user is predicted to directly hear the second user, it is automatically predicted that the second user will directly hear the first user. In a second variant, step 101 is performed twice per pair of user devices. In the latter case, it is not assumed that if the first user can directly hear the second user, the second user can also directly hear the first user. For example, the type of headphones and / or the user's hearing quality (e.g., in the case of hearing impairment) may be taken into account. One or more of user devices 71-74 may even be hearing aids.

[0060] Depending on the user device, step 103 or step 105 is performed for each pair of user devices. For example, user device 71 decides whether to mute the audio captured by user device 71 on user device 72, and user device 72 decides whether to mute the audio captured by user device 72 on user device 71. Muting the audio captured by the second user device on the first user device means preventing the first user device from reproducing the audio captured by the second user device.

[0061] exist Figure 2In the example of FIG, users 76 and 77 can directly hear each other, and therefore step 105 is performed for user device 71 with respect to the audio captured by user device 72, and step 105 is performed for user device 72 with respect to the audio captured by user device 71. Thus, user device 71 is prevented from reproducing the audio captured by user device 72, and user device 72 is prevented from reproducing the audio captured by user device 71.

[0062] exist Figure 2 In the example shown in FIG, users 78 and 79 cannot directly hear any other user, and users 76 and 77 cannot directly hear user 78 or user 79. Therefore, step 103 is performed for user devices 73 and 74 with respect to audio captured by other user devices, and step 103 is performed for user devices 71 and 72 with respect to audio captured by user devices 73 and 74. Thus, user devices 73 and 74 are enabled to reproduce audio captured by other user devices, and user devices 71 and 72 are enabled to reproduce audio captured by user devices 73 and 74.

[0063] Figure 3 A second embodiment of a method for facilitating audio conferencing is shown in FIG. Figure 3 The second embodiment is Figure 1 An extension of the first embodiment. Figure 3 In the embodiment of Figure 1 is executed before step 101, and Figure 1 Step 101 has been implemented by step 125.

[0064] Step 121 involves identifying which devices are part of a group. For example, in a museum, there are often multiple groups of people visiting the museum simultaneously. People within a group enjoy talking to each other, but not to people in other groups. Similarly, in an office, there may be multiple groups of people on different conference calls. Groups are typically managed by a central server. Users can indicate which group they want to join.

[0065] Step 123 comprises obtaining proximity and / or position data of the devices identified in step 121. The proximity and / or position data is indicative of the proximity of the first user device to at least the second user device and / or indicative of the position of at least the first user device and the position of the second user device. If the group comprises more than two devices, the proximity and / or position data is also indicative of the proximity and / or position with respect to these other devices. Steps 121 and 123 can be performed for each group of user devices, e.g. if step 123 comprises determining position data, but can alternatively be performed for each pair of user devices, e.g. if only a Bluetooth scan is used to determine the proximity of a neighbor. Steps 121 and 123 can be performed once, but are typically repeated over time (to deal with groups changing over time, such as users joining or leaving the group, or users moving around or changing position over time).

[0066] The proximity can be determined in different ways. The proximity can be determined using wireless probe signals. For example, the second user device can transmit an RF (e.g. Bluetooth or Wi-Fi), (ultra)sound or infrared signal, and the first user device can receive the signal in the same time interval, i.e. listen for the signal and determine whether it is received and whether it is received in proximity with sufficient strength. Alternatively, the wireless probe signal can for example trigger a response, similar to a "ping" on the internet, where one user device requests a response from the other user device by sending a wireless signal, and the other user device responds to this signal. In this way, both user devices can be able to determine the proximity from a single request-response exchange, or both user devices can send both the request and the response.

[0067] Step 125 comprises predicting whether the first user of the first user device is able to directly hear the second user of the second user device based on the proximity and / or position data obtained in step 123. Next, steps 102, 103 and 105 are performed as described with respect to Figure 1 For example, if the method is performed by a central audio device, steps 121 and 123 and steps 101-105 can be performed for multiple groups.

[0068] In Figure 3 embodiments, when people are in physical proximity, they do not hear each other through the audio conferencing system, but they only talk to each other directly. When the distance is larger, people are still able to talk to each other and hear each other through the audio conferencing system.

[0069] Figure 4Three types of audio conferencing systems are shown: a peer-to-peer (P2P) audio conferencing system 86, an audio conferencing system 87 using a stream forwarding unit (SFU) 81, and an audio conferencing system 88 using a multipoint control unit (MCU) 83. In the P2P audio conferencing system 86, audio streams are transmitted directly from user devices to other user devices, not to a central audio device.

[0070] The MCU 83 mixes the audio captured by multiple devices in a single stream tailored for a single user device and transmits this single stream to this user device. The MCU 83 does this for every user device. Thus, every user device transmits only one audio stream and receives only one audio stream.

[0071] The SFU 81 is a central server that has a separate signaling connection with each of the user devices in the session. In contrast to the MCU, the SFU will not decode and mix media streams; it will simply forward media packets arriving on one input connection to all output connections that have requested a specific media stream.

[0072] As such, the SFU is a signaling endpoint for each of the user devices. In SIP (Session Initiation Protocol) terms, this would be called a B2BUA, i.e. a Back-to-Back User Agent. This means that the server acts as a regular user agent towards each user device and internally connects the user agents to the various user devices. The same concept applies if other protocols are used.

[0073] To allow this to work, similar to in P2P systems but unlike in MCU-based systems, each user device needs to support the media codecs used by the other user devices to encode / format their streams. For example, generally available codecs can be used. For specific applications, specific codecs can be used if all user devices of the users in the group are identical or run the same application.

[0074] On the stream forwarding level, the SFU 41 also has a role to play. Even though it does not decode and mix media streams like the MCU 43, it will have to ensure that each user will receive the correct media stream. Thus, the SFU 41 can be configured to rewrite the RTP header, such as the SSRC number (for the identification of the stream) and the sequence number. A more comprehensive description of this aspect can be found in, for example, IETF RFC 7667 on RTP Topologies.

[0075] Typically, when signaling for the media streams to be exchanged is performed using SDP (Session Description Protocol) offer / answer (as for example in WebRTC), the user devices will describe (i.e. offer) the streams they have available to the SFU 81 and the SFU 81 will describe (i.e. offer) the streams it has available to the user devices. While each user device typically has only one or two streams to offer, e.g. an audio stream and possibly a video stream, the SFU 81 will have many streams to offer: one or possibly two streams for each input stream from each other user device. If a SFU is used, the following additional concepts can be used specifically for video: - Since not all receiving user devices have the same available bandwidth, some user devices can want streams of higher quality and thus higher bandwidth than others. This can be achieved either by simulcasting (i.e. each user device offers streams of various qualities to the SFU) or by using a scalable video codec (i.e. each user device offers the content in a layered fashion, where the base layer offers a basic quality, while the addition of other layers will improve the quality and will require more bandwidth). - Some media transport mechanisms include retransmission of lost packets. When a packet is lost only by some receiving user device, it should not be retransmitted to all user devices. This requires either buffering and retransmission by the SFU or careful state management to forward retransmitted packets only to the correct user devices.

[0076] Figures 5-7 are block diagrams of a first embodiment, a second embodiment and a third embodiment of an audio conferencing system, respectively. In the Figure 5 embodiment, the audio conferencing system 91 is a P2P system comprising three user devices 11-13. In the Figure 6 embodiment, the audio conferencing system 92 comprises three user devices 11-13 and a SFU 31. In the Figure 7 embodiment, the audio conferencing system 93 comprises three user devices 11-13 and a MCU 33. The three types of audio conferencing systems have been described with respect to Figure 4

[0077] Each user device comprises a receiver 3, a transmitter 4, a processor 5 and a memory 7. The processor 5 is configured to predict or determine whether a first user of a first user device is able to directly hear a second user of a second user device and, if it is predicted or determined that the first user is not able to directly hear the second user, enable the first user device to reproduce audio captured by the second user device when the first user device reproduces audio captured by a third user device, or, if it is predicted or determined that the first user is able to directly hear the second user, prevent the first user device from reproducing audio captured by the second user device when the first user device reproduces audio captured by the third user device. ​

[0078] The processor 5 is configured to predict or determine whether the first user of the first user device is able to directly hear the third user of the third user device, and if it is predicted or determined that the first user cannot directly hear the third user, enable the first user device to reproduce the audio captured by the second user device and the audio captured by the third user device if it is predicted or determined that the first user cannot directly hear the second user, or if it is predicted or determined that the first user is able to directly hear the second user, prevent the first user device from reproducing the audio captured by the second user device when the first user device reproduces the audio captured by the third user device.

[0079] exist Figures 5-9 In an embodiment, the processor 5 is configured to obtain proximity and / or location data and, based on the proximity and / or location data, predict whether a first user of a first user device can directly hear a second user of a second user device. The proximity and / or location data indicates the proximity of the first user device to at least the second user device and / or indicates the location of at least the first user device and the location of the second user device. Well-known techniques can be used to determine proximity and / or location.

[0080] For example, each user device may obtain / determine proximity data based (only) on its own signal strength measurements. In this case, each user device determines which other user devices are nearby based on the RF signal strength of the received RF signal (e.g., Bluetooth or Wi-Fi signal). Location data may be used instead of or in addition to proximity data. For example, each user device may use beacons 21-23 to determine its own location and share its location with other user devices or with server 25. Server 25 may then share the location data it has received with the user devices. Server 25 may have an API or equivalent interface through which user devices can share and obtain this location data. The location data may be relative location data, i.e., location data indicating where the user is within a specific location, or it may be absolute location data, i.e., describing the exact physical location on the Earth.

[0081] Proximity can be determined based on the location of the user devices and a certain threshold, for example, users being less than 2 meters apart. However, there are some obvious situations where users in such close proximity still cannot hear each other properly: when there are walls or windows between them. For example, in a museum, two users may be in completely different rooms but still be physically close to each other.

[0082] The benefit of using location data is that it makes it possible to better determine whether two users are able to hear each other directly. By using location data and a building map indicating walls, windows and ceilings, a user device can be able to determine whether other user devices are in the same room. If the user devices obtain location data from the server 25, they can be able to obtain the map from the same server 25. If the user devices obtain location data from other user devices, the user devices can be able to obtain the map from the server 25. The server 25 can also act as a beacon.

[0083] The user devices 11-13 can determine which of the received proximity and / or location data is relevant by using group information. As part of group management, the devices can have exchanged the hardware addresses of the interfaces used, or the devices can broadcast their identity together with a group identification, or the devices can be paired or directly connected, etc.

[0084] In a first implementation of the user devices 11-13, the processor 5 of a user device 11, for example, is configured to predict or determine whether a user of the user device is able to hear a first other user of a first other user device (e.g. user device 12) directly, determine whether the user of the user device is able to hear a second other user of a second other user device (e.g. user device 13) directly, and if it is predicted or determined that the user is not able to hear the second other user directly, enable the user device to reproduce audio captured by the first other user device and audio captured by the second other user device if it is predicted or determined that the user is not able to hear the first other user directly, or prevent the user device from reproducing audio captured by the first other user device when the user device reproduces audio captured by the second other user device if it is predicted or determined that the user is able to hear the first other user directly. Thus, in this first implementation, it is the receiving user device that decides whether to mute audio captured by a transmitting user device on the receiving user device.

[0085] In Figure 2 In the example of Fig. 7, if the user devices 71-74 are implemented in this way, the user device 71 will prevent audio captured by the user device 72 from being reproduced on the user device 71, and the user device 72 will prevent audio captured by the user device 71 from being reproduced on the user device 72. The user device 71 will enable audio captured by the user devices 73 and 74 to be reproduced on the user device 71.

[0086] In Figures 5-7In embodiments of the first implementation of the user devices 11-13, the processor 5 of the first user device can be configured to prevent the first user device from reproducing audio captured by the second user device by adjusting its handling of received audio packets such that the audio captured by the second user device is not reproduced, or by instructing the second user device to prevent the first user device from reproducing the audio captured by the second user device. The second user device can prevent this in a similar way as will be described in relation to the second implementation, in which the second user device decides whether to mute the audio captured by the second user device on the first user device.

[0087] As an example of instructing the second user device, in embodiments of the first implementation of the user devices 11-13, the processor 5 of the first user device can be configured to prevent the first user device from reproducing audio captured by the second user device by adjusting its handling of received audio packets such that the audio captured by the second user device is not reproduced, or by instructing the second user device to prevent the first user device from reproducing the audio captured by the second user device. The second user device can prevent this in a similar way as will be described in relation to the second implementation, in which the second user device decides whether to mute the audio captured by the second user device on the first user device. Figure 5

[0088] In embodiments of the first implementation of the user devices 11-13, the processor 5 of the first user device can alternatively be configured to prevent the first user device from reproducing audio captured by the second user device by instructing the central audio system, Figures 5-7 Figure 6 the SFU 31 of the system 30, Figure 7 the MCU 33 of the system 30. Figure 6 The SFU 31 of the system 30 can prevent this in a similar way as the SFU 41 of the system 40. Figure 8 The MCU 33 of the system 30 can prevent this in a similar way as the MCU 43 of the system 40. Figure 7 Figure 9

[0089] ​​​​In the second implementation of the user devices 11-13, the processor 5 of a user device 12, for example, is configured to predict or determine whether a first other user of a first other user device, e.g. user device 11, can directly hear the user of the user device, and if it is predicted or determined that the first other user cannot directly hear the user, to enable the first other user device to reproduce audio captured by the user device when the first other user device reproduces audio captured by a second other user device, e.g. user device 13, or if it is predicted or determined that the first other user can directly hear the user, to prevent the first other user device from reproducing audio captured by the user device when the first other user device reproduces audio captured by a second other user device. Thus, in this second implementation, it is the transmitting user device that decides whether audio captured by the transmitting user device is muted on the receiving user device.

[0090] In Figure 2 the example of Figs. 7A-7D, if the user devices 71-74 were implemented in this way, user device 71 would prevent audio captured by user device 71 from being reproduced on user device 72, and user device 72 would prevent audio captured by user device 72 from being reproduced on user device 71. User device 73 would enable audio captured by user device 73 to be reproduced on user device 71. User device 74 would enable audio captured by user device 74 to be reproduced on user device 71.

[0091] In Figures 5-7 the example of Figs. 7A-7D, if the user devices 71-74 were implemented in this way, user device 71 would prevent audio captured by user device 71 from being reproduced on user device 72, and user device 72 would prevent audio captured by user device 72 from being reproduced on user device 71. User device 73 would enable audio captured by user device 73 to be reproduced on user device 71. User device 74 would enable audio captured by user device 74 to be reproduced on user device 71.

[0092] In Figure 5 the example of Figs. 7A-7D, if the user devices 71-74 were implemented in this way, user device 71 would prevent audio captured by user device 71 from being reproduced on user device 72, and user device 72 would prevent audio captured by user device 72 from being reproduced on user device 71. User device 73 would enable audio captured by user device 73 to be reproduced on user device 71. User device 74 would enable audio captured by user device 74 to be reproduced on user device 71. Figure 5 and 6 In the example of Figs. 7A-7D, the instruction can be included as metadata in the audio stream, or can be signaled in the example of Figs. 7A-7D. Muting can be implemented in WebRTC, for example, by (the second user device) updating the media direction to "recvonly", which effectively causes the SDP exchange of the WebRTC client (of the second user device) to stop sending audio. This effectively signals to the far end that the microphone is muted. Since in WebRTC the media direction is always "sendonly" or "recvonly", the media direction can be updated to "recvonly" to signal that the microphone is muted. Figure 7In embodiments of the application, the user devices receive only one audio stream, so the signaling would be useless in this embodiment.

[0093] In Figure 5 embodiments of the application, in the second implementation of the user devices 11-13, the processor 5 of the second user device can alternatively be configured to prevent the first user device from reproducing the audio captured by the second user device by not transmitting the audio captured by the second user device to the first user device, for example by not transmitting any audio stream from the second user device to the first user device (also called "send_none"), or by transmitting a silent audio stream from the second user device to the first user device (also called "send_silent").

[0094] This is not possible in Figure 6 and 7 embodiments of the application, because each user device transmits only one audio stream with audio captured by this user device, and the user device not transmitting its captured audio would mean that none of the user devices would be able to reproduce the audio captured by this user device. Send_none can be implemented by just stopping sending audio packets. Send_silent can be an action of the audio component itself. It can pass audio packets that do not contain audio, instead of encoding the audio received from the microphone.

[0095] In alternative implementations of the user devices 11-13, the first user device can prevent the audio captured by the second user device from being reproduced on the first user device, and also prevent the audio captured by the first user device from being reproduced on the second user device. In Figure 2 the example of Fig. 7, if the user devices 71-74 would be implemented in this way, the user device 71 or the user device 72 would prevent the audio captured by the user device 72 from being reproduced on the user device 72, and prevent the audio captured by the user device 71 from being reproduced on the user device 71.

[0096] In the embodiments shown in Figures 5 to 7 Fig. 1, the user devices 11-13 comprise one processor 5. In alternative embodiments, one or more of the user devices 11-13 comprise multiple processors. The processor 5 can be a general purpose processor, for example an ARM or a Qualcomm processor, or a special purpose processor. For example, the processor 5 can run a Unix-based operating system (for example, Google Android) or Apple iOS as operating system. For example, the processor can comprise multiple cores. For example, the memory 7 can comprise solid state memory (for example one or more solid state disks (SSD) made of flash memory) or one or more hard disks.

[0097] For example, the receiver 3 and transmitter 4 of the user devices 11-13 can use one or more wireless communication technologies, such as Wi-Fi, LTE, and / or 5G New Radio, to communicate with other devices on the Internet. The receiver 3 and transmitter 4 can be combined in a transceiver. The user devices 11-13 may include other components typical of user devices, such as a battery and / or a power connector.

[0098] Figures 8-9 They are block diagrams of the fourth and fifth embodiments of the audio conferencing system respectively. Figure 8 In the embodiment of FIG. 4 , the audio conferencing system 94 includes three user devices 51-53 and an SFU 41. Figure 9 In the embodiment of the present invention, the audio conferencing system 95 includes three user devices 51-53 and an MCU 43. Figure 4 Both types of audio conferencing systems are described.

[0099] Figure 8 SFU 41 and Figure 9 The MCUs 43 of the SFUs 41 and 43 each include a receiver 3, a transmitter 4, a processor 5, and a memory 7. The processor 5 of the SFU 41 or the MCU 43 is configured to predict or determine whether a first user of the first user device can directly hear a second user of the second user device, and if it is predicted or determined that the first user cannot directly hear the second user, enable the first user device to reproduce the audio captured by the second user device when the first user device reproduces the audio captured by the third user device, or if it is predicted or determined that the first user can directly hear the second user, prevent the first user device from reproducing the audio captured by the second user device when the first user device reproduces the audio captured by the third user device.

[0100] The processor 5 of the SFU 41 or MCU 43 is also configured to predict or determine whether the first user of the first user device can directly hear the third user of the third user device, and if it is predicted or determined that the first user cannot directly hear the third user, enable the first user device to reproduce the audio captured by the second user device and the audio captured by the third user device if it is predicted or determined that the first user cannot directly hear the second user, or if it is predicted or determined that the first user can directly hear the second user, prevent the first user device from reproducing the audio captured by the second user device when the first user device reproduces the audio captured by the third user device.

[0101] exist Figure 8 and 9In embodiments of the application, the SFU 41 and the MCU 43 obtain proximity and / or location data from the user devices 51-53, or obtain location data from the server 25. The SFU 41 and / or the MCU 43 can store the location data and / or the map itself, instead of obtaining the location data from the server 25. In this case, a separate server 25 can not be necessary. Each of the user devices 51-53 can determine its location using the beacons 21-23, and share this location with the SFU 41 and / or the MCU 43.

[0102] In Figure 8 and 9 embodiments of the application, the processor 5 of the SFU 41 or the MCU 43 can be configured to prevent the first user device from reproducing audio captured by the second user device by not transmitting to the first user device audio captured by the second user device. For example, in Figure 9 embodiments of the application, the processor 5 of the MCU 43 can be configured to prevent the first user device from reproducing audio captured by the second user device by omitting audio captured by the second user device (send_silent) from a single stream transmitted to the first user device.

[0103] In Figure 8 embodiments of the application, the processor 5 of the SFU 41 can be configured to not transmit any data stream to the second user device on behalf of the first user device (send_none). Thus, the SFU 41 selectively forwards certain audio streams and not others. The SFU 41 can simply stop forwarding packets to certain users without any signaling.

[0104] In Figure 8 embodiments of the application, the processor 5 of the SFU 41 can alternatively be configured to prevent the first user device from reproducing audio captured by the second user device by sending an instruction to the first user device on the network. The instruction instructs the first user device not to reproduce audio captured by the second user device. For example, the instruction can instruct a user device receiving the instruction whether it should reproduce audio included in a certain audio stream, or can instruct a list of user devices that should or should not reproduce audio included in a certain audio stream. For example, the instruction can be signaled. For example, the SFU 41 can select the streams that a receiving user device needs, and use signaling (e.g., SDP signaling) to instruct the receiving user device which streams it should reproduce.

[0105] Thus, when using an MCU, each user device sends their stream to the MCU once. The MCU mixes the output of each individual user device containing the audio of the other users and sends this to the user device. Because the audio is modified by the MCU, it can use send_silent, but it cannot use send_none, because audio still needs to be sent for the other users. Also, since the individual streams are no longer identifiable, using signaling or metadata will not work.

[0106] When using an SFU, each user device sends their stream to the SFU once. The SFU replicates the audio stream and sends it to every other user device without modifying the audio. In this case, the signaling approach can work, and send_none can also work by simply not replicating audio packets for a certain receiving user. However, send_silent is not available, because the SFU itself does not manipulate the audio stream. Also, the use of metadata does not work, because all users receive the same audio stream (a copy of it).

[0107] In the embodiments shown in Figure 8 and 9 , the SFU 41 and the MCU 43 each comprise one processor 5. In alternative embodiments, the SFU 41 and / or the MCU 43 comprise multiple processors. The processor 5 can be a general purpose processor, such as an Intel or AMD processor, or a special purpose processor. For example, the processor 5 can run a Unix-based operating system or a Windows operating system. For example, the processor 5 can comprise multiple cores. For example, the memory 7 can comprise solid state memory, such as one or more solid state disks (SSDs) made of flash memory, or one or more hard disks. Alternatively, the SFU 41 and the MCU 43 can run in a cloud network or an edge network, typically with scalable processing and storage capabilities. In yet another embodiment, one of the user devices acts as the server 25, the SFU 41 and / or the MCU 43.

[0108] For example, the receiver 3 and the transmitter 4 can use one or more wireless communication technologies, such as Wi-Fi, LTE and / or 5G New Radio, to communicate with other devices on the Internet. The receiver 3 and the transmitter 4 can be combined in a transceiver. The SFU 41 and the MCU 43 can comprise other components typical for central audio devices, such as a power connector.

[0109] In the above description of Figures 5 to 9 , it is described how a receiving user device can decide whether to mute the audio captured by a transmitting user device on the receiving device by adjusting its handling of the received audio packets or by instructing the transmitting user device or the central audio device. In Figures 5 to 9In the above description, it is further described how a transmitting user device or a central audio device can decide whether to mute the audio captured by the transmitting user device on the receiving device, and four implementations thereof are described: 1. Send_none: Send no audio for a certain user; 2. Send_silent: Send "silent" audio for a certain user; 3. Signaling: Indicate in the signaling that audio for a certain user should not be reproduced; 4. Metadata: Indicate in the metadata of the audio stream that the audio should not be reproduced.

[0110] If a central audio device is used, it is the central audio device that uses option 1 or option 2, not the transmitting user device. The first option is most efficient for the network, but can cause problems if the receiving user device stops working when it no longer receives audio packets. The second option prevents this problem, but requires modification of the audio, and thus more processing. The use of signaling or metadata prevents this audio processing and allows muting to be cancelled more instantaneously when users leave each other's proximity (as the audio is still being delivered). Furthermore, a combination of these options can be applied, such as for example signaling and also no longer sending audio.

[0111] Table 1 below gives an overview of the various options that have been described with respect to the embodiments of Figures 5 to 9 for embodiments or implementations in which a transmitting user device or a central audio device decides whether to mute the audio captured by the transmitting user device on the receiving device. Other options can also be possible, for example, the SFU can be adjusted to be able to modify the output metadata in the RTP header and thus support the metadata option.

[0112] Table 1 Send_none Send_silent Signaling Metadata P2P X X X X SFU X X MCU X Optionally, the user devices 11-13 and / or 51-53 of the above description can use earplugs or headphones that are able to reproduce spatial audio. The audio of a remote user speaking can then sound as if it comes from the physical direction in which the person actually is. Spatial audio can work in P2P and SFU embodiments using object-based spatial audio, or in MCU embodiments using omnidirectional audio. Figures 5-9

[0113] Optionally, the above audio conference can be combined with a video conference. When the users are local, they see and hear each other. When the users are far apart, they see a video projection or some other representation of the other user (possibly in the direction in which the user physically is), and also hear each other through the audio conference system.

[0114] ​In case of a (shared) audio tour, if the live guide is not part of the group, then Figures 5-9 The user devices 11-13 and / or 51-53 described in the method of

[0115] Figure 1 The method described in the method of

[0116] This double capture problem does not always occur. In fact, in case all users have a modern smartphone, and when a single communication service with the same settings (e.g. used codec, buffering settings, etc.) is used for each user device, the delay differences will most likely be negligible, and the echo problem will most likely not occur. Furthermore, modern headsets and earplugs are very good at picking up the user's voice and filtering out ambient noise, and can thereby help prevent the echo problem by not picking up the other user's voice.

[0117] If the double capture problem does occur, it can be very annoying. There are several ways to solve this problem. First, by means of the synchronized capture by multiple microphones, the captured audio can also be played back synchronously. Such capture synchronization is well known. The most common approach to this is to synchronize the clocks on the capturing devices, e.g. using NTP, GPS or cellular clock synchronization, and then timestamp the captured audio packets in the media stream with these synchronized clocks. Timestamping can be done by inserting the timestamp in the header of the audio packet, or by signaling the relationship between the timestamp of the audio packet and the clock timestamp (e.g. in case of an RTP timestamp, which is usually a random offset). This should be done periodically, taking into account clock drift. Next, inter-media synchronization should be applied.

[0118] An alternative to playback synchronization is to use echo cancellation. The working principle of echo cancellation is to find a certain potentially modified form of an audio pattern in one piece of audio from another piece of audio, usually within a certain delay of the first pattern. Typically, this works on a single device: incoming audio is played back through a loudspeaker, and then captured by a microphone on the same device (acoustic echo). The played back audio picked up by the microphone is then filtered out (i.e. cancelled) from the outgoing audio.

[0119] Figure 10 is a flowchart of a method of solving the double capture problem by means of echo cancellation. Figure 10 The method of comprises steps 141, 143, 145 and 147. Step 141 comprises receiving audio captured by a third user device. Step 143 comprises determining whether the audio captured by the third user device comprises first audio information that is as much originating from the same source as second audio information comprised in the audio captured by the second user device.

[0120] Step 145 comprises checking whether it was determined in step 143 that the audio captured by the third user device comprises first audio information that is as much originating from the same source as the second audio information. If so, step 147 is performed. Step 147 comprises removing the second audio information from the audio captured by the second user device, or removing the first audio information from the audio captured by the third user device, if it is determined that the first audio information and the second audio information originate from the same source.

[0121] Removing audio information can comprise audio processing such as cancelling the audio information in the captured audio or subtracting the audio information from the captured audio. Such cancelling / subtracting typically uses techniques comparable to (regular) echo cancellation. The audio captured by the second and third user devices minus any removed audio information is reproduced on the first user device (and vice versa). Figure 10 (not shown in the figure).

[0122] Figure 10 The working of the echo cancellation of differs from the standard echo cancellation scenario described above in that the audio that is picked up twice is picked up by two different devices. In the example of proximity between the first and second user, the voice of the first user can be picked up by both microphones, as will be the voice of the second user. Although it is more difficult to determine which audio signal in which stream to cancel compared to classic echo cancellation, the audio volume can be used to distinguish the own user of a device (according to the assumption that said user is closest to his or her own microphone and thus has a higher volume) from the other user.

[0123] Furthermore, techniques such as those disclosed in EP3175456 A1 can be used to implement echo cancellation. In EP3175456 A1, instead of playing the device, the user of the second communication device creates a sound signal that is recorded by the first communication device. Because the second communication device also records the sound signal created by its user, the second communication device can provide noise suppression data.

[0124] Figure 11 A third embodiment of a method for facilitating audio conferencing is shown in , which solves the problem of double capture. Figure 11 The third embodiment combines Figure 1 The first embodiment and Figure 10 Method. Figure 11 In the embodiment of the present invention, if it is determined in step 101 that the first user can directly hear the second user, then after step 102, an additional Figure 10 If it is determined in step 143 that the audio captured by the third user device includes first audio information originating from the same source as the second audio information, step 147 is performed after step 145. Figure 10 Steps 141-147 are described.

[0125] Figure 12 A fourth embodiment of a method for facilitating audio conferencing is shown in , which also solves the problem of double capture. Figure 12 The fourth embodiment is Figure 1 An extension of the first embodiment. Figure 12 In an embodiment, if it is determined in step 101 that the first user can directly hear the second user, step 161 is additionally performed after step 102. Step 161 includes deactivating the capture and / or transmission of audio by the first user device or the capture and / or transmission of audio by the second user device.

[0126] exist Figure 12 In the embodiment of the present invention, if the first user can directly hear the second user, then the audio captured by the first user device and the audio captured by the second user device are determined or expected to include audio information from the same source. In alternative embodiments, whether the audio captured by the first user device and the audio captured by the second user device are determined or expected to include audio information from the same source is determined in a different manner, and step 161 is performed depending on the result of this determination.

[0127] Figure 13 Depicted diagram can be executed reference Figure 1 、 3 10-12. For the audio conferencing facilitation device and / or any user device disclosed herein, such asFigures 8-9 The user devices 51-53 and the data processing system are also exemplary.

[0128] like Figure 13 , data processing system 300 may include at least one processor 302 coupled to memory element 304 via system bus 306. As such, the data processing system may store program code within memory element 304. In addition, processor 302 may execute program code accessed from memory element 304 via system bus 306. In one aspect, the data processing system may be implemented as a computer suitable for storing and / or executing program code. However, it will be appreciated that data processing system 300 may be implemented in the form of any system including a processor and memory capable of performing the functions described in this specification.

[0129] The memory element 304 may include one or more physical memory devices, such as, for example, local memory 308 and one or more mass storage devices 310. Local memory may refer to random access memory or other (one or more) non-persistent memory devices that are typically used during the actual execution of program code. The mass storage device may be implemented as a hard drive or other persistent data storage device. The processing system 300 may also include one or more cache memories (not shown) that provide temporary storage of at least some program code in order to reduce the number of times program code must be retrieved from the mass storage device 310 during execution.

[0130] Input / output (I / O) devices, depicted as input device 312 and output device 314, may optionally be coupled to the data processing system. Examples of input devices may include, but are not limited to, a keyboard, a pointing device such as a mouse, and the like. Examples of output devices may include, but are not limited to, a monitor or display, speakers, and the like. Input and / or output devices may be coupled to the data processing system directly or through intervening I / O controllers.

[0131] In one embodiment, the input and output devices may be implemented as a combined input / output device (in Figure 13 314). An example of such a combined device is a touch-sensitive display, sometimes also referred to as a "touch screen display" or simply a "touch screen." In such an embodiment, input to the device can be provided by moving a physical object, such as, for example, a user's finger or a stylus, on or near the touch screen display.

[0132] A network adapter 316 may also be coupled to the data processing system to enable it to couple to other systems, computer systems, remote network devices, and / or remote storage devices through intervening private or public networks. The network adapter may include a data receiver for receiving data transmitted to the data processing system 300 by the system, device, and / or network, as well as a data transmitter for transmitting data from the data processing system 300 to the system, device, and / or network. Modems, cable modems, and Ethernet cards are examples of different types of network adapters that can be used with data processing system 300. Network adapter 316 may support one or more wired networks and / or one or more wireless networks (e.g., Wi-Fi and / or Bluetooth).

[0133] like Figure 13 As depicted in FIG, memory element 304 may store applications 318. In various embodiments, applications 318 may be stored in local memory 308, one or more mass storage devices 310, or separate from local memory and mass storage devices. It should be appreciated that data processing system 300 may also execute an operating system ( Figure 13 318. The application 318, implemented in the form of executable program code, may be executed by the data processing system 300, for example, by the processor 302. In response to executing the application, the data processing system 300 may be configured to perform one or more operations or method steps described herein.

[0134] Various embodiments of the present invention may be implemented as a program product for use with a computer system, wherein the program(s) of the program product define the functionality of the embodiments (including the methods described herein). In one embodiment, the program(s) may be embodied on various non-transitory computer-readable storage media, where, as used herein, the term "non-transitory computer-readable storage medium" includes all computer-readable media with the sole exception of transitory propagating signals. In another embodiment, the program(s) may be embodied on various transitory computer-readable storage media. Illustrative computer-readable storage media include, but are not limited to: (i) non-writable storage media on which information is permanently stored (e.g., a read-only memory device within a computer, such as a CD-ROM disk readable by a CD-ROM drive, a ROM chip, or any type of solid-state non-volatile semiconductor memory); and (ii) writable storage media on which variable information is stored (e.g., a flash memory, a floppy disk in a floppy disk drive or hard drive, or any type of solid-state random-access semiconductor memory). The computer program(s) may be executed on the processor 302 described herein.

[0135] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0136] The corresponding structure, material, acts, and equivalents of all means or step plus function elements in the claims that follow are intended to include any structure, material, or act for performing the function in combination with other claimed The description of embodiments of the application has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the application forms disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the application. The described embodiments were chosen and described in order to best explain the principles of the application and the practical application and to enable others skilled in the art to understand the application for various embodiments with various modifications as are suited to the particular use contemplated.

Claims

1. An audio conference promotion device (11-12, 41, 43), wherein: The audio conference facilitation device (11-12, 41, 43) includes at least one processor (5), and the at least one processor (5) is configured to: - predicting or determining whether a first user (76) of a first user device (11, 51, 71) is able to directly hear a second user (77) of a second user device (12, 52, 72), - if it is predicted or determined that the first user cannot directly hear the second user, enabling the first user device (11, 51, 71) to reproduce audio captured by the second user device (12, 52, 72) when the first user device (11, 51, 71) reproduces audio captured by a third user device (13, 53, 73), or if it is predicted or determined that the first user can directly hear the second user, preventing the first user device (11, 51, 71) from reproducing audio captured by the second user device (12, 52, 72) when the first user device (11, 51, 71) reproduces audio captured by the third user device (13, 53, 73).

2. The audio conference promotion device (11) according to claim 1, wherein: The audio conferencing facilitation device is the first user device (11).

3. The audio conference promotion device (11) according to claim 2, wherein: The at least one processor (5) is configured to prevent the first user device (11) from reproducing the audio captured by the second user device (12) by adjusting its processing of received audio packets so that the audio captured by the second user device (12) is not reproduced.

4. The audio conference promotion device (11) according to claim 2, wherein: The at least one processor (5) is configured to prevent the first user device (11) from reproducing audio captured by the second user device (12) by instructing a central audio system (31, 33) or the second user device (12) to prevent the first user device (11) from reproducing audio captured by the second user device (12).

5. The audio conference promotion device (12, 41, 43) according to claim 1, wherein: The audio conferencing facilitation device is the second user device (12) or a central audio system (41, 43).

6. The audio conference promotion device (12, 41, 43) according to claim 5, wherein: The at least one processor (5) is configured to prevent the first user device (11, 51, 71) from reproducing audio captured by the second user device (12, 52, 72) by sending an instruction to the first user device (11, 51, 71) over a network, the instruction instructing the first user device (11, 51, 71) not to reproduce the audio captured by the second user device (12, 52, 72).

7. The audio conference promotion device (11-12, 41, 43) according to claim 5, wherein: The at least one processor (5) is configured to prevent the first user device (11, 51, 71) from reproducing audio captured by the second user device (12, 52, 72) by not transmitting the audio captured by the second user device (12, 52, 72) to the first user device (11, 51, 71).

8. The audio conference promotion device (43) according to claim 5, wherein: The audio conferencing facilitation device is the central audio system (43), and the at least one processor (5) is configured to mix audio captured by multiple devices (52-54) into a single stream customized for the first user device (51) and transmit the single stream to the first user device (51), and wherein the at least one processor (5) is configured to prevent the first user device (51) from reproducing the audio captured by the second user device (52) by omitting the audio captured by the second user device (52) from the single stream transmitted to the first user device (51).

9. The audio conferencing facilitation device (11-12, 43) according to any one of the preceding claims, wherein: The at least one processor is configured to: - receiving audio captured by a third user device (13, 53, 73), - determining whether the audio captured by the third user device (13, 53, 73) includes first audio information originating from the same source as second audio information included in the audio captured by the second user device (12, 52, 72), and - if it is determined that the first audio information and the second audio information originate from the same source, removing the second audio information from the audio captured by the second user device (12, 52, 72) or removing the first audio information from the audio captured by the third user device (13, 53, 73).

10. The audio conferencing facilitation device (11-12, 41, 43) according to any one of the preceding claims, wherein: The at least one processor is configured to: - obtaining proximity and / or location data indicating proximity of the first user device (11, 51, 71) to at least the second user device (12, 52, 72) and / or indicating a location of at least the first user device (11, 51, 71) and a location of the second user device (12, 52, 72), and - predicting whether the first user (76) of the first user device (11, 51, 71) is able to directly hear the second user (77) of the second user device (12, 52, 72) based on the proximity and / or position data.

11. An audio conferencing system (91-95), comprising the audio conferencing facilitation device (11-12, 41, 43) according to any one of claims 1 to 10.

12. The audio conferencing system (91-95) according to claim 11, wherein: The audio conferencing facilitation device is a central audio system (41, 43), and the audio conferencing system further comprises the first user device (51), the second user device (52) and the third user device (53).

13. The audio conferencing system (91-95) according to claim 11, wherein: The audio conferencing facilitation device is one of the first user device (11) and the second user device (12), and the audio conferencing system further includes the other of the first user device (11) and the second user device (12), and further includes the third user device (13).

14. The audio conferencing system (91-95) according to any one of claims 11 to 13, wherein: The at least one processor (5) of the audio conferencing device (11-12, 41, 43) or at least one other processor of the audio conferencing system (91-95) is configured to: - predicting or determining whether the first user (76) of the first user device (11, 51, 71) is able to directly hear the third user (78) of the third user device (13, 53, 73), and - if it is predicted or determined that the first user (76) cannot directly hear the third user (78), enabling the first user device (11, 51, 71) to reproduce audio captured by the second user device (12, 52, 72) and audio captured by the third user device (13, 53, 73) if it is predicted or determined that the first user (76) cannot directly hear the second user (77), or preventing the first user device (11, 51, 71) from reproducing audio captured by the second user device (12, 52, 72) while the first user device (11, 51, 71) is reproducing audio captured by the third user device (13, 53, 73) if it is predicted or determined that the first user (76) can directly hear the second user (77).

15. A method for facilitating audio conferencing, the method comprising: - predicting or determining (101) whether a first user of a first user device is able to directly hear a second user of a second user device, and - if it is predicted or determined that the first user cannot directly hear the second user, enabling (103) the first user device to reproduce the audio captured by the second user device when the first user device reproduces the audio captured by the third user device, or if it is predicted or determined that the first user can directly hear the second user, preventing (105) the first user device from reproducing the audio captured by the second user device when the first user device reproduces the audio captured by the third user device.

16. A computer program or computer program suite comprising at least one software code portion, or a computer program product storing at least one software code portion, configured for performing the method of claim 15 when run on a computer system.

Citation Information

Patent Citations

  • Noise suppression system and method

    EP3175456A1