Spatial audio for video telephony
By identifying and adjusting the acoustic position of participants' devices in a video call, selecting the main device, and separating the audio signal, the problem of inaccurate spatial audio rendering caused by audio leakage was solved, achieving more accurate alignment of the audio signal direction with the image position and improving the audio rendering effect.
Patent Information
- Application Number
- CN202511115571.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-13
- Filing Date
- 2025-08-11
- Publication Date
- 2026-02-13
AI Technical Summary
In video calls, audio leakage between participants using different devices causes a mismatch between the direction of the audio signal and the position of the image in the composite image, affecting the spatial audio rendering effect.
By identifying the acoustic location of the participant's device, selecting the main participant's device, separating the audio signal into parts, associating them with the image, adjusting the direction of the audio signal to compensate for leakage, and finally rendering it to the corresponding location in the composite image.
It improves the alignment between the audio signal and the image position in the synthesized image, reduces the impact of audio leakage, and improves the spatial audio rendering effect.
Smart Images

Figure CN121531291A_ABST
Abstract
Description
Technical Field
[0001] Examples of this disclosure relate to spatial audio for video telephony. Some examples relate to spatial audio for video telephony in which two or more participants use different devices to participate in the video call, but are in the same acoustic location. Background Technology
[0002] Spatial audio can be used in video calls. Using spatial audio allows the spatial characteristics of the sending participants to be rendered, enabling the receiving participant to perceive the different locations of the sending participants. Summary of the Invention
[0003] According to various, but not necessarily all, examples of this disclosure, an apparatus for enabling video telephony can be provided, including components for the following operations:
[0004] Identify two or more participant devices in a video call at an acoustic location, wherein the participant devices provide images for one or more receivers in the video call to synthesize an image, and at least one identified participant device provides an audio signal to the video call such that the audio signal corresponds to the synthesized image;
[0005] Determine the location of images from the identified participant's device within the synthesized image;
[0006] Associate audio signals from two or more participant devices with images from the identified participant devices in a synthetic image;
[0007] The direction of the audio signals from the identified participant devices is adjusted to increase the angle with the center in order to compensate for audio leakage between the identified participant devices; and
[0008] Render the audio from the identified participant's device to the adjusted orientation.
[0009] At least one participant device can be used by multiple users.
[0010] The degree of adjustment in the direction of the audio signal can depend on the amount of audio leakage.
[0011] Audio leakage can be determined based on the correlation of audio signals from the identified participant's devices.
[0012] Acoustic echo cancellation can be performed on audio signals from the identified participant's device.
[0013] The component can be used for:
[0014] Identify two or more participant devices in a video call at an acoustic location, wherein the participant devices provide images for a composite image for one or more receivers in the video call;
[0015] Select a participant device from the identified participant devices at the acoustic location to be used as the primary participant device for the acoustic location;
[0016] Separate the audio signal from the main participant's device into parts;
[0017] The separated portions of the audio signal are correlated with images from the identified participant's device;
[0018] Determine the location of images from the identified participant's device within the synthesized image; and
[0019] The orientation of the separated portions of the audio signal is rendered to the location of the image from the identified participant's device in the synthetic image.
[0020] Audio signals from participant devices other than the primary participant device can be used to enhance the spatial rendering of audio signals from the primary participant device.
[0021] Audio signals from participant devices other than the primary participant device can be muted in one or more receiving participant devices.
[0022] Selecting participant devices from the identified participant devices may include selecting the participant device that provides the audio signal with the highest signal-to-noise ratio.
[0023] Different parts of an audio signal may include at least one of the following:
[0024] Different objects,
[0025] Different time frames, or
[0026] Different time-frequency blocks.
[0027] At least one of the following can be used to separate an audio signal into parts:
[0028] Blind source separation; or
[0029] Time-frequency transformation.
[0030] Associating a separated portion of the audio signal with an image from the identified participant's device may include at least one of the following:
[0031] Lip-sync detection
[0032] Speaker recognition,
[0033] The correlation between various audio signals,
[0034] Sound energy in the desired direction, or
[0035] Image classification.
[0036] A portion of the audio signal from the primary participant's device can be associated with each identified participant's device at the acoustic location.
[0037] A portion of the audio signal from the primary participant's device can be associated with each identified participant's device that provides video for the synthesized image.
[0038] Rendering of a direction may include at least one of the following:
[0039] Translation;
[0040] Diaurization; or
[0041] Panning of Ambisonics surround sound.
[0042] The device may be provided with at least one of the following:
[0043] Receive participant's device; or
[0044] Server equipment.
[0045] Based on various, but not necessarily all, examples of this disclosure, a method may be provided, including:
[0046] Identify two or more participant devices in a video call at an acoustic location, wherein the participant devices provide images for one or more receivers in the video call to synthesize an image, and at least one identified participant device provides an audio signal to the video call such that the audio signal corresponds to the synthesized image;
[0047] Determine the location of images from the identified participant's device within the synthesized image;
[0048] Associate audio signals from two or more participant devices with images from the identified participant devices in a synthetic image;
[0049] The direction of the audio signals from the identified participant devices is adjusted to increase the angle with the center in order to compensate for audio leakage between the identified participant devices; and
[0050] Render the audio from the identified participant's device to the adjusted orientation.
[0051] According to various, but not necessarily all, examples of this disclosure, a computer program comprising instructions that, when executed by a device, cause the device to perform:
[0052] Identify two or more participant devices in a video call at an acoustic location, wherein the participant devices provide images for one or more receivers in the video call to synthesize an image, and at least one identified participant device provides an audio signal to the video call such that the audio signal corresponds to the synthesized image;
[0053] Determine the location of images from the identified participant's device within the synthesized image;
[0054] Associate audio signals from two or more participant devices with images from the identified participant devices in a synthetic image;
[0055] The direction of the audio signals from the identified participant devices is adjusted to increase the angle with the center in order to compensate for audio leakage between the identified participant devices; and
[0056] Render the audio from the identified participant's device to the adjusted orientation.
[0057] According to various, but not necessarily all, examples of this disclosure, an apparatus for enabling video telephony can be provided, the apparatus including components for the following operations:
[0058] Identify two or more participant devices in a video call at an acoustic location, wherein the participant devices provide images for a composite image for one or more receivers in the video call;
[0059] Select a participant device from the identified participant devices at the acoustic location to be used as the primary participant device for the acoustic location;
[0060] Separate the audio signal from the main participant's device into parts;
[0061] The separated portions of the audio signal are correlated with images from the identified participant's device;
[0062] Determine the location of images from the identified participant's device within the synthesized image; and
[0063] The orientation of the separated portions of the audio signal is rendered to the location of the image from the identified participant's device in the synthetic image.
[0064] According to various, but not necessarily all, embodiments, an apparatus is provided, comprising:
[0065] At least one processor; and
[0066] At least one memory including computer program code;
[0067] At least one memory storage instruction, when executed by at least one processor, causes the device to perform at least a portion of one or more methods described herein.
[0068] According to various, but not necessarily all, embodiments, an apparatus is provided that includes components for performing at least a portion of one or more methods described herein. Furthermore, the description of functions and / or actions should be considered as also disclosing any means suitable for performing those functions and / or actions. The functions and / or actions described herein can be performed using any suitable method and in any suitable manner.
[0069] Examples as claimed in the appended claims are provided according to various, but not necessarily all, embodiments.
[0070] Although the examples and optional features of this disclosure have been described separately, it will be understood that their provision in all possible combinations and permutations is included within this disclosure. It will be understood that the various examples of this disclosure may include any or all of the features described with respect to other examples of this disclosure, and vice versa. Furthermore, it will be understood that any one or more features (in any combination) may be implemented / included / performed by means of apparatus, method, and / or computer program instructions as needed and suitably. Moreover, the description of the function should be considered as also disclosing any means suitable for performing that function. Attached Figure Description
[0071] Some examples will now be described with reference to the accompanying drawings, in which:
[0072] Figure 1 An example system is shown;
[0073] Figure 2 Another example system is shown;
[0074] Figure 3 Example methods are shown;
[0075] Figure 4A and Figure 4B An example use case scenario is shown;
[0076] Figure 5 Example methods are shown;
[0077] Figures 6A to 6E Example use case scenarios are shown; and
[0078] Figure 7 An example device is shown.
[0079] The accompanying drawings are not necessarily to scale. For clarity and brevity, some features and views in the drawings may be shown schematically or enlarged to scale. For example, the dimensions of some elements in the drawings may be enlarged relative to other elements to aid illustration. Corresponding reference numerals are used in the drawings to identify corresponding features. For clarity, not all reference numerals may be shown in all drawings. Detailed Implementation
[0080] Figure 1 An example system 100 that can be used in the examples of this disclosure is shown. System 100 can be used for video telephony, in which both audio and video are transmitted between respective participant devices 102.
[0081] System 100 includes multiple participant devices 102. Participant devices 102 may include any device configured to enable users of the device to participate in video calls. Participant devices 102 may be telephones, tablets, soundbars, microphone arrays, cameras, computing devices, teleconferencing equipment, televisions, virtual reality (VR) / augmented reality (AR) devices, or any other suitable type of device.
[0082] exist Figure 1 In the diagram, participant device 102 is shown as three transmitting participant devices 102A, 102B, and 102C, and one receiving participant device 102D. Thus, system 100 is shown to indicate how multiple signals can reach a single participant device 102. In implementations of this disclosure, transmitting participant devices 102A, 102B, and 102C can also receive signals from other participant devices, and receiving participant device 102D can also transmit signals to other participant devices.
[0083] Transmitting participant devices 102A, 102B, and 102C are configured to transmit audio signals 106A, 106B, and 106C and video signals 108A, 108B, and 108C to receiving participant device 102D. Transmitting participant devices 102A, 102B, and 102C may include any suitable components for capturing the audio signals 106A, 106B, and 106C and the video signals 108A, 108B, and 108C, such as microphones and cameras.
[0084] Receiving participant device 102D is configured to receive audio signals 106A, 106B, 106C and video signals 108A, 108B, 108C from transmitting participant devices 102A, 102B, 102C. Receiving participant device 102D is configured to process the received signals and render them for playback to a user of participant device 102D. The rendered audio signals can be played using a speaker, headphones, or any other suitable component. The video signals can be played using one or more displays.
[0085] The teleconferencing system 100 may include means for processing relevant signals. Figure 7 Example device 700 is shown in the figure. Figure 1 In the example, device 700 can be located within receiving participant device 102D.
[0086] Figure 2 Another example system 100 is shown, which also includes three transmitting participant devices 102A, 102B, and 102C and one receiving participant device 102D. Figure 2 The system 100 shown is with Figure 1 The difference in system 100 is that it includes a teleconferencing server 200. The teleconferencing server 200 is configured to receive audio signals 106A, 106B, 106C and video signals 108A, 108B, 108C from sending participant devices 102A, 102B, 102C. The teleconferencing server 200 is configured to process the received signals and transmit the processed audio signal 202A and the processed video signal 202B to receiving participant device 102D. A means 700 for processing the signals may be provided within the teleconferencing server 200.
[0087] The processing performed on the video signal 108 may include combining the corresponding video signals to generate a composite image. The composite image may be a larger image composed of component images, wherein the component images correspond to the video signals 108 from the respective transmitting participant devices 102A, 102B, and 102C. This processing may be performed by, for example... Figure 1 The receiving participant device 102D in the system 100 shown performs the action, or in a manner such as Figure 2 It may be executed in the teleconference server 200 of the system shown, or in any other suitable part of the system 100.
[0088] The processing performed on audio signal 106 may include spatial audio rendering such that the audio associated with each transmitting participant device 102A, 102B, 102C is rendered in an orientation corresponding to the position of the relevant image within the composite image. For example, if the image from the third participant device 102C is located on the right side of the composite image, the audio signal 106C from the third participant device 102C will be rendered so that it is perceived as originating from the right. This aligns the orientation of the audio signal 106 with the position of the relevant image within the composite image. This processing may be performed by, for example... Figure 1 The receiving participant device 102D in the system 100 shown performs the action, or in a manner such as Figure 2 It may be executed in the teleconference server 200 of the system shown, or in any other suitable part of the system 100.
[0089] exist Figure 1 and Figure 2 In the system 100 shown, a first transmitting participant device 102A is located at a first acoustic position 104A, and a second participant device 102B and a third participant device 102C are located at a second acoustic position 104B. The acoustic position 104 includes the area or environment surrounding the participant device 102, from which the participant device 102 can detect acoustic signals. The acoustic position 104 can be a room or other enclosed space or any other environment.
[0090] exist Figure 1 and Figure 2 In this configuration, two participant devices 102B and 102C are present at the same second acoustic location 104B. This can lead to audio leakage, where audio from a user of the second participant device 102B is captured by a third participant device 102C and / or audio from a user of the third participant device 102C is captured by the second participant device 102B.
[0091] Audio leakage can cause errors in spatial audio rendering, resulting in a mismatch between the direction of the audio signal and the image orientation in the synthesized image. For example, the position of an image within the synthesized image may not be the same as its relative position within the acoustic location. The second participant device 102B may be located to the right of the third participant device 102C at acoustic location 104, but the synthesized image may be arranged such that the image from the second participant device 102B is to the left of the image from the third participant device 102C. Therefore, any audio from the user captured by the third participant device 102C from the second participant device 102B will be rendered in an orientation not aligned with the image.
[0092] Examples of this disclosure provide methods and apparatus for rendering audio signals to improve alignment with images in a composite image and reduce the effects of audio leakage.
[0093] Figure 3 An example method is shown. This method can be used in, for example... Figure 1 or Figure 2 The method may be implemented in system 100 as shown or in any other suitable system 100. The method may be implemented by device 700 or any other suitable component. The component used to implement the method may be located within receiving participant device 102, server device (e.g., teleconference server 200), or any other suitable device or combination of devices.
[0094] In box 300, the method includes: identifying two or more participant devices 102 of a video call at an acoustic location. The video call can be a conference call between multiple participants or any other suitable type of video call.
[0095] Acoustic location 104 is a real-world location where audio from a user of the first participant device 102 can leak into an audio signal from the second participant device 102. Acoustic location 104 can be a room or other enclosed space or any other environment where participant device 102 may be close enough to detect acoustic signals from other users.
[0096] Any suitable means can be used to identify two or more participant devices 102 within the same acoustic location 104. For example, the correlation between audio signals from participant devices 102 can be calculated. If the correlation is high or exceeds a threshold, it can be assumed that the participant devices 102 are in the same acoustic location.
[0097] Participant device 102 provides images for synthesizing an image for one or more receivers in a video call. The synthesized image includes images from multiple participant devices 102. The multiple participant devices 102 may include participant devices 102 located at the same acoustic location 104, and one or more other participant devices 102 that may be located at different acoustic locations. Images may be located at different locations within the synthesized image. For example, an image from a first participant device 102 may be provided on the right side of the synthesized image, and an image from a second participant device 102 may be provided on the left side of the synthesized image. The positions of the images within the synthesized image do not need to correspond to or be determined by the relative positions of the participant devices 102 within the acoustic location 104.
[0098] In block 302, the method includes: selecting a participant device 102 from the identified participant devices 102 at acoustic location 104 to serve as the primary participant device 102 at acoustic location 104. The primary participant device 102 may be a participant device 102 used to provide audio signals from acoustic location 104 to other participant devices 102 in a video call but not at the same acoustic location 104. Participant devices 102 identified as being at the same acoustic location 104 but not selected as the primary participant device 102 may be muted during the video call.
[0099] Any suitable criterion may be used to select the primary participant device 102. In some examples, the participant device 102 with the highest signal-to-noise ratio or the participant device 102 that provides the loudest audio signal 106 may be selected, or any other suitable criterion or combination of criteria may be used.
[0100] In block 304, the method includes separating an audio signal from the primary participant device 102 into portions. The different portions of the audio signal may be based on different users or participants within acoustic location 104. The different portions of the audio signal may include different objects, different time frames, or different time-frequency blocks, or any other suitable different portions.
[0101] Any suitable means may be used to separate the audio signal from the primary participant device 102 into parts. In some examples, separating the audio signal into parts may be performed using blind source separation, time-frequency transformation, or any other suitable process.
[0102] In block 306, the method includes associating separated portions of an audio signal with images from the identified participant device 102. This association can determine which / which separated portions of the audio signal correspond to the user of the respective participant device 102. For example, it can determine which voice comes from the user of a first participant device 102 and which voice comes from the user of a second participant device 102.
[0103] Any suitable means may be used to correlate a separated portion of the audio signal with an image from the identified participant device 102, such as lip-sync detection, speaker recognition, correlation of the corresponding audio signal, acoustic energy in a desired direction, or image classification. The desired direction may be a frontal orientation or any other suitable orientation.
[0104] In some examples, a portion of the audio signal from the primary participant device 102 is associated with each of the identified participant devices 102 at acoustic location 104. In some examples, a portion of the audio signal from the primary participant device 102 is associated with each of the identified participant devices 102 at acoustic location 104 that provides video for the component images. This allows a particular audio signal to be associated with each image in the component images.
[0105] In box 308, the method includes determining the position of an image from the identified participant device within the synthesized image. For example, this may include determining whether the image from participant device 102 is located at the center of the synthesized image, or oriented to the left or right. In some examples, this may include determining the angular position of the corresponding image within the synthesized image.
[0106] In block 310, the method includes rendering the orientation of a separated portion of the audio signal to the position of the image from the identified participant device in the synthesized image. The orientation to which the separated portion of the audio signal is rendered is determined by the position of the corresponding image in the synthesized image, rather than the position of the participant device 102 at acoustic position 104. This can result in the separated portion of the audio signal being rendered in an orientation different from the real-world orientation. For example, a non-primary participant device 102 may be located to the right of the primary participant device 102 at acoustic position 104, but the synthesized image may be arranged such that the image from the non-primary participant device 102 is to the left of the image from the primary participant device 102. In the examples of this disclosure, portions of the audio signal associated with the non-primary participant device 102 are rendered to the left so that they are aligned with the associated image, rather than due to the real-world position.
[0107] Directional rendering can include: panning, binaural panning, panoramic surround sound panning, or any other suitable rendering.
[0108] In some examples, audio signals from participant device 102 other than the primary participant device can be muted. Audio signals can be muted in the receiving participant device. In some examples, audio signals from participant device 102 other than the primary participant device 102 can be used to enhance the spatial rendering of audio signals from the primary participant device.
[0109] Figure 4A and Figure 4B An example use case scenario is shown. This can be used... Figure 3 The method shown and / or any suitable variation thereof.
[0110] In this scenario, the video call is configured to prevent audio from multiple participant devices 102 at the same acoustic location 104. The reason for blocking audio from multiple participant devices 102 at the same acoustic location 104 could be because acoustic echo cancellation is not used or is ineffective, or to limit the number of channels to be manipulated and processed by system 100, or for any other reason. In this case, a primary participant device 102 is selected, and this primary participant device 102 is used to provide audio from all participants at the same acoustic location.
[0111] Figure 4A An example video call is shown. In this case, the video call is between multiple family members. Grandmother 400 is located at a first acoustic location 104A. The first acoustic location can be a room in Grandmother's house or any other suitable location. Grandmother 400 is participating in the video call using participant device 102A.
[0112] The other participants in the video call are the first child 408, the second child 410, and the mother 412. The two children, 408 and 410, are sharing participant device 102B to participate in the video call. The mother 412 is using her own participant device 102C to participate in the video call. The children's participant device 102B and the mother's participant device 102C are separate and independent devices. No connection is required between the children's participant device 102B and the mother's participant device 102C, as long as they are in the same video call and at the same second acoustic location 104B.
[0113] Two children 408, 410 and mother 412 are located at the same second acoustic position 104B. This could be the same room in their home or any other suitable location. In this example, in the real world, mother 412 is located to the right of children 408, 410 (from the viewer's perspective of the second acoustic position 104B). Audio from mother 412, captured by the children's participant device 102B, will reach the children's participant device 102B from the right side of the participant device 102B (from the viewer's perspective of the acoustic position 104B).
[0114] Participant devices 102B and 102C for children 408, 410 and mother 412 capture video signals for use in a video call. In this example, participant devices 102B and 102C can be arranged to capture images of the users of participant devices 102B and 102C such that the video signal from the child's participant device 102B includes images of the children 408 and 410, and the image from the mother's participant device 102C includes an image of the mother 412.
[0115] The video signal is processed to generate a composite image 402. The composite image 402 is used for display on the receiving participant device 102A. In this example, the composite image 402 is displayed on the screen of the grandmother's participant device 102A.
[0116] The composite image 402 includes multiple component images 404. The different component images 404 are based on video signals from the respective transmitting participant devices 102B and 102C. In this use case scenario, the composite image 402 includes a first component image 404A and a second component image 404B. The first component image 404A includes an image from the mother's participant device 102C, and the second component image 404B includes an image from the child's participant device 102B.
[0117] The corresponding component images 404A and 404B are displayed at different locations within the composite image. Figure 4AIn the example, the first component image 404A is displayed on the left side of the composite image 402, and the second component image 404B is displayed on the right side of the composite image 402.
[0118] The participant devices 102B and 102C for children 408, 410, and mother 412 also capture audio signals. However, since children 408, 410, and mother 412 are located at the same second acoustic position 104B, the child's participant device 102B is used to capture audio from mother 412. The audio signal from the mother's participant device 102C can be muted or used to enhance the rendering of the audio from the child's participant device 102B.
[0119] The audio signal is processed and rendered for playback by the grandmother's participant device 102A. The grandmother's participant device 102A includes multiple speakers 406A, 406B that can be used to play spatial audio.
[0120] Figure 4A Spatial audio rendering 414 is shown. When the audio from the child participant device 102B is spatially rendered, the audio from the mother 412 appears to be to the right of the children 408, 410, because in the actual second acoustic position 104B, the mother 412 is to the right of the children 408, 410. Therefore, spatial audio rendering 414 is misaligned with the composite image 402 (where the mother 412 is to the left of the children 408, 410).
[0121] Figure 4B It shows Figure 3 This describes how the method, or other similar methods, can be applied to solve the problem. The process can be implemented by the receiving participant device 102A (in this case, the grandmother's participant device 102A). In some examples, the process can be implemented by a teleconference server 200 (which can be positioned between the sending participant devices 102B, 102C and the receiving participant device 102A). In other examples, any other suitable device or combination of devices can be used.
[0122] In this case, the participant devices 102B and 102C of children 408, 410 and mother 412 are identified as being at the same second acoustic position 104B. The correlation between the audio signals from the participant devices 102B and 102C of children 408, 410 and mother 412 can be used to determine that they are at the same second acoustic position 104B, and / or any other suitable process can be used.
[0123] The child's participant device 102B can be selected as the primary participant device 102 of the second acoustic position 104B. The child's participant device 102B can be selected because it has the best signal-to-noise ratio, because it has the loudest audio signal, or based on any other criterion or combination of criteria. The audio signal from the mother's participant device 102C can be muted, or it can be used to improve the spatial rendering of the audio signal from the child's participant device 102B.
[0124] In box 420, the audio signal 106 from the child participant device 102B is separated into portions. The corresponding portions can be audio objects. Audio objects can represent audio from different users; for example, a first audio object could correspond to audio from children 408 and 410, and a second audio object could correspond to audio from the mother 412. In other examples, other methods of separating the audio signal 106 into portions can be used.
[0125] In box 422, video signals 108 from the child's participant device 102B and the mother's participant device 102C are processed to find the corresponding image for each separate portion of the audio signal. In this case, an image is found for each audio object.
[0126] For example, video signal 108 from the child participant device 102B can provide a first image, and video signal from the mother participant device 102C can provide a second image. In box 422, the audio object corresponding to the respective image is identified. In some examples, lip-sync detection can be used to determine the image corresponding to the audio object. In this case, the image can be analyzed to determine whether the child's lips are moving or the mother's lips are moving, and the audio can be matched with the lip movement. In some examples, the image corresponding to the audio object can be determined based on speaker recognition. For example, the speech in the audio object can be matched with the child 408, 410, or the mother 412. In other examples, other processes can be used, such as the correlation of the corresponding audio signals, the sound energy in the corresponding direction, image classification, or any other suitable means. Image classification can include identifying objects in the image (e.g., people or animals within the image) and matching them with associated sounds.
[0127] An image of each audio object can be found using any suitable means. In some examples, machine learning models can be trained and used to find images of audio objects.
[0128] In box 424, the separated audio objects are rendered in the direction corresponding to the images in the composite image. In this case, the second component image 404B from the child participant device 102B is located to the right of the composite image 402. Therefore, the audio object found to correspond to the second component image 404B from the child participant device 102B is also rendered to the right. The first component image 404A from the mother participant device 102C is located to the left of the composite image 402. Therefore, the audio object found to correspond to the first component image 404A from the mother participant device 102C is also rendered to the left. Thus, the spatial audio rendering 414 is adjusted to account for audio leakage, so that the direction of the perceived audio corresponds to the position of the corresponding image in the composite image.
[0129] Figure 5 An example method is shown, which can be used in, for example Figure 1 or Figure 2 The method can be implemented in the system 100 shown or any other suitable system 100. The method can be implemented by the device 700 or any other suitable component. The component used to implement the method can be located within the receiving participant device 102, a server device (e.g., a teleconference server 200), or any other suitable device or combination of devices.
[0130] In box 500, the method includes identifying two or more participant devices 102 in a video call at acoustic location 104. The video call can be a conference call between multiple participants or any other suitable type of video call. The two or more participant devices 102 in the same location can be identified using any suitable means, such as the correlation between audio signals from each participant device 102. If the correlation is high or exceeds a threshold, it can be assumed that the participant devices 102 are in the same acoustic location.
[0131] Participant device 102 provides audio signals for the video call and also provides images for one or more receivers in the video call to synthesize the image. In this example, audio signals from two or more participant devices 102 at acoustic location 104 can be used for audio to the receivers. The participant devices 102 within acoustic location 104 do not need to be muted.
[0132] The synthesized image includes images from multiple participant devices 102. The multiple participant devices 102 may include participant devices 102 located at the same acoustic location 104, and one or more other participant devices 102 that may be located at different acoustic locations. The images may be located at different positions within the synthesized image. For example, an image from a first participant device 102 may be provided on the right side of the synthesized image, and an image from a second participant device 102 may be provided on the left side of the synthesized image. The position of the image within the synthesized image does not need to correspond to or be determined by the relative position of the participant devices 102 within the acoustic location 104.
[0133] The audio signal provided by one or more participant devices 102 corresponds to the synthesized image. The audio signal and the image may be provided together in the combined signal. The image may include an image of the acoustic location.
[0134] In box 502, the method includes determining the position of images from the identified participant device 102 within the synthesized image. This may include determining whether the images from the participant device 102 are located at the center of the synthesized image or oriented to the left or right. In some examples, this may include determining the angular position of each image within the synthesized image.
[0135] In block 504, the method includes associating an audio signal from a participant device with an image from the participant device in a synthesized image. The image may be associated with its corresponding audio signal such that the audio signal originating from the first participant device is associated with the orientation of the corresponding image in the synthesized image.
[0136] In block 506, the method includes adjusting the orientation of audio signals from the identified participant devices 102 to increase the angle with respect to the center in order to counteract audio leakage between the identified participant devices 102. For example, if the audio signal from the first participant device 102 is rendered to align with the position of the image from the first participant device 102 within the synthesized image, the effect of audio leakage could cause the perceived sound to be incorrectly aligned with the image. Increasing the angle with respect to the center can direct the audio signal away from the center or the frontal direction.
[0137] In box 508, the method includes: rendering audio from the identified participant's device to the adjusted orientation.
[0138] Figures 6A to 6E Another example use case scenario is shown. This can be used... Figure 5 The method shown and / or any suitable variation thereof.
[0139] In this scenario, the video call is configured to allow audio from multiple participant devices 102 located at the same acoustic position 104. Acoustic echo cancellation or other noise reduction techniques can be employed effectively in this case.
[0140] Figure 6A An example video call is shown. In this case, the video call is between multiple family members. Grandmother 400 is located at a first acoustic location 104A. The first acoustic location can be a room in Grandmother's house or any other suitable location. Grandmother 400 is participating in the video call using participant device 102A.
[0141] The other participants in the video call are the first child 408, the second child 410, and the mother 412. The two children 408 and 410 are sharing participant device 102B to participate in the video call. The mother 412 is using her own participant device 102C to participate in the video call. The children's participant device 102B and the mother's participant device 102C are separate, independent devices. No connection is required between the children's participant device 102B and the mother's participant device 102C, as long as they are in the same video call and at the same second acoustic location 104B.
[0142] The two children 408 and 410 and the mother 412 are located in the same second acoustic position 104B. This could be the same room in their home or any other suitable location. There may be an acoustic leakage 600 between the children's participant device 102B and the mother's participant device 102C. That is, audio from the children 408 and 410 can be detected by the mother's participant device 102C, and / or audio from the mother 412 can be detected by the children's participant device 102B.
[0143] In this example, in the real world, the mother 412 is located to the right of the children 408, 410 (from the perspective of a viewer looking at the second acoustic position 104B). Audio from the mother 412, captured by the child's participant device 102B, will arrive at the child's participant device 102B from the right (from the perspective of a viewer looking at the second acoustic position 104B). Similarly, audio from the children 408, 410, captured by the mother's participant device 102C, will arrive at the mother's participant device 102C from the left (from the perspective of a viewer looking at the second acoustic position 104B).
[0144] The participant devices 102B and 102C for children 408, 410 and mother 412 also capture video signals for use in video calls. In this example, the participant devices 102B and 102C can be arranged to capture images of the users of the participant devices 102B and 102C, such that the video signal from the child's participant device 102B includes images of the children 408 and 410, and the image from the mother's participant device 102C includes an image of the mother 412.
[0145] The video signal is processed to generate a composite image 402. The composite image 402 is used for display on the receiving participant device 102A. In this example, the composite image 402 is displayed on the screen of the grandmother's participant device 102A.
[0146] The composite image 402 includes multiple component images 404. The different component images 404 are based on video signals from the respective transmitting participant devices 102B and 102C. In this use case scenario, the composite image 402 includes a first component image 404A and a second component image 404B. The first component image 404A includes an image from the mother's participant device 102C, and the second component image 404B includes an image from the child's participant device 102B.
[0147] The component images 404A and 404B are displayed at different locations within the composite image. Figure 6A In the example, the first component image 404A is displayed on the left side of the composite image 402, and the second component image 404B is displayed on the right side of the composite image 402.
[0148] Figure 6B The diagram illustrates how audio from the sending participant devices 102B and 102C should be rendered for the grandmother's participant device 102A. Figure 6B The scene shown illustrates how audio will be rendered without audio leakage. The audio from the mother participant device 102C can be mono and should be rendered to a sector aligned with the image from the mother participant device 102C. This is determined by… Figure 6B Point 602 is indicated in the text.
[0149] The audio from the child participant device 102B can be spatial audio and should be rendered to a sector aligned with the image from the child participant device 102B. This is determined by... Figure 6B Points 604 and 606 are indicated in the text.
[0150] Figure 6CAudio leakage is shown. Points 608 and 610 show the audio leaked into the audio captured by the mother's participant device 102C from the children 408 and 410. These points 608 and 610 are smaller than the points representing the audio from the mother 412 because the children 408 and 410 are quieter to the mother's participant device 102C (because they are farther away).
[0151] Point 612 shows the audio from the mother 412 leaked into the audio captured by the child's participant device 102B. Similarly, point 612 is smaller than points 604 and 606, which represent the audio from the children 408 and 410, because the mother 412 is quieter to the child's participant device 102B (because she is farther away).
[0152] Figure 6D This illustrates the effect of audio leakage on the actual direction to which audio is rendered. Audio leakage shifts the angle at which audio is rendered towards the center. This results in a perceived sound image that is narrower than expected, and the direction of the audio is not correctly aligned with the direction of the corresponding object in the image.
[0153] To solve this problem, the following can be implemented: Figure 5 The method or any other suitable method is used to determine the position of the images from each transmitting participant device 102B, 102C in the composite image 402. Then, the direction of the audio signal from the transmitting participant device 102C is adjusted to compensate for audio leakage.
[0154] In some examples, adjusting the direction of the audio signals may include determining a horizontal average of the direction of the audio signals from the transmitting participant device 102C. Audio signals determined to be to the right of the average are adjusted so that they are even further to the right. Similarly, audio signals determined to be to the left of the average are adjusted so that they are even further to the left. The average position may be a center position or any other suitable angle.
[0155] As an example, if a video call includes three transmitting participant devices 102 at the same acoustic location 104, the orientation of their respective images in the composite image 402 can be represented as α. i Let i = 1, …, N. The adjusted direction of the audio signal will be:
[0156] ,
[0157] The factor G is an adjustment factor. G can be fixed and can be within a defined range, such as 0.1, ..., 0.5. In some examples, the adjustment factor G can be variable and can depend on the amount of audio leakage. If there is more audio leakage, the adjustment factor G can be larger, and if there is less audio leakage, the adjustment factor G can be smaller. The amount of audio leakage can be estimated based on the correlation between audio signals from participant devices 102 at the same acoustic location 104.
[0158] Figure 6E The effect of adjusting the direction of the audio signals is illustrated. This shows that the audio from the mother's participant device 102C (as indicated by point 602) shifts to the left of the composite image 402, and the audio from the child's participant device 102B (as indicated by points 604 and 606) shifts to the right of the composite image 402. This expands the perceived sound image and counteracts the effects of audio leakage. This results in improved perceived audio for the grandmother 402 because the audio is better aligned with the component images 404A, 404B.
[0159] Figure 5 and Figures 6A to 6E The method shown can be used in situations where there is no feedback between participant devices 102 at the same acoustic location 104. This may occur if acoustic echo cancellation is used and effective. In some cases, the presence of feedback can be detected. If feedback is detected, it can be used... Figure 3 and Figures 4A to 4B Instead of using Figure 5 and Figures 6A to 6E The method.
[0160] Figure 7 An apparatus 700 that can be used to implement an example of this disclosure is schematically shown. In this example, apparatus 700 includes a controller 702. The controller 702 may be a chip or a chipset. In some examples, the controller 702 may be located within participant device 102 or teleconference server 200 or any other suitable type of device.
[0161] exist Figure 7 In the examples, controller 702 can be implemented as a controller circuit. In some examples, controller 702 can be implemented solely in hardware, solely in certain aspects of software (including firmware), or can be a combination of hardware and software (including firmware).
[0162] like Figure 7As shown, the controller 702 can be implemented using instructions that enable hardware functions, for example by using executable instructions of a computer program 708 in a general-purpose or special-purpose processor 704, which can be stored on a computer-readable storage medium (disk, memory, etc.) for execution by such processor 704.
[0163] Processor 704 is configured to read from and write to memory 706. Processor 704 may also include an output interface through which data and / or commands are output by processor 704; and an input interface through which data and / or commands are input to processor 704.
[0164] Memory 706 is configured to store a computer program 708, including computer program instructions (computer program code 710), which controls the operation of controller 702 when loaded into processor 704. The computer program instructions of computer program 708 provide logic and routines that enable controller 702 to perform the methods shown in the figures. By reading memory 706, processor 704 is able to load and execute computer program 708.
[0165] Therefore, the device 700 includes: at least one processor 704; and at least one memory 706 including computer program code 710, the at least one memory 706 and the computer program code 710 being configured together with the at least one processor 704 to cause the device 700 to perform at least the following:
[0166] Two or more participant devices 102 are identified in a video call at acoustic location 104, wherein the participant devices 102 provide images for one or more receivers in the video call to synthesize an image, and at least one identified participant device 102 provides an audio signal to the video call such that the audio signal corresponds to the synthesized image.
[0167] Determine the location of the image from the identified participant device 102 in the 502 composite image;
[0168] 504. Associate audio signals from two or more participant devices 102 with images from the identified participant devices 102 in the synthetic image.
[0169] Adjusting the direction of the audio signal from the identified participant device 102 to increase the angle with the center in order to compensate for audio leakage between the identified participant devices 102; and
[0170] The audio from the identified participant's device 102 is rendered 508 to the adjusted orientation.
[0171] like Figure 7As shown, computer program 708 can reach controller 702 via any suitable transmission mechanism 712. Transmission mechanism 712 can be, for example, a machine-readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a memory device, a recording medium (e.g., an optical disc read-only memory (CD-ROM) or a digital versatile optical disc (DVD) or solid-state memory), or an article of manufacture that includes or tangibly embodies computer program 708. Transmission mechanism 712 can be a signal configured to reliably transmit computer program 708. Controller 702 can propagate or transmit computer program 708 as a computer data signal. In some examples, wireless protocols (e.g., Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWpanan (Low Power Personal Area Network IP)) can be used. v 6) The computer program 708 is sent to the controller 702 via ZigBee, ANT+, Near Field Communication (NFC), Radio Frequency Identification, Wireless Local Area Network (Wireless LAN) or any other suitable protocol.
[0172] Computer program 708 includes computer program instructions that, when executed by device 700, cause device 700 to perform at least the following operations:
[0173] Two or more participant devices 102 are identified in a video call at acoustic location 104, wherein the participant devices 102 provide images for one or more receivers in the video call to synthesize an image, and at least one identified participant device 102 provides an audio signal to the video call such that the audio signal corresponds to the synthesized image.
[0174] Determine the location of the image from the identified participant device 102 in the 502 composite image;
[0175] 504. Associate audio signals from two or more participant devices 102 with images from the identified participant devices 102 in the synthetic image.
[0176] Adjusting the direction of the audio signal from the identified participant device 102 to increase the angle with the center in order to compensate for audio leakage between the identified participant devices 102; and
[0177] The audio from the identified participant's device 102 is rendered 508 to the adjusted orientation.
[0178] Computer program instructions may be included in computer program 708, non-transitory computer-readable medium, computer program product, or machine-readable medium. In some, but not necessarily all, examples, computer program instructions may be distributed across multiple computer programs 708.
[0179] Although memory 706 is shown as a single component / circuit, it can be implemented as one or more separate components / circuits, some or all of which can be integrated / removable, and / or can provide permanent / semi-permanent / dynamic / cached storage.
[0180] Although processor 704 is shown as a single component / circuit, it can be implemented as one or more separate components / circuits, some or all of which can be integrated / removable. Processor 704 can be a single-core or multi-core processor.
[0181] References to “computer-readable storage medium,” “computer program product,” “computer program tangibly embodied,” or “controller,” “computer,” “processor,” etc., should be understood to include not only computers with different architectures (e.g., single / multiprocessor architectures and sequential (von Neumann) / parallel architectures) but also special-purpose circuits such as field-programmable gate arrays (FPGAs), special-purpose circuits (ASICs), signal processing devices, and other processing circuits. References to computer programs, instructions, code, etc., should be understood to include software for programmable processors or firmware, such as programmable content for hardware devices, whether instructions for processors or configuration settings for fixed-function devices, gate arrays, or programmable logic devices.
[0182] As used in this application, the term "circuit" may refer to one or more or all of the following:
[0183] (a) Hardware circuit implementation only (e.g., implementation in analog and / or digital circuits only), and
[0184] (b) A combination of hardware circuitry and software, such as (if applicable):
[0185] (i) A combination of analog and / or digital hardware circuitry with software / firmware, and
[0186] (ii) Any part of a hardware processor (including a digital signal processor), software, and memory that works together to enable a device (e.g., a mobile phone or a server) to perform various functions, and
[0187] (c) Hardware circuitry and / or processors, such as microprocessors or a portion thereof, that require software (e.g. firmware) to operate, but may be absent when operation does not require it.
[0188] The definition of "circuit" applies to all uses of the term in this application (including in any claim). As a further example, as used in this application, the term "circuit" also covers an implementation of only one hardware circuit or processor and its accompanying software and / or firmware. The term "circuit" also covers (e.g., and if applicable to a particular claim element) baseband integrated circuits for mobile devices or similar integrated circuits in servers, cellular network devices, or other computing or networking devices.
[0189] like Figure 7 The device 700 shown can be housed in any suitable device. In some examples, the device 700 can be housed in an electronic device, such as a mobile phone, teleconferencing equipment, camera, computing device, server, or any other suitable device.
[0190] The boxes shown in the accompanying drawings may represent steps in the method and / or code portions in computer program 708. A specific sequence diagram of the boxes does not necessarily imply a required or preferred order, and the order and arrangement of the boxes may be changed. Furthermore, some boxes may be omitted.
[0191] The above examples apply to components that are enabled as follows:
[0192] Automotive systems; telecommunications systems; electronic systems, including consumer electronics; distributed computing systems; media systems for generating or rendering media content (including audio, visual, and audiovisual content, as well as mixed, mediated, virtual, and / or augmented reality); personal systems, including personal health systems or personal fitness systems; navigation systems; user interfaces, also known as human-computer interfaces; networks, including cellular networks, non-cellular networks, and optical networks; self-organizing networks; the Internet of Things; the Internet of Things; virtualized networks; and related software and services.
[0193] According to the examples of this disclosure, the device can be located in an electronic device, such as a mobile terminal. However, it should be understood that a mobile terminal is merely an example of an electronic device that will benefit from the implementation examples of this disclosure, and therefore should not be construed as limiting the scope of this disclosure. Although in some implementation examples the device can be located in a mobile terminal, other types of electronic devices can readily adopt the examples of this disclosure, such as, but not limited to: mobile communication devices, handheld portable electronic devices, wearable computing devices, portable digital assistants (PDAs), pagers, mobile computers, desktop computers, televisions, gaming devices, laptop computers, cameras, video recorders, GPS devices, and other types of electronic systems. Furthermore, devices can readily adopt the examples of this disclosure regardless of their intention to provide mobility.
[0194] The term "includes" as used in this document has an inclusive meaning, not an exclusive meaning. That is, any reference to X that includes Y indicates that X may include only one Y or may include multiple Ys. If the exclusive meaning of "includes" is intended to be used, it will be made clear in the context by referring to "includes only one..." or by using "comprises".
[0195] In this description, the terms “connection,” “coupling,” and “communication,” and their derivatives, mean operational connection / coupling / communication. It should be understood that any number or combination of intermediate components (including no intermediate components) may exist to provide direct or indirect connection / coupling / communication. Any such intermediate component may include hardware and / or software components.
[0196] As used herein, the term "determine" (and its grammatical variations) can include, but is not limited to: calculation, estimation, processing, derivation, measurement, investigation, identification, lookup (e.g., searching in a table, database, or another data structure), ascertainment, etc. Furthermore, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), obtaining, etc. Additionally, "determine" can include parsing, selecting, picking, building, etc.
[0197] Various examples have been referenced in this description. Descriptions of features or functions relative to examples indicate that those features or functions exist in that example. The use of the terms "example," "for example," "may," or "possibly" in the text (whether explicitly stated or not) indicates that such a feature or function exists at least in the described example (whether or not it is described as an example), and that it may, but not necessarily, exist in some or all other examples. Therefore, "example," "for example," "may," or "possibly" refers to a specific instance within a class of examples. The characteristics of an instance can be characteristics of only that instance, or characteristics of the class, or characteristics of a subclass of the class (which includes some, but not all, instances of that class). Therefore, it is implicitly disclosed that a feature described with reference to one example but not to another may (if possible) be used as part of a working combination in that other example, but is not necessarily required to be used in that other example.
[0198] As used herein, “at least one of the following:” and “at least one of the following” and similar wording (where a list of two or more elements is connected by “and” or “or”) means at least any one element, or at least any two or more elements, or at least all elements.
[0199] Although examples have been described in the preceding paragraphs with reference to various examples, it should be understood that modifications can be made to the given examples without departing from the scope of the claims.
[0200] The features described above can be used in combinations other than those explicitly described above.
[0201] Although some features have been described with reference to certain features, these features can be performed by other features (whether or not they are described).
[0202] A description of a feature (e.g., a device or component of a device) configured to perform a function or for performing a function should also be considered to disclose a method for performing that function. For example, a description of a device configured to perform one or more actions or for performing one or more actions should also be considered to disclose a method for performing those one or more actions with or without that device.
[0203] Although features have been described with reference to some examples, these features may also exist in other examples (whether or not they are described).
[0204] The terms “a” or “the” as used in this document have an inclusive rather than an exclusive meaning. That is, any reference to X that includes one / the Y indicates that X may include only one Y or may include multiple Ys, unless the context clearly indicates the opposite. If “a” or “the” with an exclusive meaning is intended to be used, it will be clarified in the context. In some cases, “at least one” or “one or more” may be used to emphasize the inclusive meaning, but no exclusive meaning should be inferred from the absence of these terms.
[0205] The feature (or combination of features) in the claims refers to the feature (or combination of features) itself, and also to features that achieve substantially the same technical effect (equivalent features). Equivalent features include, for example, features that are variations and achieve substantially the same result in substantially the same manner. Equivalent features include, for example, features that perform substantially the same function in substantially the same manner to achieve substantially the same result.
[0206] In this description, references have been made to various examples that use adjectives or adjective phrases to describe the characteristics of the example. Such a description of a characteristic relative to the example indicates that the characteristic exists exactly as described in some examples, and substantially as described in others.
[0207] The foregoing description illustrates some examples of this disclosure, but those skilled in the art will recognize possible alternative structures and methodological features that provide functionality equivalent to specific examples of such structures and features described above, and which have been omitted from the foregoing description for the sake of brevity and clarity. However, the foregoing description should be understood to implicitly include references to such alternative structures and methodological features that provide equivalent functionality, unless such alternative structures or methodological features are explicitly excluded in the foregoing description of the examples of this disclosure.
[0208] When focusing on features deemed important in the foregoing specification, the applicant may seek protection by means of the claims for any patentable features or combinations thereof (whether emphasized or not) mentioned above and / or shown in the figures.
Claims
1. An apparatus for enabling a video phone, comprising means for: identifying two or more participant devices of a video phone at an acoustic location, wherein the participant devices providing images for a composite image for one or more recipients in the video phone, and at least one identified participant device providing an audio signal for the video phone such that the audio signal corresponds to the composite image; determining locations of images from the identified participant devices in the composite image; associating the audio signals from the two or more participant devices with the images from the identified participant devices in the composite image; adjusting directions of audio signals from the identified participant devices to increase an angle from a center so as to cancel audio leakage between the identified participant devices; and rendering audio from the identified participant devices to the adjusted directions. At least one of the participant devices is used by multiple users.
2. The apparatus of claim 1, wherein, An amount of adjustment of the directions of the audio signals depends on an amount of audio leakage.
3. The apparatus of any preceding claim, wherein, The amount of audio leakage is determined based on a correlation of audio signals from the identified participant devices.
4. The apparatus of claim 3, wherein, Acoustic echo cancellation is performed on the audio signals from the identified participant devices.
5. The apparatus of any preceding claim, wherein, The means are for:
6. The apparatus of any one of claims 1-2, wherein, identifying two or more participant devices of the video phone at an acoustic location, wherein the participant devices provide images for a composite image for one or more recipients in the video phone; selecting a participant device from the identified participant devices at the acoustic location to use as a primary participant device for the acoustic location; separating an audio signal from the primary participant device into portions; associating the separated portions of the audio signal with images from the identified participant devices; determining locations of images from the identified participant devices in a composite image; and rendering directions of the separated portions of the audio signal to the locations of images from the identified participant devices in the composite image. Audio signals from participant devices other than the primary participant device are used to enhance spatial rendering of the audio signal from the primary participant device.
7. The apparatus of claim 6, wherein, Audio signals from participant devices other than the primary participant device are muted in one or more recipient participant devices.
8. The apparatus of any one of claims 6-7, wherein, Selecting a participant device from the identified participant devices includes selecting a participant device that provides an audio signal with a highest signal-to-noise ratio.
9. The apparatus of any one of claims 6-8, wherein, Different portions of the audio signal include at least one of:
10. The apparatus of any one of claims 6-9, wherein, different objects, different time frames, or different time-frequency bins. Separating the audio signal into portions is performed using at least one of:
11. The apparatus of any one of claims 6-10, wherein, blind source separation; or a time-frequency transform. Associating the separated portions of the audio signal with images from the identified participant devices includes at least one of:
12. The apparatus of any one of claims 6-11, wherein, lip-sync detection, speaker identification, a correlation of the audio signals, sound energy in a desired direction, or a classification of the images. A portion of the audio signal from the primary participant device is associated with each of the identified participant devices at the acoustic location.
13. The apparatus of any one of claims 6-12, wherein, 14. The apparatus of any one of claims 6-13, wherein, A portion of the audio signals from the primary participant devices is associated with each identified participant device providing video for the composite image.
15. The apparatus of any one of claims 6-14, wherein, Rendering the directions includes at least one of: panning; binauralization; or ambisonic panning.
16. The apparatus of any preceding claim, wherein, The apparatus is disposed within at least one of: a receiving participant device; or a server device.
17. A method comprising: identifying two or more participant devices of a video telephony at an acoustic location, wherein the participant devices provide images for a composite image for one or more recipients in the video telephony, and at least one identified participant device provides audio signals for the video telephony such that the audio signals correspond to the composite image; determining locations of images from the identified participant devices in the composite image; associating audio signals from the two or more participant devices with the images from the identified participant devices in the composite image; adjusting directions of the audio signals from the identified participant devices to increase an angle from a center so as to cancel audio leakage between the identified participant devices; and rendering the audio from the identified participant devices to the adjusted directions.
18. A computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform: identifying two or more participant devices of a video phone at an acoustic location, wherein identifying two or more participant devices of a video telephony at an acoustic location, wherein the participant devices provide images for a composite image for one or more recipients in the video telephony, and at least one identified participant device provides audio signals for the video telephony such that the audio signals correspond to the composite image; determining locations of images from the identified participant devices in the composite image; associating audio signals from the two or more participant devices with the images from the identified participant devices in the composite image; adjusting directions of the audio signals from the identified participant devices to increase an angle from a center so as to cancel audio leakage between the identified participant devices; and rendering the audio from the identified participant devices to the adjusted directions.