Audio processing method, related apparatus and communication system

By selectively processing audio signals based on the physical distance and latency conditions between devices in virtual reality or augmented reality virtual meeting systems, the problem of repetitive sound interference is solved, thus improving the user's meeting experience.

CN119601025BActive Publication Date: 2026-03-27HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-11
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In virtual reality or augmented reality virtual meeting systems, users may experience interference from repeated audio due to attendees being too close together in the same physical space, which can negatively impact their meeting experience.

Method used

When the physical distance between the first and second devices meets certain conditions, audio signals are selectively played or processed to avoid repetitive sound interference, including playing transparent or transmitted sound, and filtering out human voices when the time delay is large.

Benefits of technology

It effectively avoids interference from repetitive sounds and improves the user's auditory experience in virtual meetings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119601025B_ABST
    Figure CN119601025B_ABST
Patent Text Reader

Abstract

The application provides an audio processing method, related devices and a communication system. When accessing a virtual conference, a first device can determine whether there are devices in the same physical space as the first device among other devices accessing the virtual conference, and the physical distance between the devices in the same physical space as the first device and the first device, and then determine whether the transmission sound from the devices in the same physical space as the first device will cause repeated sound interference. If the transmission sound will cause repeated sound interference, the first device can not play the transmission sound or mix the transmission sound with the transparent sound and then play it. This can solve the problem of repeated sound interference when multiple conference participants are in the same physical space during a virtual conference, and improve the user experience of participating in a virtual conference.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of terminals, and in particular to an audio processing method, related apparatus and a communication system. BACKGROUND

[0002] A current virtual conference system based on augmented reality (AR) or virtual reality (VR) can simulate a virtual conference room. Participants can adopt virtual images and sit around a table in the virtual conference room to interact. Each participant can wear an AR / VR device to watch the scene in the virtual conference room. In addition, the AR / VR device can perform spatial audio modeling according to sound field environment information of the virtual conference room and position relationship information of each participant in the virtual conference room, so that when a user hears the voice of a participant, the direction of the participant heard by the user is consistent with the direction of the participant seen in the virtual conference room. That is, the above spatial audio modeling can make the sound heard by the user match the virtual conference room scene seen by the user.

[0003] However, in the case where multiple participants are located in the same physical space, especially when the participants are close to each other, when a participant near the user speaks, the user will hear the voice of the participant from both the ambient sound collected by the AR / VR device and the transmission sound sent by the AR / VR device of the participant. That is, the user may hear the voice of the participant twice. This will cause a repeated sound interference problem for the user and affect the conference experience of the user. SUMMARY

[0004] The present application provides an audio processing method, related apparatus and a communication system. The present application selects and processes the audio generated in a virtual conference to solve the problem of repeated sound interference in the virtual conference and improve the conference experience of the user.

[0005] In a first aspect, the present application provides an audio processing method. A first device receives first audio from a second device and generates second audio according to a collected sound signal. In the case where the first device and the second device are in the same physical space and the physical distance between the first device and the second device meets a first condition, the first device plays the second audio and does not play the first audio, or the first device detects a first time delay of the first audio and the second audio, and in the case where the first time delay is greater than a time delay threshold, filters a vocal part in the first audio to obtain third audio, and plays the second audio and the third audio.

[0006] The first audio can be collected by the second device from the ambient sound. The first audio can be referred to as a transmission sound. The second audio can be generated by the first device using a pass-through playback technology on the collected sound signal. The second audio can be referred to as a pass-through sound. The third audio can be generated by the first device after performing crosstalk filtering on the first audio.

[0007] The first condition can be satisfied when the physical distance between the first device and the second device is small enough to indicate that the first device and the second device are close to each other.

[0008] It can be understood that, when the first device and the second device are in the same physical space and close to each other, the first device and the second device can both collect the sound of the user wearing the second device speaking. That is, if the user wearing the second device speaks, the first audio and the second audio can both contain the sound of the user wearing the second device speaking. In addition, the first device and the second device can both collect the sound of the user wearing the first device speaking. That is, if the user wearing the first device speaks, the first audio and the second audio can both contain the sound of the user wearing the first device speaking. Therefore, when the first device and the second device are in the same physical space and close to each other, whether the user wearing the first device speaks or the user wearing the second device speaks, the first audio and the second audio will cause repeated sound interference.

[0009] It can be seen that, when the first device and the second device are in the same physical space and close to each other, the first device plays the second audio without playing the first audio, which can ensure that the user wearing the first device can hear the sound of himself speaking and the sound of the user wearing the second device speaking, while avoiding repeated sound interference. Alternatively, the first device can filter out the human voice part in the first audio to avoid repeated sound interference when the time delay of the first audio and the second audio is large. That is, the above method can first determine whether the pass-through sound and the transmission sound in the first device will cause repeated sound interference, and in the case of repeated sound interference, the audio is selected and processed so that the finally played audio will not cause repeated sound interference to the user, thereby improving the user's experience of participating in a virtual meeting.

[0010] In combination with the first aspect, in some embodiments, the first device and the second device are in the same physical space, specifically including: the IP address of the first device and the IP address of the second device indicating that the first device and the second device access the same network; and the physical distance between the first device and the second device satisfying the first condition, specifically including: the physical distance between the first device and the second device being less than a first distance, and / or the sound signal related to the second audio in the first audio having a loudness greater than a first loudness.

[0011] The first distance can also be referred to as a distance threshold. The first distance can be preset. For example, the first distance can be 3 meters, or 4 meters, or 5 meters, or the like.

[0012] In combination with the first aspect, in some embodiments, the first device and the second device are devices accessing a same virtual conference, wherein the first device and the second device are worn by different users, the device for playing audio in the first device adopts a block-ear design, and the second audio is a pass-through sound.

[0013] In combination with the first aspect, in some embodiments, when the first device plays the second audio and the third audio, the first device can specifically take the second audio as a main sound source, take the third audio as a non-main sound source, mix the second audio and the third audio, and play the mixed audio.

[0014] It can be understood that, in a case where the first device and the second device are in a same physical space and are close to each other, the sound quality of the sound of the user wearing the first device or the user wearing the second device in the second audio is good. Moreover, compared with the first audio from the first device, the user wearing the first device will first hear the content of the speech of the user wearing the second device through the second audio. Therefore, when the first device mixes the second audio and the third audio, the first device takes the second audio as a main sound source and takes the third audio as a non-main sound source. In the mixed audio, the information proportion of the main sound source is higher than that of the non-main sound source. That is, the content heard by the user when listening to the mixed audio is more from the main sound source.

[0015] In combination with the first aspect, in some embodiments, in a case where the first device and the second device are in a same physical space, the physical distance between the first device and the second device satisfies the first condition, and the first time delay is less than or equal to a time delay threshold, the first device takes the second audio as a main sound source, takes the first audio as a non-main sound source, mixes the first audio and the second audio, and plays the mixed audio.

[0016] It can be seen that, in a case where the time delay between the first audio and the second audio is small, the first audio and the second audio will not cause repeated sound interference to the user. Therefore, the first device can directly mix the first audio and the second audio and play the mixed audio.

[0017] In combination with the first aspect, in some embodiments, in a case where the first device and the second device are in a same physical space, and the physical distance between the first device and the second device does not satisfy the first condition, the first device plays the first audio and does not play the second audio, or in a case where the first time delay is greater than the time delay threshold, the first electronic device filters out a human voice part in the second audio to obtain a fourth audio, and plays the first audio and the fourth audio.

[0018] The physical distance between the first device and the second device does not satisfy the first condition, specifically including that the physical distance between the first device and the second device is greater than the first distance, or the sound signal loudness of the second audio related to the first audio is less than the first loudness. The physical distance between the first device and the second device not satisfying the first condition can represent that the first device and the second device are far apart.

[0019] When the first device and the second device are in the same physical space but far apart, the first device has poor pickup effect on the speech of the user wearing the second device. That is to say, the sound quality of the voice of the user wearing the second device in the second audio is poor. The first device playing the first audio without playing the second audio can ensure that the user wearing the first device can clearly hear the speech content of the user wearing the second device, while avoiding repeated sound interference. Or, the first device filters out the human voice part in the second audio to avoid repeated sound interference in the case that the first audio and the second audio have large time delay.

[0020] In combination with the first aspect, in some embodiments, the first device plays the first audio and the fourth audio, specifically taking the first audio as the main sound source and the fourth audio as the non-main sound source, mixes and plays the first audio and the fourth audio.

[0021] In combination with the first aspect, in some embodiments, in the case that the first device and the second device are in the same physical space, the physical distance between the first device and the second device does not satisfy the first condition, and the first time delay is less than or equal to the time delay threshold, the first device takes the first audio as the main sound source and the second audio as the non-main sound source, mixes and plays the first audio and the second audio.

[0022] In combination with the first aspect, in some embodiments, in the case that the first device and the second device are not in the same physical space, the first device plays the first audio and the second audio.

[0023] The second aspect provides an audio processing method. Specifically, a first device receives first audio from a second device; in the case that the first device and the second device are in the same physical space and the physical distance between the first device and the second device satisfies a first condition, the first device does not play the first audio.

[0024] In combination with the second aspect, in some embodiments, the first device and the second device are in the same physical space, specifically including that the IP address of the first device and the IP address of the second device indicate that the first device and the second device access the same network; the physical distance between the first device and the second device satisfies the first condition, specifically including that the physical distance between the first device and the second device is less than the first distance, and / or the sound signal loudness of the second audio related to the first audio is greater than the first loudness.

[0025] In some embodiments of the second aspect, the first device and the second device are devices accessing a same virtual conference, wherein the first device and the second device are worn by different users, and the means for playing audio in the first device adopts a non-occluded ear design.

[0026] It can be seen that, in the case that the first device adopts a non-occluded ear design, the first device can detect whether there are other devices in the same physical space as the first device, and among the devices in the same physical space as the first device, a device that is closer to the first device. The user wearing the first device can quickly and relatively clearly hear the sound of the closer conference participant speaking through the ambient sound. Therefore, in the case that the second device is closer to the first device, the first device can no longer play the first audio sent by the second device, so as to avoid the case that the user repeatedly listens to the same speech of the closer conference participant. In this way, the problem of repeated audio interference can be solved.

[0027] In some embodiments of the second aspect, in the case that the first device and the second device are in the same physical space, and the physical distance between the first device and the second device does not satisfy the first condition, the first device plays the first audio; in the case that the first device and the second device are not in the same physical space, the first device plays the first audio.

[0028] In a third aspect, the present application provides an electronic device, which can include an audio input module, an audio output module, a memory and a processor. The audio input module can be used to collect sound signals. The audio output module can be used to play audio. The memory can be used to store a computer program. The processor can be used to call the computer program, so that the electronic device executes any possible implementation method of the first aspect or the second aspect.

[0029] In a fourth aspect, the present application provides a computer readable storage medium, which includes instructions, when the instructions run on an electronic device, make the electronic device execute any possible implementation method of the first aspect or the second aspect.

[0030] In a fifth aspect, the present application provides a computer program product, which can include computer instructions, when the computer instructions run on an electronic device, make the electronic device execute any possible implementation method of the first aspect or the second aspect.

[0031] In a sixth aspect, the present application provides a chip, which is applied to an electronic device, and the chip includes one or more processors, which are used to call computer instructions to make the electronic device execute any possible implementation method of the first aspect or the second aspect.

[0032] It can be understood that the electronic device provided in the third aspect, the computer readable storage medium provided in the fourth aspect, the computer program product provided in the fifth aspect, and the chip provided in the sixth aspect are all used to execute the method provided in the embodiments of the present application. Therefore, the beneficial effects that can be achieved thereby can refer to the beneficial effects in the corresponding method, which will not be described here again. BRIEF DESCRIPTION OF DRAWINGS

[0033] Figure 1 is a schematic diagram of a virtual conference scenario provided by an embodiment of the present application;

[0034] Figure 2 is a schematic diagram of an architecture of a communication system 10 provided by an embodiment of the present application;

[0035] Figure 3 is a schematic diagram of a hardware structure of an electronic device 100 provided by an embodiment of the present application;

[0036] Figure 4A is a flowchart of an audio processing method provided by an embodiment of the present application;

[0037] Figure 4B is a schematic diagram of a time delay between acoustic signals provided by an embodiment of the present application;

[0038] Figure 5 is a flowchart of another audio processing method provided by an embodiment of the present application;

[0039] Figure 6 is a schematic diagram of a structure of an electronic device 100 provided by an embodiment of the present application. DETAILED DESCRIPTION

[0040] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing the specific embodiments of the present application, and are not intended to be limiting on the present application. As used in the specification and the appended claims of the present application, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that “at least one” and “one or more” refer to one or two or more (including two) in the following embodiments of the present application. The term “and / or” is used to describe the association relationship of the associated objects, which means that there can be three relationships; for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character “ / ” generally represents an “or” relationship between the associated objects.

[0041] Reference within the specification to "one embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" or "in some embodiments" in various places within specified

[0042] In the embodiments of the present application, the words "exemplary" and "for example" are used to mean serving as an example, instance, or illustration. Any implementation described herein as "exemplary" or as an "example" is not necessarily to be construed as preferred or advantageous over other implementations. Rather, the

[0043] Here, a scenario of a virtual meeting using AR / VR technology is first introduced.

[0044] Figure 1 An exemplary schematic diagram of a virtual meeting is shown.

[0045] As shown in Figure 1 A user can participate in a virtual meeting by wearing an electronic device 100. The electronic device 100 can be worn on the head of the user. For example, the electronic device 100 can be AR / VR glasses, an AR / VR head-mounted display (HMD), an AR / VR all-in-one machine, or the like, which supports AR / VR technology. The embodiments of the present application do not limit the specific type of the electronic device 100.

[0046] All participants in the virtual meeting can wear the same type of device as the electronic device 100 to access the virtual meeting. Some participants in the virtual meeting can be located in the same physical space in the real environment, and some participants in the virtual meeting can also be located in different physical spaces in the real environment. For example, in the real environment, some participants can be in a company conference room, and some participants can be at home. The embodiments of the present application do not limit the location of the participants in the virtual meeting in the real environment.

[0047] The electronic device 100 can displayFigure 1 The user interface 210 shown. As can be seen from the user interface 210, the electronic device 100 can provide a virtual conference room by using AR / VR technology, and display one or more virtual images of participants in the virtual conference room. For example, virtual image 211, virtual image 212, virtual image 213, virtual image 214, and so on. One virtual image can represent one participant in the virtual conference.

[0048] In some embodiments, the character features of a virtual image (such as the gender, hairstyle, clothing, and so on of the virtual image) can be set by the participant corresponding to the virtual image. The sitting position of a virtual image in the virtual conference room can also be selected by the user. For example, the sitting positions of the virtual image 211 and the virtual image 212 in the virtual conference room are adjacent. The sitting position of the virtual image 213 is opposite to the virtual image 211.

[0049] In some embodiments, the electronic device 100 can present the scene of the virtual conference according to the perspective of the virtual image of the user wearing the electronic device 100 in the virtual conference room. For example, the sitting position selected by the user wearing the electronic device 100 in the virtual conference room is opposite to the virtual image 214. The user interface 210 presents the scene of the virtual conference room that can be seen from the perspective of the position opposite to the virtual image 214. This can give the participants an immersive experience when participating in the virtual conference. Figure 1

[0050] In some embodiments, the electronic device 100 can perform spatial audio modeling according to the sound field environment information of the virtual conference room and the positional relationship of each participant in the virtual conference room, so that when the user hears the voice of a participant, the direction of the participant in the hearing sense is consistent with the direction of the participant seen in the virtual conference room. The sound field environment of the virtual conference room can be set by the participants. For example, before the virtual conference, the participants can set the room size, the room reflection material, the reverberation time, the number of seats, and so on of the virtual conference room. The electronic device 100 can determine the sound field environment information of the virtual conference room according to the above user-set parameters. The electronic device 100 performs spatial audio modeling, which can determine the audio playback parameters for playing the sound of each participant in the virtual conference room, so as to achieve the effect that the sound heard by the user matches the virtual conference room scene seen by the user.

[0051] ​Exemplarily, the electronic device 100 receives audio 1 sent from the electronic device of the virtual image 211. The audio 1 can be the audio of the speaker corresponding to the virtual image 211. The electronic device 100 can play the audio 1 using the audio playing parameter of the virtual image 211 determined based on the spatial audio modeling, so that the user can distinguish that the sound of the speaker corresponding to the virtual image 211 is emitted from the left front in the hearing sense. The electronic device 100 receives audio 2 sent from the electronic device of the virtual image 213. The audio 2 can be the audio of the speaker corresponding to the virtual image 213. The electronic device 100 can play the audio 2 using the audio playing parameter of the virtual image 213 determined based on the spatial audio modeling, so that the user can distinguish that the sound of the speaker corresponding to the virtual image 213 is emitted from the right front in the hearing sense.

[0052] In some embodiments, the electronic device 100 adopts a plug-ear design. The plug-ear design can mean that the device for playing audio in the electronic device 100 can plug the user's ears and isolate part of the environmental sound. For example, the electronic device 100 can play audio through an in-ear earphone. In the case of adopting the plug-ear design, the electronic device 100 can collect environmental sound and process the environmental sound through the pass-through playback technology to generate pass-through sound. The electronic device 100 can play the pass-through sound. The pass-through sound can also be called pass-through audio and the like. It can be understood that in the case of the user's ears being plugged, the user's hearing of his own speaking sound can be affected. The above-mentioned environmental sound can include the user's own speaking sound. The electronic device 100 can facilitate the user to normally hear his own speaking sound by playing the above-mentioned pass-through sound. Moreover, the electronic device 100 can send the collected environmental sound to the electronic device worn by other participants in the virtual conference to enable other participants to hear the user's speech. In addition, in the virtual conference scenario, when other participants speak, the electronic device worn by the speaker can send the collected sound signal to the electronic device 100. The electronic device 100 can receive and play the transmission sound of the speech of other participants. The transmission sound can also be called transmission audio.

[0053] It can be seen that in the case of the above-mentioned plug-ear design, the electronic device 100 will play both the pass-through sound and the transmission sound.

[0054] In the case where there are other participants near the electronic device 100, the ambient sound collected by the electronic device 100 can contain the speech of the participant near the electronic device 100. That is, when the participant near the electronic device 100 speaks, the user wearing the electronic device 100 will hear the speech of the participant from both the pass-through sound played by the electronic device 100 and the transmission sound played by the electronic device 100. The pass-through sound and the transmission sound can have a time delay. Therefore, in the case where the participant near the electronic device 100 speaks, the user wearing the electronic device 100 will hear the speech of the participant twice in succession. This can cause a double sound interference problem for the user, affecting the user's conference experience.

[0055] In some embodiments, the electronic device 100 adopts a non-occlusion design. The non-occlusion design can mean that the device for playing audio in the electronic device 100 does not occlude the user's ear and does not isolate ambient sound. For example, the electronic device 100 can play audio through a bone conduction earphone. In the case of the non-occlusion design, because the ambient sound is not isolated, the user can directly hear the ambient sound, and therefore the electronic device 100 can not need to play pass-through sound using the pass-through playback technology. The electronic device 100 can send the collected ambient sound to the electronic devices worn by other participants attending the virtual conference, so that the other participants can hear the speech of the user. In the virtual conference scenario, when the other participants speak, the electronic device worn by the speaker can send the collected sound signal to the electronic device 100. The electronic device 100 can receive and play the transmission sound of the speech of the other participants.

[0056] As can be seen, in the case of the non-occlusion design described above, the electronic device 100 can only play transmission sound.

[0057] In the case where the other participants near the electronic device 100 speak, the speech of the other participants can be directly transmitted to the ears of the user wearing the electronic device 100 through a medium such as air and heard by the user wearing the electronic device 100. In addition, the user wearing the electronic device 100 will also hear the speech of the participant from the transmission sound played by the electronic device 100. The time at which the user hears the speech of the speaker through the transmission sound is usually later than the time at which the user directly hears the speech of the speaker through the ambient sound. Therefore, in the case of the non-occlusion design and the other participants near the electronic device 100 speaking, the user wearing the electronic device 100 will also hear the speech of the participant twice in succession. This can cause a double sound interference problem for the user, affecting the user's conference experience.

[0058] For ease of understanding, the AR / VR technology is introduced here.

[0059] Electronic device 100 can utilize AR / VR technology to provide a virtual environment. Specifically, electronic device 100 can render and display one or more virtual objects using AR / VR technology. These virtual objects can be generated by electronic device 100 itself using computer graphics technology, computer simulation technology, etc., or they can be generated by other electronic devices using computer graphics technology, computer simulation technology, etc., and then sent to electronic device 100. Other electronic devices can be, for example, servers, or devices such as mobile phones and computers connected to or paired with electronic device 100. For example, the aforementioned virtual environment can include virtual conference rooms, virtual cinemas, virtual supermarkets, etc. These virtual objects can also be referred to as virtual images, virtual elements, or virtual figures. Virtual objects can be two-dimensional or three-dimensional. Virtual objects are fake and not real objects existing in the physical world. Virtual objects can be virtual objects that mimic real objects existing in the physical world, thereby providing users with an immersive experience. Types of virtual objects can include virtual animals, virtual characters, virtual trees, virtual buildings, virtual labels, icons, images, or videos, etc.

[0060] Different virtual objects can be located in different positions within a virtual environment. For example, virtual objects representing attendees can be seated in different positions within a virtual meeting room. Electronic device 100 can determine the positional relationships of the virtual objects representing attendees within the virtual meeting room.

[0061] A virtual environment is the opposite of a real environment. A real environment can refer to the actual physical environment or space where the user and electronic device 100 are currently located. Objects existing in the real environment are real objects. For example, real objects can include participants in a virtual meeting.

[0062] Based on AR / VR technology, electronic device 100 can display virtual objects on a screen or present a virtual environment to the user by projecting optical signals to the user's retina. The user can experience a three-dimensional virtual environment by wearing electronic device 100. In this three-dimensional virtual environment, virtual objects are essentially mapped onto the actual real environment, and the user feels or perceives that the virtual object exists in the actual real environment. For example, in the above... Figure 1 In the scenario shown, a user wearing electronic device 100 can view user interface 210. User interface 210 presents the user with a three-dimensional virtual conference room and virtual objects participating in the meeting within the virtual conference room. Each virtual object participating in the meeting can correspond to a real object, i.e., a participant in a real environment. In some embodiments, electronic device 100 can map virtual objects onto the real environment. In this way, the user can see both real objects in the real environment and virtual objects.

[0063] Figure 2An exemplary architecture diagram of the communication system 10 provided by the present application is shown.

[0064] As shown in Figure 2 , the communication system 10 can include an electronic device 100, an electronic device 101, an electronic device 102, and the like, which are all AR / VR technology-enabled electronic devices. Among them, the electronic device 100, the electronic device 101, and the electronic device 102 can all be devices for wearing and providing virtual environments to users using AR / VR technology. For example, in the virtual meeting scenario shown in the above Figure 1 , multiple participants can wear the electronic device 100, the electronic device 101, and the electronic device 102 to attend the virtual meeting.

[0065] Communication connections can be established between each of the electronic devices in the communication system 10. For example, the electronic device 100, the electronic device 101, and the electronic device 102 can communicate with each other in pairs. The present application does not limit the method by which the electronic devices in the communication system 10 establish communication connections. For example, the communication connections can include, but are not limited to, wireless fidelity (Wi-Fi) connections, Bluetooth connections, and the like.

[0066] In the scenario in which the electronic device 100, the electronic device 101, and the electronic device 102 access the same virtual meeting, these electronic devices can provide the same virtual meeting room. The electronic device 100 can send data such as virtual avatar information set by the user wearing the electronic device 100 and selected position information to other electronic devices accessing the virtual meeting, such as the electronic device 101 and the electronic device 102. The virtual avatar information can be used to display a virtual avatar. For example, the virtual avatar information can include, but is not limited to, information such as the gender, hairstyle, and clothing of the virtual avatar. After receiving the information sent by the electronic device 100, the other electronic devices can render the virtual avatar of the user wearing the electronic device 100 on the corresponding seat in the virtual meeting room. The electronic device 100 can also receive virtual avatar information and selected position information sent by other electronic devices accessing the virtual meeting. For example, the electronic device 100 receives virtual avatar information and selected position information sent by the electronic device 101. The electronic device 100 can render the virtual avatar of the user wearing the electronic device 101 on the corresponding seat in the virtual meeting room. In this way, the user wearing the electronic device 100 can view the scenario in which the virtual avatars of the participants are seated at corresponding positions in the virtual meeting room to conduct a virtual meeting.

[0067] The electronic device 100 can also send the collected ambient sound to other electronic devices accessing the virtual conference, so that the other participants can hear the speech of the user wearing the electronic device 100. The electronic device 100 can receive the transmitted sound sent by other electronic devices accessing the virtual conference. For example, the electronic device 100 receives the transmitted sound sent by the electronic device 101. The transmitted sound can include the speech of the participant wearing the electronic device 101. In this way, the user wearing the electronic device 100 can hear the speech of the other participants in the virtual conference in real time.

[0068] In some embodiments, the user can control the motion of the virtual image through somatosensory interaction such as hand / arm motion, head motion, and / or eye rotation. For example, the electronic device 100 can detect the body state of the user wearing the electronic device 100, and adjust the virtual image of the user according to the change of the body state of the user. For example, when the user wearing the electronic device 100 rotates the head to the left, the electronic device 100 can adjust the head of the virtual image of the user to rotate to the left. The electronic device 100 can send information about the left rotation of the head of the virtual image to other electronic devices accessing the virtual conference, so that the other electronic devices adjust the virtual image of the user wearing the electronic device 100.

[0069] Optionally, each electronic device in the communication system 10 can be used with a handheld device. For example, the handheld device can include a controller, a gyro mouse, a stylus, or other handheld computing device. The handheld device can be configured with various sensors, such as an acceleration sensor, a gyroscope sensor, a magnetic sensor, etc. The electronic device 100 and other electronic devices can detect and track the change of the user's posture and other information through the handheld device.

[0070] The structure of the electronic device 100 is described below. The structures of other electronic devices (such as the electronic device 101, the electronic device 102, etc.) in the communication system 10 can refer to the structure of the electronic device 100.

[0071] Figure 3 An exemplary schematic diagram of the hardware structure of the electronic device 100 is shown.

[0072] As shown in Figure 3 , the electronic device 100 can include a processor 110, a memory 120, a communication module 130, a sensor module 140, a key 150, an input / output interface 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset interface 170D, a display device 180, a camera 190, and a battery 1100, etc.

[0073] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device can include more or fewer components than shown. For example, it can also include infrared transceiver devices, ultrasonic transceiver devices, motors, and flashlights, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0074] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a video processing unit (VPU) controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices, or can be integrated in one or more processors.

[0075] Among them, the controller can be the nerve center and command center of the electronic device. The controller can generate operation control signals according to instruction operation codes and timing signals to complete the control of fetching instructions and executing instructions.

[0076] The memory in the processor 110 can also be provided for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can be directly called from the memory. Avoiding repeated access reduces the waiting time of the processor 110, thus improving the efficiency of the system.

[0077] In the present application, the memory stores a computer program for enabling the controller or the processor to implement the audio processing method of the present application through an interface or a protocol.

[0078] The electronic device can realize wireless communication functions through the communication module 130. The communication module 130 can include an antenna, a wireless communication module, a mobile communication module, a modem processor, and a baseband processor, etc. The communication module 130 can also include more or fewer devices.

[0079] Antennas are used to transmit and receive electromagnetic wave signals. Multiple antennas can be included in an electronic device, each of which can be used to cover a single or multiple communication bands. Different antennas can also be multiplexed to improve the utilization of the antennas. For example, a certain antenna can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, antennas can be used in combination with tuning switches.

[0080] A mobile communication module can provide a solution for wireless communication including 2G / 3G / 4G / 5G, etc. applied on an electronic device. The mobile communication module can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module can receive electromagnetic waves from an antenna, and perform filtering, amplification, etc. on the received electromagnetic waves, and transmit the processed signals to a modem processor for demodulation. The mobile communication module can also amplify signals modulated by the modem processor, and convert the signals into electromagnetic waves radiated through an antenna.

[0081] A modem processor can include a modulator and a demodulator. The modulator is configured to modulate a low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is configured to demodulate a received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to a baseband processor for processing. The low-frequency baseband signal processed by the baseband processor is transmitted to an application processor. The application processor outputs a sound signal through an audio device (not limited to a speaker, etc.), or displays an image or a video through a display device 180.

[0082] A wireless communication module can provide a solution for wireless communication including wireless local area networks (WLAN) (such as a Wi-Fi network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied on an electronic device. The wireless communication module can be one or more devices integrated with at least one communication processing module. The wireless communication module receives electromagnetic waves via an antenna, performs frequency modulation and filtering on the electromagnetic wave signals, and transmits the processed signals to a processor 110. The wireless communication module can also receive signals to be transmitted from the processor 110, perform frequency modulation and amplification, and convert the signals into electromagnetic waves radiated through an antenna.

[0083] The electronic device implements display functionality through a GPU, display device 180, and application processor, among other components. The GPU is a microprocessor for image processing, connected to the display device 180 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.

[0084] In embodiments of the application, the display device 180 can be used to present one or more virtual objects, thereby enabling the electronic device 100 to provide a user with a scene of a virtual environment. The manner in which the display device 180 presents the virtual objects can include one or more of the following:

[0085] 1. In some embodiments, the display device 180 is a display screen, which can include a display panel. The display panel of the display device 180 can be used to display virtual objects, thereby presenting a user with a stereoscopic virtual environment. The user can see the virtual objects from the display panel and experience a virtual reality scene. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diodes (QLED), among other display panels.

[0086] 2. In some embodiments, the display device 180 can include an optical device for projecting optical signals (e.g., light beams) directly onto a user’s retina. The user can see the virtual objects and experience a stereoscopic virtual environment directly through the optical signals projected by the optical device. The optical device can be a micro-projector, among other optical devices.

[0087] The number of display devices 180 in the electronic device can be two, corresponding to the two eyeballs of a user. The content displayed on the two display devices can be independently displayed. Different images can be displayed on the two display devices to improve the stereoscopic effect of the images. In some possible embodiments, the number of display devices 180 in the electronic device can also be one, corresponding to the two eyeballs of a user.

[0088] The electronic device can implement a photographing function through an ISP, camera 190, video codec, GPU, display device 180, and application processor, among other components.

[0089] ISP is used to process the data feedback by the camera 190. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing, and converts it into a visible image.

[0090] The camera 190 is used to capture still images or videos. In some embodiments, the electronic device can include 1 or N cameras 190, N being a positive integer greater than 1.

[0091] In some embodiments, the camera 190 can capture images of the user's hands or body, and the processor 110 can be used to analyze the images captured by the camera 190 to identify the user's input hand gestures or body gestures.

[0092] In some embodiments, the camera 190 can be used in conjunction with an infrared device (such as an infrared emitter) to detect eye movements of the user, such as eye gaze direction, blinking operation, gaze operation, etc., to achieve eye tracking.

[0093] The digital signal processor is used to process digital signals, in addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0094] The NPU is a neural network (NN) computing processor, which learns from the structure of biological neural networks, such as the transmission mode between human brain neurons, and quickly processes input information, and can also continuously self-learn. Through the NPU, the electronic device can achieve intelligent cognition applications such as image recognition, face recognition, speech recognition, text understanding, etc.

[0095] The memory 120 can be used to store computer-executable program code including instructions. The processor 110 performs various functional applications of the electronic device and data processing by executing the instructions stored in the memory 120. The memory 120 can include a program area storing programs and a data area storing data. The program area can store an operating system, application programs (e.g., VR / AR / MR applications) required for at least one function (e.g., a sound play function, an image play function, etc.), and the like. The data area can store data (e.g., audio data, virtual avatar data, etc.) created during use of the electronic device, and the like. In addition, the memory 120 can include a high-speed random access memory, and can further include a non-volatile memory such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like.

[0096] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the microphone 170C, the earphone interface 170D, and the application processor, etc. For example, the electronic device 100 can collect ambient sound, play audio, etc.

[0097] The audio module 170 is used to convert digital audio information into an analog audio signal output, and is also used to convert an analog audio input into a digital audio signal. The audio module can also be used to encode and decode audio signals. The speaker 170A, also referred to as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. In some embodiments, the electronic device 100 can include a speaker array composed of multiple speakers. The microphone 170C, also referred to as a "microphone", "microphone", is used to convert a sound signal into an electrical signal. The electronic device can be provided with at least one microphone 140. In some embodiments, the electronic device can be provided with multiple microphones 170C to implement sound source positioning and other functions. The earphone interface 170D is used to connect a wired earphone.

[0098] In some embodiments, the electronic device 100 can use a device with an ear-plugging design to play audio. For example, the electronic device 100 includes an earphone. The earphone plugs the user's ear when worn. In other embodiments, the electronic device 100 can use a device with a non-ear-plugging design to play audio. For example, the electronic device 100 includes an earphone. The earphone does not plug the user's ear when worn.

[0099] In some embodiments, the electronic device can include one or more keys 150, which can control the electronic device and provide the user with access to functions on the electronic device. In some embodiments, the electronic device can include an input / output interface 160, which can connect other devices to the electronic device through suitable components. The components can include, for example, audio / video jacks, data connectors, and the like.

[0100] The sensor module 140 can include various sensors, such as a proximity light sensor, a distance sensor, a gyroscope sensor, an ambient light sensor, an acceleration sensor, a temperature sensor, a magnetic sensor, a bone conduction sensor, a fingerprint sensor, and the like.

[0101] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture, and the like, which is not limited in the present application. For example, the electronic device 100 in the embodiments of the present application can carry an OS, or other operating systems.

[0102] Based on the above virtual conference scenario and the electronic device 100, the present application provides an audio processing method to solve the problem of repeated sound interference in virtual conferences and improve the conference experience of users.

[0103] When accessing a virtual conference, the electronic device 100 can determine whether the electronic device 100 is in the same physical space as other electronic devices accessing the virtual conference. If the electronic device 100 is not in the same physical space as other electronic devices accessing the virtual conference, the electronic device 100 can play the transmission sound after receiving the transmission sound sent by other electronic devices. If the electronic device 100 is in the same physical space as other electronic devices accessing the virtual conference, the electronic device 100 can not play the transmission sound, or mix the transmission sound and the transparent sound and then play the mixed audio.

[0104] In the case where the electronic device 100 adopts the earplug design, the electronic device 100 can generate the transparent sound according to the collected environmental sound. The electronic device 100 can detect the distance between the electronic device 100 and other electronic devices accessing the virtual conference that are in the same physical space as the electronic device 100. If the distance is far, the electronic device 100 can mix the transmission sound from the electronic device at the far distance as the main sound source and the transparent sound as the non-main sound source. If the distance is close, the electronic device 100 can mix the transparent sound as the main sound source and the transmission sound from the electronic device at the close distance as the non-main sound source. In addition, before mixing the main sound source and the non-main sound source, the electronic device 100 can detect the time delay between the main sound source and the non-main sound source. If the time delay is large, the electronic device 100 can perform crosstalk filtering cancellation on the non-main sound source to filter out the effective components in the non-main sound source. The effective components can include the speech part in the non-main sound source. Then, the electronic device 100 can mix the main sound source and the filtered non-main sound source and play the mixed audio. If the time delay is small, the electronic device 100 directly mixes the main sound source and the non-main sound source.

[0105] It can be understood that if other electronic devices are in the same physical space as the electronic device 100, the speech of the participants wearing other electronic devices may be included in the pass-through sound played by the electronic device 100. In this application, the electronic device 100 can mix the pass-through sound and the transmission sound sent from the electronic devices in the same physical space as the electronic device 100 to ensure that the user only hears the sound of one participant speaking in the user's surroundings when the participant speaks. Moreover, when the transmission sound and the pass-through sound have a large time delay, the electronic device 100 performs crosstalk filtering and elimination, which can better avoid the same participant's speech appearing twice in the mixed audio. This can solve the problem of repeated audio interference when multiple participants are in the same physical space in a virtual conference scenario.

[0106] In the case of the electronic device 100 adopting a non-occluded ear design, the electronic device 100 can detect the distance between the electronic device 100 and other electronic devices accessing the virtual conference and located in the same physical space as the electronic device 100. If the distance is far, the electronic device 100 can play the transmission sound sent from the electronic device at a far distance. If the distance is close, the electronic device 100 can not play the transmission sound sent from the electronic device at a close distance.

[0107] It can be understood that if other electronic devices are in the same physical space as the electronic device 100 and are far away, the user wearing the electronic device 100 can have difficulty clearly hearing the speech of the participants at a far distance through the ambient sound. Therefore, the electronic device 100 can play the transmission sound sent from the electronic device at a far distance to enable the user to accurately hear the content of the speech of the participants at a far distance in the same space. If other electronic devices are in the same physical space as the electronic device 100 and are close, the user wearing the electronic device 100 can quickly and relatively clearly hear the sound of the speech of the participants at a close distance through the ambient sound. Therefore, the electronic device 100 can no longer play the transmission sound sent from the electronic device at a close distance to avoid the user repeatedly hearing the speech of the participants at a close distance twice. This can solve the problem of repeated audio interference.

[0108] Figure 4A An exemplary flowchart of an audio processing method provided by the present application is shown.

[0109] Figure 4A The method shown can be an audio processing method for the electronic device 100 adopting an occluded ear design. As shown in the figure, Figure 4A The method can include steps S411-S421. Among them:

[0110] S411. The electronic device 100 determines the pass-through sound according to the collected sound signal, and receives the transmission sound from the electronic device 101.

[0111] In the case that the electronic device 100 adopts the ear-plugging design, the electronic device 100 can collect the ambient sound (or referred to as the ambient sound signal). The ambient sound can include a background sound part. When there is a person speaking in the environment where the electronic device 100 is located, the ambient sound can also include a human voice part. The human voice part can also be referred to as speech. The electronic device 100 can process the collected ambient sound using the pass-through playback technology. The audio obtained after the pass-through playback processing described above can be referred to as the pass-through sound.

[0112] In some embodiments, the electronic device 100 can filter out part or all of the background sound part in the ambient sound to generate the pass-through sound.

[0113] The pass-through sound described above can facilitate the user to hear the sound emitted by the nearby person or object in the case that the ear is plugged by the electronic device 100.

[0114] The transmission sound from the electronic device 101 can be the audio collected by the electronic device 101 and transmitted through the communication connection between the electronic device 100 and the electronic device 101. Among them, in the case that the user wearing the electronic device 101 speaks, the transmission sound from the electronic device 101 can include the sound of the user wearing the electronic device 101 speaking.

[0115] Here, the electronic device 100 and the electronic device 101 are taken as an example to access the same virtual meeting. Among them, more electronic devices can also be accessed in the virtual meeting. That is, not limited to the participants wearing the electronic device 100 and the electronic device 101, the virtual meeting can also have more participants participating. In addition to receiving the transmission sound from the electronic device 101, the electronic device 100 can also receive the transmission sound sent by more electronic devices accessing the virtual meeting. The case that the electronic device 100 receives the transmission sound sent by other electronic devices can refer to the case that the electronic device 100 receives the transmission sound sent by the electronic device 101.

[0116] S412. The electronic device 100 determines whether the electronic device 100 and the electronic device 101 are in the same physical space.

[0117] Since the electronic device 100 and the electronic device 101 are both worn by users, the position of the electronic device 100 can represent the position of the user wearing the electronic device 100 in the real environment, and the position of the electronic device 101 can represent the position of the user wearing the electronic device 101 in the real environment. That is, the electronic device 100 and the electronic device 101 being in the same physical space can represent the user wearing the electronic device 100 and the user wearing the electronic device 101 being in the same physical space. The electronic device 100 and the electronic device 101 not being in the same physical space can represent the user wearing the electronic device 100 and the user wearing the electronic device 101 not being in the same physical space.

[0118] When the two users are in the same physical space, especially in a relatively close physical distance, the two users can hear the voice of the other user without transmitting sound.

[0119] In a possible implementation, the electronic device 100 and the electronic device 101 can determine whether the electronic device 100 and the electronic device 101 are in the same physical space through an internet protocol (IP) address. The IP address of the electronic device 100 and the IP address of the electronic device 101 can reflect whether the electronic device 100 and the electronic device 101 are connected to the same network. If the electronic device 100 and the electronic device 101 are connected to the same network, the electronic device 100 can determine that the electronic device 100 and the electronic device 101 are in the same physical space. If the electronic device 100 and the electronic device 101 are connected to different networks, the electronic device 100 can determine that the electronic device 100 and the electronic device 101 are not in the same physical space.

[0120] For example, when the electronic device 100 and the electronic device 101 are both connected to the network through the same router, the electronic device 100 and the electronic device 101 are connected to the same network. Even if the user wearing the electronic device 100 and the user wearing the electronic device 101 are in different rooms, the electronic device 100 determines that the electronic device 100 and the electronic device 101 are in the same physical space. It can be understood that the signal coverage range of a router is limited. Even if the user wearing the electronic device 100 and the user wearing the electronic device 101 are in different rooms, the actual physical distance between the two users is usually not too far apart. For example, the user wearing the electronic device 100 can be in the bedroom. The user wearing the electronic device 101 can be in the living room. The electronic device 100 and the electronic device 101 are both connected to the network through the router located in the living room.

[0121] When the electronic device 100 and the electronic device 101 are connected to the network through different routers, the electronic device 100 and the electronic device 101 are connected to different networks. At this time, the electronic device 100 can determine that it is not in the same physical space as the electronic device 101.

[0122] The above method of determining whether the electronic devices are in the same physical space through the IP address is only an example of the present application and should not be construed as limiting the present application. The electronic device 100 can also determine whether the electronic device 100 and the electronic device 101 are in the same physical space through other methods.

[0123] If the electronic device 100 and the electronic device 101 are not in the same physical space, the electronic device 100 can perform the following step S413.

[0124] If the electronic device 100 and the electronic device 101 are in the same physical space, the electronic device 100 can perform the following step S414.

[0125] S413, the electronic device 100 plays the transparent sound and the transmission sound.

[0126] In the case where the electronic device 100 and the electronic device 101 are not in the same physical space, the user wearing the electronic device 100 cannot hear the voice of the user wearing the electronic device 101 speaking in the environment in which the user wearing the electronic device 100 is located through the environmental sound. That is, the transparent sound determined by the electronic device 100 does not contain the voice of the user wearing the electronic device 101 speaking. The voice of the user wearing the electronic device 101 speaking does not cause the problem of repeated sound interference to the user wearing the electronic device 100. Therefore, the electronic device 100 can directly play the transmission sound from the electronic device 101. In addition, since the transparent sound and the transmission sound do not cause the problem of repeated sound interference, the electronic device 100 can also play the transparent sound.

[0127] The electronic device 100 can superimpose the transparent sound and the transmission sound and then play them. For example, the transparent sound is a1(t). The transmission sound from the electronic device 101 is a2(t). The electronic device 100 superimposes a1(t) and a2(t) to obtain a3(t). a3(t) = a1(t) + a2(t). The electronic device 100 can play the audio corresponding to a3(t). The present application does not limit the method of superimposing the transparent sound and the transmission sound.

[0128] S414, the electronic device 100 determines whether the physical distance between the electronic device 100 and the electronic device 101 is greater than a distance threshold.

[0129] In the case that the electronic device 100 and the electronic device 101 are in the same physical space, the user wearing the electronic device 100 can hear the voice of the user wearing the electronic device 101 speaking in the environment in which the user wearing the electronic device 100 is located through the ambient sound. That is, the pass-through sound determined by the electronic device 100 can contain the voice of the user wearing the electronic device 101 speaking. The voice of the user wearing the electronic device 101 speaking can cause the problem of repeated sound interference to the user wearing the electronic device 100. Therefore, the electronic device 100 can detect the physical distance between the electronic device 100 and the electronic device 101 to determine the quality of the sound effect of the user wearing the electronic device 100 hearing the voice of the user wearing the electronic device 101 speaking through the pass-through sound. The better the quality of the sound effect of the user wearing the electronic device 100 hearing the voice of the user wearing the electronic device 101 speaking through the pass-through sound, the more likely the pass-through sound and the transmission sound cause repeated sound interference to the user.

[0130] In the case that the physical distance between the electronic device 100 and the electronic device 101 is greater than the distance threshold, the electronic device 100 can perform the following step S415.

[0131] In the case that the physical distance between the electronic device 100 and the electronic device 101 is less than or equal to the distance threshold, the electronic device 100 can perform the following step S416.

[0132] The distance threshold can be preset. The present application does not limit the value of the distance threshold. For example, the distance threshold can be 3 meters or 4 meters or 5 meters or the like.

[0133] The distance threshold can be preset. The present application does not limit the value of the distance threshold. For example, the distance threshold can be 3 meters or 4 meters or 5 meters or the like.

[0134] In a possible implementation, the electronic device 100 can determine the physical distance between the electronic device 100 and the electronic device 101 according to a signal (such as a Wi-Fi signal or a Bluetooth signal, etc.) communicated between the electronic device 100 and the electronic device 101. In this case, the electronic device 100 and the electronic device 101 can send signals to each other. The electronic device 100 can determine the physical distance between the electronic device 100 and the electronic device 101 according to a time taken for the signal to be transmitted between the electronic device 100 and the electronic device 101.

[0135] In a possible implementation, the electronic device 100 can determine the physical distance between the electronic device 100 and the electronic device 101 using a sound source positioning technology. In this case, the electronic device 100 can collect ambient sound based on an audio collection device such as a microphone array, and identify a physical distance between one or more sound sources in the ambient sound and the electronic device 100. For example, the electronic device 100 collects ambient sound containing human voice, and the electronic device 100 can determine a physical distance between a user who emits the human voice and the electronic device 100 using the sound source positioning technology. In the case where the electronic device 100 collects ambient sound containing human voice, if the electronic device 100 receives a transmission sound from the electronic device 101, the electronic device 100 can determine whether the human voice in the ambient sound is related to the transmission sound from the electronic device 101. The human voice in the ambient sound being related to the transmission sound from the electronic device 101 can mean that the human voice in the ambient sound is emitted by a user wearing the electronic device 101. Therefore, the physical distance between the sound source of the human voice determined by the electronic device 100 using the sound source positioning technology and the electronic device 100 is the physical distance between the electronic device 100 and the electronic device 101.

[0136] In a possible implementation, the electronic device 100 can also determine the physical distance between the electronic device 100 and the electronic device 101 according to an image captured by a camera. In this case, the electronic device 100 can store face feature data of a user wearing the electronic device 101. In the case where the user wearing the electronic device 100 and the user wearing the electronic device 101 are in the same physical space, the electronic device 100 can capture an image containing the user wearing the electronic device 101. The electronic device 100 can determine a position of the user wearing the electronic device 101 in the image captured by the electronic device 100 according to the face feature data of the user, and further calculate a physical distance between the user (i.e., the electronic device 101) and the electronic device 100.

[0137] The above method is only an example of the present application, and should not be construed as limiting the present application. The electronic device 100 can also determine the physical distance between the electronic device 100 and the electronic device 101 using other methods.

[0138] In some embodiments, the quality of the sound heard by the user wearing the electronic device 100 through the pass-through sound of the speech of the nearby participant is not limited to being determined by the physical distance between the electronic devices. The electronic device 100 can also determine the quality of the sound heard by the user wearing the electronic device 100 through the pass-through sound of the speech of the nearby participant according to whether the loudness or the sound pickup quality of the human voice in the environmental sound (or the pass-through sound) meets a preset condition. The sound pickup quality can include the equalization effect of different frequency signals collected in the environmental sound, and can be used to reflect the subjective listening experience of the environmental sound. The better the sound pickup quality, the clearer the sound in the environmental sound heard by the user.

[0139] For example, when the environmental sound is collected, the electronic device 100 can separate the human voice from the environmental sound and detect the loudness of the human voice. If the loudness of the human voice is greater than a preset loudness, it can be indicated that the user who emits the human voice is relatively close to the user wearing the electronic device 100. If the loudness of the human voice is less than or equal to the preset loudness, it can be indicated that the user who emits the human voice is relatively far from the user wearing the electronic device 100. The electronic device 100 can perform correlation determination on the human voice in the environmental sound and the transmission sound from the electronic device 101 to determine whether the human voice in the environmental sound is correlated with the transmission sound of the electronic device 101. If the human voice in the environmental sound is correlated with the transmission sound of the electronic device 101 and the loudness of the human voice is greater than the preset loudness, the electronic device 100 can perform the following step S416. If the human voice in the environmental sound is correlated with the transmission sound of the electronic device 101 and the loudness of the human voice is less than or equal to the preset loudness, the electronic device 100 can perform the following step S415.

[0140] S415, the electronic device 100 determines the transmission sound as the main sound source and the pass-through sound as the non-main sound source.

[0141] When the physical distance between the electronic device 100 and the electronic device 101 is greater than the distance threshold, the quality of the sound of the user wearing the electronic device 101 in the pass-through sound determined by the electronic device 100 is poor. Therefore, the electronic device 100 mixes the transmission sound and the pass-through sound by taking the transmission sound as the main sound source and the pass-through sound as the non-main sound source. In this way, the user wearing the electronic device 100 can better hear the content of the speech of the user wearing the electronic device 101.

[0142] In the mixed sound, the information proportion of the main sound source is higher than that of the non-main sound source. That is, the content heard by the user when listening to the mixed sound is more from the main sound source.

[0143] S416, the electronic device 100 determines the pass-through sound as the main sound source and the transmission sound as the non-main sound source.

[0144] When the physical distance between the electronic device 100 and the electronic device 101 is less than or equal to the distance threshold, the electronic device 100 determines that the sound quality of the sound of the user wearing the electronic device 101 in the through transmission sound is better. In addition, compared with the transmission sound from the electronic device 101, the user wearing the electronic device 100 will usually hear the content of the speech of the user wearing the electronic device 101 through the through transmission sound first. Therefore, the electronic device 100 takes the through transmission sound as the main sound source and the transmission sound as the non-main sound source when mixing the through transmission sound and the transmission sound.

[0145] In S417, the electronic device 100 determines the time delay between the through transmission sound and the transmission sound.

[0146] The through transmission sound and the transmission sound are obtained by the electronic device 100 through different channels. There is usually a time delay between the through transmission sound and the transmission sound. The time delay between the through transmission sound and the transmission sound can represent the time difference between the user hearing the speech of the user wearing the electronic device 101 through the through transmission sound and the transmission sound. The above-mentioned through transmission sound is determined by the electronic device 100 after collecting the environmental sound and played to the user. The above-mentioned transmission sound is collected by the electronic device 101, transmitted to the electronic device 100 through the communication network, and then played to the user by the electronic device 100. The transmission of the above-mentioned transmission sound in the communication network takes a long time. Therefore, under normal circumstances, the user wearing the electronic device 100 will hear the speech of the user wearing the electronic device 101 through the through transmission sound earlier than through the transmission sound.

[0147] In order to avoid the case that the mixed sound audio obtained by mixing the through transmission sound and the transmission sound has repeated sound, the electronic device 100 can detect the time delay between the through transmission sound and the transmission sound.

[0148] Figure 4B An exemplary schematic diagram of the time delay between the through transmission sound and the transmission sound is shown.

[0149] As shown in Figure 4B , the sound signal 431 can be the sound signal of the through transmission sound. The sound signal 432 can be the sound signal of the transmission sound. The electronic device 100 can perform correlation judgment on the sound signal 431 and the sound signal 432 to determine the sound signal corresponding to the speech of the user wearing the electronic device 101 in the through transmission sound and the transmission sound. Among them, the electronic device 100 can compare the correlation between the signals starting at different time points of the sound signal 431 and the sound signal 432. For example, the electronic device 100 detects that the part of the sound signal 431 starting at t1 has the highest correlation with the part of the sound signal 432 starting at t1+Δt1. Then, the electronic device 100 can determine that the time delay between the through transmission sound and the transmission sound is Δt1.

[0150] The method for estimating the time delay of the electronic device 100 is not limited in the embodiments of the present application. For example, the electronic device 100 can use a time delay estimation method combining fractional low-order covariance (FLOC) and linear prediction (LPC).

[0151] In S418, the electronic device 100 determines whether the time delay is greater than a time delay threshold.

[0152] According to the Haas effect, if the time difference between the sound waves of two same sound sources reaching the listener is within 35 milliseconds (ms), the listener cannot distinguish the two sound sources, and the two same sound sources will not cause the problem of repeated sound interference for the listener. Therefore, the electronic device 100 can determine whether the time delay between the main sound source and the non-main sound source is greater than the time delay threshold. The time delay threshold can be 35 ms or 34 ms or 36 ms, etc. The specific value of the time delay threshold is not limited in the embodiments of the present application.

[0153] If the time delay between the main sound source and the non-main sound source is less than or equal to the time delay threshold, the electronic device 100 can perform the following step S419. If the time delay between the main sound source and the non-main sound source is greater than the time delay threshold, the electronic device 100 can perform the following step S420.

[0154] In S419, the electronic device 100 mixes the main sound source and the non-main sound source to obtain and play mixed audio.

[0155] In some embodiments, the electronic device 100 can use an adaptive mixing weighting algorithm or an automatic alignment algorithm to mix the main sound source and the non-main sound source. The mixing method used by the electronic device 100 is not limited in the embodiments of the present application.

[0156] In S420, the electronic device 100 performs crosstalk filtering cancellation on the non-main sound source to filter out the effective components in the non-main sound source.

[0157] In S421, the electronic device 100 mixes the main sound source and the filtered non-main sound source to obtain and play mixed audio.

[0158] When the time delay between the main sound source and the non-main sound source is greater than the time delay threshold, the electronic device 100 can first filter out the effective components in the non-main sound source and then mix the main sound source and the non-main sound source. The effective components in the non-main sound source can include the vocal part of the non-main sound source. In this way, the same speech of a participant can not appear twice in the mixed audio, which can avoid repeated sound interference for the user.

[0159] In some embodiments, the electronic device 100 stores an audio playback parameter for playing the audio of the speech of the user wearing the electronic device 101. The audio playback parameter can be determined by the electronic device 100 according to the positional relationship between the user wearing the electronic device 101 and the user wearing the electronic device 100 in the virtual conference room for spatial audio modeling. The electronic device 100 plays the mixed audio determined in the step S419 or the step S421 using the audio parameter, which can bring an immersive experience to the user participating in the virtual conference. Wherein, when the user hears the speech of one participant, the source of the sound is determined from the hearing sense, which can match the position of the virtual image of the one participant in the virtual conference room.

[0160] In some embodiments, when the physical distance between the electronic device 100 and the electronic device 101 is greater than the distance threshold, the electronic device 100 can not need to distinguish the main sound source and the non-main sound source. Wherein, the electronic device 100 can superimpose and play the pass-through sound and the transmitted sound. For details, please refer to the aforementioned step S413. It can be understood that, if the physical distance between the electronic device 100 and the electronic device 101 is greater than the distance threshold, the electronic device 100 can only collect a small amount of sound of the speech of the user wearing the electronic device 101, or even can not collect the sound of the speech of the user wearing the electronic device 101. Therefore, the pass-through sound almost does not contain the sound of the speech of the user wearing the electronic device 101. Therefore, the pass-through sound and the transmitted sound from the electronic device 101 will not cause repeated sound interference to the user wearing the electronic device 100. The electronic device 100 can play the audio directly superimposed by the pass-through sound and the transmitted sound.

[0161] In some embodiments, the steps S415-S421 are optional. Wherein, when the physical distance between the electronic device 100 and the electronic device 101 is greater than the distance threshold, the electronic device 100 can only play the transmitted sound from the electronic device 101 when receiving the transmitted sound, without playing the pass-through sound. When the physical distance between the electronic device 100 and the electronic device 101 is less than or equal to the distance threshold, the electronic device 100 can play the pass-through sound, without playing the transmitted sound from the electronic device 101. In this way, repeated sound interference to the user caused by playing the pass-through sound and the transmitted sound at the same time can also be avoided.

[0162] In some embodiments, when the electronic device 100 and the electronic device 101 are located in the same physical space, especially when the physical distance between the electronic device 100 and the electronic device 101 is less than a distance threshold, if the user wearing the electronic device 100 speaks, the sound of the user wearing the electronic device 100 speaking is included in the transparent sound generated by the electronic device 100, and the electronic device 101 can also collect the sound of the user wearing the electronic device 100 speaking. That is, the sound of the user wearing the electronic device 100 speaking is also included in the transmission sound from the electronic device 101. The time delay of the transparent sound and the transmission sound in step S417 described above can also represent the time difference between the user wearing the electronic device 100 hearing his own voice through the transparent sound and the transmission sound. The electronic device 100 plays audio according to the method shown in Figure 4A the above method not only can reduce the situation that the user repeatedly hears other users speaking, but also can reduce the situation that the user repeatedly hears himself speaking.

[0163] In some embodiments, if the electronic device 100 and the electronic device 101 are in the same physical space and the physical distance between them is greater than the distance threshold, the electronic device 100 can determine whether the loudness of the human voice part in the pass-through sound is greater than the loudness of the human voice part in the transmitted sound, or the electronic device 100 can determine which of the relevant parts in the pass-through sound and the transmitted sound has greater loudness. For example, if the loudness of the human voice part in the transmitted sound is greater, the electronic device 100 can perform the above step S415 to determine the transmitted sound as the main sound source and the pass-through sound as the non-main sound source. If the loudness of the human voice part in the pass-through sound is greater, the electronic device 100 can perform the above step S416 to determine the pass-through sound as the main sound source and the transmitted sound as the non-main sound source. It can be understood that the repeated sound interference in the virtual conference can be that the user repeatedly hears his own voice, or that the user hears the voices of other users. When the electronic device 100 and the electronic device 101 are in the same physical space and are far apart, if the user wearing the electronic device 101 speaks, the voice of the user wearing the electronic device 101 can be included in both the transmitted sound and the pass-through sound. However, the sound quality of the voice of the user wearing the electronic device 101 in the transmitted sound is better. Therefore, the electronic device 100 takes the transmitted sound as the main sound source and the pass-through sound as the non-main sound source, which can allow the user wearing the electronic device 100 to clearly hear the voice of the user wearing the electronic device 101 and reduce the repeated sound interference. If the user wearing the electronic device 100 speaks, the voice of the user wearing the electronic device 100 can be included in both the transmitted sound and the pass-through sound. However, the sound quality of the voice of the user wearing the electronic device 100 in the pass-through sound is better. Therefore, the electronic device 100 takes the pass-through sound as the main sound source and the transmitted sound as the non-main sound source, which can allow the user wearing the electronic device 100 to clearly hear his own voice and reduce the repeated sound interference. It can be seen that in the virtual conference, each participant usually speaks alternately. In the case that the electronic device 100 and the electronic device 101 are in the same physical space and are far apart, the electronic device 100 can adjust the main sound source and the non-main sound source in real time according to the change of the speaker, so as to ensure that the user wearing the electronic device 100 can clearly hear the voice of the speaker while reducing the repeated sound interference and improving the conference experience of the user.

[0164] From the above Figure 4AThe method shown can know that in the case that the audio to be played in the electronic device 100 includes the pass-through sound and the transmission sound, the electronic device 100 can determine whether the pass-through sound and the transmission sound will cause repeated sound interference to the user wearing the electronic device 100, and the degree of interference, to determine the processing and playing method of the pass-through sound and the transmission sound. In the case that the pass-through sound and the transmission sound will cause repeated sound interference to the user (i.e., the electronic device 100 and the electronic device 101 are in the same physical space), the electronic device 100 can determine one of the pass-through sound and the transmission sound as the main sound source and the other as the non-main sound source according to the degree of interference, and then mix the sound. For example, if the physical distance between the electronic device 100 and the electronic device 101 is greater than the distance threshold, the degree of interference is small. If the physical distance between the electronic device 100 and the electronic device 101 is less than or equal to the distance threshold, the degree of interference is large. In addition, the electronic device 100 can also perform crosstalk filtering elimination on the non-main sound source when mixing the sound. This can better reduce the repeated sound interference and improve the user's experience of participating in the virtual meeting.

[0165] Figure 5 An exemplary flowchart of another audio processing method provided by the present application is shown.

[0166] Figure 5 The method shown can be an audio processing method for an electronic device 100 adopting a non-occlusion design. As shown in the method, Figure 5 The method can include steps S511-S516. Among them:

[0167] S511, the electronic device 100 receives the transmission sound from the electronic device 101.

[0168] In the case that the electronic device 100 adopts a non-occlusion design, the ear of the user wearing the electronic device 100 is not occluded, and the user can directly hear the sound in the environment (such as his own voice, the voice of others speaking nearby, etc.). Therefore, the electronic device 100 does not need to generate and play the pass-through sound. The transmission sound from the electronic device 101 can refer to the introduction of the aforementioned Figure 4A The step S411 shown.

[0169] S512, the electronic device 100 determines whether the electronic device 100 and the electronic device 101 are in the same physical space.

[0170] Step S512 can refer to the introduction of the aforementioned Figure 4A The step S412 shown.

[0171] S513, the electronic device 100 plays the transmission sound.

[0172] In the case that the electronic device 100 and the electronic device 101 are not in the same physical space, the user wearing the electronic device 100 cannot hear the voice of the user wearing the electronic device 101 speaking in the environment in which the user wearing the electronic device 100 is located through the environmental sound. That is, the user wearing the electronic device 100 can only hear the voice of the user wearing the electronic device 101 speaking through the transmission sound. Therefore, the electronic device 100 can directly play the transmission sound from the electronic device 101.

[0173] S514, the electronic device 100 determines whether the physical distance between the electronic device 100 and the electronic device 101 is greater than the distance threshold.

[0174] Step S514 can refer to the foregoing Figure 4A description of step S414.

[0175] S515, the electronic device 100 plays the transmission sound.

[0176] In the case that the physical distance between the electronic device 100 and the electronic device 101 is greater than the distance threshold, the user wearing the electronic device 100 is far away from the user wearing the electronic device 101. When the voice of the user wearing the electronic device 101 speaking is transmitted to the user wearing the electronic device 100 through the medium such as air, it is greatly attenuated. The user wearing the electronic device 100 can hardly hear the content of the user wearing the electronic device 101 speaking from the environmental sound. Therefore, the electronic device 100 can play the transmission sound.

[0177] S516, the electronic device 100 does not play the transmission sound.

[0178] In the case that the physical distance between the electronic device 100 and the electronic device 101 is less than or equal to the distance threshold, the user wearing the electronic device 100 is close to the user wearing the electronic device 101. When the voice of the user wearing the electronic device 101 speaking is transmitted to the user wearing the electronic device 100 through the medium such as air, it is less attenuated. The user wearing the electronic device 100 can clearly hear the content of the user wearing the electronic device 101 speaking from the environmental sound. At this time, if the electronic device 100 plays the transmission sound, the user can hear the user wearing the electronic device 101 speaking from the environmental sound and the transmission sound again. That is, the user wearing the electronic device 100 can be disturbed by the repeated sound. Therefore, after receiving the transmission sound from the electronic device 101, the electronic device 100 can not play the transmission sound.

[0179] from the foregoing Figure 5As shown in the method, in a case where a user can hear speeches of other users in the vicinity through ambient sound, the electronic device 100 can determine whether the received transmission sound will cause repeated sound interference to the user wearing the electronic device 100, to determine whether to play the transmission sound. For example, in a case where the electronic device 100 and the electronic device 101 are in the same physical space, and the physical distance between the two is relatively close, the transmission sound from the electronic device 101 will cause relatively large repeated sound interference to the user. The electronic device 100 can not play the transmission sound from the electronic device 101. In this way, the problem of repeated sound interference can be avoided, and the experience of the user participating in the virtual conference can be improved.

[0180] Figure 6 An example shows a structure of an electronic device 100 provided by an embodiment of the present application.

[0181] As shown in the method, in a case where a user can hear speeches of other users in the vicinity through ambient sound, the electronic device 100 can determine whether the received transmission sound will cause repeated sound interference to the user wearing the electronic device 100, to determine whether to play the transmission sound. For example, in a case where the electronic device 100 and the electronic device 101 are in the same physical space, and the physical distance between the two is relatively close, the transmission sound from the electronic device 101 will cause relatively large repeated sound interference to the user. The electronic device 100 can not play the transmission sound from the electronic device 101. In this way, the problem of repeated sound interference can be avoided, and the experience of the user participating in the virtual conference can be improved. Figure 6 The audio input module 611 can be used to collect ambient sound. For example, the audio input module 611 can include a microphone.

[0182] The communication module 612 can be used for the electronic device 100 to communicate with other electronic devices (such as the electronic device 101, etc.). The communication module 612 can refer to the communication module 130 described in the foregoing

[0183] Figure 3 The communication module 612 can be used for the electronic device 100 to communicate with other electronic devices (such as the electronic device 101, etc.). The communication module 612 can refer to the communication module 130 described in the foregoing

[0184] The audio output module 613 can be used to play audio. For example, the audio output module 613 can include a loudspeaker.

[0185] The audio processing module 614 can be used to perform crosstalk filtering cancellation, mixing, etc. on audio. In some embodiments, the audio processing module 614 can also be used to process the ambient sound collected by the audio input module 611 described above, to generate a transparent sound.

[0186] The audio processing module 614 can be used to perform crosstalk filtering cancellation, mixing, etc. on audio. In some embodiments, the audio processing module 614 can also be used to process the ambient sound collected by the audio input module 611 described above, to generate a transparent sound.​

[0187] The time delay estimation module 616 can be configured to estimate the time delay between two acoustic signals.

[0188] In some embodiments, the electronic device 100 adopts the occluded ear design. The electronic device 100 contains the pass-through sound and the transmission sound. Here, the transmission sound from the electronic device 101 is taken as an example for illustration. The acoustic signal selection module 615 can determine whether the electronic device 100 and the electronic device 101 are in the same physical space. The specific method is described in the foregoing Figure 4A The step S412 shown. If the electronic device 100 and the electronic device 101 are not in the same physical space, the pass-through sound and the transmission sound will not cause repeated sound interference to the user when playing. Therefore, the acoustic signal selection module 615 can instruct the audio output module 613 to play the pass-through sound and the transmission sound. In this way, the user wearing the electronic device 100 can hear the content of the user wearing the electronic device 101 speaking through the transmission sound, and can hear his own voice and other sounds nearby through the pass-through sound. Since the transmission sound does not contain the voice of the user speaking from the electronic device 101, the audio output module 613 playing the pass-through sound and the transmission sound from the electronic device 101 will not cause repeated sound interference to the user.

[0189] If the electronic device 100 and the electronic device 101 are in the same physical space, the pass-through sound and the transmission sound may cause repeated sound interference to the user when playing, and the degree of interference depends on the physical distance between the electronic device 100 and the electronic device 101. The greater the physical distance between the electronic device 100 and the electronic device 101, the smaller the degree of repeated sound interference. Conversely, the greater the degree of repeated sound interference. The acoustic signal selection module 615 can determine whether the physical distance between the electronic device 100 and the electronic device 101 is greater than the distance threshold. The specific method is described in the foregoing Figure 4A The step S414 shown.

[0190] If the physical distance between the electronic device 100 and the electronic device 101 is greater than the distance threshold, the acoustic signal selection module 615 can determine the transmission sound as the main sound source and the pass-through sound as the non-main sound source. If the physical distance between the electronic device 100 and the electronic device 101 is less than or equal to the distance threshold, the acoustic signal selection module 615 can determine the pass-through sound as the main sound source and the transmission sound as the non-main sound source. Then, the acoustic signal selection module 615 can instruct the audio processing module 614 to mix the main sound source and the non-main sound source.

[0191] The audio processing module 614 can first instruct the delay estimation module 616 to determine the time delay between the main sound source and the non-main sound source (i.e., the transmitted sound and the transmitted sound). If the time delay between the main sound source and the non-main sound source is greater than the time delay threshold, the audio processing module 614 can perform crosstalk filtering to eliminate the non-main sound source, filtering out the effective components in the non-main sound source, and then mix the main sound source and the filtered non-main sound source to obtain the mixed audio (refer to the foregoing). Figure 4A (See steps S420 and S421). If the time delay between the main sound source and the non-main sound source is less than the time delay threshold, the audio processing module 614 can mix the main sound source and the non-main sound source to obtain mixed audio.

[0192] The audio processing module 614 can instruct the audio output module 613 to play the above-mentioned mixed audio.

[0193] In some embodiments, if the physical distance between electronic device 100 and electronic device 101 is greater than a distance threshold, the sound signal selection module 615 can select to play the transmitted sound. Specifically, the sound signal selection module 615 can instruct the audio output module 613 to play the transmitted sound instead of the transmitted sound. If the physical distance between electronic device 100 and electronic device 101 is less than a distance threshold, the sound signal selection module 615 can select to play the transmitted sound. Again, the sound signal selection module 615 can instruct the audio output module 613 to play the transmitted sound instead of the transmitted sound.

[0194] In some embodiments, electronic device 100 employs a non-opaque design. Electronic device 100 includes transmitted sound but not transparent sound. The example here uses transmitted sound from electronic device 101. The sound signal selection module 615 can determine whether electronic device 100 and electronic device 101 are in the same physical space. For a detailed description of the method, please refer to the foregoing. Figure 4A The step S412 is shown. If electronic device 100 and electronic device 101 are not in the same physical space, the user wearing electronic device 100 can only hear the content spoken by the user wearing electronic device 101 through the transmitted sound. Therefore, the sound signal selection module 615 can instruct the audio output module 613 to play the transmitted sound.

[0195] If electronic devices 100 and 101 are in the same physical space, the transmitted sound and the ambient sound transmitted through media such as air, which is heard by the user, may cause repetitive sound interference. The degree of interference depends on the physical distance between electronic devices 100 and 101. The greater the physical distance between electronic devices 100 and 101, the less severe the repetitive sound interference. Conversely, the smaller the physical distance, the greater the repetitive sound interference. The sound signal selection module 615 can determine whether the physical distance between electronic devices 100 and 101 is greater than a distance threshold. For a detailed description of the method, please refer to the foregoing.Figure 4A The step S414 is shown.

[0196] If the physical distance between the electronic device 100 and the electronic device 101 is greater than the distance threshold, the sound signal selection module 615 can select the transmission sound to be played. Wherein, the sound signal selection module 615 can instruct the audio output module 613 to play the transmission sound. If the physical distance between the electronic device 100 and the electronic device 101 is less than or equal to the distance threshold, the transmission sound from the electronic device 101 does not need to be played. That is, the audio output module 613 can not play the transmission sound from the electronic device 101.

[0197] It should be noted that, Figure 6 The structure shown is only an exemplary description of the present application, and should not be construed as limiting the present application. The electronic device 100 can also include more or less modules, or combine certain modules, or split certain modules.

[0198] From the structure shown, Figure 6 As can be seen from the structure shown, the electronic device 100 can select appropriate audio for playing based on the sound signal selection module, which can solve the problem of hearing the same content sound from the speaker twice when multiple conference participants are in the same physical space for a virtual conference, and improve the user's experience of participating in a virtual conference.

[0199] It should be noted that, without causing contradiction or conflict, any feature in any embodiment of the present application, or any part of any feature, can be combined, and the combined technical solution is also within the scope of the embodiments of the present application.

[0200] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. An audio processing method, characterized by, The method comprises: The first device receives first audio from a second device and generates second audio according to collected sound signals, and a device for playing audio in the first device adopts a deaf-ear design; In a case where the first device and the second device are in the same physical space and a physical distance between the first device and the second device satisfies a first condition, the first device detects a first time delay of the first audio and the second audio; In a case where the first time delay is greater than a time delay threshold, the first device filters out a vocal part in the first audio to obtain third audio, and plays mixed audio of the second audio and the third audio; In a case where the first time delay is less than or equal to the time delay threshold, the first device plays mixed audio of the first audio and the second audio.

2. The method of claim 1, wherein, The first device and the second device are in the same physical space, specifically comprising: IP addresses of the first device and the second device indicate that the first device and the second device access the same network; The physical distance between the first device and the second device satisfies the first condition, specifically comprising: the physical distance between the first device and the second device is less than a first distance, and / or a sound signal related to the second audio in the first audio has a loudness greater than a first loudness.

3. The method according to claim 1 or 2, characterized in that, The first device and the second device are devices accessing the same virtual conference, wherein the first device and the second device are worn by different users, and the second audio is a pass-through sound.

4. The method according to claim 1 or 2, characterized in that, The playing of the mixed audio of the second audio and the third audio specifically comprises: The first device takes the second audio as a main sound source, takes the third audio as a non-main sound source, mixes the second audio and the third audio, and plays the mixed audio.

5. The method of claim 3, wherein, The playing of the mixed audio of the second audio and the third audio specifically comprises: The first device takes the second audio as a main sound source, takes the third audio as a non-main sound source, mixes the second audio and the third audio, and plays the mixed audio.

6. The method of any one of claims 1, 2, 5, wherein, The playing of the mixed audio of the first audio and the second audio by the first device specifically comprises: The first device takes the second audio as a main sound source, takes the first audio as a non-main sound source, mixes the first audio and the second audio, and plays the mixed audio.

7. The method of claim 3, wherein, The playing of the mixed audio of the first audio and the second audio by the first device specifically comprises: The first device takes the second audio as a main sound source, takes the first audio as a non-main sound source, mixes the first audio and the second audio, and plays the mixed audio.

8. The method of claim 4, wherein, The playing of the mixed audio of the first audio and the second audio by the first device specifically comprises: The first device takes the second audio as a main sound source, takes the first audio as a non-main sound source, mixes the first audio and the second audio, and plays the mixed audio.

9. The method of any one of claims 1, 2, 5, 7, 8, wherein, The method further comprises: In a case where the first device and the second device are in the same physical space and a physical distance between the first device and the second device does not satisfy the first condition, the first device plays the first audio without playing the second audio or detects the first time delay of the first audio and the second audio. In a case where the first time delay is greater than the time delay threshold, the first device filters a vocal part in the second audio to obtain fourth audio, and plays a mixed audio of the first audio and the fourth audio.

10. The method of claim 3, wherein, The method further includes: In a case where the first device and the second device are in the same physical space and a physical distance between the first device and the second device does not satisfy the first condition, the first device plays the first audio without playing the second audio or detects the first time delay of the first audio and the second audio. In a case where the first time delay is greater than the time delay threshold, the first device filters a vocal part in the second audio to obtain fourth audio, and plays a mixed audio of the first audio and the fourth audio.

11. The method of claim 4, wherein, The method further includes: In a case where the first device and the second device are in the same physical space and a physical distance between the first device and the second device does not satisfy the first condition, the first device plays the first audio without playing the second audio or detects the first time delay of the first audio and the second audio. In a case where the first time delay is greater than the time delay threshold, the first device filters a vocal part in the second audio to obtain fourth audio, and plays a mixed audio of the first audio and the fourth audio.

12. The method of claim 6, wherein, The method further includes: In a case where the first device and the second device are in the same physical space and a physical distance between the first device and the second device does not satisfy the first condition, the first device plays the first audio without playing the second audio or detects the first time delay of the first audio and the second audio. In a case where the first time delay is greater than the time delay threshold, the first device filters a vocal part in the second audio to obtain fourth audio, and plays a mixed audio of the first audio and the fourth audio.

13. The method of claim 9, wherein, The playing of the mixed audio of the first audio and the fourth audio specifically includes: The first device takes the first audio as a main sound source, takes the fourth audio as a non-main sound source, mixes the first audio and the fourth audio, and plays the mixed audio.

14. The method according to any one of claims 10-12, characterized in that, The playing of the mixed audio of the first audio and the fourth audio specifically includes: The first device takes the first audio as a main sound source, takes the fourth audio as a non-main sound source, mixes the first audio and the fourth audio, and plays the mixed audio.

15. The method of claim 9, wherein, The method further includes: In a case that the first device and the second device are in the same physical space, a physical distance between the first device and the second device does not satisfy the first condition, and the first time delay is less than or equal to the time delay threshold, the first device takes the first audio as a main sound source, takes the second audio as a non-main sound source, mixes the first audio and the second audio, and plays the mixed audio.

16. The method of any one of claims 10-13, wherein, The method further includes: In a case that the first device and the second device are in the same physical space, a physical distance between the first device and the second device does not satisfy the first condition, and the first time delay is less than or equal to the time delay threshold, the first device takes the first audio as a main sound source, takes the second audio as a non-main sound source, mixes the first audio and the second audio, and plays the mixed audio.

17. The method of claim 14, wherein, The method further includes: In a case that the first device and the second device are in the same physical space, a physical distance between the first device and the second device does not satisfy the first condition, and the first time delay is less than or equal to the time delay threshold, the first device takes the first audio as a main sound source, takes the second audio as a non-main sound source, mixes the first audio and the second audio, and plays the mixed audio.

18. The method of any one of claims 1, 2, 5, 7, 8, 10-13, 15, 17, wherein, The method further includes: In a case that the first device and the second device are not in the same physical space, the first device plays the first audio and the second audio.

19. The method of claim 3, wherein, The method further includes: In a case that the first device and the second device are not in the same physical space, the first device plays the first audio and the second audio.

20. The method of claim 4, wherein, The method further includes: In a case that the first device and the second device are not in the same physical space, the first device plays the first audio and the second audio.

21. The method of claim 6, wherein, The method further includes: In a case that the first device and the second device are not in the same physical space, the first device plays the first audio and the second audio.

22. The method of claim 9, wherein, The method further includes: In a case that the first device and the second device are not in the same physical space, the first device plays the first audio and the second audio.

23. The method of claim 14, wherein, The method further includes: In a case that the first device and the second device are not in the same physical space, the first device plays the first audio and the second audio.

24. The method of claim 16, wherein, The method further includes: In a case that the first device and the second device are not in the same physical space, the first device plays the first audio and the second audio.

25. An electronic device, comprising: The method further includes:

26. A computer-readable storage medium comprising instructions, wherein: In a case that the first device and the second device are not in the same physical space, the first device plays the first audio and the second audio. The electronic device includes an audio input module, an audio output module, a memory, and a processor, wherein the audio input module is configured to collect sound signals; the audio output module is configured to play audio; the memory is configured to store a computer program; and the processor is configured to invoke the computer program, so that the electronic device executes the method in any one of claims 1-24. When the instructions are run on the electronic device, the electronic device executes the method in any one of claims 1-24.

27. A computer program product, characterised in that, The computer program product comprises computer instructions which, when run on an electronic device, cause the electronic device to perform the method of any one of claims 1-24.

Citation Information

Patent Citations

  • Audio processing method, device and equipment applied to multi-party call

    CN111756723A

  • Echo suppression method, terminal, system, storage device and computer program product

    CN113206972A

  • Headset apparatus, teleconference system, user device and teleconferencing method

    EP4184507A1