A method and apparatus for assisting in listening
By using microphones and HRTF technology to generate virtual noise and superimpose it onto non-focused sound sources, the problem of user attention being distracted in multi-sound-source environments is solved, enabling focused listening to the sound sources of interest.
Patent Information
- Application Number
- CN202180004382.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-26
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2041-02-26
AI Technical Summary
In a multi-source environment, users may find it difficult to focus on the audio content of the source, while audio content that is not of interest may distract them.
External audio signals are acquired through a microphone, the coordinates of the sound source are determined, and virtual noise is generated using the head correlation transfer function (HRTF). This virtual noise is then superimposed on the audio signal of the non-interesting sound source to reduce its clarity, thereby allowing the user to focus on the sound source.
It effectively reduces the clarity of audio signals from non-interesting sound sources, helping users focus more on the audio content of the sound source they are interested in, thus aiding in listening.
Smart Images

Figure CN115250646B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart terminal technology, and in particular to an auxiliary listening method and device. Background Technology
[0002] When multiple sound sources are present around a user, these sources include the one the user wants to hear, such as... Figure 1 The sound source S1 shown, and the sound sources that the user does not want to listen to, such as Figure 1 The sound source S2 shown is noise to the user. How to enable the user to listen more attentively to the audio content of the sound source S1 is the technical problem to be solved by the embodiments of this application. Summary of the Invention
[0003] This application provides an auxiliary listening method and device, which enables users to focus their attention more on the audio content of the sound source they are interested in, and assists users in selecting sounds.
[0004] In a first aspect, an assisted listening method is provided, comprising: determining the coordinates of a first sound source based on an external audio signal collected by a microphone and a first direction, wherein the first direction is a direction determined by user detection, and the first sound source is a sound source in a direction other than the first direction; determining a first HRTF corresponding to the first sound source based on the coordinates of the first sound source and a preset head correlation transfer function (HRTF); obtaining a noise signal corresponding to the first sound source based on the first HRTF and a preset virtual noise, and playing the noise signal.
[0005] Taking a first direction as the user's focus direction, and the sound source in the first direction (i.e., the sound source in the user's focus direction) as S1, and the aforementioned first sound source as a sound source S2 in a direction other than the user's focus direction as an example, through the design of the first aspect mentioned above, noise N can be superimposed on the audio content of the sound source S2 that the user is not focused on / uninterested in, thereby reducing the clarity of the audio content of the sound source S2, making it impossible for the user to clearly hear the audio content of the sound source S2, thus making the user's attention more focused on the audio content of the sound source S1, assisting in the selection of listening.
[0006] In one possible design, determining the coordinates of the first sound source based on the external audio signal collected by the microphone and the first direction includes: determining the coordinates of at least one sound source near the user based on the external audio signal collected by the microphone; detecting the user and determining the first direction; and determining the coordinates of the first sound source from the coordinates of at least one sound source near the user based on the first direction.
[0007] In one possible design, determining the coordinates of at least one sound source near the user based on the external audio signals collected by the microphones includes: having multiple microphones, each microphone collecting external audio signals, with a time delay between the external audio signals collected by different microphones; and determining the coordinates of at least one sound source near the user based on the time delay of the external audio signals collected by different microphones.
[0008] In one possible design, detecting the user and determining the first direction includes: detecting the user's gaze direction; or, detecting the difference in the user's binocular currents and determining the user's gaze direction based on the correspondence between the difference in binocular currents and the gaze direction, wherein the user's gaze direction is the first direction.
[0009] In one possible design, determining the coordinates of the first sound source based on the first direction, within the coordinates of at least one sound source near the user, includes: analyzing the coordinates of the at least one sound source near the user to determine the directional relationship between each sound source and the user; determining the deviation between each sound source and the first direction based on the first direction and the directional relationship between the sound source and the user; and selecting, from the at least one sound source near the user, a sound source whose deviation from the first direction is greater than a threshold, and the coordinates of that sound source are the coordinates of the first sound source.
[0010] In one possible design, the external audio signal collected by the microphone is a mixed audio signal, which includes audio signals output from multiple sound sources; the external audio signal collected by the microphone is separated to obtain a first audio signal output from the first sound source.
[0011] In one possible design, the method further includes: analyzing the separated first audio signal to determine the content of the first audio signal; and determining the type of virtual noise to be added based on the content of the first audio signal.
[0012] Through the above design, different types of noise are added to different user-unwanted sound sources, i.e., the first sound source, according to the different content of the audio signal output by the first sound source. This helps to mask the content of the audio signal of the first sound source and helps the user listen to the audio signal of the sound source they are interested in.
[0013] In one possible design, if the content of the first audio signal is human conversation, then the type of virtual noise to be added is multi-person conversation babble noise.
[0014] In one possible design, the method further includes: determining the energy of the separated first audio signal; and determining the energy of the virtual noise to be added based on the energy of the first audio signal.
[0015] Through the above design, virtual noise with corresponding energy can be added according to the different energy of the first audio signal output by the first sound source, which can avoid adding too much virtual noise and reduce the power consumption of electronic devices.
[0016] Secondly, an auxiliary listening device is provided, comprising corresponding functional modules or units for implementing the functions described in the first aspect or any of the designs described in the first aspect. These functions can be implemented by hardware or by hardware executing corresponding software, wherein the hardware or software includes one or more modules or units corresponding to the aforementioned functions.
[0017] Thirdly, an auxiliary listening device is provided, comprising a processor and a memory. The memory stores computational programs or instructions, and the processor is coupled to the memory; when the processor executes the computer program or instructions, the device performs the method described in the first aspect or any of the designs in the first aspect.
[0018] Fourthly, an electronic device is provided for performing the method described in the first aspect or any of the designs in the first aspect. Optionally, the electronic device may be an earphone (including wired or wireless earphones, etc.), a smartphone, an in-vehicle device, or a wearable device, etc. Wireless earphones include, but are not limited to, Bluetooth earphones, and wearable devices may be smart glasses, smartwatches, or smart bracelets, etc.
[0019] Fifthly, a computer-readable storage medium is provided that stores a computer program or instructions, which, when executed by a device, cause the device to perform the method described in the first aspect or any one of the designs of the first aspect.
[0020] In a sixth aspect, a computer program product is provided, comprising a computer program or instructions that, when executed by a device, cause the device to perform the method described in the first aspect or any of the designs in the first aspect. Attached Figure Description
[0021] Figure 1 A schematic diagram illustrating user selection of audio input method provided in an embodiment of this application;
[0022] Figure 2 A schematic diagram illustrating the principle of assisted listening provided in the embodiments of this application;
[0023] Figure 3 A schematic diagram of an electronic device provided in an embodiment of this application;
[0024] Figure 4 A flowchart illustrating the assisted listening method provided in the embodiments of this application;
[0025] Figure 5 A schematic diagram of the coordinate system provided in the embodiments of this application;
[0026] Figure 6 A schematic diagram of a microphone and sound source provided in an embodiment of this application;
[0027] Figure 7 This is a functional schematic diagram of the electronic device provided in the embodiments of this application;
[0028] Figure 8 This is a schematic diagram of HRTF rendering provided for an embodiment of this application;
[0029] Figure 9 This is a schematic diagram of the apparatus provided in an embodiment of this application. Detailed Implementation
[0030] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.
[0031] There are various sound sources in our environment, such as Figure 1 As shown, sound sources S1 and S2 exist near the user. For people with normal hearing, we can still consciously choose the audio content we want to listen to in a scenario with multiple sound sources, such as the audio content of sound source S1. This ability is called the cocktail party effect.
[0032] The cocktail party effect refers to a human auditory selective ability. In this case, a person's attention is focused on the conversation of one person while ignoring other conversations or noises in the background. This effect reveals the surprising ability of the human auditory system to allow us to converse even in noisy environments. The cocktail party effect is the auditory version of the figure-ground phenomenon. Here, the "figure" is the sound that we notice or that attracts our attention, and the "ground" is the other sounds.
[0033] In a noisy indoor environment, such as a cocktail party, many different sound sources exist simultaneously, such as the voices of multiple people speaking at the same time, the clinking of cutlery, music, and reflected sounds from walls and objects in the room. During the transmission of sound, sound waves from different sources, as well as direct and reflected sound, superimpose in the propagation medium (usually air) to form complex mixed sound waves. Therefore, the individual sound waves no longer exist in the mixed sound sources reaching the listener's ear canal. Yet, in this acoustic environment, the listener can understand the target statement to a considerable extent. How does the listener separate the speech signals of different speakers from the received mixed sound waves and thus understand the target statement? This is the famous "cocktail party problem" proposed in 1953. The industry typically uses the following principles of auditory attention and acoustics to explain and illustrate the "cocktail party effect."
[0034] The principle of auditory attention: This is because when a person's auditory attention is focused on a certain thing, the conscious mind will exclude some irrelevant sound stimuli, while the unconscious mind is constantly monitoring external stimuli. Once there are some special stimuli that are relevant to oneself, they can immediately attract attention. This effect is actually an adaptive ability of the auditory system. Simply put, our brain makes a certain judgment about the sound before deciding whether to listen or not.
[0035] Acoustic principle: The cocktail party effect in acoustics refers to the masking effect of the human ear. In the noisy crowd of a cocktail party, two people can have a smooth conversation. Despite the high level of ambient noise, they only hear each other's voices, seemingly oblivious to other noises. This is because people have focused their attention (this is selective attention) on the topic of the conversation.
[0036] This cocktail party effect allows us to influence people's listening choices by controlling the clarity of audio objects in the environment. Generally speaking, the clearer the content of an audio object, the more attention it receives and the more attention it garners. Currently, solutions to assist users in making listening choices mainly include the following two types:
[0037] The first approach, in its basic process, is as follows:
[0038] 1) The spectrogram of the mixed audio is fed into multiple deep neural networks (DNNs), each of which is trained to separate specific speech sources.
[0039] 2) When a user is listening to one of the speech sources, the neurons in their brain can record a spectrogram used to reconstruct that speech source.
[0040] 3) Compare the reconstructed spectrogram with the output of each DNN. If they match, amplify the sound source.
[0041] This solution, used in hearing aids, enhances the playback of audio objects that the user is interested in, allowing them to be heard more clearly. However, because this method relies on brainwave technology to identify external sound source signals that the user is currently interested in (reconstructing and recovering them from brainwaves) and then enhancing and playing the corresponding sound source signals, it presents significant challenges. Technically, this technology requires speech separation before matching with the brainwave decoding signal. Its effectiveness depends heavily on the accuracy of the front-end speech signal separation, making it a highly complex and difficult-to-implement solution.
[0042] The second approach includes: Step 1, detecting the user's focus or interest direction, such as based on the user's gaze or head orientation, to determine the user's focus or interest direction. Step 2, focusing sound pickup to capture the audio signal s(n) in the user's focus direction; optionally, sound pickup refers to the specific action of acquiring audio signals. Methods for focusing sound pickup to capture the user's focus audio signal may include: capturing all audio signals around the user and then separating the audio signal the user is interested in; or, only capturing the audio signal the user is interested in, for example, using beamforming methods to acquire only the audio signal the user is interested in. Step 3, binaural playback, performing binaural rendering playback on the sound in the pickup direction, and combining it with the head related transfer function (HRTF) to enhance the sense of presence of s(n).
[0043] The second approach essentially improves the clarity of sound source S1 by focusing on its pickup, thus enhancing the user's attention to S1. However, it doesn't fundamentally change the signal-to-noise ratio of S2 relative to N, meaning S2 can still be perceived, potentially affecting the user's effective attention to S1. For example, a user is conversing with person A, but person B is also conversing nearby; person B is noise to the user. In the second approach, focused pickup increases the volume or clarity of person A's conversation. However, it doesn't process person B's conversation. Therefore, if person B's conversation contains words the user is interested in, it might still attract the user's attention, preventing them from focusing on person A. For instance, if a user is interested in "salary increase," then if person B mentions "salary increase" during their conversation with person A, it will distract the user, preventing them from focusing on person A.
[0044] This application provides an auxiliary listening method and device, such as Figure 2 As shown, the principle of this scheme is to superimpose noise (N) onto the audio content of sound source S2, which is not of interest to the user, thereby reducing the clarity of the audio content of sound source S2. This prevents the user from clearly hearing the audio content of sound source S2, thus making the user's attention more focused on the audio content of sound source S1, assisting in selective listening. Following the above example, using the scheme in this application embodiment, noise N can be superimposed on the conversation content of the aforementioned subject B, thereby reducing the signal-to-noise ratio of the conversation content of subject B. Even if the conversation content of subject B involves words such as "salary increase" that the user is more concerned about, through the superposition of the aforementioned noise N, the user can no longer clearly hear the conversation content of subject B, and the conversation content of subject B will no longer attract the user's attention, making the user more focused on listening to the conversation content of subject A.
[0045] The assisted listening method provided in this application can be applied to electronic devices, including but not limited to headphones, smartphones, in-vehicle devices, or wearable devices (such as smart glasses, smartwatches, or wristbands). Taking headphones as an example, an application scenario could be that the headphones have an assisted listening button. When the user activates this button, the headphones can detect the user and determine the sound source of interest. Noise N is added to the audio signal of the sound source of interest, thereby reducing the signal-to-noise ratio of the audio signal of the non-interested sound source, allowing the user to focus more on listening to the content of the sound source of interest, thus assisting in selective listening.
[0046] like Figure 3 As shown, this application embodiment provides an electronic device that can be used to implement the auxiliary ringing method provided in this application embodiment. The electronic device includes at least: a processor 301, a memory 302, at least one speaker 304, at least one microphone 305, and a power supply 306, etc.
[0047] The memory 302 can store program code, which may include program code for implementing assisted listening. The processor 301 can execute the above program code to implement the assisted listening function in this embodiment. For example, the processor 301 can execute the program code in the memory 302 to implement the following functions: determining the coordinates of a first sound source in a direction other than the user's attention based on the external audio signal collected by the microphone and the detected first direction of user focus; determining a first HRTF corresponding to the first sound source based on the coordinates of the first sound source and the HRTF of predetermined different locations; obtaining a noise signal corresponding to the first sound source based on the first HRTF and preset virtual noise, etc.
[0048] The speaker 304 can be used to convert audio electrical signals into sound and play them. For example, the speaker 304 can be used to play noise signals corresponding to the first sound source mentioned above.
[0049] Microphone 305, also known as a microphone, transducer, etc., is used to convert sound signals into audio electrical signals. For example, microphone 305 can collect sound signals near the user and convert them into audio electrical signals. It should be understood that the audio electrical signal is the audio signal in the embodiments of this application.
[0050] The power supply 306 can be used to supply power to the various components included in an electronic device. In some embodiments, the power supply 306 may be a battery, such as a rechargeable battery.
[0051] Optionally, if the electronic device 300 is a wireless headset, the electronic device 300 may also include: a sensor 303 and a wireless communication module 307, etc.
[0052] Sensor 303 can be a proximity sensor. Processor 301 can use sensor 303 to determine whether the earphones are being worn by the user. For example, processor 301 can use the proximity sensor to detect whether there is an object near the earphones, thereby determining whether the earphones are being worn. Alternatively, if the earphones are charging in the charging case, processor 301 can use the proximity sensor to determine whether the charging case is open, thereby determining whether to control the earphones in pairing mode. In addition, in some embodiments, the earphones may also include a bone conduction sensor, which can acquire the vibration signal of the sound-emitting bone, and processor 301 can parse the voice signal to realize the control function corresponding to the voice signal. In other embodiments, the earphones may also include a touch sensor or a pressure sensor, which are used to detect the user's touch operation and pressing operation on the earphones, respectively. In other embodiments, the earphones may also include a fingerprint sensor, which is used to detect the user's fingerprint and identify the user's identity.
[0053] The wireless communication module 307 is used to establish a wireless connection with other electronic devices, enabling the headphones to interact with other electronic devices. In some embodiments, the wireless communication module 307 can be a near field communication (NFC) module, allowing the headphones to communicate with other electronic devices equipped with NFC modules. The NFC module can store relevant information about the headphones, such as the headphone's name, address information, or unique identifier. Other electronic devices equipped with NFC modules can then establish an NFC connection with the headphones based on this information and transmit data via the NFC connection. In other embodiments, the wireless communication module 307 can also be a Bluetooth module, storing the headphones' Bluetooth address. This allows other electronic devices to establish a Bluetooth connection with the headphones based on the Bluetooth address and transmit audio data via the Bluetooth connection. In this application embodiment, the Bluetooth module can simultaneously support multiple Bluetooth connection types, such as the traditional Bluetooth serial port profile (SPP) or Bluetooth Low Energy (BLE) generic attribute profile (GAP), etc., without limitation.
[0054] In other embodiments, the wireless communication module 307 may also be an infrared module or wireless fidelity (WIFI), etc. The specific implementation of the wireless communication module 307 is not limited here.
[0055] Furthermore, in this embodiment, only one wireless communication module 307 may be provided, or multiple modules may be provided as needed. For example, two wireless communication modules may be provided in the headset, one of which is a Bluetooth module and the other is an NFC module, etc. In this way, the headset can conduct data communication through these two wireless communication modules respectively, and the number of wireless communication modules 307 is not limited here.
[0056] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 300. It may have more than Figure 3 The more or fewer components shown can be combined into two or more components, or they can have different component configurations. For example, the electronic device 300 may also include components such as indicator lights (which can indicate status such as battery level) and dust filters (which can be used with a handset). Figure 3 The various components shown may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing or application-specific integrated circuits.
[0057] like Figure 4 As shown, a process for providing an assisted listening method is provided, which includes at least the following:
[0058] Step 401: The electronic device determines the coordinates of the first sound source based on the external audio signal collected by the microphone and the first direction. The first direction is determined by detecting the user, and the first sound source is a sound source in a direction other than the first direction. For example, the first direction can be the direction the user is interested in, and the first sound source can be a sound source in a direction the user is not interested in. For example, sound sources S1, S2, and S3 exist near the user. If the user is interested in sound source S1, for example, listening to the audio output of sound source S1, then S1 can be a sound source in the direction the user is interested in, and sound sources S2 and S3 are sound sources in directions the user is not interested in.
[0059] Optionally, the implementation process of step 401 above may include: the electronic device determining the coordinates of at least one sound source near the user based on the external audio signals collected by the microphone. For example, the electronic device is equipped with a microphone array, which includes at least one microphone, each microphone collecting external audio signals, with a time delay between the external audio signals collected by different microphones. The electronic device can determine the coordinates of at least one sound source near the user based on the time delay between the external audio signals collected by different microphones. Optionally, the microphone can be a vector microphone, etc. For example, there are sound sources S1 and S2 near the user. When the user wears the electronic device, the microphone array of the electronic device can collect the external sound signals and convert them into audio signals, which include the audio signals corresponding to sound source S1 and the audio signals corresponding to sound source S2. The electronic device can determine the coordinates of sound source S1 and sound source S2 based on the time delay between the audio signals collected by different microphones in the microphone array.
[0060] In one possible implementation, the model described in the industry paper "Research on Sound Source Localization Algorithm Based on Microphone Array" (Master's Thesis, Nanjing University, Liu Chao) can be used to implement the above-mentioned method of determining the coordinates of each sound source by utilizing the time delay of different microphones in the embodiments of this application.
[0061] Specifically, such as Figure 6 As shown, suppose the microphone array consists of N+1 microphones, where the N+1 microphones are (M0, M1, ..., M...). n ), etc. Let the spatial coordinates of the i-th microphone be represented as r. i =(x i ,y i ,z i The coordinates of microphone M0 are the origin of the spatial coordinate system, i.e., r0 = (0, 0, 0), and the spatial coordinates of sound source S are r0 = (0, 0, 0). s= (x, y, z), then:
[0062] a. The distance between the sound source S and the i-th microphone is:
[0063]
[0064] Where, d i Let (x, y, z) represent the distance from sound source S to the i-th microphone, and let (x, y, z) represent the coordinates of sound source S. i ,y i ,z i ) represents the coordinates of the i-th microphone.
[0065] b. The distance difference between the sound source S and each microphone is:
[0066] d ij =d i -d j i, j = 0, 1, ..., N
[0067] Where, d ij d represents the distance d from the sound source to the i-th microphone. i The distance d from the sound source to the j-th microphone j The difference d between ij c. The relative distance between the i-th microphone and the j-th microphone is:
[0068]
[0069] in, The distance between the i-th microphone and the j-th microphone is represented by c, where c represents the speed of sound, and τ represents the relative distance between them. ij This represents the time delay between the i-th microphone and the j-th microphone.
[0070] In the above expressions, the distances between the microphones and the speed of sound are known. By combining and solving these expressions, we can obtain the spatial location of the sound source S: s = (x... s ,y s ,z s The specific algorithm can be the maximum likelihood estimation method or the least squares method, etc., without limitation.
[0071] Based on the above description, the coordinates of at least one sound source near the user can be obtained. Next, the process of detecting the first direction of the user's attention and how to filter out the first non-attentive sound source from the at least one sound source based on the first direction of the user's attention are described.
[0072] For example, a user can wear an electronic device. For instance, if the electronic device is headphones, the user can wear it around their ears. The electronic device can determine the user's primary focus direction by detecting the user. For example, the electronic device can detect the user's gaze direction and use that gaze direction as the user's primary focus direction. In one possible implementation, the electronic device can be equipped with an inertial measurement unit (IMU). The electronic device can use the IMU to determine the user's head orientation and, based on that, determine the user's gaze direction. For example, if the IMU detects that the user's head is facing forward, then the user's gaze direction, i.e., the user's primary focus direction, is forward. Alternatively, the electronic device can be equipped with a camera that can capture images of the user's head and determine the user's gaze direction based on those images. Alternatively, the electronic device can be equipped with a brainwave sensor that can detect the difference in electrical current between the user's two ears and, based on the correspondence between this difference and the gaze direction, determine the user's gaze direction.
[0073] Then, the electronic device can determine the coordinates of the first sound source based on the first direction and the coordinates of at least one sound source near the user. This process can be implemented in two specific ways:
[0074] In the first method, the electronic device, based on the aforementioned first direction, identifies the sound source of interest to the user from at least one sound source near the user; then, it excludes the sound source of interest from the at least one sound source near the user, and the remaining sound source is the first sound source that the user is not interested in. For example, if there are 5 sound sources near the user, and sound source A is identified as the sound source of interest to the user based on the aforementioned first direction of interest, then the remaining 4 sound sources (excluding sound source A) are the first sound sources that the user is not interested in. It should be understood that in the embodiments of this application, the first sound source that the user is not interested in may include one sound source or multiple sound sources, etc., without limitation.
[0075] The second method involves the electronic device directly identifying a first sound source that is not of interest to the user from at least one sound source near the user, based on the aforementioned first direction.
[0076] For example, an electronic device can analyze the coordinates of at least one sound source near the user to determine the positional relationship between each sound source and the user. Optionally, the coordinates of the sound sources can be relative to the user. Figure 5As shown, a coordinate system is established. Optionally, this coordinate system can be a three-dimensional coordinate system. The origin of this coordinate system is the user's head position, the X-axis represents the user's left / right direction, the Y-axis represents the user's front / back direction, and the Z-axis represents the user's up / down direction. For example, in one possible implementation, the positive direction of the X-axis is the user's right-hand direction, the negative direction of the X-axis is the user's left-hand direction, the positive direction of the Y-axis is the user's front-hand direction, the negative direction of the Y-axis is the user's back-hand direction, the positive direction of the Z-axis is the user's up-hand direction, and the negative direction of the Z-axis is the user's down-hand direction. It can be understood that the positive directions of the X, Y, and Z axes all represent directions greater than 0, while the negative directions of the X, Y, and Z axes all represent directions less than 0. By analyzing the coordinates of each sound source relative to the user, the electronic device can determine the positional relationship between each sound source and the user. For example, if the coordinates of a sound source relative to the user are (1, 2, 0), it means that the sound source's position relative to the user is: 1m directly to the right of the user, 2m directly in front of the user, and on the same plane as the user. Using the aforementioned coordinates of the sound source, the location of the sound source can be accurately determined, thereby establishing the relationship between that location and the user's orientation.
[0077] Based on the detected first direction and the directional relationship between the sound source and the user, the deviation between each sound source and the first direction is determined. For example, the detected first direction of user focus is 15 degrees away from directly in front of the user. If a sound source's directional relationship with the user is 20 degrees away from directly in front of the user, then the deviation between this sound source and the first direction is 5 degrees. Among at least one sound source near the user, the sound source with a deviation greater than a threshold from the first direction is selected; this sound source is the first sound source that the user is not focused on. Alternatively, among at least one sound source near the user, the sound source with a deviation less than or equal to the threshold from the first direction is selected; this sound source can then be the sound source that the user is focused on. Afterward, among all sound sources near the user, the sound sources that the user is focused on are excluded; these are the sound sources that the user is not focused on.
[0078] The deviation between each sound source and the first direction can be determined using the above method. When the deviation of a sound source from the first direction is greater than a threshold, the sound source is considered an uninteresting sound source for the user. Conversely, when the deviation of a sound source from the first direction is less than or equal to the threshold, the sound source is considered an interested sound source for the user. Continuing with the above... Figure 1In the example shown, sound sources S1 and S2 exist near the user. If the deviation of sound source S1 from the first direction is less than or equal to the aforementioned threshold, then sound source S1 is considered a sound source of interest to the user. If the deviation of sound source S2 from the first direction is greater than the aforementioned threshold, then sound source S2 is considered a sound source of non-interest to the user, and so on. Of course, the aforementioned threshold can be set at the factory of the electronic device, or it can be synchronized or notified to the electronic device through the server of the electronic device later, etc., without limitation. It should be noted that in the description of the embodiments of this application, "interest" and "interested" are not distinguished and can be used interchangeably. For example, the aforementioned sound source of interest to the user can also be replaced with a sound source of interest to the user; the sound source of non-interest to the user can also be replaced with a sound source of non-interest to the user, etc. In one understanding, a sound source of interest to the user or a sound source of interest to the user can be a sound source related to the user's current activity. For example, if the user is currently watching TV, and the TV is playing audio signals as a sound source, then the TV can be a sound source of interest or interest to the user. Technically, sound sources that a user is interested in or focused on can be those whose deviation from the user's gaze direction is less than or equal to a threshold. Similarly, sound sources that a user is not interested in or focused on can be those unrelated to the user's current activity; these sources would be considered noise to the user. Continuing with the example above, if a user is watching television, watching television is their current activity, while listening to music is not. If a music player is playing music nearby, then that music player can be considered a sound source that is not the user's focus. Technically, sound sources that are not the user's focus can be those whose deviation from the user's gaze direction is greater than a threshold.
[0079] Step 402: The electronic device determines the first HRTF corresponding to the first sound source based on the coordinates of the first sound source and the preset HRTF. Optionally, the first HRTF may include two HRTFs, such as... Figure 7 As shown, these correspond to the HRTF of the left ear and the HRTF of the right ear, respectively.
[0080] Step 403: The electronic device obtains the noise signal corresponding to the first sound source based on the first HRTF and the preset virtual noise, and plays the noise signal.
[0081] Optionally, in the frequency domain, the electronic device can multiply the first HRTF with a preset virtual noise frequency domain to obtain the noise signal corresponding to the first sound source. For example, continuing the above example, the first HRTF includes the HRTF of the left ear and the HRTF of the right ear. Multiplying the HRTF of the left ear with the preset virtual noise frequency domain yields the noise signal of the left ear, and multiplying the HRTF of the right ear with the preset virtual noise frequency domain yields the noise signal of the right ear. The noise signals of the left and right ears can be referred to as binaural noise signals. HRTF is an algorithm for sound effect localization, corresponding to the head-related inpulse response (HRIR) in the time domain. The process of multiplying the HRTF with the preset virtual noise in the frequency domain described above is, in the time domain, manifested as the process of convolving the HRIR with the preset virtual noise.
[0082] To facilitate understanding, let's first introduce HRTF. The concept and significance of HRTF are as follows: Humans have two ears, yet they can locate sounds from three-dimensional space thanks to the human ear's sound signal analysis system. HRTF can simulate this system. Essentially, HRTF contains spatial location information of the sound source; different spatial locations correspond to different HRTFs. For any mono audio signal, multiplying it by the HRTF of the left and right ears in the frequency domain yields the corresponding audio signals for both ears. Playing these signals through headphones allows you to experience 3D audio.
[0083] In one possible implementation, multiple HRTFs for different locations can be pre-stored. Using these pre-stored HRTFs (e.g., by interpolating the HRTFs at these locations), the first HRTF corresponding to the first sound source can be obtained. Then, the first HRTF is multiplied by a preset virtual noise frequency domain to obtain the noise signal. It should be noted that the HRTF is location-dependent, and the first HRTF is obtained based on the coordinates of the first sound source. By multiplying the first HRTF by the preset virtual noise frequency domain, the virtual noise signal can be superimposed on the audio signal of the first sound source, reducing the signal-to-noise ratio of the audio signal of the first sound source. This allows the user to focus more intently on listening to the audio signal of the sound source they are interested in, thus assisting the user's listening experience. For example, continuing with the above... Figure 1 For example, based on the coordinates of the sound source S2, the first HRTF is obtained, and then the first HRTF is multiplied by a preset virtual noise frequency domain to obtain a noise signal. By playing the noise signal, the noise signal can be superimposed on the sound source S2, thereby reducing the signal-to-noise ratio of the audio signal of the sound source S2, so that the user can listen to the audio signal of the sound source S1 more attentively.
[0084] In this embodiment of the application, the process of determining the first HRTF corresponding to the first sound source based on the coordinates of the first sound source and the preset HRTF in step 402 above includes, but is not limited to, the following two implementation methods:
[0085] Method 1: Accurately locate the coordinates of each sound source relative to the user and pre-store a large number of HRTFs.
[0086] In this method, the electronic device precisely locates the position of each sound source relative to the user. Continuing with the example above, if the coordinates of a sound source relative to the user are (1, 2, 0), it means the sound source is located 1 meter to the right of the user, 2 meters in front of the user, and on the same horizontal plane as the user. Since HRTF is position-dependent, in this case, the electronic device may need to pre-calculate a relatively large number of HRTFs to determine the HRTF corresponding to the sound source's position. The advantage of this method is that it can determine a more accurate HRTF, and the noise signal determined based on this HRTF can be more accurately superimposed on non-interesting sound sources.
[0087] Method 2: Roughly locate the coordinates of each sound source relative to the user, and pre-store a small number of HRTFs.
[0088] In this method, the electronic device can roughly determine the direction of the sound source relative to the user, without precisely locating the exact position of each sound source. For example, using this method, the coordinates of a sound source relative to the user could be (1, 1, 0), representing that the sound source is located in the user's right front direction, without calculating the exact position in the right front direction. The electronic device can store the HRTF (Head-To-Face Values) for four directions: front, back, left, and right, and calculate the HRTF for the right direction based on these four HRTFs. The advantage of this approach is that it reduces the storage space required by the electronic device, simplifies the calculation process, and saves power consumption.
[0089] Optionally, the virtual noise added in step 403 above can be white noise. Alternatively, the virtual noise added in step 403 above can also be noise that matches the content of the first audio signal from the first sound source. The electronic device can analyze the first audio signal from the first sound source to determine the content of the first audio signal, and determine the type of virtual noise to be added based on the content of the first audio signal. For example, if the content of the first audio signal is human conversation, the electronic device determines that the type of virtual noise to be added is multi-person conversation babble noise, etc. In one possible implementation, the electronic device can detect whether the content of the first audio signal contains human speech. If the content of the first audio signal contains human speech, the virtual noise added to the first audio signal can be multi-person conversation babble noise, etc. Regarding the detection of whether the content of the first audio signal from the first sound source contains human speech, voice activity detection (VAD) technology can be used, such as short time energy (STE) and zero cross-counter (ZCC) detection methods.
[0090] In one possible implementation, the STE and ZCC of the first audio signal from the first sound source can be detected. Since the STE of speech segments is relatively large and the ZCC is relatively small, while the STE of non-speech segments is relatively small and the ZCC is relatively large, this is mainly because the energy of speech signals is mostly contained in the low-frequency band, while noise signals usually have lower energy and contain information from higher frequency bands. Therefore, a certain threshold can be set. When the STE of the first audio signal from the first sound source is greater than or equal to a first threshold and the ZCC is less than or equal to a second threshold, it can be considered that the first audio signal from the first sound source includes human speech. Conversely, when the STE of the first audio signal from the first sound source is less than the first threshold and the ZCC is greater than the second threshold, it can be considered that the first audio signal from the first sound source does not include human speech. ZCC refers to the rate of sign change of a signal, that is, the number of times a frame of speech time-domain signal crosses the time axis. It is calculated by shifting all signals within the frame by 1, then multiplying the corresponding points. A negative sign indicates a zero-crossing. The zero-crossing rate of the frame is obtained by calculating the product of all negative numbers within the frame. STE refers to the energy of a single frame of speech signal.
[0091] Optionally, the electronic device can also determine the energy of the first audio signal; and based on the energy of the first audio signal, determine the energy of the virtual noise to be added, etc. Essentially, this method is used to control the signal-to-noise ratio of the first audio signal after adding virtual noise. For example, the signal-to-noise ratio of the first audio signal to the added virtual noise can be preset to 50%. For instance, if the energy of the first audio signal is W1, then the energy of the virtual noise to be added can be half the energy of the first audio signal, i.e., 0.5W1.
[0092] Optionally, in this embodiment, the external audio signal collected by the microphone may be a mixed audio signal, which includes audio signals output from multiple sound sources. The electronic device can separate the mixed audio signal collected by the microphone to obtain a first audio signal corresponding to a first sound source. Then, the above method is used to determine the audio content in the first audio signal, and / or the energy of the first audio signal, etc.
[0093] In one possible implementation, a simple frequency domain speech separation algorithm can be adopted, as described in the industry paper "Research on Key Technologies in Multi-channel Speech Signal Processing - Sound Field Reconstruction and Speech Separation" (Doctoral Dissertation of Dalian University of Technology, Wang Lin). This algorithm can separate the mixed audio signal collected by the microphone into multiple independent audio signals, among which the multiple independent audio signals include the first audio signal corresponding to the first sound source.
[0094] Given N independent sound sources and M microphones, the sound source vector is (n) = [s1(n),...s2(n)]. N (n)] T The observation vector is x(n) = [x1(n),...x M (n)] T If the length of the mixing filter is P, then the convolution mixing process can be expressed as:
[0095]
[0096] The hybrid network H(n) is an M*N matrix sequence composed of the impulse responses of the hybrid filters. Let the length of the separating filter be L, and the estimated sound source vector be y(n) = [y1(n),...y...]. N (n)] T Its expression is:
[0097]
[0098] Here, the separation network W(n) is an N*M matrix, which is composed of the impulse responses of the separation filters, and * denotes matrix convolution operation.
[0099] The separation network W(n) can be obtained using a frequency-domain blind source separation algorithm. After undergoing an L-point short-time Fourier transform (STFT), the time-domain convolution is transformed into a frequency-domain multiplication, i.e.
[0100] X(m,f)=H(f)S(m,f)
[0101] Y(m,f)=W(f)X(m,f)
[0102] Where m is obtained by downsampling the time index value n by L points, X(m,f) and Y(m,f) are obtained by performing STFT on x(n) and y(n) respectively, H(f) and W(f) are the Fourier transform forms of H(n) and W(n) respectively, and f∈[f0,...f L / 2 [ ] represents frequency.
[0103] By inversely transforming Y(m,f) obtained after blind source separation back to the time domain, the estimated sound source signals y1(n),...y are obtained. N (n).
[0104] It should be noted that in the embodiments of this application, the first sound source that the user is not interested in may include one or more sound sources. When multiple sound sources are included, each sound source is processed according to steps 402 and 403 above, that is, the HRTF corresponding to each sound source is obtained, and each HRTF is multiplied by a preset virtual noise frequency domain to obtain the noise signal corresponding to each sound source. By playing the noise signal of each sound source, the noise signal can be superimposed on the corresponding sound source, thereby reducing the signal-to-noise ratio of the sound sources that the user is not interested in, allowing the user to focus more on listening to the audio signal of the sound source that is of interest, and assisting in the selection of listening.
[0105] For example, this application embodiment provides an electronic device for implementing the assisted listening method provided in this application embodiment. The electronic device at least performs the following two functions:
[0106] Part 1: Detection and Recognition.
[0107] Modules in electronic devices that implement detection and recognition functions may include an environmental information acquisition module, an environmental information processing module, and an attention detection module. The environmental information acquisition module is used to collect audio signals from the environment and can be environmental information acquisition sensors such as microphones and cameras deployed in the electronic device. The environmental information processing module can determine the location of the sound source corresponding to the collected audio signals. The attention detection module is used to determine the direction of the user's attention, such as the direction of the user's head or the direction of the user's gaze, and can be an IMU or camera deployed in the electronic device.
[0108] Part Two: Functional Processing.
[0109] The modules in an electronic device that implement this function may include an audio processing module and an audio playback module. The audio processing module, used to add noise to audio signals that are not of interest to the user, can be a processor or similar component deployed in the electronic device. The audio playback module, used to play back the added noise signal, can be a speaker or similar component deployed in the electronic device.
[0110] This application provides a specific example of an assisted listening method, which includes at least the following steps:
[0111] Step 1: Direction recognition.
[0112] Detecting user focus areas. Methods for detecting user focus areas include, but are not limited to, the following two:
[0113] The first method involves deploying a camera on the electronic device to detect the direction of the user's gaze. This detected direction of the user's gaze is the direction of the user's attention.
[0114] The second method involves deploying a brainwave detection sensor on the electronic device. This sensor detects the difference in electrical current between the user's two ears. Based on the difference in current between the two ears in different gaze directions, the user's gaze direction is determined. Similarly, this gaze direction is the user's focus direction.
[0115] Step 2: Location identification.
[0116] Detect and identify sound sources in directions not of interest to the user. For example, an electronic device may have a microphone array deployed on it, which can be used to detect the coordinates of all sound sources near the user. Based on the identified user-interested direction, select the sound sources not of interest to the user from all sound sources near the user. Optionally, there can be one or more sound sources not of interest to the user, without limitation. For example, there can be n sound sources not of interest to the user, and the coordinates of these n sound sources can be p1(x1,y1,z1),...,pn(xn,yn,zn), where n is a positive integer greater than 1.
[0117] Step 3: Binocular rendering, see [link / reference] Figure 8 As shown.
[0118] 1. Obtain the binaural HRTF based on the coordinates of the user's non-interested sound sources. Optionally, the electronic device can pre-store multiple HRTFs for different locations. By interpolating these multiple HRTFs, the HRTF corresponding to the coordinates of the user's non-interested sound sources can be obtained. Continuing with the above example, if the coordinates of the above n non-interested sound sources are p1(x1,y1,z1),...,pn(xn,yn,zn), then in this embodiment, the binaural HRTFs of the above n non-interested sound sources can be obtained respectively.
[0119] 2. Process the virtual noise with the aforementioned binaural HRTF, such as through temporal convolution or frequency multiplication, to obtain the binaural audio signal, which is then played back in real time. This playback method can be acoustic, bone conduction, or other methods; there are no limitations.
[0120] The virtual noise mentioned above can be a noise audio file, which can be stored on the electronic device or in the cloud. When needed, the electronic device downloads these noise audio files from the cloud and renders and plays them. Continuing with the example above, the n non-interesting sound sources correspond to n binaural HRTFs, and the n binaural HRTFs can correspond to n virtual noises. The n binaural HRTFs and their corresponding n virtual noises can be processed to obtain n binaural noise signals, which are then played. Optionally, the n virtual noises can be n1(n), ..., nn(n), and these n virtual noises can be the same or different.
[0121] By superimposing virtual noise onto the audio signal of the uninterested sound source, the clarity of the audio signal of the uninterested sound source is reduced, thereby improving the perceived clarity of the audio signal of the sound source that the user is interested in.
[0122] The methods provided in the embodiments of this application above are described from the perspective of an electronic device as the executing entity. To implement the functions of the methods provided in the embodiments of this application above, the electronic device may include hardware structures and / or software modules, implementing the above functions in the form of hardware structures, software modules, or a combination of hardware structures and software modules. Whether a particular function is executed in the form of hardware structures, software modules, or a combination of hardware structures and software modules depends on the specific application and design constraints of the technical solution.
[0123] like Figure 9 As shown in the figure, this application provides an auxiliary listening device, which includes at least a processing unit 901 and a playback unit 902.
[0124] The processing unit 901 is configured to determine the coordinates of a first sound source based on the external audio signal collected by the microphone and a first direction, wherein the first direction is a direction determined by user detection and the first sound source is a sound source in a direction other than the first direction; the processing unit 901 is also configured to determine a first HRTF corresponding to the coordinates of the first sound source based on the coordinates of the first sound source and a preset HRTF; the processing unit 901 is also configured to obtain a noise signal corresponding to the first sound source based on the first HRTF and a preset virtual noise; and the playback unit 902 is configured to play the noise signal.
[0125] In one possible implementation, determining the coordinates of the first sound source based on the external audio signal collected by the microphone and the first direction includes: determining the coordinates of at least one sound source near the user based on the external audio signal collected by the microphone; detecting the user and determining the first direction; and determining the coordinates of the first sound source based on the first direction and the coordinates of the at least one sound source near the user.
[0126] In one possible implementation, determining the coordinates of at least one sound source near the user based on the external audio signals collected by the microphones includes: having multiple microphones, each microphone collecting external audio signals, with a time delay between the external audio signals collected by different microphones; and determining the coordinates of at least one sound source near the user based on the time delay of the external audio signals collected by different microphones.
[0127] In one possible implementation, detecting the user and determining the first direction includes: detecting the user's gaze direction; or, detecting the difference in the user's binocular currents and determining the user's gaze direction based on the correspondence between the difference in binocular currents and the gaze direction, wherein the user's gaze direction is the first direction.
[0128] In one possible implementation, determining the coordinates of the first sound source based on the first direction, within the coordinates of at least one sound source near the user, includes: analyzing the coordinates of the at least one sound source near the user to determine the directional relationship between each sound source and the user; determining the deviation between each sound source and the first direction based on the first direction and the directional relationship between the sound source and the user; and selecting, from the at least one sound source near the user, a sound source whose deviation from the first direction is greater than a threshold, and the coordinates of that sound source are the coordinates of the first sound source.
[0129] In one possible implementation, the processing unit 901 is further configured to: the external audio signal collected by the microphone is a mixed audio signal, which includes audio signals output by multiple sound sources; and to separate the external audio signal collected by the microphone to obtain a first audio signal output by a first sound source.
[0130] In one possible implementation, the processing unit 901 is further configured to: analyze the separated first audio signal to determine the content of the first audio signal; and determine the type of virtual noise to be added based on the content of the first audio signal.
[0131] In one possible implementation, if the content of the first audio signal is human conversation, then the type of virtual noise to be added is multi-person conversation babble noise.
[0132] In one possible implementation, the processing unit 901 is further configured to: determine the energy of the separated first audio signal; and determine the energy of the virtual noise to be added based on the energy of the first audio signal.
[0133] This application also provides a computer-readable storage medium including a program, which, when run by a processor, executes the methods described in the above method embodiments.
[0134] A computer program product comprising computer program code, which, when run on a computer, causes the computer to implement the methods described in the above method embodiments.
[0135] A chip includes: a processor coupled to a memory for storing programs or instructions that, when executed by the processor, cause a device to perform the methods described in the above method embodiments.
[0136] In the description of this application, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can represent A or B. "And / or" in this application merely describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. Additionally, to facilitate a clear description of the technical solutions of the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or the order of execution, and that the words "first" and "second" do not necessarily imply that they are different.
[0137] In this application embodiment, the processor can be a general-purpose processor, digital signal processor, application-specific integrated circuit, field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in this application embodiment. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0138] In the embodiments of this application, the memory can be non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or it can be volatile memory, such as random-access memory (RAM). Memory is any other medium capable of carrying or storing desired program code in the form of instructions or data structures, and accessible by a computer, but is not limited thereto. The memory in the embodiments of this application can also be a circuit or any other device capable of implementing storage functions, used to store program instructions and / or data.
[0139] The methods provided in this application can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented in software, they can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media (e.g., SSDs), etc.
[0140] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
[0141] It should be noted that a portion of this patent application contains copyrighted material. The copyright holder retains all rights except for making copies of the contents of patent documents or records from the patent office.
Claims
1. An assisted listening method, characterized in that, include: Based on the external audio signal collected by the microphone and the first direction, the coordinates of the first sound source are determined. The first direction is the direction determined by user detection, and the first sound source is a sound source in a direction other than the first direction. Based on the coordinates of the first sound source and the preset head-related transfer function (HRTF), determine the first HRTF corresponding to the first sound source; Based on the first HRTF and the preset virtual noise, the noise signal corresponding to the first sound source is obtained, and the noise signal is played. Determining the coordinates of the first sound source based on the external audio signal collected by the microphone and the first direction includes: Based on the external audio signals collected by the microphone, determine the coordinates of at least one sound source near the user; The user is detected to determine the first direction; Based on the first direction, determine the coordinates of the first sound source in the coordinates of at least one sound source near the user; Determining the coordinates of the first sound source in the coordinates of at least one sound source near the user according to the first direction includes: The coordinates of at least one sound source near the user are analyzed to determine the directional relationship between each sound source and the user. Based on the first direction and the directional relationship between the sound source and the user, determine the deviation between each sound source and the first direction; Among at least one sound source near the user, a sound source with a deviation greater than a threshold from the first direction is selected, and the coordinates of the sound source are the coordinates of the first sound source.
2. The method as described in claim 1, characterized in that, Determining the coordinates of at least one sound source near the user based on the external audio signal collected by the microphone includes: There are multiple microphones, each of which collects external audio signals, and there is a time delay between the external audio signals collected by different microphones; Based on the time delay of the external audio signals collected by different microphones, determine the coordinates of at least one sound source near the user.
3. The method as described in claim 1 or 2, characterized in that, The step of detecting the user and determining the first direction includes: Detect the user's gaze direction; or, The difference in current between the user's two ears is detected, and the user's gaze direction is determined based on the correspondence between the difference in current between the two ears and the gaze direction. The user's gaze direction is the first direction.
4. The method according to any one of claims 1 to 3, characterized in that, Also includes: The external audio signal collected by the microphone is a mixed audio signal, which includes audio signals output from multiple sound sources; The external audio signal collected by the microphone is separated to obtain the first audio signal output by the first sound source.
5. The method as described in claim 4, characterized in that, Also includes: The separated first audio signal is analyzed to determine its content; Based on the content of the first audio signal, determine the type of virtual noise that needs to be added.
6. The method as described in claim 5, characterized in that, If the content of the first audio signal is human conversation, then the type of virtual noise to be added is multi-person conversation babble noise.
7. The method according to any one of claims 4 to 6, characterized in that, Also includes: Determine the energy of the separated first audio signal; The energy of the virtual noise to be added is determined based on the energy of the first audio signal.
8. An auxiliary listening device, characterized in that, include: The processing unit is used to determine the coordinates of a first sound source based on the external audio signal collected by the microphone and a first direction, wherein the first direction is a direction determined by user detection and the first sound source is a sound source in a direction other than the first direction. The processing unit is further configured to determine the first HRTF corresponding to the coordinates of the first sound source based on the coordinates of the first sound source and the preset HRTF; The processing unit is further configured to obtain the noise signal corresponding to the first sound source based on the first HRTF and the predetermined virtual noise. A playback unit is used to play the noise signal; Determining the coordinates of the first sound source based on the external audio signal collected by the microphone and the first direction includes: Based on the external audio signals collected by the microphone, determine the coordinates of at least one sound source near the user; The user is detected to determine the first direction; Based on the first direction, determine the coordinates of the first sound source in the coordinates of at least one sound source near the user; Determining the coordinates of the first sound source in the coordinates of at least one sound source near the user according to the first direction includes: The coordinates of at least one sound source near the user are analyzed to determine the directional relationship between each sound source and the user. Based on the first direction and the directional relationship between the sound source and the user, determine the deviation between each sound source and the first direction; Among at least one sound source near the user, a sound source whose deviation from the first direction is greater than a threshold is selected, and the coordinates of the sound source are the coordinates of the first sound source.
9. The apparatus as claimed in claim 8, characterized in that, Determining the coordinates of at least one sound source near the user based on the external audio signal collected by the microphone includes: There are multiple microphones, each of which collects external audio signals, and there is a time delay between the external audio signals collected by different microphones; Based on the time delay of the external audio signals collected by different microphones, determine the coordinates of at least one sound source near the user.
10. The apparatus as claimed in claim 8 or 9, characterized in that, The step of detecting the user and determining the first direction includes: Detect the user's gaze direction; or, The difference in current between the user's two ears is detected, and the user's gaze direction is determined based on the correspondence between the difference in current between the two ears and the gaze direction. The user's gaze direction is the first direction.
11. The apparatus as claimed in any one of claims 8 to 10, characterized in that, The processing unit is also used for: The external audio signal collected by the microphone is a mixed audio signal, which includes audio signals output from multiple sound sources; The external audio signal collected by the microphone is separated to obtain the first audio signal output by the first sound source.
12. The apparatus as claimed in claim 11, characterized in that, The processing unit is also used for: The separated first audio signal is analyzed to determine its content; Based on the content of the first audio signal, determine the type of virtual noise that needs to be added.
13. The apparatus as claimed in claim 12, characterized in that, If the content of the first audio signal is human conversation, then the type of virtual noise to be added is multi-person conversation babble noise.
14. The apparatus as claimed in any one of claims 11 to 13, characterized in that, The processing unit is also used for: Determine the energy of the separated first audio signal; The energy of the virtual noise to be added is determined based on the energy of the first audio signal.
15. An electronic device, characterized in that, The electronic device includes a memory and one or more processors, wherein the memory is used to store computer program code, the computer program code including computer instructions; when the computer instructions are executed by the processor, the electronic device performs the method as described in any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that, Includes a program or instructions that, when run on a computer, cause the method as described in any one of claims 1 to 7 to be performed.
Citation Information
Patent Citations
Personalized, real-time audio processing
CN108601519A
Generating a modified audio experience for an audio system
US10638248B1