Binaural sound pickup method and apparatus
By employing blind source separation and signal adjustment technologies in terminal devices, the problem of discrepancies between human voices and ambient sounds in TWS earphone recordings has been solved, resulting in a more natural recording effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-02
- Publication Date
- 2026-04-07
AI Technical Summary
When recording a person's voice using TWS earphones, if the person recording is too close to the microphone and the ambient sound source is too far from the microphone, the person's voice will be louder than the ambient sound, affecting the recording quality.
By using a microphone to acquire audio signals through a terminal device, blind source separation is performed to determine the target signal and non-target signal. After weight adjustment, they are fused to achieve energy adjustment and fusion of the target human voice and non-target human voice, resulting in a more natural binaural pickup result.
It improves the naturalness of the recording, balances the recorded human voice with the ambient sound, and enhances the recording effect.
Smart Images

Figure CN116781817B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to a binaural sound pickup method and device. Background Technology
[0002] With the popularization and development of terminal devices, true wireless stereo (TWS) earbuds have become increasingly common devices for audio or video recording. When using TWS earbuds for audio recording, the binaural pickup technology of TWS earbuds can simulate the way human ears hear, making the sound recorded by TWS earbuds have a better sense of space and presence.
[0003] However, when using TWS earphones to record a person's voice, the recording results in a situation where the person's voice is too loud and the ambient sound is too quiet because the person is too close to the microphone in the TWS earphone and other sound sources in the environment are too far away from the microphone in the TWS earphone, thus affecting the recording quality. Summary of the Invention
[0004] This application provides a binaural sound pickup method and apparatus, enabling a terminal device to extract target human voices and non-target human voices in a scene, and to obtain a more natural binaural sound pickup result by adjusting and fusing the energy of the target human voices and non-target human voices.
[0005] In a first aspect, embodiments of this application provide a binaural pickup method applied to a terminal device. The method includes: the terminal device acquiring audio signals using a microphone; the terminal device performing blind source separation on the audio signals to obtain N audio signals; N being an integer greater than or equal to 2; the terminal device determining a target signal from the N audio signals; the terminal device adjusting the target signal using a first weight and adjusting the non-target signals using a second weight to obtain an adjusted target signal and an adjusted non-target signal; the non-target signal being any signal other than the target signal in the N audio signals; and the terminal device fusing the adjusted target signal and the adjusted non-target signal into a binaural pickup result. This allows the terminal device to extract target human voices and non-target human voices from a scene separately, and through energy adjustment and fusion of the target human voices and non-target human voices, obtain a binaural pickup result with a more natural sound perception.
[0006] The target signal can be a signal containing the target human voice in the embodiments of this application; the non-target signal can be a signal other than the target signal in the audio signal in the embodiments of this application.
[0007] In one possible implementation, the terminal device determines the target signal from N audio signals by selecting one signal from the N audio signals that satisfies a first preset direction and contains the target human voice. This allows the terminal device to determine the speaker's voice signal during recording based on the first preset direction and the target human voice.
[0008] The first preset direction can be directly in front of the microphone of the terminal device.
[0009] In one possible implementation, the target human voice is a sound that satisfies a preset frequency and / or a preset harmonic. The preset frequency can be a frequency range, for example, 50-4000 Hz.
[0010] In one possible implementation, the microphones include two microphones in an earphone connected to the terminal device. The direction of the signal is estimated by the terminal device from the direction of arrival (DOA) of the reconstructed signal corresponding to one of the N audio signals. The reconstructed signal is obtained by reconstructing one of the N audio signals, and this reconstruction process maps one of the N audio signals to the two microphones in the earphone. This allows the terminal device to more accurately determine the direction of the signal based on the DOA estimation of the reconstructed signal.
[0011] In one possible implementation, the method further includes: the terminal device calculating the energy difference between the target signal and the non-target signal; the terminal device adjusting the target signal using a first weight and adjusting the non-target signal using a second weight to obtain the adjusted target signal and the adjusted non-target signal, including: when the terminal device determines that the difference is greater than a first threshold, the terminal device adjusting the target signal using the first weight and adjusting the non-target signal using the second weight to obtain the adjusted target signal and the adjusted non-target signal. Thus, when multiple target signals that satisfy both a first preset direction and the target human voice appear in the terminal device, the terminal device can more accurately identify the target signal through the energy difference between the target signal and the non-target signal.
[0012] In one possible implementation, the microphone includes two microphones in an earphone connected to the terminal device. The terminal device uses the microphones to acquire audio signals. The implementation includes: the terminal device displaying a first interface, which includes a first option for recording using the two microphones in the earphone; when the terminal device receives an operation to select the first option, the terminal device uses the two microphones in the earphone to acquire audio signals. This allows users to select different microphones for audio recording according to their recording needs, thereby improving the user experience of using the recording function.
[0013] The first interface can be an interface for starting recording, and the first option can be the option corresponding to the standard recording (AI recording) in this embodiment of the application.
[0014] In one possible implementation, the method further includes: when the terminal device receives an operation to set a recording mode, the terminal device displays a second interface; the second interface includes: a first control for setting the audio signal to be acquired using the two microphones in the headphones during recording; when the terminal device receives an operation to select the first option, the terminal device acquires the audio signal using the two microphones in the headphones, including: when the first control is enabled, when the terminal device receives an operation to select the first option, the terminal device acquires the audio signal using the two microphones in the headphones. This allows users to select a suitable recording mode according to their recording needs, enabling audio recording using different microphones corresponding to that recording mode, thereby improving the user experience of using the recording function.
[0015] The second interface can be used to set recording permissions, and the first control can be the control corresponding to AI recording in the recording mode. AI recording can be understood as using the microphone in the headphones to obtain audio signals when the user is wearing headphones.
[0016] In one possible implementation, the method further includes: when the terminal device receives an operation to select the first option while the first control is in a closed state, the terminal device uses its microphone to acquire an audio signal. This allows the user to select a suitable recording mode according to their recording needs, enabling audio recording using different microphones corresponding to that recording mode, thereby improving the user experience of the recording function.
[0017] In one possible implementation, the microphone includes: at least one microphone in the terminal device, and two microphones in an earphone connected to the terminal device. The method further includes: the terminal device calculating a forward beam corresponding to the audio signal; the forward beam being used to suppress audio signals located not directly in front of the microphone, and to preserve audio signals located directly in front of the microphone; the terminal device calculating the correlation values between the forward beam and N audio signals respectively; and the terminal device determining a target signal from the N audio signals, including: the terminal device selecting one signal from the N audio signals whose correlation value is greater than a second threshold as the target signal. This allows the terminal device to more accurately identify the target signal based on the correlation between the forward beam and the N audio signals respectively.
[0018] In one possible implementation, the audio signal includes a first audio signal, a second audio signal, and a third audio signal. The terminal device calculates the forward beam corresponding to the audio signal, including: the terminal device acquiring filter coefficients corresponding to a second direction; the second direction being the direction directly in front of the microphone; the terminal device using the filter coefficients corresponding to the second direction, combined with the first audio signal, the second audio signal, and the third audio signal, to obtain the forward beam corresponding to the audio signal. This allows the terminal device to accurately calculate the forward beam using the audio signals corresponding to the three microphones respectively.
[0019] In one possible implementation, the method further includes: the terminal device displaying a first interface, the first interface including a second option for recording using two microphones in the headset and at least one microphone in the terminal device; when the terminal device receives an operation to select the second option, the terminal device uses the two microphones in the headset and at least one microphone in the terminal device to acquire audio signals. This allows users to select different microphones for audio recording according to their recording needs, thereby improving the user experience of using the recording function.
[0020] In one possible implementation, the method further includes: when the terminal device receives an operation to end recording, the terminal device encodes the binaural pickup results into a first voice and stores the first voice; when the terminal device receives an operation to start the recording application, the terminal device displays a third interface; wherein the third interface includes the first voice and a first identifier corresponding to the first voice; the first identifier indicates that the first voice was recorded using two microphones in the earphones, or using two microphones in the earphones and at least one microphone in the terminal device. This allows the user to accurately determine which microphone recorded the voice based on the first identifier corresponding to the voice.
[0021] In one possible implementation, the method further includes: the terminal device performing a Fourier transform on the audio signal to obtain the Fourier-transformed audio signal; and the terminal device performing blind source separation on the audio signal to obtain N audio signals, including: the terminal device performing blind source separation on the Fourier-transformed audio signal to obtain N audio signals. This allows the terminal device to convert the time-domain audio signal into a frequency-domain audio signal using a Fourier transform, facilitating subsequent signal processing.
[0022] In one possible implementation, the terminal device fuses the adjusted target signal and the adjusted non-target signal into a binaural pickup result, including: the terminal device fuses the adjusted target signal and the adjusted non-target signal into a fourth audio signal; the terminal device performs an inverse Fourier transform on the fourth audio signal to obtain the binaural pickup result. This allows the terminal device to convert the frequency domain audio signal into a time domain audio signal using the inverse Fourier transform, facilitating subsequent signal processing.
[0023] Secondly, embodiments of this application provide a binaural pickup device, comprising a processing unit for acquiring audio signals using a microphone; the processing unit is further configured to perform blind source separation on the audio signals to obtain N audio signals; N is an integer greater than or equal to 2; the processing unit is further configured to determine a target signal among the N audio signals; the processing unit is further configured to adjust the target signal using a first weight and adjust the non-target signal using a second weight to obtain an adjusted target signal and an adjusted non-target signal; the non-target signal is any signal other than the target signal among the N audio signals; the processing unit is further configured to fuse the adjusted target signal and the adjusted non-target signal into a binaural pickup result.
[0024] In one possible implementation, the processing unit is specifically used by the terminal device to select one of the N audio signals that satisfies a first preset direction and contains the target human voice as the target signal.
[0025] In one possible implementation, the target human voice is a sound that satisfies a preset frequency and / or a preset harmonic.
[0026] In one possible implementation, the microphone includes two microphones in an earphone connected to a terminal device, wherein the direction of the signal is estimated by the terminal device from the direction of arrival (DOA) of the reconstructed signal corresponding to one of the N audio signals; wherein the reconstructed signal is obtained by reconstructing one of the N audio signals, and the reconstructing process is used to map one of the N audio signals to the two microphones in the earphone.
[0027] In one possible implementation, the processing unit is further configured to calculate the difference between the energy of the target signal and the energy of the non-target signal; when the terminal device determines that the difference is greater than a first threshold, the processing unit is further configured to adjust the target signal using a first weight and adjust the non-target signal using a second weight to obtain the adjusted target signal and the adjusted non-target signal.
[0028] In one possible implementation, the microphone includes: two microphones in an earpiece connected to a terminal device; a display unit for displaying a first interface, the first interface including: a first option for recording using the two microphones in the earpiece; and a processing unit for acquiring audio signals using the two microphones in the earpiece when the terminal device receives an operation to select the first option.
[0029] In one possible implementation, when the terminal device receives an operation to set the recording mode, the display unit is further configured to display a second interface; the second interface includes: a first control for setting the acquisition of audio signals using two microphones in the headphones during recording; when the terminal device receives an operation to select the first option while the first control is in the enabled state, the processing unit is further configured to acquire audio signals using the two microphones in the headphones.
[0030] In one possible implementation, when the terminal device receives an operation to select the first option while the first control is in the closed state, the processing unit is also used to acquire audio signals using the microphone in the terminal device.
[0031] In one possible implementation, the microphone includes: at least one microphone in the terminal device, and two microphones in the earphone connected to the terminal device; the processing unit is further configured to calculate a forward beam corresponding to the audio signal; the forward beam is used to suppress audio signals located not directly in front of the microphone, and to retain audio signals located directly in front of the microphone; the processing unit is further configured to calculate the correlation values between the forward beam and N audio signals respectively; the processing unit is further configured to select one signal among the N audio signals whose correlation value is greater than a second threshold as the target signal.
[0032] In one possible implementation, the audio signal includes: a first audio signal, a second audio signal, and a third audio signal. The processing unit is specifically used to obtain filter coefficients corresponding to a second direction; the second direction is the direction directly in front of the microphone. The processing unit is specifically used to use the filter coefficients corresponding to the second direction, combined with the first audio signal, the second audio signal, and the third audio signal, to obtain the forward beam corresponding to the audio signal.
[0033] In one possible implementation, the display unit is further configured to display a first interface, the first interface including: a second option for recording using two microphones in the headset and at least one microphone in the terminal device; when the terminal device receives an operation to select the second option, the processing unit is further configured to acquire audio signals using the two microphones in the headset and at least one microphone in the terminal device.
[0034] In one possible implementation, when the terminal device receives an operation to end recording, the processing unit is further configured to encode the binaural pickup result into a first voice and store the first voice; when the terminal device receives an operation to start the recording application, the display unit is further configured to display a third interface; wherein the third interface includes the first voice and a first identifier corresponding to the first voice; the first identifier is used to indicate that the first voice was recorded based on two microphones in the earphone, or based on two microphones in the earphone and at least one microphone in the terminal device.
[0035] In one possible implementation, the processing unit is further configured to perform a Fourier transform on the audio signal to obtain a Fourier-transformed audio signal; the processing unit is further configured to perform blind source separation on the Fourier-transformed audio signal to obtain N audio signals.
[0036] In one possible implementation, the processing unit is specifically used to fuse the adjusted target signal and the adjusted non-target signal into a fourth audio signal; the processing unit is also specifically used to perform an inverse Fourier transform on the fourth audio signal to obtain the binaural pickup result.
[0037] Thirdly, embodiments of this application provide a binaural pickup device, including a processor and a memory, the memory being used to store code instructions; the processor being used to run the code instructions, causing the electronic device to perform the binaural pickup method as described in the first aspect or any implementation thereof.
[0038] Fourthly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed, cause a computer to perform the binaural pickup method as described in the first aspect or any implementation thereof.
[0039] Fifthly, a computer program product comprising a computer program that, when run, causes the computer to perform the binaural pickup method as described in the first aspect or any implementation thereof.
[0040] It should be understood that the second to fifth aspects of this application correspond to the technical solutions of the first aspect of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be repeated here. Attached Figure Description
[0041] Figure 1 A scenario diagram provided for an embodiment of this application;
[0042] Figure 2 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application;
[0043] Figure 3A flowchart illustrating a binaural pickup method provided in an embodiment of this application;
[0044] Figure 4 A schematic diagram of an interface for starting recording provided in an embodiment of this application;
[0045] Figure 5 A schematic diagram illustrating the principle of DOA estimation provided in this application embodiment;
[0046] Figure 6 A flowchart illustrating another binaural pickup method provided in an embodiment of this application;
[0047] Figure 7 A schematic diagram of another interface for starting recording provided in an embodiment of this application;
[0048] Figure 8 A schematic diagram of a process for generating a forward beam is provided for an embodiment of this application;
[0049] Figure 9 A schematic diagram of the orientation provided for an embodiment of this application;
[0050] Figure 10 A beamforming pattern provided in an embodiment of this application;
[0051] Figure 11 A schematic diagram of an interface for enabling AI recording provided in an embodiment of this application;
[0052] Figure 12 A schematic diagram of an interface for displaying a recording identifier is provided in an embodiment of this application;
[0053] Figure 13 This is a schematic diagram of the structure of a binaural pickup device provided in an embodiment of this application;
[0054] Figure 14 This is a schematic diagram of the hardware structure of another terminal device provided in an embodiment of this application;
[0055] Figure 15 This is a schematic diagram of the structure of a chip provided in an embodiment of this application. Detailed Implementation
[0056] The following description of the terminology used in the embodiments of this application is provided. It should be understood that this description is intended to provide a clearer explanation of the embodiments of this application and does not necessarily constitute a limitation thereof.
[0057] (1) Blind source separation (BSS)
[0058] In this embodiment of the application, blind source separation can also be called blind signal separation, which is a method for separating the source signal from the received mixed signal without knowing the source signal and the signal mixing parameters.
[0059] The blind source separation method may include: Independent Vector Analysis (IVA), Independent Component Analysis (ICA), or Non-negative Matrix Factorization (NMF), etc.
[0060] (2) Fundamental wave and harmonics
[0061] In this embodiment, the fundamental wave can be a sinusoidal component that is equal to the longest period of the oscillation in a complex periodic oscillation, and the frequency corresponding to the above period is called the fundamental frequency; the harmonic can be a sinusoidal component that is an integer multiple of the fundamental frequency.
[0062] (3) Direction of arrival (DOA) estimation
[0063] In this embodiment of the application, DOA estimation can be a method for obtaining the target's distance and orientation information by processing the received echo signal of the target. The echo signal can be the audio signal of a sound source.
[0064] (4) Inhibition
[0065] In this embodiment, suppression refers to reducing the energy of an audio signal so that it becomes quieter or even inaudible. Suppression of the audio signal can be achieved by reducing its amplitude.
[0066] The amplitude value represents the voltage level of the audio signal; it can also represent the energy level of the audio signal; or the decibel level.
[0067] (5) Beamforming and Gain Coefficient
[0068] In this embodiment, beamforming can be used to describe the correspondence between the audio captured by the microphone of the terminal device and the audio transmitted to the speaker for playback. This correspondence is a set of gain coefficients used to represent the degree of suppression of the audio signals captured by the microphone in each direction.
[0069] In audio signal reduction, the energy of the audio signal is reduced to make it quieter or even inaudible. The degree of reduction describes the extent to which the audio signal is reduced. A higher degree of reduction means a greater reduction in the audio signal's energy. For example, a gain of 0.0 indicates complete removal of the audio signal, while a gain of 1.0 indicates no reduction. A gain closer to 0.0 indicates a higher degree of reduction, and a gain closer to 1.0 indicates a lower degree of reduction.
[0070] To facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. For example, the first value and the second value are only used to distinguish different values and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0071] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0072] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0073] It is understandable that microphone pickup refers to the process of collecting sound, while binaural pickup refers to the process of collecting sound by using two microphones to simulate the way human ears hear. Since the sound waves emitted by the sound source can be reflected by the environment, the human torso, and the auricle before finally reaching the human ear, allowing the human ear to hear sound with a sense of stereo, the binaural pickup method can preserve the spatial location information of the sound source and the environment to the greatest extent.
[0074] For example, Figure 1This is a schematic diagram of a scenario provided for an embodiment of this application. For example... Figure 1 As shown, this scenario may include: user 101, TWS earphones 102 worn by user 101, and sound source 103. The TWS earphones 102 may include a left earphone and a right earphone; the left earphone contains at least one microphone, and the right earphone contains one microphone. The sound source 103 may include the voices of other users and ambient sounds.
[0075] During the process of user 101 using TWS earphone 102 to record audio (or video), the TWS earphone 102 can record not only user 101's own voice, but also the voices of other users in the sound source 103, as well as ambient sounds.
[0076] However, because user 101 is closer to TWS earphone 102 and sound source 103 is farther away from TWS earphone 102, the recording result shows that user 101's voice is louder and is not in harmony with other sounds in the current environment, such as sound source 103, which affects the audio recording effect.
[0077] In view of this, embodiments of this application provide a binaural sound pickup method, enabling a terminal device to extract target human voice and non-target human voice separately, and to obtain a recording with a more natural sound quality by adjusting and fusing the energy of the target human voice and non-target human voice. The target human voice can be the voice of the person recording, and the non-target human voice can be the voice of someone other than the person recording or ambient sound.
[0078] It is understood that the binaural pickup method provided in the embodiments of this application can be applied not only to, but also to, other applications. Figure 1 The recording scenario shown can also be used in video recording scenarios or live streaming scenarios, etc., which involve sound pickup. This application embodiment does not specifically limit this.
[0079] It is understood that the aforementioned terminal devices can also be referred to as terminals, user equipment (UE), mobile stations (MS), mobile terminals (MT), etc. Terminal devices can be mobile phones with microphones, smart TVs, wearable devices, tablets, computers with wireless transceiver capabilities, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, and so on. The embodiments of this application do not limit the specific technologies or device forms used in the terminal devices.
[0080] Therefore, in order to better understand the embodiments of this application, the structure of the terminal device of the embodiments of this application will be described below. For example, Figure 2 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application.
[0081] The terminal device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, an indicator 192, a camera 193, and a display screen 194, etc.
[0082] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the terminal device. In other embodiments of this application, the terminal device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0083] The processor 110 may include one or more processing units. These processing units may be independent devices or integrated within one or more processors. The processor 110 may also include memory for storing instructions and data.
[0084] USB port 130 is a USB standard compliant interface, which can be a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge terminal devices, and can also be used for data transfer between terminal devices and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other terminal devices, such as AR devices.
[0085] The charging management module 140 receives charging input from the charger. The charger can be a wireless charger or a wired charger. The power management module 141 connects the charging management module 140 to the processor 110.
[0086] The wireless communication function of the terminal device can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor.
[0087] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Antennas in terminal equipment can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.
[0088] The mobile communication module 150 can provide solutions for wireless communication applications including 2G / 3G / 4G / 5G on terminal devices. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation.
[0089] The wireless communication module 160 can provide solutions for wireless communication applications on terminal devices, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), etc.
[0090] The terminal device implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering.
[0091] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the terminal device may include one or N display screens 194, where N is a positive integer greater than 1.
[0092] Terminal devices can achieve shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0093] Camera 193 is used to capture still images or videos. In some embodiments, the terminal device may include one or N cameras 193, where N is a positive integer greater than 1.
[0094] The external storage interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the terminal device. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.
[0095] Internal memory 121 can be used to store executable program code, including instructions. Internal memory 121 may include a program storage area and a data storage area.
[0096] The terminal device can implement audio functions such as music playback and recording through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, and an application processor.
[0097] Audio module 170 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signal. Speaker 170A, also called a "loudspeaker," is used to convert audio electrical signals into sound signals. Terminal devices can listen to music or make hands-free calls through speaker 170A. Receiver 170B, also called a "handset," is used to convert audio electrical signals into sound signals. When the terminal device answers a phone call or voice message, it can listen to the voice by bringing the receiver 170B close to the user's ear. Headphone jack 170D is used to connect wired headphones.
[0098] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. In this embodiment, the terminal device can receive sound signals based on microphone 170C and convert the sound signals into electrical signals that can be further processed. The terminal device can have at least one microphone 170C.
[0099] Sensor module 180 may include one or more of the following sensors: pressure sensor, gyroscope sensor, barometric pressure sensor, magnetic sensor, accelerometer, distance sensor, proximity sensor, fingerprint sensor, temperature sensor, touch sensor, ambient light sensor, or bone conduction sensor, etc. Figure 2 (Not shown in the image).
[0100] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. The terminal device can receive button input and generate key signal inputs related to user settings and function control of the terminal device. Indicator 192 can be an indicator light, used to indicate charging status, power level changes, messages, missed calls, notifications, etc.
[0101] The software system of terminal devices can adopt layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture, etc., which will not be elaborated here.
[0102] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be implemented independently or in combination with each other. The same or similar concepts or processes may not be described again in some embodiments.
[0103] It is understandable that terminal devices can use the microphone in TWS earbuds for audio recording (e.g., Figure 3 (Corresponding embodiments), or, the terminal device can also use the MIC in the TWS earphone and the MIC in the terminal device for audio recording (e.g. Figure 6 (Corresponding embodiments) In the embodiments of this application, the specific recording method is not limited.
[0104] In possible implementations, the terminal device may also adopt the following... Figure 3 or Figure 6 The binaural pickup method in the corresponding embodiment processes the sound received by the MIC during video recording or live streaming, but this application embodiment does not limit this process.
[0105] For example, Figure 3 This is a flowchart illustrating a binaural sound pickup method provided in an embodiment of this application. Figure 3In the corresponding embodiments, the example is given by taking a mobile phone as an example, where the terminal device can use the MIC in the TWS earphone to record audio. This example does not constitute a limitation on the embodiments of this application.
[0106] like Figure 3 As shown, this binaural pickup method may include the following steps:
[0107] S301. When the terminal device receives a user's operation to start recording, the terminal device uses the TWS earphone to obtain the audio signal.
[0108] In this embodiment of the application, the operation of starting the recording can be a trigger operation for the recording function, a voice operation, or other gesture operations, etc., and this embodiment of the application does not limit it.
[0109] The TWS earbuds can establish a communication connection with the terminal device in advance, so that the terminal device can use the MIC in the TWS earbuds to obtain the audio signal when recording begins.
[0110] For example, Figure 4 This is a schematic diagram of an interface for starting recording, provided as an embodiment of this application. When the terminal device receives an operation from the user to open the recording application, the terminal device can display as shown below. Figure 4 The interface shown in Figure 'a' may include: more controls for enabling more functions, a speaker control, an input box for searching for recording files, multiple recording files, and a control 401 for starting recording. The multiple recording files include Recording 1, Recording 2, Recording 3, and Recording 4, etc., and the recording time, recording duration, and playback control are displayed around each recording file.
[0111] like Figure 4 In the interface shown in Figure 'a', when the terminal device receives the user's operation on the control 401 for starting recording, and the terminal device detects that the TWS earphones are in the wearing state, the terminal device can use the MIC in the TWS earphones to obtain the audio signal and display it as shown in Figure 'a'. Figure 4 The interface shown in b. Figure 4 The interface shown in b can include: the waveform of the recorded sound, the recording time, a control 402 for stopping recording, a marker control, and a pause control, etc.
[0112] In possible implementations, the terminal device can also be based on Figure 7 as well as Figure 11 The corresponding implementation enables recording; see details below. Figure 7 as well as Figure 11 The corresponding description.
[0113] In this embodiment of the application, the TWS earphone may include at least two microphones, so the audio signal may include at least the audio signal obtained by the microphone in the left earphone and the audio signal obtained by the microphone in the right earphone.
[0114] For example, if the current environment includes the voice of the person recording (or target voice, target user's voice, target signal) and the voice of the person not recording (or non-target voice, or non-target signal), and the person recording is wearing TWS earphones, the TWS earphones can use the microphones corresponding to the two earphones to obtain audio signals. For example, the microphone of the left earphone can obtain the voice of the person recording and the voice of the person not recording received by the left earphone; the microphone of the right earphone can obtain the voice of the person recording and the voice of the person not recording received by the right earphone. The non-recorded voices may include the voices of other users and ambient sounds.
[0115] S302. The terminal device performs a Fourier transform on the audio signal to obtain the Fourier transformed audio signal.
[0116] Understandably, this Fourier transform is used to convert audio signals in the time domain (or time domain) into audio signals in the frequency domain (or frequency domain).
[0117] For example, if the number of sound sources in the current environment is N, the time-domain sequence of the sound sources can be represented as: s 1(t) s 2(t) , ..., s N(t) ; s represents the source signal, t represents the sampling sequence over time; assuming there are M microphones collecting audio signals, the audio signal corresponding to the sound source obtained by the terminal device can be represented as x 1(t) x 2(t) , ..., x M(t) .
[0118] It is understandable that sound waves need to travel a transmission path from the sound source to the microphone (such as time delay, reflection, and mixing caused by different sound sources entering the same microphone). Therefore, the audio signal x collected by the microphone... m(t) With source signal s n(t) The relationship in the time domain is expressed as follows:
[0119]
[0120] Where τ is the time delay, L is the maximum time delay, and h can be understood as the source signal s n(t) The signal x collected by the MIC m(t) The transmission path between them.
[0121] Furthermore, the terminal device can perform a Fourier transform on the aforementioned time-domain audio signal to obtain the frequency domain relationship between the source signal s and the signal x acquired by the MIC:
[0122] x(ω,t)=A(ω,t)s(ω,t) Formula (2)
[0123] Where N is the number of sound sources and M is the number of microphones, then x is a vector of length M, s is a vector of length N, ω is the frequency, t is the number of frames, and A is an M-row N-column matrix, or can be understood as the transmission path between the source signal s and the signal x acquired by the microphone.
[0124] S303. The terminal device performs blind source separation on the Fourier transform audio signal to obtain the first separated signal and the second separated signal.
[0125] In this embodiment of the application, the blind source separation can be used to separate the original signal from the received mixed signal. The number of original signals separated by the blind source separation method is not limited to the two paths mentioned above, and this embodiment of the application does not limit this.
[0126] For example, if the current environment includes target human voices, such as user A's voice, non-target human voices, such as user B's voice, and ambient sounds, the terminal device can use a blind source separation method to separate multiple audio signals from the acquired audio signal. In one implementation, the terminal device can separate two audio signals, such as separating user A's audio signal and other audio signals, wherein the other audio signals may include user B's audio signal and ambient audio signals. In another implementation, the terminal device can separate three audio signals, such as separating user A's audio signal, user B's audio signal, and ambient audio signals.
[0127] For example, if the current environment includes non-target human voices such as user B's voice, as well as ambient sounds, the terminal device can use a blind source separation method to separate user B's audio signal and ambient audio signal from the acquired audio signal.
[0128] S304. The terminal equipment performs voice search on the first and second separated signals respectively.
[0129] In this embodiment of the application, the terminal device can determine whether a human voice is detected in the audio signal based on a preset frequency and / or preset harmonics.
[0130] For example, when the terminal device finds an audio signal that satisfies a preset frequency and / or harmonic laws in the first separation signal (or the second separation signal), it can be understood that the first separation signal (or the second separation signal) contains human voice. The preset frequency can be the frequency range corresponding to human voice, such as 50-4000 Hz; the harmonics can be sinusoidal components that are integer multiples of the fundamental frequency.
[0131] S305. The terminal device reconstructs the first separation information and the second separation signal respectively to obtain the first reconstructed signal and the second reconstructed signal.
[0132] In this embodiment of the application, the reconstruction can be used to map audio signals to the left and right microphones of the TWS earphone; the first reconstruction signal includes: the signal corresponding to the left microphone in the TWS earphone in the first separation signal, and the signal corresponding to the right microphone in the TWS earphone in the first separation signal; the second reconstruction signal includes: the signal corresponding to the left microphone in the TWS earphone in the second separation signal, and the signal corresponding to the right microphone in the TWS earphone in the second separation signal.
[0133] For example, formula (2) can also be:
[0134] W(ω,t)x(ω,t)=s(ω,t) Formula (3)
[0135] Where W can be an N-row M-column matrix, the other parameters in formula (3) can be found in the description in formula (2), and will not be repeated here.
[0136] It is understandable that the terminal device can use the independence between the source signals s in formula (3) to solve W, and obtain each source signal s through W. The source signal s can be a single-channel signal, or it can be understood as the original sound wave signal emitted by the sound source. Furthermore, after obtaining each source signal s as in formula (3), the reconstructed signal containing the specific sound source can be obtained by keeping one sound source, setting the other sound sources to 0, and multiplying by matrix A in formula (2).
[0137] S306. The terminal equipment performs DOA estimation on the first reconstructed signal and the second reconstructed signal respectively.
[0138] For example, Figure 5 This is a schematic diagram illustrating the principle of DOA estimation provided in an embodiment of this application.
[0139] like Figure 5As shown in diagram 'a', this scenario can include TWS earbuds, which include a left microphone in the left earbud and a right microphone in the right earbud. Sound waves emitted by the sound source reach both the left and right microphones via the environment. Because the sound source is closer to the right microphone, the time it takes for the sound source to reach the right microphone is shorter than the time it takes to reach the left microphone. Simultaneously, the sound source reaches point Q near the right microphone. The angle of arrival of the sound source relative to the TWS earbuds is θ, which can range from 0 degrees (°) to 180°. For example... Figure 5 As shown in a, when the distance between the left and right microphones is d, the delay of the sound source reaching the left and right microphones can be dcosθ.
[0140] Understandably, terminal devices can use the generalized cross-correlation phase transformation (GCC-PHAT) method in DOA estimation to determine the angle of the sound source (e.g., the first reconstructed signal and the second reconstructed signal).
[0141] Specifically, the cross-correlation function between the left MIC and the right MIC. for:
[0142]
[0143] Where IDFI can be represented as a Fourier transform, X a X can be the first reconstructed signal after Fourier transform (this first reconstructed signal can be a frequency domain signal), b It can be the second reconstructed signal after Fourier transform (this second reconstructed signal can be a frequency domain signal), t can be time, which can be identified as the frame number here, f can be the frequency point, and * can be represented as conjugate.
[0144] The delay between the sound source reaching the left and right microphones It can be:
[0145]
[0146] The arrival angle θ of the sound source to the left and right microphones can be:
[0147]
[0148] It is understandable that, such as Figure 5As shown in b, since the user's voice usually comes from directly in front when recording audio while wearing TWS earphones, when θ is in the range of 60°-120°, it can be understood that the sound source is located directly in front of the TWS earphones. However, when the sound source is directly in front of the TWS earphones, the range of θ is not limited to the above 60°-120°, and the value of θ can also be different in different coordinate systems. This embodiment does not impose such limitations.
[0149] Understandably, compared to the terminal device directly performing DOA estimation using the audio signal obtained through the TWS earphone in step S301, using the reconstructed signal for DOA estimation can avoid the influence of complex sound sources in the environment on the accuracy of DOA estimation, thus obtaining a more accurate location.
[0150] S307. The terminal device determines whether the sound source is from the target human voice.
[0151] In this embodiment of the application, when the terminal device determines that a human voice can be detected in the sound source (as shown in step S304, where a human voice is detected based on a preset frequency and / or preset harmonics), and the sound source is located directly in front of the TWS earphones (as shown in step S306, where the angle of arrival θ of the sound source satisfies 60°-120°), then the terminal device can determine that the sound source is from the target human voice. Here, the target human voice can be understood as the voice of the person recording audio while wearing the TWS earphones.
[0152] For example, if the first separated signal (or the first reconstructed signal) in the current scenario includes the audio signal of the person recording audio while wearing TWS earphones; and the second separated signal (or the second reconstructed signal) includes the audio signal of other users and the audio signal corresponding to the car horn in the environment, then there are three scenarios in the process of the terminal device determining whether the sound source is from the target human voice.
[0153] In one implementation, when the first separated signal includes the audio signal of the person recording, and the energy of the audio signal of a car horn in the environment where the second separated signal (or the second reconstructed signal) is located is greater than the energy of the audio signals of other users in that environment, the terminal device determines that a human voice can be found in the first separated signal but not in the second separated signal. Furthermore, the terminal device determines that the first reconstructed signal is located directly in front of the TWS earphones, and the car in the second reconstructed signal is also located directly in front of the TWS earphones. Therefore, the terminal device can determine that the sound source in the first separated signal (or the first reconstructed signal) originates from the target human voice, while the sound source in the second separated signal (or the second reconstructed signal) does not belong to the target human voice.
[0154] In another implementation, when the first separated signal includes the audio signal of the person recording, and the energy of the audio signal of a car horn in the environment where the second separated signal (or the second reconstructed signal) is located is less than the energy of the audio signals of other users in that environment, the terminal device determines that a human voice can be found in both the first and second separated signals. Furthermore, the terminal device determines that the first reconstructed signal is located directly in front of the TWS earphone, while the other users in the second reconstructed signal are located in other directions from the TWS earphone. Therefore, the terminal device can determine that the sound source in the first separated signal (or the first reconstructed signal) originates from the target human voice, while the sound source in the second separated signal (or the second reconstructed signal) does not belong to the target human voice.
[0155] In another implementation, when performing a voice search, the first separated signal includes the audio signal of the person recording the voice, and the energy of the audio signal of a car horn in the environment where the second separated signal (or the second reconstructed signal) is located is less than the energy of the audio signals of other users in that environment. Furthermore, when performing DOA (Directional Aspect) estimation, the energy of the audio signal of a car horn in the environment of the second separated signal (or the second reconstructed signal) is greater than the energy of the audio signals of other users in that environment. In this scenario, the terminal device can determine that a human voice is found in the first separated signal, and it can also find the voice of another user in the second separated signal. Furthermore, the terminal device can determine that the first reconstructed signal is located directly in front of the TWS earphone, and the second reconstructed signal is also located directly in front of the TWS earphone. Therefore, the terminal device can determine that the sound source in the first separated signal (or the first reconstructed signal) originates from the target human voice, and that the sound source in the second separated signal (or the second reconstructed signal) also belongs to the target human voice.
[0156] At this point, the terminal device can further distinguish between the first separated signal (or the first reconstructed signal) and the second separated signal (or the second reconstructed signal) to determine whether the sound source belongs to the target human voice. For example, the terminal device can calculate the energy of the first reconstructed signal and the second reconstructed signal respectively, and determine that the sound source in the reconstructed signal with the higher energy belongs to the target human voice.
[0157] Understandably, sound waves weaken significantly with increasing distance, and the closer a signal is to the terminal device, the higher its energy, and the higher the probability that the sound source in that signal belongs to the target human voice. Therefore, the terminal device can further determine which reconstructed signal has higher energy and thus belongs to the target human voice by judging the energy levels of the two reconstructed signals.
[0158] In this embodiment of the application, when the terminal device determines that a scene from which a sound source is detected is the target human voice (or understands that the target human voice exists in the current scene), the terminal device can execute the step shown in S308; or, when the terminal device determines that a scene from which a sound source is detected is not the target human voice (or understands that the target human voice does not exist in the current scene), the terminal device can execute the step shown in S310.
[0159] S308. The terminal device determines whether the energy difference is greater than the threshold based on the reconstructed signal.
[0160] For example, the terminal device can calculate the energy of the first reconstructed signal, the energy of the second reconstructed signal, and the energy difference between the first and second reconstructed signals, respectively. Further, when the terminal device determines that the energy difference is greater than (or greater than or equal to) an energy threshold, the terminal device can execute the step shown in S309; or when the terminal device determines that the energy difference is less than or equal to (or less than) an energy threshold, the terminal device can execute the step shown in S310.
[0161] S309. The terminal device adjusts the energy of the signal corresponding to the target human voice and the signal corresponding to the non-target human voice based on preset weights.
[0162] In this embodiment, the preset weight range of the signal corresponding to the target human voice can be 0.4-0.7; the preset threshold range of the signal corresponding to the non-target human voice can be 0.8-1.2, and this embodiment does not impose specific limitations on this.
[0163] Specifically, when the first reconstructed signal contains the target human voice and the second reconstructed signal contains a non-target human voice, the terminal device can multiply the audio signals corresponding to the left MIC and the right MIC in the first reconstructed signal by the preset weight corresponding to the target human voice, and multiply the audio signals corresponding to the left MIC and the right MIC in the second reconstructed signal by the preset weight corresponding to the non-target human voice, to obtain the first adjustment signal corresponding to the first reconstructed signal and the second adjustment signal corresponding to the second reconstructed signal.
[0164] S310, terminal equipment performs sound source mixing.
[0165] In this embodiment of the application, the terminal device can perform sound source mixing on the audio signals that have not undergone energy adjustment in step S307 (or S308), such as the first reconstruction signal and the second reconstruction signal; or, the terminal device can also perform sound source mixing on the audio signals that have undergone energy adjustment in step S309, such as the first adjustment signal and the second adjustment signal.
[0166] In one possible implementation, the terminal device can superimpose sound sources based on the number of microphones outputting audio.
[0167] In one implementation, when the number of microphones for output audio is 2, the terminal device can superimpose the audio signal corresponding to the left microphone in the first reconstruction signal (or the first adjustment signal) and the audio signal corresponding to the left microphone in the second reconstruction signal (or the second adjustment signal) to obtain the audio signal corresponding to the left microphone. It can also superimpose the audio signal corresponding to the right microphone in the first reconstruction signal (or the first adjustment signal) and the audio signal corresponding to the right microphone in the second reconstruction signal (or the second adjustment signal) to obtain the audio signal corresponding to the left microphone and the right microphone respectively.
[0168] In another implementation, when the number of microphones outputting audio is 1, the terminal device can superimpose the output audio of the left microphone and the output audio of the right microphone again and divide by 2 to obtain a mixed audio signal, which can then be used as the output audio of a single microphone.
[0169] S311. The terminal device performs an inverse Fourier transform on the mixed audio signal to obtain the binaural pickup result.
[0170] The inverse Fourier transform is the reverse of the Fourier transform in step S302, and it is used to convert the audio signal in the frequency domain into the audio signal in the time domain.
[0171] In possible implementations, during audio recording, the terminal device can process the audio signal acquired by the microphone in real time to obtain binaural pickup results, store the binaural pickup results in real time, and encode the stored binaural pickup results into a recording result when it receives a user's operation to end recording; alternatively, the terminal device can also store the audio signal acquired in step S301 in real time, and when it receives a user's operation to end recording, perform the audio processing steps shown in S302-S311 on the stored audio signal to obtain binaural pickup results and encode them into a recording result. This embodiment of the application does not limit this.
[0172] Understandably, the method of storing audio signals or binaural pickup results in real time by the terminal device can meet the user's audio recording needs for the terminal device.
[0173] In possible implementations, in scenarios such as live streaming or video calls, the terminal device can encode the obtained binaural sound pickup results and the video content acquired by the camera in chronological order and store them as video recording results.
[0174] Understandably, the terminal device encodes the binaural sound pickup results and video content according to time, enabling the terminal device to meet audio or video recording needs in scenarios such as live streaming (or video call scenarios).
[0175] It is understood that the subsequent processing procedure for the binaural pickup results is not specifically limited in the embodiments of this application.
[0176] Based on this, the terminal device can extract the target human voice and the non-target human voice separately, and by adjusting and blending the energy of the target human voice and the non-target human voice, a recording with a more natural sound can be obtained.
[0177] Among the possible implementations, in Figure 3 Based on the corresponding embodiments, the terminal device can also use the MIC in the TWS earphone and the MIC in the terminal device for binaural sound pickup.
[0178] For example, Figure 6 This is a flowchart illustrating another binaural sound pickup method provided in an embodiment of this application. Figure 6 In the corresponding embodiment, the example is given by taking a mobile phone as an example, where the terminal device can use the MIC in the TWS earphone and the MIC in the terminal device to perform binaural sound pickup. This example does not constitute a limitation on the embodiments of this application.
[0179] like Figure 6 As shown, this binaural pickup method may include the following steps:
[0180] S601. When the terminal device receives an operation from the user to trigger advanced recording, the terminal device uses the TWS earphone and the terminal device to obtain the audio signal.
[0181] In this embodiment of the application, the advanced recording can be a recording mode in a recording application. The advanced recording can be understood as using two microphones in the TWS earphone and one microphone in the terminal device to simultaneously acquire audio signals.
[0182] It is understood that a terminal device may include one, two, or three microphones, etc. Therefore, during recording, the terminal device may use one (or two, or three) microphones in the device to obtain audio signals. This application embodiment does not limit this.
[0183] For example, Figure 7 This is a schematic diagram of another interface for starting recording provided in an embodiment of this application. For example... Figure 7 As shown in Figure 'a', when the terminal device receives an operation from the user that triggers the control 701 used to start recording, the terminal device can display the following: Figure 7 The interface shown in b. Figure 7 The interface shown in b includes a prompt box 702, which may contain text labels indicating standard recording or artificial intelligence (AI) recording, an activation control 703 for standard recording, a text label for advanced recording, and an activation control 704 for advanced recording. Standard recording can be done using a mobile phone or headphones; advanced recording can be done using both a mobile phone and headphones. Figure 7 a (or Figure 7 Other content displayed in the interface shown in b) is similar to Figure 4 The interface shown in 'a' is similar and will not be described again here.
[0184] like Figure 7 In the interface shown in b, when the terminal device receives an operation from the user that triggers the start control 704 corresponding to the advanced recording, the terminal device can use the microphone in the TWS earphone and the microphone in the terminal device to obtain audio signals and display them as shown in the image. Figure 7 The interface shown in 'c'. Figure 7 The interface shown in 'c' may include: an identifier 705 indicating the recording mode, for example, identifier 705 could be "Advanced Recording". Figure 7 The other content displayed in the interface shown by 'c' is similar to... Figure 4 The interface shown in b is similar, so it will not be described again here.
[0185] Among the possible implementations, in such Figure 7 In the interface shown in b, when the terminal device receives the user's operation to trigger the start control 703 corresponding to the standard recording, and the terminal device detects that the TWS earphones are in the wearing state, the terminal device can use the MIC in the TWS earphones to obtain audio signals.
[0186] Among the possible implementations, in such Figure 7 In the interface shown in b, when the terminal device receives the user's operation to trigger the start control 703 corresponding to the standard recording, and the terminal device detects that the TWS earphones do not meet the wearing status, the terminal device can use the MIC in the terminal device to obtain the audio signal.
[0187] S602. The terminal device performs a Fourier transform on the audio signal to obtain the Fourier transformed audio signal.
[0188] S603, The terminal device performs blind source separation on the Fourier transform audio signal to obtain three separated signals.
[0189] The function and method of blind source separation can be found in the steps shown in S303, and will not be repeated here.
[0190] S604. The terminal device obtains the three reconstructed signals corresponding to the three separated signals.
[0191] It is understandable that the process by which the terminal device obtains the reconstructed signals corresponding to the three separated signals can be referred to the steps shown in S304, and will not be repeated here.
[0192] In this embodiment, the terminal device can reconstruct the three separated signals respectively to obtain three reconstructed signals. Further, the terminal device can calculate the energy of each of the three reconstructed signals and obtain the two signals with the highest energy as the reconstructed signals for correlation calculation in step S606. The two reconstructed signals with the highest energy may include the first and second reconstructed signals; the reconstructed signal with the lower energy may be the third reconstructed signal.
[0193] S605. The terminal device calculates the forward beam corresponding to the audio signal after Fourier transform.
[0194] For example, the terminal device can use a fixed beam method to acquire the forward beam. It can be understood that when the number of microphones during audio recording is 3, the terminal device can acquire three input audio signals, and the Fourier transform audio signal can include: the first audio signal acquired by the first microphone, the second audio signal acquired by the second microphone, and the third audio signal acquired by the third microphone.
[0195] Understandably, the terminal device can suppress non-forward signals in the first, second, and third audio signals by calculating the forward beam of the audio signal, while keeping the forward signal unchanged.
[0196] For example, Figure 8 This is a schematic diagram illustrating a process for generating a forward beam, provided as an embodiment of this application. Figure 8 In a corresponding embodiment, the forward direction can be... Figure 8 The second direction described in the text.
[0197] like Figure 8 As shown, the method for generating a forward beam may include the following steps:
[0198] S801, The terminal device obtains the filter coefficients corresponding to the second direction.
[0199] For a detailed description of the second direction in this embodiment, please refer to [link to relevant documentation]. Figure 9 The corresponding description is as follows: The filter coefficients for the second direction are pre-configured in the terminal device before it leaves the factory. Alternatively, the filter coefficients for the second direction can also be generated by the terminal device; this embodiment does not limit this.
[0200] For example, Figure 9 This is a schematic diagram illustrating a direction for an embodiment of this application. For example... Figure 9 As shown, the three microphones can include one microphone from the terminal device and two microphones from the TWS earphones. The forward orientation of the terminal device and the TWS earphones can be in the range of 0°-180°.
[0201] like Figure 9 As shown, the first direction can be a 135° direction (or the first direction can also be any direction within the range of 10° clockwise from the front to 70° clockwise from the front of the electronic device), the second direction can be a 90° direction (or the second direction can also be any direction within the range of 10° counterclockwise from the front to 10° clockwise from the front), and the third direction can be a 45° direction (or the third direction can also be any direction within the range of 10° counterclockwise from the front to 70° counterclockwise from the front).
[0202] It is understood that the angles mentioned above are merely examples and can be adjusted to other angles as needed; this application does not limit them.
[0203] Furthermore, the filter coefficients corresponding to the second direction include: filter coefficients corresponding to the first microphone in the second direction, filter coefficients corresponding to the second microphone in the second direction, and filter coefficients corresponding to the third microphone in the second direction. Specifically, the filter coefficients corresponding to the first microphone in the second direction can be used to retain the audio signal acquired directly in front of the terminal device in the first audio signal, while suppressing the audio signals acquired to the left and right. The filter coefficients corresponding to the second microphone in the second direction can be used to retain the audio signal acquired directly in front of the terminal device in the second audio signal, while suppressing the audio signals acquired to the left and right. The filter coefficients corresponding to the first microphone in the third direction can be used to retain the audio signal acquired directly in front of the terminal device in the third audio signal, while suppressing the audio signals acquired to the left and right.
[0204] The formula for the terminal device to generate the filter coefficients corresponding to the second direction is as follows (7):
[0205]
[0206] Here, w2(ω) represents the filter coefficients, which consist of three elements, where the i-th element can be represented as w 2i (ω), w 2i(ω) represents the filter coefficients corresponding to the i-th MIC in the second direction, H1(ω) represents the first test audio signal, H2(ω) represents the second test audio signal, and H3(ω) represents the third test audio signal. G(H1(ω),H2(ω),H3(ω)) represents the processing of the first, second, and third test audio signals through the device-related transfer function, which can be used to describe the correlation between the first, second, and third test audio signals. H2 represents the forward beam corresponding to the second direction, w2 represents the filter coefficients that can be obtained in the second direction, and argmin represents the filter coefficients corresponding to the second direction obtained by using the least squares frequency-invariant fixed beamforming method.
[0207] The first test audio signal is a collection of input audio signals collected by the first microphone of the terminal device at different distances in multiple directions. The second test audio signal is a collection of input audio signals collected by the second microphone of the electronic device at different distances in multiple directions. The third test audio signal is a collection of input audio signals collected by the third microphone of the electronic device at different distances in multiple directions.
[0208] The forward beam is used by the terminal device to generate a second filter corresponding to the second direction, which describes the degree of filtering by the terminal device in multiple directions.
[0209] In some embodiments, when there are 36 directions, the forward beam has 36 gain coefficients. The i-th gain coefficient represents the filtering degree in the i-th direction, and each direction corresponds to a gain coefficient. Specifically, the gain coefficient corresponding to the second direction is 1. Then, for each direction that differs from the second direction by 10°, the gain coefficient is successively reduced by 1 / 36. Therefore, the closer the direction is to the second direction, the closer the element is to 1, and the farther the direction is from the second direction, the closer the element is to 0.
[0210] S802, The terminal device uses the filter coefficients corresponding to the second direction, combined with the first audio signal, the second audio signal and the third audio signal, to generate the forward beam corresponding to the second direction.
[0211] The forward beam corresponding to the second direction is the audio signal synthesized by the terminal device from the first, second, and third audio signals. During the synthesis process, the terminal device can retain the audio signals collected directly in front of the terminal device from the first, second, and third audio signals, while suppressing the audio signals collected to the left and right.
[0212] Specifically, the terminal device uses the filter coefficients corresponding to the second direction, combined with the first input audio signal, the second input audio signal and the third audio input signal, to generate the forward beam corresponding to the second direction. The formulas involved are as follows: formula (8)-formula (10).
[0213]
[0214] Where y2 represents the forward beam corresponding to the second direction, which includes N elements. Each element represents a frequency point. The number of frequency points corresponding to this forward beam is the same as the number of frequency points corresponding to the first audio signal, the second audio signal, and the third audio signal.
[0215] In the formula w 2i (ω) represents the filter coefficients corresponding to the i-th MIC in the second direction, w 2i The j-th element in (ω) represents the degree of suppression of the audio signal corresponding to the j-th frequency point in the audio signal. i (ω) represents the audio signal corresponding to the i-th MIC, x i The j-th element in (ω) represents the complex field of the j-th frequency point, which represents the amplitude and phase information of the sound signal corresponding to that frequency point.
[0216] For example, the j-th element in the filter coefficients corresponding to the i-th MIC in the second direction is denoted as c. ji Let b be the j-th element in the audio signal corresponding to the i-th microphone. ij Then the above formula (8) can be expressed as the following formula (9):
[0217]
[0218] The forward beam corresponding to the second direction can be specifically expressed as the following formula (10):
[0219]
[0220] According to the above formula (10), it can be seen that the terminal device synthesizes the first audio signal, the second audio signal and the third audio signal. The audio signal collected relative to the front of the terminal device is retained, while the audio signals collected to the left and to the right are suppressed.
[0221] It should be understood that when the j-th element of the filter coefficients corresponding to the M microphones in the second direction is equal to or close to 1, the terminal device does not suppress the audio signal corresponding to the frequency point multiplied by the j-th element, i.e., it retains it, and the direction of the audio signal corresponding to the j-th frequency point is considered to be close to the second direction. In other cases, the audio signal corresponding to the frequency point multiplied by the j-th element is suppressed. For example, when the j-th element is equal to or close to 0, the greater the degree of suppression by the terminal device, the further the direction of the audio signal corresponding to the j-th frequency point is considered to be away from the second direction.
[0222] To more clearly illustrate the suppression of different audio signals by the forward beam, the embodiments of this application combine... Figure 10 The corresponding embodiments are explained and illustrated. For example, Figure 10 A beamforming pattern is provided for an embodiment of this application.
[0223] like Figure 10 As shown, the sound signal is represented by a solid line. The shooting scene includes user 1001, user 1002, and car 1003. User 1001 is located at a 90° angle to the terminal device and TWS earphones, user 1002 is located at a 60° angle, and car 1003 is located at a 150° angle. Since user 1001's voice is the target human voice, it does not need to be suppressed. However, the voices of user 1002 and car 1003 are non-target human voices, therefore their voices are suppressed, and the degree of suppression may differ between them.
[0224] like Figure 10 As shown, when performing audio processing on a terminal device, it is possible to utilize, for example... Figure 10The beamforming pattern of the mono channel shown generates a forward beam. The line of symmetry of this beamforming pattern is in the 90° direction. The terminal device can use this mono beamforming pattern to generate mono audio. From the beamforming pattern, it can be seen that the gain coefficient corresponding to the direction of user 1001 is 1 (or close to 1), therefore the terminal device will not suppress the sound of user 1001. However, the gain coefficients corresponding to the direction of user 1002 are all 0.4 (or close to 0.4), therefore the terminal device can suppress the sound of user 1002. The gain coefficients corresponding to the direction of car 1003 are all 0 (or close to 0), therefore the terminal device can suppress the sound of car 1003. The audio signal collected by the terminal device includes the sounds of user 1001, user 1002 and car 1003. However, in the audio played based on the forward beam, the sounds of user 1002 and car 1003 are suppressed. Aurally, the sounds of user 1002 and car 1003 are either inaudible or sound quieter.
[0225] Understandable, Figure 10 The corresponding beamforming pattern is only an example and is not limited in this embodiment.
[0226] S606, the terminal equipment calculates the correlation between the reconstructed signal and the forward beam.
[0227] In this embodiment, the reconstructed signal can be the first reconstructed signal and the second reconstructed signal shown in step S604. The correlation between the first reconstructed signal and the second reconstructed signal and the forward beam can be used to characterize the similarity between any reconstructed signal and the forward beam. For example, when the correlation value is 0, it can be understood that the reconstructed signal is completely unrelated to the forward beam, and when the correlation value is 1, it can be understood that the reconstructed signal is completely correlated with the forward beam.
[0228] Understandably, terminal devices can determine the target human voice in the reconstructed signal by calculating the correlation. For example, when calculating the correlation with the forward beam, the reconstructed signal with the higher correlation value can be the signal corresponding to the target human voice.
[0229] Specifically, the formula for calculating correlation can be:
[0230]
[0231] Where γ represents correlation, a represents forward beam, b represents any reconstructed signal, t represents the number of frames in the Fourier transform, which is equivalent to time, and f represents frequency.
[0232]
[0233] Where * denotes conjugate computation, and E denotes mathematical expectation.
[0234] S607. The terminal device determines whether all correlation values are greater than the correlation threshold.
[0235] In this embodiment of the application, the correlation value may include: the correlation value α1 calculated between the first reconstructed signal and the forward beam, and the correlation value α2 calculated between the second reconstructed signal and the forward beam.
[0236] In one implementation, when the terminal device determines that both α1 and α2 are greater than (or greater than or equal to) the correlation threshold, and α1 is greater than α2, the terminal device can determine that the first reconstructed signal corresponding to α1 contains the target human voice, and the second reconstructed signal corresponding to α2 contains the non-target human voice; further, the terminal device can perform the step shown in S608 to perform energy adjustment based on the weight corresponding to the target human voice and the weight corresponding to the non-target human voice.
[0237] In another implementation, when the terminal device determines that α1 is greater than (or greater than or equal to) the correlation threshold and α2 is less than or equal to (or less than) the correlation threshold, the terminal device can determine that the first reconstructed signal corresponding to α1 contains the target human voice and the second reconstructed signal corresponding to α2 contains the non-target human voice; further, the terminal device can perform the steps shown in S608 to adjust the energy based on the weights corresponding to the target human voice and the non-target human voice.
[0238] In another implementation, when the terminal device determines that both α1 and α2 are less than or equal to (or less than) the correlation threshold, the terminal device can determine that there is currently no target human voice; further, the terminal device can perform the steps shown in S609 to perform sound source fusion on the first reconstructed signal and the second reconstructed signal.
[0239] It is understood that the value of the correlation threshold can be 0.8 or other values, and this application embodiment does not specifically limit it.
[0240] S608, the terminal equipment adjusts the energy of the three reconstructed signals based on preset weights.
[0241] For example, the three reconstructed signals include a first reconstructed signal, a second reconstructed signal, and a third reconstructed signal. The first reconstructed signal may correspond to the target human voice described in step S607, and the second and third reconstructed signals may correspond to the non-target human voice described in step S607. Furthermore, the terminal device may adjust the energy of the three reconstructed signals based on the preset weights corresponding to the target human voice and the preset weights corresponding to the non-target human voice.
[0242] It is understandable that, since the energy of the third reconstructed signal is relatively low, it can be assumed to be a non-target human voice. The preset weights for the target human voice, the preset weights for the non-target human voice, and the energy adjustment method can be found in the steps shown in S309, and will not be repeated here.
[0243] S609, the terminal equipment performs sound source mixing.
[0244] In the scenario where the terminal device does not detect the reconstruction signal of the target human voice in step S607 and performs sound source fusion, the terminal device can directly perform sound source fusion on the first reconstruction signal, the second reconstruction signal, and the third reconstruction signal. For the specific fusion method, please refer to the step S310.
[0245] S610: The terminal device performs an inverse Fourier transform on the mixed audio signal to obtain the binaural pickup result.
[0246] It is understood that the steps shown in S609-S610 can be referred to the steps shown in S310-S311, and will not be repeated here.
[0247] Based on this, terminal devices can use terminal devices and TWS earphones to acquire target human voices and non-target human voices, and by adjusting and blending the energy of the target human voices and non-target human voices, a recording with a more natural sound can be obtained.
[0248] Among the possible implementations, in Figure 3 (or Figure 6 Based on the corresponding embodiments, the terminal device can execute the steps shown in S301-S311 (or S601-S610) in the device; or, it can execute the binaural pickup method in the TWS earphone. For example, after the TWS earphone acquires the audio signal, it can directly execute the steps shown in S301-S311 (or S601-S610) to obtain the binaural pickup result and send the obtained binaural pickup result to the terminal device; or, the terminal device can execute the binaural pickup method in the server. For example, after the terminal device acquires the audio signal in S301 (or S601), it can send the audio signal to the server, so that the server can execute the steps shown in S302-S311 (or S602-S610) to obtain the binaural pickup result, and the server can send the binaural pickup result to the terminal device.
[0249] It is understood that the processing device for the binaural pickup method is not specifically limited in the embodiments of this application.
[0250] Among the possible implementations, in Figure 3 or Figure 6Based on the corresponding embodiments, the terminal device can also support the acquisition of audio signals using different recording modes in scenarios such as recording (or video recording, live streaming).
[0251] For example, Figure 11 This is a schematic diagram of an interface for enabling AI recording, provided as an embodiment of this application. The example illustrates setting the recording mode in a recording application, but this example does not constitute a limitation on the embodiments of this application.
[0252] When the terminal device receives a user's request to grant permission to use the recorder, the terminal device can display something like this: Figure 11 The interface shown in 'a' displays functional controls for the storage function, such as controls for reading location information from your media collection, controls for reading content from the memory card, and controls for modifying or deleting content from the secure digital memory card (SD). This interface can also display functional controls for the MIC function, such as controls for setting recording audio and controls 1101 for setting the recording mode.
[0253] When the terminal device receives the user's Figure 11 In the interface shown in Figure 'a', when operating the control 1101 used to set the recording mode, the terminal device can display as follows: Figure 11 The interface shown in b. Figure 11 The interface shown in b above displays the function controls corresponding to the recording mode, such as the AI recording function control 1102. AI recording can be understood as acquiring audio signals using the microphone in the headphones when the system detects that the user is wearing headphones.
[0254] It is understandable that when the AI recording function control 1102 is enabled, and the terminal device receives a recording from the user, such as... Figure 7 When the interface shown in b is operated on the activation control 703 corresponding to the standard recording, the terminal device can use the microphone in the headphones to obtain audio data; or, when the AI recording function control 1102 is in the off state, when the terminal device receives the user's input as shown in the image, the terminal device can use the microphone in the headphones to obtain audio data. Figure 7 When the interface shown in b operates on the opening control 703 corresponding to the standard recording, the terminal device can use the MIC in this device to obtain audio data.
[0255] Based on this, terminal devices can flexibly set the recording mode according to their own needs, enhancing the user experience of using the recording function.
[0256] Among the possible implementations, in Figure 11Based on the corresponding embodiments, when the terminal device uses AI recording (or advanced recording) for recording, the terminal device can also display the identifier corresponding to the AI recording (or advanced recording). For example, Figure 12 This is a schematic diagram of an interface for displaying a recording identifier, provided as an embodiment of this application.
[0257] When the terminal device receives a message from the user indicating that AI recording has been enabled, and then receives a message from the user indicating that recording has ended, the terminal device can display something like this: Figure 12 The interface shown. Compared to Figure 7 The interface shown in 'a' is... Figure 12 The interface shown can display, using, as Figure 4 The recording 5 obtained in the AI recording mode in the corresponding embodiment, and the AI recording identifier 1201 can be displayed around the recording 5.
[0258] Alternatively, when the terminal device receives a message from the user indicating that advanced recording has been enabled and that recording has ended, the terminal device can display something like this: Figure 12 The interface shown. Compared to Figure 7 The interface shown in 'a' is... Figure 12 The interface shown can display, using, as Figure 6 The recording 6 obtained in the advanced recording mode in the corresponding embodiment, and the advanced recording identifier 1202 can be displayed around the recording 6.
[0259] Based on this, terminal devices can intuitively see which mode is being used for recording based on the label, thereby enhancing the user experience of using the recording function.
[0260] It is understood that the interface provided in the embodiments of this application is only an example and does not constitute a limitation on the embodiments of this application.
[0261] The above combination Figures 3-12 The methods provided in the embodiments of this application have been described. The apparatus for executing the above methods, provided in the embodiments of this application, is described below. Figure 13 As shown, Figure 13 This is a schematic diagram of a binaural pickup device provided in an embodiment of this application. The binaural pickup device can be a terminal device in the embodiment of this application, or a chip or chip system within the terminal device.
[0262] like Figure 13 As shown, the binaural pickup device 130 can be used in communication equipment, circuits, hardware components, or chips. The binaural pickup device includes a display unit 1301 and a processing unit 1302. The display unit 1301 is used to support the display steps performed by the binaural pickup device 130; the processing unit 1302 is used to support the information processing steps performed by the binaural pickup device 130.
[0263] This application provides a binaural pickup device 130, a processing unit 1302, for acquiring audio signals using a microphone; the processing unit 1302 is further configured to perform blind source separation on the audio signals to obtain N audio signals; N is an integer greater than or equal to 2; the processing unit 1302 is further configured to determine a target signal among the N audio signals; the processing unit 1302 is further configured to adjust the target signal using a first weight and adjust the non-target signal using a second weight to obtain an adjusted target signal and an adjusted non-target signal; the non-target signal is any signal other than the target signal among the N audio signals; the processing unit 1302 is further configured to fuse the adjusted target signal and the adjusted non-target signal into a binaural pickup result.
[0264] In one possible implementation, the processing unit 1302 is specifically used by the terminal device to select one of the N audio signals that satisfies a first preset direction and contains the target human voice as the target signal.
[0265] In one possible implementation, the target human voice is a sound that satisfies a preset frequency and / or a preset harmonic.
[0266] In one possible implementation, the microphone includes two microphones in an earphone connected to a terminal device, wherein the direction of the signal is estimated by the terminal device from the direction of arrival (DOA) of the reconstructed signal corresponding to one of the N audio signals; wherein the reconstructed signal is obtained by reconstructing one of the N audio signals, and the reconstructing process is used to map one of the N audio signals to the two microphones in the earphone.
[0267] In one possible implementation, the processing unit 1302 is further configured to calculate the difference between the energy of the target signal and the energy of the non-target signal using the terminal device; when the terminal device determines that the difference is greater than a first threshold, the processing unit 1302 is further configured to adjust the target signal using a first weight and adjust the non-target signal using a second weight to obtain the adjusted target signal and the adjusted non-target signal.
[0268] In one possible implementation, the microphone includes: two microphones in an earphone connected to the terminal device; a display unit 1301 for displaying a first interface, the first interface including: a first option for recording using the two microphones in the earphone; when the terminal device receives an operation to select the first option, a processing unit 1302 for acquiring audio signals using the two microphones in the earphone.
[0269] In one possible implementation, when the terminal device receives an operation to set the recording mode, the display unit 1301 is further configured to display a second interface; the second interface includes: a first control for setting the acquisition of audio signals using two microphones in the headphones during recording; when the terminal device receives an operation to select the first option while the first control is in the enabled state, the processing unit 1302 is further configured to acquire audio signals using the two microphones in the headphones.
[0270] In one possible implementation, when the terminal device receives an operation to select the first option while the first control is in the closed state, the processing unit 1302 is also used to acquire an audio signal using the microphone in the terminal device.
[0271] In one possible implementation, the microphone includes: at least one microphone in the terminal device, and two microphones in the headset connected to the terminal device. The processing unit 1302 is further configured to calculate the forward beam corresponding to the audio signal; the forward beam is used to suppress audio signals located not directly in front of the microphone and to retain audio signals located directly in front of the microphone; the processing unit 1302 is further configured to calculate the correlation values between the forward beam and N audio signals respectively; the processing unit 1302 is further configured to select one signal with a correlation value greater than a second threshold from the N audio signals as the target signal.
[0272] In one possible implementation, the audio signal includes a first audio signal, a second audio signal, and a third audio signal. The processing unit 1302 is specifically used to obtain the filter coefficients corresponding to a second direction; the second direction is the direction directly in front of the microphone; the processing unit 1302 is specifically used to use the filter coefficients corresponding to the second direction, combined with the first audio signal, the second audio signal, and the third audio signal, to obtain the forward beam corresponding to the audio signal.
[0273] In one possible implementation, the display unit 1301 is further configured to display a first interface, the first interface including: a second option for recording using two microphones in the headset and at least one microphone in the terminal device; when the terminal device receives an operation to select the second option, the processing unit 1302 is further configured to acquire audio signals using the two microphones in the headset and at least one microphone in the terminal device.
[0274] In one possible implementation, when the terminal device receives an operation to end recording, the processing unit 1302 is further configured to encode the binaural pickup result into a first voice and store the first voice; when the terminal device receives an operation to start the recording application, the display unit 1301 is further configured to display a third interface; wherein the third interface includes the first voice and a first identifier corresponding to the first voice; the first identifier is used to indicate that the first voice was recorded based on two microphones in the earphone, or based on two microphones in the earphone and at least one microphone in the terminal device.
[0275] In one possible implementation, the processing unit 1302 is further configured to perform a Fourier transform on the audio signal to obtain a Fourier transformed audio signal; the processing unit 1302 is further configured to perform blind source separation on the Fourier transformed audio signal to obtain N audio signals.
[0276] In one possible implementation, the processing unit 1302 is specifically used to fuse the adjusted target signal and the adjusted non-target signal into a fourth audio signal; the processing unit 1302 is also specifically used to perform an inverse Fourier transform on the fourth audio signal to obtain a binaural pickup result.
[0277] In a possible implementation, the binaural pickup device 130 may also include a communication unit 1303. Specifically, the communication unit supports the binaural pickup device 130 in performing data transmission and data reception steps. The communication unit 1303 may be an input or output interface, pins, or circuitry, etc.
[0278] In a possible embodiment, the binaural pickup device may further include a storage unit 1304. The processing unit 1302 and the storage unit 1304 are connected via a line. The storage unit 1304 may include one or more memories, which may be devices in one or more devices or circuits used for storing programs or data. The storage unit 1304 may exist independently and be connected to the processing unit 1302 of the binaural pickup device via a communication line. Alternatively, the storage unit 1304 may be integrated with the processing unit 1302.
[0279] Storage unit 1304 may store computer-executable instructions for the methods in the terminal device, so that processing unit 1302 executes the methods in the above embodiments. Storage unit 1304 may be a register, cache, or RAM, etc., and storage unit 1304 may be integrated with processing unit 1302. Storage unit 1304 may be a read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, and storage unit 1304 may be independent of processing unit 1302.
[0280] Figure 14 This is a schematic diagram of the hardware structure of another terminal device provided in an embodiment of this application, such as... Figure 14 As shown, the terminal device includes a processor 1401, a communication line 1404, and at least one communication interface. Figure 14 (The example described uses communication interface 1403 as an example).
[0281] The processor 1401 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present application.
[0282] Communication line 1404 may include circuitry for transmitting information between the aforementioned components.
[0283] Communication interface 1403 uses any transceiver-like device for communicating with other devices or communication networks, such as Ethernet, wireless local area networks (WLAN), etc.
[0284] The terminal device may also include a memory 1402.
[0285] The memory 1402 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory may exist independently and be connected to the processor via communication line 1404. The memory may also be integrated with the processor.
[0286] The memory 1402 stores computer execution instructions for implementing the scheme of this application, and the processor 1401 controls the execution. The processor 1401 executes the computer execution instructions stored in the memory 1402, thereby implementing the binaural pickup method provided in the embodiments of this application.
[0287] It is possible that the computer execution instructions in the embodiments of this application may also be referred to as application code, and the embodiments of this application do not specifically limit this.
[0288] In a specific implementation, as one embodiment, the processor 1401 may include one or more CPUs, for example... Figure 14 CPU0 and CPU1 in the CPU.
[0289] In a specific implementation, as one example, the terminal device may include multiple processors, for example... Figure 14 Processors 1401 and 1405 are mentioned. Each of these processors can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. A processor here can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).
[0290] For example, Figure 15 This is a schematic diagram of a chip structure provided in an embodiment of this application. The chip 150 includes one or more processors 1520 and a communication interface 1530.
[0291] In some implementations, memory 1540 stores elements such as executable modules or data structures, or subsets thereof, or extended sets thereof.
[0292] In this embodiment, memory 1540 may include read-only memory and random access memory, and provides instructions and data to processor 1520. A portion of memory 1540 may also include non-volatile random access memory (NVRAM).
[0293] In this embodiment, the memory 1540, the communication interface 1530, and the memory 1540 are coupled together via a bus system 1510. The bus system 1510 includes a data bus, and may also include a power bus, a control bus, and a status signal bus, etc. For ease of description, in... Figure 15 The general labeled all buses as Bus System 1510.
[0294] The methods described in the embodiments of this application can be applied to, or implemented by, processor 1520. Processor 1520 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the hardware of processor 1520 or by instructions in software form. Processor 1520 may be a general-purpose processor (e.g., a microprocessor or conventional processor), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, transistor logic devices, or discrete hardware components. Processor 1520 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention.
[0295] The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in mature storage media in the art, such as random access memory, read-only memory, programmable read-only memory, or electrically erasable programmable read-only memory (EEPROM). This storage medium is located in memory 1540, and processor 1520 reads information from memory 1540 and, in conjunction with its hardware, completes the steps of the above method.
[0296] In the above embodiments, the instructions stored in the memory for execution by the processor can be implemented in the form of a computer program product. This computer program product can be pre-written into the memory, or it can be downloaded and installed into the memory as software.
[0297] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. For example, available media may include magnetic media (e.g., floppy disks, hard disks, or magnetic tapes), optical media (e.g., digital versatile discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).
[0298] This application also provides a computer-readable storage medium. The methods described in the above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. The computer-readable medium may include computer storage media and communication media, and may also include any medium capable of transferring a computer program from one place to another. The storage medium can be any target medium accessible by a computer.
[0299] As one possible design, computer-readable media may include compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM, or other optical disc storage; computer-readable media may also include disk storage or other disk storage devices. Furthermore, any connecting cable may also be appropriately referred to as computer-readable media. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. As used herein, disks and optical discs include optical discs (CD), laser discs, optical discs, digital versatile discs (DVD), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs optically reproduce data using lasers.
[0300] The above combinations should also be included within the scope of computer-readable media. The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A binaural pickup method, characterized in that, Applied to a terminal device, the method includes: The terminal device uses a microphone to acquire audio signals; The terminal device performs blind source separation on the audio signal to obtain N audio signals; N is an integer greater than or equal to 2. The terminal device determines the target signal from the N audio signals; The terminal device adjusts the target signal using a first weight and adjusts the non-target signal using a second weight to obtain the adjusted target signal and the adjusted non-target signal; the non-target signal is any signal other than the target signal among the N audio signals. The terminal device fuses the adjusted target signal and the adjusted non-target signal into a binaural pickup result; The microphone includes: at least one microphone in the terminal device, and two microphones in an earphone connected to the terminal device; the method further includes: The terminal device calculates the forward beam corresponding to the audio signal; the forward beam is used to suppress audio signals located not directly in front of the microphone, and to retain audio signals located directly in front of the microphone. The terminal device calculates the correlation values between the forward beam and the N audio signals respectively; The terminal device determines the target signal from the N audio signals by selecting one signal from the N audio signals whose correlation value is greater than a second threshold as the target signal.
2. The method according to claim 1, characterized in that, The terminal device determines the target signal from the N audio signals, including: The terminal device selects one of the N audio signals that satisfies a first preset direction and contains the target human voice as the target signal.
3. The method according to claim 2, characterized in that, The target human voice is a sound that meets a preset frequency and / or a preset harmonic.
4. The method according to claim 2, characterized in that, The microphone includes two microphones in an earphone connected to the terminal device, wherein the direction of the signal is obtained by the terminal device from the direction of arrival (DOA) estimation of the reconstructed signal corresponding to one of the N audio signals; wherein the reconstructed signal is obtained by reconstructing one of the N audio signals, and the reconstructing process is used to map one of the N audio signals to the two microphones in the earphone.
5. The method according to any one of claims 1-4, characterized in that, The method further includes: The terminal device calculates the difference between the energy of the target signal and the energy of the non-target signal; The terminal device adjusts the target signal using a first weight and adjusts the non-target signal using a second weight to obtain the adjusted target signal and the adjusted non-target signal, including: when the terminal device determines that the difference is greater than a first threshold, the terminal device adjusts the target signal using the first weight and adjusts the non-target signal using the second weight to obtain the adjusted target signal and the adjusted non-target signal.
6. The method according to any one of claims 1-4, characterized in that, The microphone includes two microphones in an earphone connected to the terminal device. The terminal device uses the microphones to acquire audio signals, including: The terminal device displays a first interface, which includes a first option for recording using the two microphones in the headset. When the terminal device receives an operation to select the first option, the terminal device uses the two microphones in the headset to acquire audio signals.
7. The method according to claim 6, characterized in that, The method further includes: When the terminal device receives an operation to set the recording mode, the terminal device displays a second interface; the second interface includes: a first control for setting the audio signal to be acquired using the two microphones in the headphones during recording; When the terminal device receives an operation to select the first option, the terminal device uses the two microphones in the headset to acquire audio signals, including: when the first control is in the enabled state, when the terminal device receives an operation to select the first option, the terminal device uses the two microphones in the headset to acquire audio signals.
8. The method according to claim 7, characterized in that, The method further includes: When the first control is closed, when the terminal device receives an operation to select the first option, the terminal device uses the microphone in the terminal device to obtain an audio signal.
9. The method according to claim 1, characterized in that, The number of microphones is 3, and the audio signal includes: a first audio signal acquired by a first microphone, a second audio signal acquired by a second microphone, and a third audio signal acquired by a third microphone. The terminal device calculates the forward beam corresponding to the audio signal, including: The terminal device acquires the filter coefficients corresponding to the second direction; the second direction is the direction directly in front of the microphone. The terminal device uses the filter coefficients corresponding to the second direction to combine the first audio signal, the second audio signal, and the third audio signal to obtain the forward beam corresponding to the audio signal.
10. The method according to claim 1 or 9, characterized in that, The method further includes: The terminal device displays a first interface, which includes a second option for recording using two microphones in the earphone and at least one microphone in the terminal device. When the terminal device receives an operation to select the second option, the terminal device uses the two microphones in the headset and at least one microphone in the terminal device to acquire audio signals.
11. The method according to claim 10, characterized in that, The method further includes: When the terminal device receives an operation to end the recording, the terminal device encodes the binaural pickup result into a first speech and stores the first speech. When the terminal device receives an operation to start the recording application, the terminal device displays a third interface; wherein, the third interface includes the first voice and a first identifier corresponding to the first voice; the first identifier is used to indicate that the first voice was recorded based on the two microphones in the earphone, or based on the two microphones in the earphone and at least one microphone in the terminal device.
12. The method according to any one of claims 1-4, 7-9, and 11, characterized in that, The method further includes: The terminal device performs a Fourier transform on the audio signal to obtain the Fourier transformed audio signal. The terminal device performs blind source separation on the audio signal to obtain N audio signals, including: the terminal device performs blind source separation on the Fourier transform audio signal to obtain the N audio signals.
13. The method according to claim 12, characterized in that, The terminal device fuses the adjusted target signal and the adjusted non-target signal into a binaural pickup result, including: The terminal device fuses the adjusted target signal and the adjusted non-target signal into a fourth audio signal; The terminal device performs an inverse Fourier transform on the fourth audio signal to obtain the binaural pickup result.
14. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it causes the terminal device to perform the method as described in any one of claims 1 to 13.
15. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it causes the computer to perform the method as described in any one of claims 1 to 13.
16. A computer program product, characterized in that, Includes a computer program that, when run, causes a computer to perform the method as described in any one of claims 1 to 13.
Citation Information
Patent Citations
Audio data processing method, electronic equipment and medium
CN111370018A
Video recording method and electronic equipment
CN113473057A
Method for estimating direction of arrival of sound source, electronic equipment and chip system
CN113889135A
Signal processing device, signal processing method, and computer program product
US20180061433A1