An audio signal processing method, apparatus, device and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-24
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本发明的目的在于,针对上述现有技术中的不足,本申请提供了一种动力电池监测方法、装置、设备及存储介质,以解决现有技术中无法精准地获取语音数据等问题
[0037]本申请提供一种音频信号处理方法、装置、设备及存储介质,该方法通过对双音频信号进行缓存,双音频信号对应双麦克风;将双音频信号进行时频变换,得到频域双音频信号;根据频域双音频信号,确定声音入射角度;根据声音入射角度,对频域双音频信号进行合并处理,得到合并后的频域音频信号;将合并后的频域音频信号进行时频逆变换,得到目标音频信号。从而,通过确定声音入射角度,以对双麦克风的音频信号进行合并处理,极大程度消除了非目标方向上的干扰声音,使目标方向上的声音得到增强,提高了获取高清语音音频信号的精准度。
Smart Images

Figure CN115881158B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and more specifically, to an audio signal processing method, apparatus, device, and storage medium. Background Technology
[0002] With the development of audio processing technology, multiple microphones are typically used to acquire high-quality audio data. However, this involves the issue of merging the audio data acquired by multiple microphones.
[0003] Existing multi-microphone merging technologies primarily target sampling signals at 8kHz and 16kHz. Current multi-microphone merging methods fall into two categories: additive and differential. Additive methods do not sufficiently attenuate sound from non-target directions, while differential methods cannot adjust the target speech direction. Insufficient spectral resolution leads to noise, especially affecting high-definition speech (48kHz). Furthermore, they cannot accurately acquire speech data. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of the prior art by providing a power battery monitoring method, apparatus, device, and storage medium to solve problems such as the inability to accurately acquire voice data in the prior art.
[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows:
[0006] In a first aspect, embodiments of this application provide an audio signal processing method, the method comprising:
[0007] The dual audio signals, corresponding to two microphones, are buffered.
[0008] The dual-tone signal is subjected to time-frequency transformation to obtain a frequency-domain dual-tone signal;
[0009] The sound incident angle is determined based on the frequency domain dual-tone signal;
[0010] Based on the sound incident angle, the frequency domain dual-tone signals are merged to obtain the merged frequency domain audio signal;
[0011] The merged frequency domain audio signal is subjected to inverse time-frequency transformation to obtain the target audio signal.
[0012] Optionally, buffering the dual-tone signals includes:
[0013] The audio signal of the current frame and the audio signal of the previous frame are buffered for each microphone, and each microphone corresponds to one audio signal.
[0014] Optionally, determining the sound incident angle based on the frequency domain dual-tone signal includes:
[0015] Based on the frequency domain dual-tone signal, the frequency domain dual-tone signal is combined using multiple preset incident angles to obtain the target energy of multiple preset sub-bands corresponding to the multiple preset incident angles.
[0016] Based on the target energy of the plurality of preset sub-bands corresponding to each preset incident angle, calculate the sum of the energies corresponding to each preset incident angle;
[0017] The sound incident angle is determined from the plurality of preset incident angles based on the sum of the energy corresponding to the plurality of preset incident angles.
[0018] Optionally, the step of merging the frequency domain dual-tone signals using multiple preset incident angles to obtain the target energy of multiple preset sub-bands corresponding to the multiple preset incident angles includes:
[0019] Based on each preset incident angle, the distance between the two microphones, and the preset sampling frequency, the energy coefficient corresponding to each preset incident angle of the two microphones is calculated.
[0020] Based on the energy coefficient of each preset incident angle corresponding to each microphone, the frequency domain audio signal corresponding to the dual microphones is weighted to obtain the target energy of multiple preset sub-bands corresponding to each preset incident angle.
[0021] Optionally, the step of merging the frequency-domain dual-tone signals according to the sound incident angle to obtain the merged frequency-domain audio signal includes:
[0022] If the sound incident angle is not zero, then the signal filtering coefficient corresponding to the dual microphones is calculated based on the sound incident angle, the distance between the dual microphones, and the preset sampling frequency.
[0023] Based on the signal filtering coefficients corresponding to the dual microphones, the frequency domain dual audio signals are differentially combined to obtain the combined frequency domain audio signal.
[0024] Optionally, the step of performing differential combining processing on the frequency domain dual audio signals according to the signal filtering coefficients corresponding to the dual microphones to obtain the combined frequency domain audio signal includes:
[0025] Based on the sound incident angle and the signal filtering coefficients corresponding to the dual microphones, the frequency domain dual audio signals are differentially merged to obtain the merged frequency domain audio signal.
[0026] Optionally, the step of merging the frequency domain dual-tone signals according to the sound incident angle to obtain the merged frequency domain audio signal further includes:
[0027] If the incident angle of the sound is zero, the frequency domain dual-tone signal is averaged to obtain the merged frequency domain audio signal.
[0028] Secondly, embodiments of this application provide an audio signal processing apparatus, the apparatus comprising:
[0029] A buffer module is used to buffer dual audio signals, which correspond to dual microphones;
[0030] The first transformation module is used to perform time-frequency transformation on the dual-tone signal to obtain a frequency-domain dual-tone signal;
[0031] The determining module is used to determine the sound incident angle based on the frequency domain dual-tone signal;
[0032] The merging module is used to merge the frequency domain dual-tone signals according to the sound incident angle to obtain the merged frequency domain audio signal.
[0033] The second transformation module is used to perform inverse time-frequency transformation on the merged frequency domain audio signal to obtain the target audio signal.
[0034] Thirdly, embodiments of this application provide an electronic device, including: a processor and a storage medium, wherein the processor and the storage medium are connected via a bus for communication, and the storage medium stores program instructions executable by the processor, wherein the processor calls the program stored in the storage medium to perform the steps of the audio signal processing method as described in any of the first aspects.
[0035] Fourthly, embodiments of this application provide a storage medium storing a computer program, which, when executed by a processor, performs the steps of the audio signal processing method as described in any of the first aspects.
[0036] Compared with the prior art, this application has the following beneficial effects:
[0037] This application provides an audio signal processing method, apparatus, device, and storage medium. The method buffers dual audio signals, each corresponding to a dual microphone; performs a time-frequency transformation on the dual audio signals to obtain a frequency-domain dual audio signal; determines the sound incident angle based on the frequency-domain dual audio signal; performs a merging process on the frequency-domain dual audio signals based on the sound incident angle to obtain a merged frequency-domain audio signal; and performs an inverse time-frequency transformation on the merged frequency-domain audio signal to obtain the target audio signal. Therefore, by determining the sound incident angle to merge the audio signals from the dual microphones, interference from non-target directions is greatly eliminated, enhancing the sound from the target direction and improving the accuracy of acquiring high-definition voice and audio signals. Attached Figure Description
[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 A flowchart illustrating an audio signal processing method provided in an embodiment of this application;
[0040] Figure 2 A schematic diagram illustrating the sound incident angle under a dual-microphone horizontal structure provided in this application embodiment;
[0041] Figure 3 A schematic diagram of a Hann window with a length of 1024 provided for an embodiment of this application;
[0042] Figure 4 A flowchart illustrating a method for determining the incident angle of sound according to an embodiment of this application;
[0043] Figure 5 A flowchart illustrating a method for calculating the target energy of multiple preset sub-bands, provided in an embodiment of this application;
[0044] Figure 6 A flowchart illustrating a method for merging frequency domain audio signals provided in an embodiment of this application;
[0045] Figure 7 This is a schematic diagram of an audio signal processing device provided in an embodiment of this application;
[0046] Figure 8 This is a schematic diagram of an electronic device provided in an embodiment of this application.
[0047] Icons: 701-Cache module, 702-First transformation module, 703-Determine module, 704-Merge module, 705-Second transformation module, 801-Processor, 802-Storage medium. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. The components of the embodiments of the present application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0049] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0050] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0051] Furthermore, the terms "first" and "second" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0052] It should be noted that, where there is no conflict, the features in the embodiments of the present invention can be combined with each other.
[0053] When receiving audio signals, both microphones are connected to the electronic device. Both microphones can receive audio signals, and after the electronic device acquires the dual audio signals through the dual microphones, it can process the dual audio signals for subsequent playback.
[0054] To accurately acquire voice data, this application provides an audio signal processing method, apparatus, device, and storage medium.
[0055] The following specific examples illustrate an audio signal processing method provided in this application. Figure 1 This is a flowchart illustrating an audio signal processing method provided in an embodiment of this application. The execution subject of this method is an electronic device, which can be a device with computing processing capabilities. Figure 1 As shown, the method includes:
[0056] S101, Buffer the dual-tone signal.
[0057] The dual audio signals correspond to dual microphones, with each microphone receiving an audio signal.
[0058] Dual audio signals are acquired using dual microphones and buffered for subsequent audio signal processing.
[0059] S102. Perform time-frequency transformation on the dual-tone signal to obtain a frequency-domain dual-tone signal.
[0060] When acquiring audio signals, they are obtained in a time-domain manner, typically within time frames. Each frame usually contains multiple audio sampling points. During audio signal processing, each sampling point is processed. Therefore, a time-frequency transformation is performed on the dual-audio signals to obtain frequency-domain dual-audio signals, facilitating subsequent audio signal processing.
[0061] Before performing time-frequency transformation, the dual-tone signal can be windowed to filter the time-domain signal. Then, the windowed dual-tone signal can be time-frequency transformed to obtain the frequency-domain dual-tone signal.
[0062] For example, the Hann window function can be used to window the audio signal first. The specific windowing calculation formula is shown in formula (1) below:
[0063] (1)
[0064] in, For the first time after adding a window The microphone The audio signal at each sampling point For the first One sampling point, Indicates the first One microphone, For the cached first The microphone The audio signal at each sampling point Take the square root of the Hanning window coefficients. The length of the Hanning window needs to be the same as the buffer length of the audio signal, for example, a length of 1024 samples. It should be noted that if the time-frequency transformation does not require a power of two, the Hanning window length can also be set to 960 samples.
[0065] For example, the time-frequency transformation can be performed using the windowed dual-tone signal of the Discrete Fourier Transform (DFT). The specific time-frequency transformation formula is shown in formula (2) below:
[0066] (2)
[0067] in, For the first The microphone The frequency domain audio signal of each sub-band Indicates the first A person with a belt The length of the DFT (e.g., 1024). It is in complex exponential form.
[0068] S103. Determine the sound incident angle based on the frequency domain dual-tone signal.
[0069] To make the processed audio signal more accurate, the sound direction factor is incorporated into the audio signal processing. Specifically, the sound direction is represented by the sound incidence angle.
[0070] By comprehensively referencing the dual-tone frequency domain signals, the sound incidence angle can be determined. This enables adaptive determination of the sound incidence angle, improving the efficiency of audio signal processing.
[0071] Specifically, the sound incident angle is the angle between the direction of sound propagation through the center point of the two microphones and the vertical direction of the two microphones. For example, the sound incident angle will be explained using a two-microphone structure. Figure 2 This application provides a schematic diagram of the sound incident angle under a dual-microphone horizontal structure, as shown in the embodiment. Figure 2 As shown, the two microphones are arranged side by side on the same horizontal plane. The center point of the two microphones is located between them, and the direction perpendicular to the line connecting the two microphones is the vertical direction of the two microphones. Figure 2 (The dotted line between the two microphones). When sound travels from the area to the right front of the center point, the angle of incidence of the sound is defined as positive, such as... Figure 2 30 60°, 90°. When sound propagates from the area to the left front of the center point, the angle of incidence of the sound is defined as negative, such as... Figure 2 -30°, -60°, and -90°.
[0072] In addition, the presence of external indicators can be used to determine whether the sound incidence angle needs to be redefined. External indicators are the detection of voice activity signals over a continuous period of time.
[0073] If an external indicator exists, the sound incident angle is determined based on the frequency domain dual-tone signal of the external indicator. If no external indicator exists, the current sound incident angle remains unchanged.
[0074] For example, if the primary audio signal captured by the microphone is continuous (e.g., a speaker continues to speak), or if the primary audio signal is interrupted (e.g., a speaker stops speaking), but the microphone continues to capture other audio signals (e.g., ambient noise), then it is considered that there is no external indication. If, during a period when the primary audio signal captured by the microphone is interrupted, a continuous primary audio signal is detected again (e.g., a speaker begins speaking again, or another speaker begins speaking), then it is considered that there is an external indication. This external indication may be from the same sound source as before, or it may be from a different sound source; therefore, it is necessary to redetermine the sound incidence angle.
[0075] S104. Based on the sound incident angle, the frequency domain dual-tone signals are merged to obtain the merged frequency domain audio signal.
[0076] After obtaining the sound incidence angle, the frequency-domain dual-tone signals are merged based on this angle to obtain a merged frequency-domain audio signal. The merged frequency-domain audio signal already incorporates the sound incidence angle, making it more accurate. This significantly reduces interference from non-target directions and enhances the sound from the target direction.
[0077] S105. Perform inverse time-frequency transformation on the merged frequency domain audio signal to obtain the target audio signal.
[0078] The specific time-frequency inverse transformation method is shown in the following formula (3):
[0079] (3)
[0080] in, This is the audio signal at the nth sampling point after inverse time-frequency transformation. This is the frequency domain audio signal after merging the k-th sub-band.
[0081] After completing the inverse time-frequency transform, windowing processing is required. The audio signal after windowing is only an audio signal at a certain frequency point, while the completed audio signal is continuous. Therefore, a segment of the target audio signal needs to be added to form a continuous audio signal. The specific windowing processing method is shown in the following formula (4):
[0082] (4)
[0083] in, For the target audio signal, This is the target audio signal from the previous segment.
[0084] In summary, in this embodiment, by buffering the dual-audio signals (DAT) corresponding to dual microphones, and performing time-frequency transformation on the DAT signals to obtain frequency-domain DAT signals, the sound incidence angle is determined based on the frequency-domain DAT signals. Based on the sound incidence angle, the frequency-domain DAT signals are then merged to obtain a merged frequency-domain audio signal. Finally, the merged frequency-domain audio signal undergoes an inverse time-frequency transformation to obtain the target audio signal. Therefore, by determining the sound incidence angle and merging the audio signals from the dual microphones, interference from non-target directions is significantly eliminated, enhancing the sound from the target direction and improving the accuracy of acquiring high-definition voice audio signals.
[0085] In the above Figure 1 Based on the corresponding embodiments, this application also provides an audio data caching method. Caching the dual audio signals in S101 includes:
[0086] The audio signal of the current frame and the audio signal of the previous frame are buffered for each microphone, with one audio signal corresponding to each microphone.
[0087] To improve audio signal processing efficiency, when buffering the audio signal captured by each microphone, both the current frame's audio signal and the previous frame's audio signal are buffered simultaneously. This allows for processing one frame at a time while buffering the next, avoiding interruptions in audio signal output and improving processing efficiency.
[0088] For example, to improve spectral resolution, an audio signal with 1024 samples can be buffered, processing 512 samples at a time (i.e., each frame is 512 samples long). Using a buffered 1024-sample audio signal as an example, a Hann window of length 1024 can also be explained. It should be noted that if the time-frequency transformation does not require a power of two, an audio signal with 960 samples can also be buffered, processing 480 samples at a time (i.e., each frame is 480 samples long, or 10ms at a sampling rate of 48kHz).
[0089] Figure 3 This is a schematic diagram of a Hann window with a length of 1024 provided for an embodiment of this application. (See diagram below.) Figure 3 As shown, each time a window is applied to process the audio signal of 1024 sampling points simultaneously, one frame of audio signal is processed and one frame of audio signal is buffered.
[0090] In summary, this embodiment buffers the current frame audio signal and the previous frame audio signal captured by each microphone, with each microphone corresponding to one audio signal. This improves processing efficiency.
[0091] In the above Figure 1Based on the corresponding embodiments, this application also provides a method for determining the sound incident angle. Figure 4 This is a flowchart illustrating a method for determining the sound incident angle provided in an embodiment of this application. Figure 4 As shown, in S103, determining the sound incident angle based on the frequency domain dual-tone signal includes:
[0092] S201. Based on the frequency domain dual-tone signal, the frequency domain dual-tone signal is merged using multiple preset incident angles to obtain the target energy of multiple preset sub-bands corresponding to the multiple preset incident angles.
[0093] To improve audio data processing efficiency, multiple preset incident angles can be set in advance, for example, preset incident angles. =30°, 60°, 90°, 0°, -30°, -60°, -90°. The specific number of preset incident angles can be set according to actual processing needs and is not limited here.
[0094] The frequency domain dual-tone signals are combined at each preset incident angle to obtain the target energy of multiple preset sub-bands corresponding to the preset incident angles. It should be noted that different incident angles yield different target energies, and the sound incident angle can be further determined based on the target energy.
[0095] S202. Calculate the sum of energy corresponding to each preset incident angle based on the target energy of multiple preset sub-bands corresponding to each preset incident angle.
[0096] The sum of the energies of all sub-bands corresponding to each preset incident angle can be calculated based on the target energies of all sub-bands. However, to improve the efficiency of audio signal processing, the sum of the energies corresponding to each preset incident angle can be calculated using the target energies of multiple preset sub-bands.
[0097] For example, the preset sub-band can be a low-frequency sub-band, which can already represent most of the audio information. For example, the low-frequency sub-band can be the first 64 sub-bands. The calculation method for the sum of the energy corresponding to each preset incident angle using the target energy of the first 64 sub-bands is shown in the following formula (5):
[0098] (5)
[0099] in, For the first The sum of the combined energies corresponding to each preset incident angle, This indicates taking the conjugate.
[0100] S203. Determine the sound incident angle from multiple preset incident angles based on the sum of energy corresponding to multiple preset incident angles.
[0101] Among multiple preset incident angles, the angle of incidence of the sound is determined by the sum of the energy and the maximum preset incident angle. The specific calculation method is shown in the following formula (6):
[0102] , (6)
[0103] in, For the angle of sound incidence, express When taking the maximum value, the corresponding value.
[0104] In summary, in this embodiment, based on the frequency domain dual-tone signals, multiple preset incident angles are used to merge the frequency domain dual-tone signals to obtain the target energy of multiple preset sub-bands corresponding to the multiple preset incident angles; based on the target energy of the multiple preset sub-bands corresponding to each preset incident angle, the energy sum corresponding to each preset incident angle is calculated; based on the energy sum corresponding to the multiple preset incident angles, the sound incident angle is determined from the multiple preset incident angles. Thus, the sound incident angle is determined by the energy sum.
[0105] In the above Figure 4 Based on the corresponding embodiments, this application also provides a method for calculating the target energy of multiple preset subbands. Figure 5 This is a flowchart illustrating a method for calculating the target energy of multiple preset sub-bands, provided as an embodiment of this application. Figure 5 As shown, in S201, based on the frequency domain dual-tone signal, multiple preset incident angles are used to merge the frequency domain dual-tone signal to obtain the target energy of multiple preset sub-bands corresponding to the multiple preset incident angles, including:
[0106] S301. Based on each preset incident angle, the distance between the two microphones, and the preset sampling frequency, calculate the energy coefficient corresponding to each preset incident angle of the two microphones.
[0107] Taking two microphones as an example, the specific calculation method for the energy coefficient is shown in the following formula (7):
[0108] (7)
[0109] in, The distance between the two microphones (in meters). The speed of sound propagation (340m / s). Sampling rate (48kHz, high-definition audio).
[0110] Furthermore, since the energy coefficients for calculating preset incident angles are all fixed values, in order to save computational load, the energy coefficients for each preset incident angle can be calculated in advance and stored in tabular form for easy use.
[0111] S302. Based on the energy coefficient of each preset incident angle corresponding to each microphone, the frequency domain audio signals corresponding to the dual microphones are weighted to obtain the target energy of multiple preset sub-bands corresponding to each preset incident angle.
[0112] Taking two microphones as an example, the specific weighted calculation method for the target energy can be shown in the following formula (8):
[0113] (8)
[0114] In summary, in this embodiment, based on each preset incident angle, the distance between the two microphones, and the preset sampling frequency, the energy coefficient for each preset incident angle corresponding to the two microphones is calculated. Based on the energy coefficient for each preset incident angle corresponding to each microphone, the frequency domain audio signals corresponding to the two microphones are weighted to obtain the target energy of multiple preset sub-bands corresponding to each preset incident angle. Therefore, by using energy coefficients to weight the frequency domain audio signals, the target energy is accurately calculated.
[0115] In the above Figure 1 Based on the corresponding embodiments, this application also provides a method for merging frequency domain audio signals. Figure 6 This is a flowchart illustrating a method for merging frequency domain audio signals provided in an embodiment of this application. Figure 6 As shown, in step S104, the frequency domain dual-tone signals are merged according to the sound incident angle to obtain the merged frequency domain audio signal, including:
[0116] S401. If the sound incident angle is not zero, calculate the signal filtering coefficients corresponding to the two microphones based on the sound incident angle, the distance between the two microphones, and the preset sampling frequency.
[0117] Taking two microphones as an example, the specific calculation method for the signal filtering coefficient is shown in the following formula (9):
[0118] , (9)
[0119] in, H These are the signal filtering coefficients. The default value is (e.g., -1). , express The absolute value, , Represents an imaginary number.
[0120] To ensure that the amplitude remains consistent before and after filtering, it is necessary to... , Normalization is performed, and the specific normalization method is shown in the following formula (10):
[0121] , (10)
[0122] in, After normalization , After normalization .
[0123] To save computational load, you can , , , Calculate in advance and save it as a table, according to The values are selected from the signal filtering coefficients in the table for easy calculation.
[0124] S402. Based on the signal filtering coefficients corresponding to the dual microphones, differential merging is performed on the frequency domain dual audio signals to obtain the merged frequency domain audio signal.
[0125] After determining the signal filtering coefficients corresponding to the two microphones, the frequency domain dual audio signals can be merged based on these coefficients to obtain the merged frequency domain audio signal. Since the audio signals acquired by the two microphones overlap to some extent, differential merging is used to obtain the merged frequency domain audio signal, thereby improving its accuracy.
[0126] In summary, in this embodiment, if the sound incident angle is not zero, the signal filtering coefficients corresponding to the two microphones are calculated based on the sound incident angle, the distance between the two microphones, and the preset sampling frequency. Based on these signal filtering coefficients, the frequency-domain dual audio signals are differentially combined to obtain the combined frequency-domain audio signal. This improves the accuracy of the frequency-domain audio signal.
[0127] In the above Figure 6 Based on the corresponding embodiments, this application also provides a method for obtaining a merged frequency domain audio signal. In step S402, differential merging processing is performed on the frequency domain dual audio signals according to the signal filtering coefficients corresponding to the dual microphones to obtain the merged frequency domain audio signal, including:
[0128] Based on the sound incident angle and the signal filtering coefficients corresponding to the two microphones, the frequency domain dual audio signals are differentially merged to obtain the merged frequency domain audio signal.
[0129] Taking two microphones as an example, the specific differential merging process is shown in the following formula (10):
[0130] (10)
[0131] when When the signal filtering coefficients of the first microphone and the second microphone are used, the frequency domain audio signals of the first microphone and the second microphone are weighted and then subtracted; when At that time, the frequency domain audio signals of the second microphone and the first microphone are weighted and subtracted respectively using the signal filtering coefficients of the first microphone and the signal filtering coefficients of the second microphone.
[0132] In summary, in this embodiment, differential merging is performed on the dual-audio signals in the frequency domain based on the sound incident angle and the signal filtering coefficients corresponding to the two microphones to obtain the merged frequency domain audio signal. Thus, the differentially merged frequency domain audio signal is accurately obtained.
[0133] In the above Figure 1 Based on the corresponding embodiments, this application also provides another method for merging frequency domain audio signals. In step S104, the frequency domain dual-tone signals are merged according to the sound incident angle to obtain the merged frequency domain audio signal, and the method further includes:
[0134] If the incident angle of the sound is zero, the frequency domain dual-tone signals are averaged to obtain the merged frequency domain audio signal.
[0135] Taking two microphones as an example, the specific averaging method is shown in the following formula (11):
[0136] (11)
[0137] In summary, in this embodiment, if the sound incident angle is zero, the frequency domain dual-tone signals are averaged to obtain the merged frequency domain audio signal. Thus, the differentially merged frequency domain audio signal is accurately obtained.
[0138] The following describes the audio signal processing apparatus, device, and storage medium provided in this application for implementation. The specific implementation process and technical effects are described above and will not be repeated below.
[0139] Figure 7 This is a schematic diagram of an audio signal processing device provided in an embodiment of this application. Figure 7As shown, the device includes:
[0140] The buffer module 701 is used to buffer the dual audio signals, which correspond to the dual microphones;
[0141] The first transformation module 702 is used to perform time-frequency transformation on the dual-tone signal to obtain a frequency domain dual-tone signal;
[0142] The determination module 703 is used to determine the sound incident angle based on the frequency domain dual-tone signal;
[0143] The merging module 704 is used to merge the frequency domain dual-tone signals according to the sound incident angle to obtain the merged frequency domain audio signal.
[0144] The second transformation module 705 is used to perform inverse time-frequency transformation on the merged frequency domain audio signal to obtain the target audio signal.
[0145] Furthermore, the caching module 701 is specifically used to cache the current frame audio signal and the previous frame audio signal collected by each microphone, with each microphone corresponding to one audio signal.
[0146] Furthermore, the determination module 703 is specifically used to merge the frequency domain dual-tone signals according to the frequency domain dual-tone signals using multiple preset incident angles to obtain the target energy of multiple preset sub-bands corresponding to the multiple preset incident angles; calculate the sum of energy corresponding to each preset incident angle based on the target energy of the multiple preset sub-bands corresponding to each preset incident angle; and determine the sound incident angle from the multiple preset incident angles based on the sum of energy corresponding to the multiple preset incident angles.
[0147] Furthermore, the determining module 703 is specifically used to calculate the energy coefficient of each preset incident angle corresponding to the dual microphones based on each preset incident angle, the distance between the dual microphones, and the preset sampling frequency; and to weight the frequency domain audio signal corresponding to the dual microphones based on the energy coefficient of each preset incident angle corresponding to each microphone to obtain the target energy of multiple preset sub-bands corresponding to each preset incident angle.
[0148] Furthermore, the merging module 704 is specifically used to calculate the signal filtering coefficients corresponding to the two microphones based on the sound incident angle, the distance between the two microphones, and the preset sampling frequency if the sound incident angle is not zero; and to perform differential merging processing on the frequency domain dual audio signals based on the signal filtering coefficients corresponding to the two microphones to obtain the merged frequency domain audio signal.
[0149] Furthermore, the merging module 704 is specifically used to perform differential merging processing on the frequency domain dual audio signals according to the sound incident angle and the signal filtering coefficients corresponding to the dual microphones, so as to obtain the merged frequency domain audio signal.
[0150] Furthermore, the merging module 704 is specifically used to average the frequency domain dual-tone signals if the sound incident angle is zero, to obtain the merged frequency domain audio signal.
[0151] Figure 8 This is a schematic diagram of an electronic device provided in an embodiment of this application. The electronic device may be a device with computing processing capabilities.
[0152] The electronic device includes a processor 801 and a storage medium 802. The processor 801 and the storage medium 802 are connected via a bus.
[0153] Storage medium 802 is used to store programs, and processor 801 calls the programs stored in storage medium 802 to execute the above method embodiments. The specific implementation and technical effects are similar, and will not be described in detail here.
[0154] Optionally, the present invention also provides a storage medium including a program, which, when executed by a processor, is used to perform the above-described method embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0155] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0156] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0157] The integrated units implemented as software functional units described above can be stored in a storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. An audio signal processing method, characterized in that, The method includes: The dual audio signals, corresponding to two microphones, are buffered. The dual-tone signal is subjected to time-frequency transformation to obtain a frequency-domain dual-tone signal; Based on each preset incident angle, the distance between the two microphones, and the preset sampling frequency, the energy coefficient corresponding to each preset incident angle of the two microphones is calculated. Based on the energy coefficient of each preset incident angle corresponding to each microphone, the frequency domain audio signal corresponding to the dual microphones is weighted to obtain the target energy of multiple preset sub-bands corresponding to each preset incident angle. Based on the target energy of the plurality of preset sub-bands corresponding to each preset incident angle, calculate the sum of the energies corresponding to each preset incident angle; Based on the sum of energy corresponding to the plurality of preset incident angles, the preset incident angle with the largest sum of energy is determined as the sound incident angle from the plurality of preset incident angles; Based on the sound incident angle, the frequency domain dual-tone signals are merged to obtain the merged frequency domain audio signal; The merged frequency domain audio signal is subjected to inverse time-frequency transformation to obtain the target audio signal.
2. The method according to claim 1, characterized in that, The buffering of the dual-tone signals includes: The audio signal of the current frame and the audio signal of the previous frame are buffered for each microphone, and each microphone corresponds to one audio signal.
3. The method according to claim 1, characterized in that, The step of merging the frequency domain dual-tone signals according to the sound incident angle to obtain the merged frequency domain audio signal includes: If the sound incident angle is not zero, then the signal filtering coefficient corresponding to the dual microphones is calculated based on the sound incident angle, the distance between the dual microphones, and the preset sampling frequency. Based on the signal filtering coefficients corresponding to the dual microphones, the frequency domain dual audio signals are differentially combined to obtain the combined frequency domain audio signal.
4. The method according to claim 3, characterized in that, The step of performing differential combining processing on the frequency domain dual audio signals according to the signal filtering coefficients corresponding to the dual microphones to obtain the combined frequency domain audio signal includes: Based on the sound incident angle and the signal filtering coefficients corresponding to the dual microphones, the frequency domain dual audio signals are differentially merged to obtain the merged frequency domain audio signal.
5. The method according to claim 3, characterized in that, The step of merging the frequency domain dual-tone signals according to the sound incident angle to obtain the merged frequency domain audio signal further includes: If the incident angle of the sound is zero, the frequency domain dual-tone signal is averaged to obtain the merged frequency domain audio signal.
6. An audio signal processing device, characterized in that, The device includes: A buffer module is used to buffer dual audio signals, which correspond to dual microphones; The first transformation module is used to perform time-frequency transformation on the dual-tone signal to obtain a frequency-domain dual-tone signal; The determining module is used to determine the sound incident angle based on the frequency domain dual-tone signal; The merging module is used to merge the frequency domain dual-tone signals according to the sound incident angle to obtain the merged frequency domain audio signal. The second transformation module is used to perform inverse time-frequency transformation on the merged frequency domain audio signal to obtain the target audio signal; The determining module is specifically used to merge the frequency domain dual-tone signal according to the frequency domain dual-tone signal using multiple preset incident angles to obtain the target energy of multiple preset sub-bands corresponding to the multiple preset incident angles. Based on the target energy of the plurality of preset sub-bands corresponding to each preset incident angle, calculate the sum of the energies corresponding to each preset incident angle; Based on the sum of energy corresponding to the plurality of preset incident angles, the preset incident angle with the largest sum of energy is determined as the sound incident angle from the plurality of preset incident angles; The determining module is specifically used to calculate the energy coefficient of each preset incident angle corresponding to the dual microphones based on each preset incident angle, the distance between the dual microphones, and the preset sampling frequency. Based on the energy coefficient of each preset incident angle corresponding to each microphone, the frequency domain audio signal corresponding to the dual microphones is weighted to obtain the target energy of multiple preset sub-bands corresponding to each preset incident angle.
7. An electronic device, characterized in that, include: The processor and the storage medium are connected via a bus for communication. The storage medium stores program instructions executable by the processor. The processor calls the program stored in the storage medium to perform the steps of the audio signal processing method as described in any one of claims 1 to 5.
8. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, performs the steps of the audio signal processing method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Signal generation method and voice recognition method and device based on artificial intelligence
CN110517702A
Microphone speech enhancement method and device, terminal and storage medium
CN113421582A
Dual-microphone directional pickup method and device with adjustable pickup angle range
CN113660578A