An optimization method and device of audio data, electronic equipment and storage medium
By cyclically reading and parallel processing audio frequency domain data, calculating the power spectrum and conjugate weighted cross power spectrum, the MDF algorithm is optimized, solving the time consumption problem caused by the large amount of MDF calculation and improving the efficiency of echo cancellation.
Patent Information
- Application Number
- CN202110190070.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-18
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2041-02-18
AI Technical Summary
In existing technologies, multi-delay block frequency domain adaptive filters (MDFs) are computationally expensive and inefficient in echo cancellation due to the large computational cost of matrix operators.
The operands in the audio frequency domain data of the local and remote ends are read in a loop. The power spectrum and conjugate weighted cross power spectrum are calculated. The echo component is eliminated through parallel processing. The matrix operator operation is optimized using the NEON instruction set.
This achieves a 38% reduction in echo cancellation time, improving efficiency and enhancing the efficiency of audio data processing.
Smart Images

Figure CN114974276B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a method, apparatus, electronic device, and storage medium for optimizing audio data. Background Technology
[0002] Live streaming is a popular form of entertainment, and users have increasingly higher expectations for the audio-visual experience. For example, they demand high-quality video and audio. Echoes during live streams can negatively impact the viewing experience.
[0003] In existing technologies, echo cancellation algorithms are used to suppress echoes in live streaming. The multi-delay block frequency-domain adaptive filter (MDF) is widely used in audio preprocessing algorithms for echo cancellation. MDF contains numerous matrix operators for frequency domain audio data computation. Due to the large computational load of these matrix operators, echo cancellation using MDF is time-consuming and inefficient. Summary of the Invention
[0004] This invention provides an audio data optimization method, apparatus, electronic device, and storage medium to solve the problem in the prior art that the large computational load of matrix operators in MDF leads to long time consumption and low efficiency when using MDF for echo cancellation.
[0005] In a first aspect, the present invention provides a method for optimizing audio data, comprising:
[0006] The program reads M first operands from the frequency domain data of the first audio at the local end in a loop, and reads M second operands from the frequency domain data of the second audio at the remote end in a loop, where M is a positive integer and M>1;
[0007] Based on the M first operands read in a loop, calculate the power spectrum corresponding to the frequency domain data of the first audio.
[0008] Based on the M first operands, the M second operands, and... For each first weight, calculate the weighted cross-power spectrum of the conjugate, where the... Each first weight corresponds to one of the M first operands and one of the M second operands;
[0009] Based on the power spectrum corresponding to the frequency domain data of the first audio and the conjugate weighted cross power spectrum, the echo components in the frequency domain data of the first audio that are related to the frequency domain data of the second audio are eliminated.
[0010] Preferably, the step of calculating the power spectrum corresponding to the frequency domain data of the first audio based on the M first operands read in a loop includes:
[0011] The M first operands include The first actual operation and The first dummy operand;
[0012] The power spectrum corresponding to the frequency domain data of the first audio signal is calculated using the following formula:
[0013]
[0014] Where ps is the power spectrum, X r For the The first target operand in the first operand, X i For the The first target dummy operand is the first target real operand among the first dummy operands.
[0015] Preferably, the step involves reading the M first operands, the M second operands, and... The first weight is used to calculate the weighted cross-power spectrum of the conjugate, including:
[0016] The M second operands comprise M / 2 second real operands and M / 2 second dummy operands;
[0017] The weighted cross-power spectrum of the conjugate is calculated using the following formula:
[0018] prod.r=p×w×(X r ×Y r +X i ×Y i )
[0019] prod.i=p×w×(X r ×Y i -X i ×Y r )
[0020] Where prod.r is the real part of the conjugate weighted cross-power spectrum, prod.i is the imaginary part of the conjugate weighted cross-power spectrum, and Y r For the The second target operand in the second operand, Y i For the Among the second dummy operands, the second target dummy operand corresponding to the second target real operand, w is the second target dummy operand. In the first weight, X r The X i The Y r and the Y i The corresponding first objective weight, where p is the total weight.
[0021] Preferably, the step of cyclically reading M first operands from the frequency domain data of the first audio at the local end and cyclically reading M second operands from the frequency domain data of the second audio at the remote end includes:
[0022] The system reads M first operands from the frequency domain data of the first audio signal from the local terminal's memory in a loop, and reads M second operands from the frequency domain data of the second audio signal from the remote terminal's memory in a loop.
[0023] Preferably, the frequency domain data of the first audio at the local end is obtained by sampling the first analog audio data at the local end to obtain the time domain data of the first audio, and then performing a Fourier transform on the time domain data of the first audio.
[0024] Preferably, the first analog audio data of the local device is sampled at a sampling frequency of 48 kHz to obtain the time domain data of the first audio.
[0025] Preferably, M is set to 8.
[0026] Secondly, the present invention also provides an audio data optimization apparatus, comprising:
[0027] The loop reading module is used to loop read M first operands from the frequency domain data of the first audio at the local end, and loop read M second operands from the frequency domain data of the second audio at the remote end, where M is a positive integer and M>1;
[0028] The first calculation module is used to calculate the power spectrum corresponding to the frequency domain data of the first audio based on the M first operands read in a loop.
[0029] The second calculation module is used to calculate based on the M first operands, the M second operands, and... For each first weight, calculate the weighted cross-power spectrum of the conjugate, where the... Each first weight corresponds to one of the M first operands and one of the M second operands;
[0030] The elimination module is used to eliminate echo components in the frequency domain data of the first audio that are related to the frequency domain data of the second audio, based on the power spectrum corresponding to the frequency domain data of the first audio and the conjugate weighted cross power spectrum.
[0031] Thirdly, the present invention also provides an electronic device, including a memory and a processor, wherein the processor is configured to execute a computer program stored in the memory to implement the steps of the audio data optimization method described in the first aspect.
[0032] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that: when the computer program is executed by a processor, it implements the steps of the audio data optimization method described in the first aspect.
[0033] As can be seen from the above technical solutions, the audio data optimization method, apparatus, electronic device, and storage medium provided by the embodiments of the present invention cyclically read M first operands from the frequency domain data of the first audio at the local end, and cyclically read M second operands from the frequency domain data of the second audio at the remote end, where M is a positive integer and M>1; calculate the power spectrum corresponding to the frequency domain data of the first audio based on the cyclically read M first operands; and calculate the power spectrum based on the cyclically read M first operands, the M second operands, and... For each first weight, calculate the weighted cross-power spectrum of the conjugate, where the... Each first weight corresponds to one of the M first operands and the M second operands. Based on the power spectrum corresponding to the frequency domain data of the first audio and the conjugate weighted cross-power spectrum, echo components in the frequency domain data of the first audio that are related to the frequency domain data of the second audio are eliminated. In this way, the M first operands in the frequency domain data of the first audio at the local end can be read cyclically, and the M second operands in the frequency domain data of the second audio at the remote end can be read cyclically. Furthermore, based on the cyclically read M first operands in the frequency domain data of the first audio at the local end and the M second operands in the frequency domain data of the second audio at the remote end, echo components in the frequency domain data of the first audio that are related to the frequency domain data of the second audio can be eliminated. This allows for parallel processing of multiple operands, resulting in shorter processing time and higher efficiency during echo cancellation. Attached Figure Description
[0034] Figure 1 A flowchart illustrating an audio data optimization method provided in this application embodiment;
[0035] Figure 2 A schematic diagram of the first 8 first operands included in the frequency domain data X of the first audio signal provided in an embodiment of this application;
[0036] Figure 3 This is a schematic diagram showing the first 8 first operands included in the frequency domain data X of the first audio at one end and the first 8 second operands included in the frequency domain data Y of the second audio at the other end, provided in an embodiment of this application.
[0037] Figure 4 A comparison chart showing the time consumption of echo cancellation using the original MDF algorithm and the method of this application, provided for embodiments of this application;
[0038] Figure 5 A structural diagram of an audio data optimization device provided in an embodiment of this application;
[0039] Figure 6 A schematic diagram illustrating an embodiment of an electronic device provided in this application;
[0040] Figure 7 This is a schematic diagram illustrating an embodiment of a computer-readable storage medium provided in this application. Detailed Implementation
[0041] To better understand the technical solutions provided in the embodiments of this specification, the technical solutions of the embodiments of this specification will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of this specification and the specific features in the embodiments are detailed descriptions of the technical solutions of the embodiments of this specification, rather than limitations on the technical solutions of this specification. In the absence of conflict, the embodiments of this specification and the technical features in the embodiments can be combined with each other.
[0042] In this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, without necessarily requiring or implying any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The term "two or more" includes two or more cases.
[0043] See Figure 1 , Figure 1 This is a flowchart of an audio data optimization method provided by the present invention. For example... Figure 1 As shown, it includes the following steps:
[0044] S101. Loop through M first operands in the frequency domain data of the first audio at this end, and loop through M second operands in the frequency domain data of the second audio at the other end, where M is a positive integer and M>1.
[0045] In S101, assume there are two streamers during a live stream: Streamer A and Streamer B. If Streamer B uses speakerphone during the stream, Streamer A's voice will be amplified through Streamer B's speaker. Streamer B's sound card will pick up both Streamer B's voice and Streamer A's voice amplified through Streamer B's speaker. When Streamer A's terminal receives the audio signal from Streamer B's terminal, it will hear both Streamer B's voice and its own voice, resulting in an echo. Similarly, if Streamer A uses speakerphone, Streamer B's terminal will also hear both Streamer A's voice and its own voice, again resulting in an echo. In this scenario, the live stream experience for both Streamer A and Streamer B is poor.
[0046] Therefore, in this embodiment, M first operands in the frequency domain data of the first audio at the local end can be read cyclically, and M second operands in the frequency domain data of the second audio at the remote end can be read cyclically. Here, M is a positive integer, and M > 1. It should be noted that the side of broadcaster B can be considered the local end, and the side of broadcaster A can be considered the remote end. In this case, the frequency domain data X of the first audio at the local end can contain audio data of what broadcaster B said and audio data of what broadcaster A said. The audio data of what broadcaster A said is the audio data of what broadcaster A said, which is collected and played back through broadcaster B's sound card. The frequency domain data Y of the second audio at the remote end is the audio data of what broadcaster A said. In this way, multiple operands can be processed in parallel, resulting in shorter processing time and higher efficiency during echo cancellation.
[0047] Preferably, the step of cyclically reading M first operands from the frequency domain data of the first audio at the local end and cyclically reading M second operands from the frequency domain data of the second audio at the remote end includes:
[0048] The system reads M first operands from the frequency domain data of the first audio signal from the local terminal's memory in a loop, and reads M second operands from the frequency domain data of the second audio signal from the remote terminal's memory in a loop.
[0049] Furthermore, M first operands from the frequency domain data of the first audio signal at the local terminal can be read cyclically from the local terminal's memory, and M second operands from the frequency domain data of the second audio signal at the remote terminal can be read cyclically from the local terminal's memory. As mentioned earlier, broadcaster B can be considered the local end, and broadcaster A can be considered the remote end. In this case, M first operands from the frequency domain data X of the first audio signal at the local terminal can be read cyclically from the local terminal's memory, and M second operands from the frequency domain data Y of the second audio signal at the remote terminal can be read cyclically from the remote terminal's memory.
[0050] Preferably, M is set to 8.
[0051] Furthermore, the value of M can be 8. That is, it can read 8 first operands from the frequency domain data X of the first audio at the local terminal in a loop from the local terminal's memory, and it can read 8 second operands from the frequency domain data Y of the second audio at the remote terminal in a loop from the local terminal's memory.
[0052] Preferably, the frequency domain data of the first audio at the local end is obtained by sampling the first analog audio data at the local end to obtain the time domain data of the first audio, and then performing a Fourier transform on the time domain data of the first audio.
[0053] It should be noted that the frequency domain data X of the first audio at this end is obtained by sampling the first analog audio data at this end, obtaining the time domain data of the first audio, and then performing a Fourier transform on the time domain data of the first audio.
[0054] Preferably, the first analog audio data of the local device is sampled at a sampling frequency of 48 kHz to obtain the time domain data of the first audio.
[0055] Furthermore, the first analog audio data at this end can be sampled at a sampling frequency of 48kHz to obtain the time-domain data of the first audio.
[0056] Audio sampling converts the sound wave waveform into a series of binary data for representation. For example, the sampling frequency can be 48 kHz, meaning 48,000 data points are sampled per second. That is, a 1-second audio clip can be represented as 48,000 data points; this is the time-domain representation of audio. The Fast Fourier Transform (FFT) can be used to convert the time-domain audio data into the frequency-domain audio data. Frequency-domain audio data is complex, containing real and imaginary parts. For example, the subscript 'r' can represent the real part, and the subscript 'i' can represent the imaginary part.
[0057] S102. Calculate the power spectrum corresponding to the frequency domain data of the first audio based on the M first operands read in a loop.
[0058] In S102, the power spectrum corresponding to the frequency domain data X of the first audio can be calculated based on the M first operands read in a loop. That is, the power spectrum corresponding to the frequency domain data X of the first audio can be calculated based on the 8 first operands read in a loop.
[0059] As mentioned earlier, 1 second of audio can be represented by 48,000 data points, which is the time domain representation of audio. After converting the time-domain audio data to the frequency domain, the frequency-domain audio data also contains 48,000 data points. The first time, the first to eighth operands of the 48,000 data points in the frequency domain data X of the first audio can be read; the second time, the ninth to sixteenth operands of the 48,000 data points can be read; the third time, the seventeenth to twenty-fourth operands of the 48,000 data points can be read. This process continues until all 48,000 data points have been read and processed.
[0060] Preferably, the step of calculating the power spectrum corresponding to the frequency domain data of the first audio based on the M first operands read in a loop includes:
[0061] The M first operands include The first actual operation and The first dummy operand;
[0062] The power spectrum corresponding to the frequency domain data of the first audio signal is calculated using the following formula:
[0063]
[0064] Where ps is the power spectrum, X r For the The first target operand in the first operand, X i For the The first target dummy operand is the first target real operand among the first dummy operands.
[0065] It should be noted that M first operands can contain M / 2 first real operands and M / 2 first dummy operands. That is, 8 first operands can contain 8 / 2 = 4 first real operands and 8 / 2 = 4 first dummy operands. For example, each time the first operand is read, the NEON instruction vld2q_f32 can be used to read the 8 first operands of the frequency domain data X of the first audio signal on the local end in an interleaved manner.
[0066] like Figure 2The diagram shown illustrates the first eight first operands contained in the frequency domain data X of a first audio signal from this terminal. Figure 2 In this example, the four first operands numbered 1, 3, 5, and 7 are the first real operands; the four first operands numbered 2, 4, 6, and 8 are the first dummy operands. Furthermore, the first real operand numbered 1 corresponds to the first dummy operand numbered 2; the first real operand numbered 3 corresponds to the first dummy operand numbered 4; the first real operand numbered 5 corresponds to the first dummy operand numbered 6; and the first real operand numbered 7 corresponds to the first dummy operand numbered 8.
[0067] The aforementioned "interleaving method" refers to reading the four first real operands numbered 1, 3, 5, and 7 into register Q1; and reading the four first virtual operands numbered 2, 4, 6, and 8 into register Q2. It should be noted that NEON is an extended architecture based on 128-bit Single Instruction Multiple Data (SIMD) implementation under the Advanced RISC Machine (ARM) architecture. It can copy multiple operands and pack them into a set of instructions within a large register. This allows multiple operands in a register to execute a single instruction simultaneously, achieving parallel processing and significantly improving execution efficiency. The operand types involved in MDF algorithms are typically 32-bit floating-point, while NEON typically has 16 128-bit registers. This allows four operands to be packed into one NEON register each time, enabling parallel processing of these four operands during matrix operator operations. After processing, the data in the register is written back to memory.
[0068] The power spectrum corresponding to the frequency domain data X of the first audio audio can be calculated using the following formula:
[0069]
[0070] Where ps is the power spectrum, X r for The first target operand in the first operand, X i for The first target dummy operand, which corresponds to the first target real operand in the first dummy operand. That is, X r X is the first target operand among the four first operands numbered 1, 3, 5, and 7; i This is the first destination virtual operand among the four first virtual operands numbered 2, 4, 6, and 8. For example, X r When X is the first target operand numbered 1, i That is, the first target dummy operand with the corresponding number 2; Xr When X is the first target operand numbered 3, i This corresponds to the first target dummy operand, numbered 4; X r When X is the first target operand numbered 5, i This corresponds to the first target dummy operand, numbered 6; X r When X is the first target operand numbered 7, i This is the first target dummy operand with the corresponding number 8.
[0071] The NEON instruction vmulq_f32 can be used to perform multiplication operations and calculate... Stored in the ps register. Next, the NEON instruction vmlaq_f32 can be used to perform a multiplication-then-add operation to calculate... That is, calculation Complete the calculation of the power spectrum corresponding to the frequency domain data X of the first audio audio. At this point, four [data points] can be obtained. The data results can then be written back to the memory of the local terminal using the NEON instruction vst1q_f32.
[0072] Furthermore, the data result in register ps can be written back to the local terminal's memory in the following order: calculated using the first real operand (number 1) and the first virtual operand (number 2). —Calculated using the first real operand numbered 3 and the first dummy operand numbered 4 —Calculated using the first real operand numbered 5 and the first dummy operand numbered 6 —Calculated using the first real operand numbered 7 and the first dummy operand numbered 8 This makes it easy to find the power spectrum corresponding to the frequency domain data X of the first audio signal in memory.
[0073] S103, based on the M first operands and the M second operands read in a loop... For each first weight, calculate the weighted cross-power spectrum of the conjugate, where the... Each first weight corresponds to one of the M first operands and one of the M second operands.
[0074] In S103, the M first operands, M second operands, and... can be read in a loop. The first weight is used to calculate the weighted cross-power spectrum of the conjugate. Wherein, Each first weight corresponds to M first operands and M second operands. That is, based on the 8 first operands, 8 second operands, and... The first weight is used to calculate the weighted cross-power spectrum of the conjugate. Wherein, Each first weight corresponds to eight first operands and eight second operands. It should be noted that the frequency domain data Y of the second audio signal at the other end can also contain 48,000 data points. The first time, the first to eighth second operands out of the 48,000 data points can be read; the second time, the ninth to sixteenth second operands out of the 48,000 data points can be read; the third time, the seventeenth to twenty-fourth second operands out of the 48,000 data points can be read. This process continues until all 48,000 data points have been read and processed.
[0075] Preferably, the step involves reading the M first operands, the M second operands, and... The first weight is used to calculate the weighted cross-power spectrum of the conjugate, including:
[0076] The M second operands comprise M / 2 second real operands and M / 2 second dummy operands; the weighted cross-power spectrum of the conjugate is calculated using the following formula:
[0077] prod.r=p×w×(X r ×Y r +X i ×Y i )
[0078] prod.i=p×w×(X r ×Y i -X i ×Y r )
[0079] Where prod.r is the real part of the conjugate weighted cross-power spectrum, prod.i is the imaginary part of the conjugate weighted cross-power spectrum, and Y r For the The second target operand in the second operand, Y i For the Among the second dummy operands, the second target dummy operand corresponding to the second target real operand, w is the second target dummy operand. In the first weight, X r The X i The Y r and the Y i The corresponding first objective weight, where p is the total weight.
[0080] It should be noted that M second operands can include M / 2 second real operands and M / 2 second virtual operands. That is, 8 second operands can include 8 / 2 = 4 second real operands and 8 / 2 = 4 second virtual operands.
[0081] like Figure 3 The diagram shows the first eight first operands contained in the frequency domain data X of the first audio signal at one end, and the first eight second operands contained in the frequency domain data Y of the second audio signal at the other end. Figure 3 In this example, the four first operands numbered 1, 3, 5, and 7 are the first real operands; the four first operands numbered 2, 4, 6, and 8 are the first dummy operands. Furthermore, the first real operand numbered 1 corresponds to the first dummy operand numbered 2; the first real operand numbered 3 corresponds to the first dummy operand numbered 4; the first real operand numbered 5 corresponds to the first dummy operand numbered 6; and the first real operand numbered 7 corresponds to the first dummy operand numbered 8.
[0082] The four second operands numbered 1*, 3*, 5*, and 7* are the second real operands; the four second operands numbered 2*, 4*, 6*, and 8* are the second virtual operands. Furthermore, the second real operand numbered 1* corresponds to the second virtual operand numbered 2*; the second real operand numbered 3* corresponds to the second virtual operand numbered 4*; the second real operand numbered 5* corresponds to the second virtual operand numbered 6*; and the second real operand numbered 7* corresponds to the second virtual operand numbered 8*.
[0083] Each time the first and second operands are read, the NEON instruction vld2q_f32 can be used to read the eight first operands of the frequency domain data X of the first audio and the eight second operands of the frequency domain data Y of the second audio in an interleaved manner. The aforementioned "interleaved manner" means that the four first real operands numbered 1, 3, 5, and 7 are read into register Q1; the four second real operands numbered 1*, 3*, 5*, and 7* are read into register Q2; the four first virtual operands numbered 2, 4, 6, and 8 are read into register Q3; and the four second virtual operands numbered 2*, 4*, 6*, and 8* are read into register Q4.
[0084] The weighted cross-power spectrum of the conjugate can be calculated using the following formula:
[0085] prod.r=p×w×(X r ×Y r +X i ×Y i )
[0086] prod.i=p×w×(X r×Y i -X i ×Y r )
[0087] Where prod.r is the real part of the conjugate weighted cross-power spectrum, prod.i is the imaginary part of the conjugate weighted cross-power spectrum, and Y r for The second target operand in the second operand, Y i for The second target dummy operand is the second target real operand among the second dummy operands. That is, Y. r Y is the second target operand among the four second operands numbered 1*, 3*, 5*, and 7*; i This refers to the second target virtual operand among the four second virtual operands numbered 2*, 4*, 6*, and 8*. For example, Y r When Y is the second target operand numbered 1*, i That is, the second target dummy operand with the corresponding number 2*; Y r When Y is the second target operand numbered 3*, i This corresponds to the second target dummy operand with the number 4*; Y r When Y is the second target operand numbered 5*, i This corresponds to the second target dummy operand with the number 6*; Y r When Y is the second target operand numbered 7*, i This corresponds to the second target dummy operand with the number 8*.
[0088] w can be Among the first weights, X r X i Y r and Y i The corresponding first objective weight, w, can be Among the first weights, X r X i Y r and Y i The corresponding first objective weight. p is the total weight.
[0089] The four first weights can be read sequentially using the NEON instruction vld1q_f32 and stored in register Q5. Next, the NEON instruction vmulq_f32 can be used to perform a multiplication operation, calculate the overall weight w×p, and store it in register Q6.
[0090] Then, the real part of the conjugate cross-power spectrum can be calculated using the NEON instructions vmulq_f32 and vmlaq_f32, corresponding to the operation X.r ×Y r +X i ×Y i The imaginary part of the conjugate cross-power spectrum is calculated using the NEON instructions vmulq_f32 and vmlsq_f32, corresponding to the operation X. r ×Y i -X i ×Y r .
[0091] Next, the NEON instruction vmluq_f32 can be used to multiply the real part and the imaginary part of the conjugate cross-power spectrum by the comprehensive weight w×p, respectively, to obtain the real part prod.r and the imaginary part prod.i of the conjugate weighted cross-power spectrum.
[0092] It should be noted that, in Figure 3 In this process, the first real operand numbered 1 in the frequency domain data X of the first audio at this end, the first dummy operand numbered 2 in the frequency domain data X of the first audio at this end, the second real operand numbered 1* in the frequency domain data Y of the second audio at the other end, and the second dummy operand numbered 2* in the frequency domain data Y of the second audio at the other end can be used as a set of data to calculate the real part prod.r and the imaginary part prod.i of the conjugate weighted cross power spectrum. At this time, X r This refers to the first real operand, numbered 1, in the frequency domain data X of the first audio signal on this end; X i That is, the first dummy operand numbered 2 in the frequency domain data X of the first audio signal at this end; Y r This refers to the second real operand, numbered 1*, in the frequency domain data Y of the second audio signal at the other end; Y i This refers to the second dummy operand, numbered 2*, in the frequency domain data Y of the second audio signal from the other end. At this point, the aforementioned X... r X i Y r Y i This corresponds to a first objective weight w.
[0093] Similarly, the first real operand numbered 3 in the frequency domain data X of the first audio at this end, the first dummy operand numbered 4 in the frequency domain data X of the first audio at this end, the second real operand numbered 3* in the frequency domain data Y of the second audio at the other end, and the second dummy operand numbered 4* in the frequency domain data Y of the second audio at the other end can be used as a set of data to calculate the real part prod.r and the imaginary part prod.i of the conjugate weighted cross power spectrum. At this time, X r This refers to the first real operand numbered 3 in the frequency domain data X of the first audio signal on this end; X iThis refers to the first dummy operand, numbered 4, in the frequency domain data X of the first audio signal at this end; Y r This refers to the second real operand, numbered 3*, in the frequency domain data Y of the second audio signal at the other end; Y i This refers to the second dummy operand, numbered 4*, in the frequency domain data Y of the second audio signal from the other end. At this point, X... r X i Y r Y i This also corresponds to a first objective weight w.
[0094] Alternatively, the first real operand numbered 5 in the frequency domain data X of the first audio signal at this end, the first dummy operand numbered 6 in the frequency domain data X of the first audio signal at this end, the second real operand numbered 5* in the frequency domain data Y of the second audio signal at the other end, and the second dummy operand numbered 6* in the frequency domain data Y of the second audio signal at the other end can be used as a set of data to calculate the real part prod.r and the imaginary part prod.i of the conjugate weighted cross-power spectrum. In this case, X r This refers to the first real operand, numbered 5, in the frequency domain data X of the first audio signal on this end; X i This refers to the first dummy operand, numbered 6, in the frequency domain data X of the first audio signal at this end; Y r This refers to the second real operand, numbered 5*, in the frequency domain data Y of the second audio signal at the other end; Y i This refers to the second dummy operand, numbered 6*, in the frequency domain data Y of the second audio signal from the other end. At this point, X... r X i Y r Y i This also corresponds to a first objective weight w.
[0095] Alternatively, the first real operand numbered 7 in the frequency domain data X of the first audio at this end, the first dummy operand numbered 8 in the frequency domain data X of the first audio at this end, the second real operand numbered 7* in the frequency domain data Y of the second audio at the other end, and the second dummy operand numbered 8* in the frequency domain data Y of the second audio at the other end can be used as a set of data to calculate the real part prod.r and the imaginary part prod.i of the conjugate weighted cross power spectrum. In this case, X r This refers to the first real operand, numbered 7, in the frequency domain data X of the first audio signal on this end; X i That is, the first dummy operand numbered 8 in the frequency domain data X of the first audio signal on this end; Y r This refers to the second real operand, numbered 7*, in the frequency domain data Y of the second audio signal at the other end; Y i This refers to the second dummy operand, numbered 8*, in the frequency domain data Y of the second audio signal from the other end. At this point, X...r X i Y r Y i This also corresponds to a first objective weight w.
[0096] At this point, four pairs of prod.r and prod.i can be obtained. Finally, the NEON instruction vst2q_f32 can be used to write the four pairs of prod.r and prod.i in the register back to the memory of the local terminal in an interleaved manner. That is, the NEON instruction vst2q_f32 can be used to write the four pairs of prod.r and prod.i in the register back to the memory of the terminal on the broadcaster B's side in an interleaved manner.
[0097] The interleaving method refers to the write-back order being: prod.r.1 calculated using the first real operand (numbered 1), the first virtual operand (numbered 2), the second real operand (numbered 1*), and the second virtual operand (numbered 2*) as a set of data; prod.i.1 calculated using the first real operand (numbered 1), the first virtual operand (numbered 2), the second real operand (numbered 1*), and the second virtual operand (numbered 2*) as a set of data; prod.r.2 calculated using the first real operand (numbered 3), the first virtual operand (numbered 4), the second real operand (numbered 3*), and the second virtual operand (numbered 4*) as a set of data; and prod.r.2 calculated using the first real operand (numbered 3), the first virtual operand (numbered 4), the second real operand (numbered 3*), and the second virtual operand (numbered 4*) as a set of data. prod.i.2 – prod.r.3 is calculated using the first real operand (numbered 5), the first dummy operand (numbered 6), the second real operand (numbered 5*), and the second dummy operand (numbered 6*) as a set of data. prod.i.3 is calculated using the first real operand (numbered 5), the first dummy operand (numbered 6), the second real operand (numbered 5*), and the second dummy operand (numbered 6*) as a set of data. prod.r.4 is calculated using the first real operand (numbered 7), the first dummy operand (numbered 8), the second real operand (numbered 7*), and the second dummy operand (numbered 8*) as a set of data. This allows for convenient searching of the conjugate weighted cross-power spectrum in memory.
[0098] S104. Based on the power spectrum corresponding to the frequency domain data of the first audio and the conjugate weighted cross power spectrum, eliminate the echo component in the frequency domain data of the first audio that is related to the frequency domain data of the second audio.
[0099] In step S104, echo components related to the frequency domain data Y of the second audio can be eliminated from the frequency domain data X of the first audio based on the power spectrum corresponding to the frequency domain data of the first audio and the conjugate weighted cross power spectrum. After eliminating the echo components related to the frequency domain data Y of the second audio in the frequency domain data X of the first audio, that is, after eliminating the voice of broadcaster A played back by broadcaster B's sound card, when broadcaster A's terminal receives the audio signal sent by broadcaster B's terminal, it will only hear broadcaster B's voice and not its own voice, thus eliminating the echo. This avoids the poor live streaming experience caused by echoes for broadcasters A and B.
[0100] It should be noted that a 25-second audio data segment can be selected, and the original MDF algorithm and the method of this application (i.e., the algorithm optimized using the NEON instruction set) can be executed 10 times respectively. This allows for timing accuracy down to the microsecond level. Figure 4 The image shows a comparison of the time consumption results for echo cancellation using the original MDF algorithm and the method of this application. Figure 4 The document lists the results of each of the 10 executions. Each execution result includes the time taken for echo cancellation using the original MDF algorithm, the time taken for echo cancellation using the method of this application (i.e., the algorithm optimized using the NEON instruction set), the difference (delta) between the time taken for echo cancellation using the original MDF algorithm and the time taken for echo cancellation using the method of this application, and the percentage (percent) of this difference to the time taken for echo cancellation using the original MDF algorithm. Additionally, in... Figure 4 The document also lists the average time for echo cancellation using the original MDF algorithm in 10 executions, the average time for echo cancellation using the method of this application, i.e., the algorithm optimized using the NEON instruction set, the difference between the two average times, and the ratio of the difference between the two average times to the average time for echo cancellation using the original MDF algorithm.
[0101] from Figure 4 It can be seen that the average time for echo cancellation using the method of this application, that is, the algorithm optimized using the NEON instruction set, is reduced by about 38% compared with the average time for echo cancellation using the original MDF algorithm. The method of this application is faster and more efficient.
[0102] It should be noted that in the existing technology, due to the large amount of computation required for matrix operators in MDF, the time consumed and the efficiency are low when using MDF for echo cancellation.
[0103] In this application, M first operands from the frequency domain data of the first audio at the local end can be read cyclically, and M second operands from the frequency domain data of the second audio at the remote end can be read cyclically. Then, based on the cyclically read M first operands from the frequency domain data of the first audio at the local end and the M second operands from the frequency domain data of the second audio at the remote end, echo components related to the frequency domain data of the second audio in the frequency domain data of the first audio can be eliminated. This allows for parallel processing of multiple operands, resulting in shorter processing time and higher efficiency during echo cancellation.
[0104] As can be seen from the above technical solutions, the audio data optimization method provided by the embodiments of the present invention involves: cyclically reading M first operands from the frequency domain data of a first audio source at the local end, and cyclically reading M second operands from the frequency domain data of a second audio source at the remote end, where M is a positive integer and M>1; calculating the power spectrum corresponding to the frequency domain data of the first audio source based on the cyclically read M first operands; and calculating the power spectrum based on the cyclically read M first operands, the M second operands, and... For each first weight, calculate the weighted cross-power spectrum of the conjugate, where the... Each first weight corresponds to one of the M first operands and the M second operands. Based on the power spectrum corresponding to the frequency domain data of the first audio and the conjugate weighted cross-power spectrum, echo components in the frequency domain data of the first audio that are related to the frequency domain data of the second audio are eliminated. In this way, the M first operands in the frequency domain data of the first audio at the local end can be read cyclically, and the M second operands in the frequency domain data of the second audio at the remote end can be read cyclically. Furthermore, based on the cyclically read M first operands in the frequency domain data of the first audio at the local end and the M second operands in the frequency domain data of the second audio at the remote end, echo components in the frequency domain data of the first audio that are related to the frequency domain data of the second audio can be eliminated. This allows for parallel processing of multiple operands, resulting in shorter processing time and higher efficiency during echo cancellation.
[0105] See Figure 5 , Figure 5 This is a structural diagram of an audio data optimization device provided by the present invention. Figure 5 As shown, the audio data optimization device 500 includes a loop reading module 501, a first calculation module 502, a second calculation module 503, and an elimination module 504, wherein:
[0106] The loop reading module 501 is used to loop read M first operands from the frequency domain data of the first audio at the local end, and loop read M second operands from the frequency domain data of the second audio at the remote end, where M is a positive integer and M>1;
[0107] The first calculation module 502 is used to calculate the power spectrum corresponding to the frequency domain data of the first audio based on the M first operands read in a loop.
[0108] The second calculation module 503 is used to calculate based on the M first operands, the M second operands, and... (the rest of the text is missing). For each first weight, calculate the weighted cross-power spectrum of the conjugate, where the... Each first weight corresponds to one of the M first operands and one of the M second operands;
[0109] The elimination module 504 is used to eliminate echo components in the frequency domain data of the first audio that are related to the frequency domain data of the second audio, based on the power spectrum corresponding to the frequency domain data of the first audio and the conjugate weighted cross power spectrum.
[0110] The audio data optimization device 500 can achieve Figure 1 The various processes implemented by the audio data optimization device in the method embodiment will not be described again here to avoid repetition. The audio data optimization device 500 can cyclically read M first operands from the frequency domain data of the first audio at its local end and cyclically read M second operands from the frequency domain data of the second audio at the remote end. Then, based on the cyclically read M first operands from the frequency domain data of the first audio at its local end and the M second operands from the frequency domain data of the second audio at the remote end, echo components related to the frequency domain data of the second audio in the frequency domain data of the first audio can be eliminated. That is, parallel processing of multiple operands can be achieved, resulting in shorter processing time and higher efficiency during echo cancellation.
[0111] Please see Figure 6 , Figure 6 A schematic diagram illustrating an embodiment of the electronic device provided in this application.
[0112] like Figure 6 As shown, this application embodiment provides an electronic device 600, including a memory 610, a processor 620, and a computer program 611 stored in the memory 610 and executable on the processor 620. When the processor 620 executes the computer program 611, it performs the following steps:
[0113] The program reads M first operands from the frequency domain data of the first audio at the local end in a loop, and reads M second operands from the frequency domain data of the second audio at the remote end in a loop, where M is a positive integer and M>1;
[0114] Based on the M first operands read in a loop, calculate the power spectrum corresponding to the frequency domain data of the first audio.
[0115] Based on the M first operands, the M second operands, and... For each first weight, calculate the weighted cross-power spectrum of the conjugate, where the... Each first weight corresponds to one of the M first operands and one of the M second operands;
[0116] Based on the power spectrum corresponding to the frequency domain data of the first audio and the conjugate weighted cross power spectrum, the echo components in the frequency domain data of the first audio that are related to the frequency domain data of the second audio are eliminated.
[0117] In practical implementation, when the processor 620 executes the computer program 611, it can achieve... Figure 1 Any of the corresponding implementation methods in the embodiments.
[0118] Since the electronic device described in this embodiment is a device used to implement an audio data optimization device in the embodiments of this application, those skilled in the art can understand the specific implementation method and various variations of the electronic device in this embodiment based on the method described in the embodiments of this application. Therefore, how the electronic device implements the method in the embodiments of this application will not be described in detail here. Any device used by those skilled in the art to implement the method in the embodiments of this application falls within the scope of protection of this application.
[0119] Please see Figure 7 , Figure 7 This is a schematic diagram illustrating an embodiment of a computer-readable storage medium provided in this application.
[0120] like Figure 7 As shown, this embodiment provides a computer-readable storage medium 700, on which a computer program 711 is stored. When the computer program 711 is executed by a processor, it performs the following steps:
[0121] The program reads M first operands from the frequency domain data of the first audio at the local end in a loop, and reads M second operands from the frequency domain data of the second audio at the remote end in a loop, where M is a positive integer and M>1;
[0122] Based on the M first operands read in a loop, calculate the power spectrum corresponding to the frequency domain data of the first audio.
[0123] Based on the M first operands, the M second operands, and... For each first weight, calculate the weighted cross-power spectrum of the conjugate, where the... Each first weight corresponds to one of the M first operands and one of the M second operands;
[0124] Based on the power spectrum corresponding to the frequency domain data of the first audio and the conjugate weighted cross power spectrum, the echo components in the frequency domain data of the first audio that are related to the frequency domain data of the second audio are eliminated.
[0125] In practical implementation, when the computer program 711 is executed by the processor, it can achieve the following: Figure 1 Any of the corresponding implementation methods in the embodiments.
[0126] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0127] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0128] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0129] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0130] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0131] This application also provides a computer program product, which includes computer software instructions that, when executed on a processing device, cause the processing device to perform actions such as... Figure 1 The process of optimizing audio data in the corresponding embodiment.
[0132] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0133] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0134] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0135] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0136] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0137] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0138] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for optimizing audio data, characterized in that, include: The program reads M first operands from the frequency domain data of the first audio at the local end in a loop, and reads M second operands from the frequency domain data of the second audio at the remote end in a loop, where M is a positive integer and M>1; Based on the M first operands read in a loop, calculate the power spectrum corresponding to the frequency domain data of the first audio. Based on the M first operands, the M second operands, and... For each first weight, calculate the weighted cross-power spectrum of the conjugate, where the... Each first weight corresponds to one of the M first operands and one of the M second operands; Based on the power spectrum corresponding to the frequency domain data of the first audio and the conjugate weighted cross power spectrum, the echo components in the frequency domain data of the first audio that are related to the frequency domain data of the second audio are eliminated. The step of calculating the power spectrum corresponding to the frequency domain data of the first audio based on the M first operands read in a loop includes: The M first operands include The first actual operation and The first dummy operand; The power spectrum corresponding to the frequency domain data of the first audio signal is calculated using the following formula: Where ps is the power spectrum, X r For the The first target operand in the first operand, X i For the The first target dummy operand is the first target real operand among the first dummy operands.
2. The method as described in claim 1, characterized in that, The M first operands and the M second operands read in a loop are... The first weight is used to calculate the weighted cross-power spectrum of the conjugate, including: The M second operands comprise M / 2 second real operands and M / 2 second dummy operands; The weighted cross-power spectrum of the conjugate is calculated using the following formula: prod.r=p×w×(X r ×Y r +X i ×Y i ) prod.i=p×w×(X r ×Y i -X i ×Y r ) Where prod.r is the real part of the conjugate weighted cross-power spectrum, prod.i is the imaginary part of the conjugate weighted cross-power spectrum, and Y r For the The second target operand in the second operand, Y i For the Among the second dummy operands, the second target dummy operand corresponding to the second target real operand, w is the second target dummy operand. In the first weight, X r The X i The Y r and the Y i The corresponding first objective weight, where p is the total weight.
3. The method as described in claim 2, characterized in that, The step of cyclically reading M first operands from the frequency domain data of the first audio at the local end and cyclically reading M second operands from the frequency domain data of the second audio at the remote end includes: The system reads M first operands from the frequency domain data of the first audio signal from the local terminal's memory in a loop, and reads M second operands from the frequency domain data of the second audio signal from the remote terminal's memory in a loop.
4. The method according to any one of claims 1 to 3, characterized in that, The frequency domain data of the first audio signal at the local end is obtained by sampling the first analog audio data at the local end to obtain the time domain data of the first audio signal, and then performing a Fourier transform on the time domain data of the first audio signal.
5. The method as described in claim 4, characterized in that, The first analog audio data of the local device is sampled at a sampling frequency of 48 kHz to obtain the time domain data of the first audio.
6. The method as described in claim 5, characterized in that, The value of M is 8.
7. An electronic device, comprising a memory and a processor, characterized in that, When the processor executes a computer program stored in the memory, it implements the steps of the method for optimizing audio data as described in any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the method for optimizing audio data as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Frequency-domain echo cancellation method for speech recognition front end and computer storage medium
CN109727604A