Microphone sound mixing method, system and device, communication equipment and readable storage medium

By using adaptive filters on the recording end for echo cancellation and limiting adjustment, combined with audio decoding and mixing processing at the mixing end, the problem of inefficiency of traditional multi-terminal audio mixing is solved, and efficient audio processing and synchronization is achieved.

CN120496556APending Publication Date: 2025-08-15CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510717735.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional multi-terminal audio mixing methods are inefficient, and there are problems of audio delay and audio and video aberration, especially in multi-terminal coordination, the complexity of echo cancellation and noise suppression increases.

Method used

Echo cancellation is performed through an adaptive filter on the recording end, echo cancellation and limit adjustment is performed using filter parameter configuration and mixing weights, and audio decoding and mixing are performed on the mixing end. The recording end and mixing end jointly undertake the mixing work.

Benefits of technology

Improves mixing efficiency, reduces audio delay, avoids the problem of out-of-synchronization of audio and video, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496556A_ABST
    Figure CN120496556A_ABST
Patent Text Reader

Abstract

The invention relates to a microphone sound mixing method, system and device, communication equipment and a readable storage medium. A recording audio signal of a current frame collected by a microphone is obtained, an original echo signal corresponding to the current frame sent by a sound mixing end is received, the sound mixing end is in communication connection with a recording end, the original echo signal and the recording audio signal are input, and a self-adaptive filter is configured by using filter parameters of the current frame; the method comprises the following steps: acquiring a pre-estimated echo signal corresponding to a current frame, performing echo cancellation on a recording audio signal by using a preset sound mixing weight and the pre-estimated echo signal, performing audio coding on the recording audio signal after echo cancellation, and sending the coded recording audio signal to a sound mixing end, and the sound mixing end carries out audio decoding on the coded recording audio signal and carries out sound mixing processing on the decoded recording audio signal. The recording end and the sound mixing end jointly undertake the sound mixing work, so that the sound mixing efficiency is improved, and the problem of sound and picture desynchrony caused by audio delay is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of wireless communications and terminal technologies, and in particular to a microphone mixing method, system, apparatus, communication equipment, computer-readable storage medium, and computer program product. Background Art

[0002] The microphone application in multi-terminal collaboration allows users to use the microphone hardware on their mobile phones to collaborate across devices with conference tablets or TVs, providing business functions such as video conferencing and karaoke for conference tablets or TVs that do not have microphones.

[0003] However, traditional multi-end audio mixing suffers from problems such as low mixing efficiency due to audio encoding and decoding, echo cancellation, noise suppression, and voice enhancement. Summary of the Invention

[0004] Based on this, it is necessary to provide a microphone mixing method, system, apparatus, communication device, computer-readable storage medium and computer program product that can improve mixing efficiency in response to the above technical problems.

[0005] In a first aspect, the present application provides a microphone mixing method, applied to a recording end, comprising:

[0006] Acquire the recorded audio signal of the current frame collected by the microphone, and receive the original echo signal corresponding to the current frame sent by the mixing end; the mixing end is connected to the recording end;

[0007] The original echo signal and the recorded audio signal are input into an adaptive filter configured using the filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame;

[0008] Use pre-set mixing weights and estimated echo signals to perform echo cancellation on recorded audio signals;

[0009] The audio signal after echo cancellation is audio-encoded and sent to the mixing end; the mixing end is used to decode the encoded audio signal and mix the decoded audio signal.

[0010] In combination with the first aspect, in one embodiment, before the original echo signal and the recorded audio signal are input into the adaptive filter configured using the filter parameters of the current frame, the method further includes:

[0011] Get the distance change and filter weight of the current frame, as well as the preset initial echo delay and initial echo strength;

[0012] Obtain the echo delay of the current frame according to the distance change of the current frame and the initial echo delay;

[0013] Based on the distance change of the current frame and the initial echo intensity, the echo intensity of the current frame is obtained;

[0014] The filter weight, echo delay, and echo strength of the current frame are used as filter parameters of the current frame, and the adaptive filter is configured using the filter parameters of the current frame.

[0015] In combination with the first aspect, in one embodiment, when the current frame is the first frame, obtaining the distance change and filter weight of the current frame includes:

[0016] Obtaining a preset initial distance change, and using the preset initial distance change as the distance change of the current frame;

[0017] The preset initial filter weight is obtained, and the preset initial filter weight is used as the filter weight of the current frame.

[0018] In combination with the first aspect, in one embodiment, when the current frame is not the first frame, obtaining the distance change and filter weight of the current frame includes:

[0019] Obtain the signal strength of the communication connection between the recording end and the mixing end of the current frame, and obtain the distance change of the current frame based on the signal strength;

[0020] Obtaining a preset initial step factor, a filter weight of a previous frame, a reference echo signal, and an error signal; the previous frame is the frame before the current frame;

[0021] Based on the distance change of the current frame and the initial step factor, the step factor of the current frame is obtained;

[0022] The filter weight of the current frame is obtained according to the step factor of the current frame, the filter weight of the previous frame, the reference echo signal and the error signal.

[0023] In conjunction with the first aspect, in one embodiment, after obtaining the estimated echo signal corresponding to the current frame, the method further includes:

[0024] Obtain the delayed recorded audio signal of the current frame collected by the microphone after the echo delay;

[0025] The difference between the delayed recorded audio signal and the estimated echo signal is used as the error signal of the current frame.

[0026] In conjunction with the first aspect, in one embodiment, inputting the original echo signal and the recorded audio signal into an adaptive filter configured using filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame includes:

[0027] The original echo signal and the recorded audio signal are input into the adaptive filter, and the original echo signal is updated using the distance change of the current frame to obtain the reference echo signal of the current frame;

[0028] Inputting the reference echo signal and the recorded audio signal into the feature extraction unit of the adaptive filter to obtain corresponding acoustic features;

[0029] The acoustic features are input into the intelligent control unit of the adaptive filter to obtain the estimated echo signal corresponding to the current frame.

[0030] In conjunction with the first aspect, in one embodiment, the mixing weight includes a voice weight and an echo weight. Before performing echo cancellation on the recorded audio signal using the preset mixing weight and the estimated echo signal, the method further includes:

[0031] The mixing weight sent by the mixer is used as the preset mixing weight;

[0032] Use pre-set mixing weights and estimated echo signals to perform echo cancellation on recorded audio signals, including:

[0033] Echo cancellation is performed on the recorded audio signal using the echo weight, voice weight, and estimated echo signal.

[0034] In a second aspect, the present application further provides a microphone mixing method, which is applied to a mixing end and includes:

[0035] The original echo signal corresponding to the current frame is sent to the recording end; the recording end is communicatively connected to the mixing end; the recording end is used to obtain the recorded audio signal of the current frame collected by the microphone, receive the original echo signal, input the original echo signal and the recorded audio signal into the adaptive filter configured using the filter parameters of the current frame, and obtain the estimated echo signal corresponding to the current frame; echo cancel the recorded audio signal using the preset mixing weight and the estimated echo signal; audio encode the recorded audio signal after echo cancellation and return it;

[0036] The encoded recorded audio signal is received, audio decoding is performed on the encoded recorded audio signal, and mixing processing is performed on the decoded recorded audio signal.

[0037] In conjunction with the second aspect, in one embodiment, mixing the decoded recorded audio signal includes:

[0038] Get background sounds and preset mix weights;

[0039] The decoded recorded audio signal is mixed using the background sound weight and the background sound included in the mixing weight.

[0040] In conjunction with the second aspect, in one embodiment, after mixing the decoded recorded audio signal, the method includes:

[0041] The mixing weight is sent to the recording end; the recording end is used to receive the mixing weight and use the mixing weight as the preset mixing weight.

[0042] In a third aspect, the present application further provides a microphone mixing system, the system comprising a recording end and a mixing end, the recording end and the mixing end being communicatively connected;

[0043] The mixing end is used to send the original echo signal corresponding to the current frame to the recording end;

[0044] The recording end is used to obtain the recorded audio signal of the current frame collected by the microphone and receive the original echo signal, input the original echo signal and the recorded audio signal into the adaptive filter configured with the filter parameters of the current frame, obtain the estimated echo signal corresponding to the current frame, and use the preset mixing weights and the estimated echo signal to perform echo cancellation on the recorded audio signal; perform audio encoding on the recorded audio signal after echo cancellation, and send the encoded recorded audio signal to the mixing end;

[0045] The mixing end is used to receive the encoded recorded audio signal, perform audio decoding on the encoded recorded audio signal, and perform mixing processing on the decoded recorded audio signal.

[0046] In a fourth aspect, the present application further provides a microphone mixing device, which is applied to a recording end and includes:

[0047] A first signal acquisition module is used to acquire the recorded audio signal of the current frame collected by the microphone and receive the original echo signal corresponding to the current frame sent by the mixing end; the mixing end is communicatively connected to the recording end;

[0048] A second signal acquisition module is used to input the original echo signal and the recorded audio signal into an adaptive filter configured using the filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame;

[0049] An echo cancellation module is used to perform echo cancellation on the recorded audio signal using pre-set mixing weights and estimated echo signals;

[0050] The first signal sending module is used to perform audio encoding on the recorded audio signal after echo cancellation and send the encoded recorded audio signal to the mixing end; the mixing end is used to perform audio decoding on the encoded recorded audio signal and perform mixing processing on the decoded recorded audio signal.

[0051] In a fifth aspect, the present application further provides a microphone mixing device, which is applied to a mixing end and includes:

[0052] The second signal sending module is configured to send the original echo signal corresponding to the current frame to the recording end; the recording end is communicatively connected to the mixing end; the recording end is configured to obtain the recorded audio signal of the current frame collected by the microphone, receive the original echo signal, input the original echo signal and the recorded audio signal into an adaptive filter configured using the filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame; use the preset mixing weights and the estimated echo signal to echo cancel the recorded audio signal; perform audio encoding on the recorded audio signal after echo cancellation, and return the audio signal;

[0053] The signal receiving module is used to receive the encoded recorded audio signal, perform audio decoding on the encoded recorded audio signal, and perform mixing processing on the decoded recorded audio signal.

[0054] In a sixth aspect, the present application further provides a communication device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0055] Acquire the recorded audio signal of the current frame collected by the microphone, and receive the original echo signal corresponding to the current frame sent by the mixing end; the mixing end is connected to the recording end;

[0056] The original echo signal and the recorded audio signal are input into an adaptive filter configured using the filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame;

[0057] Use pre-set mixing weights and estimated echo signals to perform echo cancellation on recorded audio signals;

[0058] The audio signal after echo cancellation is audio-encoded and sent to the mixing end; the mixing end is used to decode the encoded audio signal and mix the decoded audio signal.

[0059] In a seventh aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps:

[0060] Acquire the recorded audio signal of the current frame collected by the microphone, and receive the original echo signal corresponding to the current frame sent by the mixing end; the mixing end is connected to the recording end;

[0061] The original echo signal and the recorded audio signal are input into an adaptive filter configured using the filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame;

[0062] Use pre-set mixing weights and estimated echo signals to perform echo cancellation on recorded audio signals;

[0063] The audio signal after echo cancellation is audio-encoded and sent to the mixing end; the mixing end is used to decode the encoded audio signal and mix the decoded audio signal.

[0064] In an eighth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:

[0065] Acquire the recorded audio signal of the current frame collected by the microphone, and receive the original echo signal corresponding to the current frame sent by the mixing end; the mixing end is connected to the recording end;

[0066] The original echo signal and the recorded audio signal are input into an adaptive filter configured using the filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame;

[0067] Use pre-set mixing weights and estimated echo signals to perform echo cancellation on recorded audio signals;

[0068] The audio signal after echo cancellation is audio-encoded and sent to the mixing end; the mixing end is used to decode the encoded audio signal and mix the decoded audio signal.

[0069] The microphone mixing method, system, apparatus, communication device, computer-readable storage medium, and computer program product described above obtain a recorded audio signal of the current frame collected by a microphone, and receive an original echo signal corresponding to the current frame sent by a mixing end. The mixing end and the recording end are communicatively connected, and the original echo signal and the recorded audio signal are input into an adaptive filter configured using the filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame. The recorded audio signal is echo-cancelled using pre-set mixing weights and the estimated echo signal, and the echo-cancelled recorded audio signal is audio-encoded. The encoded recorded audio signal is then sent to the mixing end, which then decodes the encoded recorded audio signal and performs mixing processing on the decoded recorded audio signal. The recording end is responsible for echo cancellation, and simultaneously performs echo cancellation for the current frame while receiving the recorded audio. The recording end and the mixing end jointly undertake the mixing work, thereby speeding up the mixing efficiency and avoiding the problem of audio delay causing audio and video synchronization. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0071] Figure 1 A diagram illustrating an application environment of a microphone mixing method according to an embodiment;

[0072] Figure 2 1 is a flow chart of a microphone mixing method according to an embodiment;

[0073] Figure 3 is a flowchart of a microphone mixing method in another embodiment;

[0074] Figure 4 is a schematic structural diagram of a microphone mixing system in one embodiment;

[0075] Figure 5 A schematic diagram showing a comparison between the present application and the original solution in one embodiment;

[0076] Figure 6 is a system architecture diagram of a microphone mixing system in another embodiment;

[0077] Figure 7 A schematic flow chart of a microphone mixing method according to another embodiment;

[0078] Figure 8 is a schematic diagram of a model in an embodiment;

[0079] Figure 9 A schematic diagram of multi-terminal interaction of a microphone mixing method according to another embodiment;

[0080] Figure 10 is a structural block diagram of a microphone mixing device in one embodiment;

[0081] Figure 11 is a structural block diagram of a microphone mixing device in another embodiment;

[0082] Figure 12 FIG. 4 is a diagram showing the internal structure of a communication device in one embodiment. DETAILED DESCRIPTION

[0083] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0084] The microphone application in multi-device collaboration allows users to use the microphone hardware on their mobile phone to collaborate with conference tablets or TVs across devices, providing business functions such as video conferencing and karaoke for conference tablets or TVs without microphones. However, traditional multi-device audio mixing can cause audio delays or audio-video asynchrony due to audio codecs, echo cancellation, noise suppression, and voice enhancement.

[0085] In existing multi-terminal collaborative mixing scenarios, echo cancellation and voice enhancement are more complex than those processed by a single terminal, increasing mixing latency.

[0086] Audio needs to be recorded and encoded at one terminal and transmitted to another for decoding, mixing, and playback. The uncertainty of latency further increases the difficulty of echo cancellation, and the mixing performance of the mixing terminal becomes a service bottleneck.

[0087] The echo changes caused by people moving around make it difficult for the mixing terminal to distinguish between human voices and background sounds (echoes), further increasing the difficulty of voice enhancement.

[0088] To solve the above problems, this patent proposes using the signal strength change index of the two-end connection based on multi-end collaboration as a reference for the echo filter, optimizing the adaptive filtering process of the recording terminal. The separation, weighted balance and limiting of human voice and echo are completed in advance at the recording end (such as a mobile phone). The mixing end (such as a conference tablet or TV) only needs to complete dynamic range compression, equalization and reverberation, further reducing the performance pressure of the mixing end, thereby reducing the delay of service mixing and improving user experience.

[0089] The microphone mixing method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the recording end 102 is communicatively connected to the mixing end 104. The recording end 102 obtains the recording audio signal of the current frame collected by the microphone, and receives the original echo signal corresponding to the current frame sent by the mixing end, inputs the original echo signal and the recording audio signal into the adaptive filter configured with the filter parameters of the current frame, obtains the estimated echo signal corresponding to the current frame, uses the preset mixing weight and the estimated echo signal to perform echo cancellation on the recording audio signal, audio encodes the recording audio signal after echo cancellation, and sends the encoded recording audio signal to the mixing end, which performs audio decoding on the encoded recording audio signal and performs mixing processing on the decoded recording audio signal.

[0090] The mixing end can be a terminal, which can include terminal devices such as conference tablets or TVs; the recording end can also be a terminal, but this terminal is a terminal that includes a microphone, which can include terminal devices such as mobile phones and microphones.

[0091] In an exemplary embodiment, Figure 2 As shown, a microphone mixing method is provided, which is applied to Figure 1 Taking the recording terminal 102 in FIG. 1 as an example, the method includes the following steps S201 to S204.

[0092] Step S201: Acquire the recorded audio signal of the current frame collected by the microphone, and receive the original echo signal corresponding to the current frame sent by the mixing end; the mixing end is communicatively connected to the recording end.

[0093] Among them, the recorded audio signal can be understood as the sound stream recorded by the microphone, including human voice and echo sound, and the original echo signal can be understood as the echo stream signal that the mixing end synchronously transmits to the recording end while playing the mixed sound source.

[0094] Exemplarily, the recording end 102 establishes a point-to-point connection with the mixing end 104, the recording end synchronizes the mixing parameters of the mixing end 104, the microphone of the recording end 102 starts recording, the recording end 102 obtains the recording audio signal of the current frame collected by the microphone, and receives the original echo signal corresponding to the current frame sent synchronously by the mixing end 104 when playing the mixed sound source.

[0095] Based on the above implementation method, before officially starting the recording work, the recording end synchronizes the mixing parameters of the mixing end, and uses the microphone to obtain the recorded audio signal of the current frame. It receives the original echo signal corresponding to the current frame synchronously transmitted by the mixing end when playing the mixed sound source, which lays a data foundation for the subsequent calculation of the estimated echo signal and echo cancellation of the recorded audio signal, thereby accelerating the efficiency of echo cancellation.

[0096] Step S202 : Input the original echo signal and the recorded audio signal into an adaptive filter configured using the filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame.

[0097] Among them, the filter parameters may include filter weight, echo delay and echo strength, echo delay and echo strength, among which the echo delay is positively correlated with the distance between the recording end and the mixing end, and the echo strength is negatively correlated with the distance; the adaptive filter can be understood as a filter that automatically updates the filter parameters in each frame, including adaptive update of filter parameters, labeling of human voice and echo, acquisition of estimated echo signals and other functions; the estimated echo signal can be understood as the echo component received by the microphone inferred by the adaptive filter.

[0098] In an exemplary embodiment, the recording end 102 obtains the distance change between the recording end 102 and the mixing end 104 in the current frame. The first frame is directly set to 1, and the non-first frame obtains the corresponding distance change based on the signal strength of the first frame and the signal strength of the current frame. The distance change and the original echo signal are used to obtain a reference echo signal. The reference echo signal and the recorded audio signal are input into an adaptive filter configured using the filter parameters of the current frame. By analyzing the echo reference signal and the recorded audio signal, an estimated echo signal received by the microphone is obtained.

[0099] According to the aforementioned embodiment, the recording terminal 102 inputs the original echo signal and the recorded audio signal into an adaptive filter configured using the filter parameters of the current frame. The adaptive filter then infers the original echo signal and the recorded audio signal to obtain an estimated echo signal received by the microphone. Because the filter parameters of the adaptive filter are updated with each frame, the estimated echo signal output by the adaptive filter is always accurate to the current situation, laying a data foundation for subsequent echo cancellation and ensuring the accuracy of echo cancellation.

[0100] Step S203 : performing echo cancellation on the recorded audio signal using the preset mixing weights and the estimated echo signal.

[0101] The mixing weight can be understood as the weight corresponding to each component in the mix output by the mixing end, generally including human voice, background sound and echo. Echo cancellation can be understood as removing the echo signal from the received recorded audio signal and retaining the echo with the corresponding weight.

[0102] Exemplarily, the recording end 102 synchronizes the mixing parameters of the mixing end 104, including the mixing weights of the mixing end 104, where the human voice: background: echo is equal to a:b:c, the sum of a, b, and c is not equal to 1, and the weights corresponding to the three are independent of each other. The weight c corresponding to the echo is used to cancel the echo of the recorded audio signal, and at the same time, the limiting amplitude is adjusted, and the limiting amplitude is d. Specifically: d*(a*recorded audio signal-(1-c)*estimated echo signal)=recorded audio signal after echo cancellation.

[0103] Based on the aforementioned implementation, echo cancellation is performed on the recorded audio signal by utilizing pre-synchronized mixing weights and estimated echo signals. At the same time, limiting adjustment is performed to ensure the amplitude consistency of the audio signal. The microphone recording stream is limited in advance while echo cancellation is being performed, thereby reducing the mixing processing time.

[0104] Step S204: audio encode the recorded audio signal after echo cancellation, and send the encoded recorded audio signal to the mixing end; the mixing end is used to audio decode the encoded recorded audio signal and perform mixing processing on the decoded recorded audio signal.

[0105] Among them, audio coding can be understood as the encoding process of converting analog signals into digital signals.

[0106] In an exemplary embodiment, the recording end 102 performs audio encoding on the recorded audio signal after echo cancellation, converts the analog signal into a digital signal, and sends the encoded recorded audio signal to the mixing end 104. After receiving the encoded recorded audio signal, the mixing end 104 performs audio decoding on it, restores the digital signal to the corresponding analog signal, and performs mixing processing on the decoded recorded audio signal. The mixed audio after processing is played through the microphone and returned to the recording end 102 for the next estimated echo signal acquisition.

[0107] According to the above embodiment, by performing audio encoding on the recorded audio signal after echo cancellation and converting the analog signal into a digital signal, the amount of data transmitted is reduced, and at the same time, the audio signal is prevented from being contaminated by other sound sources during transmission, thereby ensuring the integrity and security of the recorded audio signal. At the same time, the mixing end performs mixing processing after decoding, and one copy is played and the other is returned to the recording end for the next mixing, which also ensures the information synchronization of the interaction between the two ends, thereby improving the processing efficiency of the mixing processing.

[0108] In the above-mentioned microphone mixing method, the recorded audio signal of the current frame collected by the microphone and the original echo signal corresponding to the current frame sent by the receiving mixing end are obtained. The mixing end and the recording end are communicatively connected, and the original echo signal and the recorded audio signal are input into the adaptive filter configured using the filter parameters of the current frame to obtain the estimated echo signal corresponding to the current frame. The recorded audio signal is echo-cancelled using the pre-set mixing weights and the estimated echo signal, and the recorded audio signal after echo cancellation is audio-encoded, and the encoded recorded audio signal is sent to the mixing end. The mixing end performs audio decoding on the encoded recorded audio signal and performs mixing processing on the decoded recorded audio signal. The recording end is responsible for the echo cancellation work, and the echo cancellation work of the current frame is synchronously performed while receiving the recorded audio. The recording end and the mixing end jointly undertake the mixing work, thereby speeding up the mixing efficiency and avoiding the problem of audio delay causing audio and video synchronization.

[0109] In one embodiment, before the original echo signal and the recorded audio signal are input into the adaptive filter configured using the filter parameters of the current frame, the process further includes: obtaining the distance change and filter weight of the current frame, as well as a preset initial echo delay and initial echo intensity; obtaining the echo delay of the current frame based on the distance change and initial echo delay of the current frame; obtaining the echo intensity of the current frame based on the distance change and initial echo intensity of the current frame; using the filter weight, echo delay, and echo intensity of the current frame as the filter parameters of the current frame, and configuring the adaptive filter using the filter parameters of the current frame.

[0110] The distance change can be understood as the ratio of the distance between the recording end and the mixing end of the current frame to the distance between the recording end and the mixing end of the first frame. The initial echo delay can be understood as the default time delay for the microphone to continue collecting audio in the current frame.

[0111] Exemplarily, the recording end 102 obtains the distance change and filter weight of the current frame, as well as the pre-set initial echo delay and initial echo intensity, and obtains the echo delay of the current frame based on the distance change and initial echo delay of the current frame. Based on the distance change and initial echo intensity of the current frame, the echo intensity of the current frame is obtained, and the filter weight, echo delay and echo intensity of the current frame are used as the filter parameters of the current frame, and the filter parameters of the current frame are used to configure the adaptive filter.

[0112] Based on the aforementioned implementation, the initial echo delay and initial echo strength are updated by utilizing the change in the distance between the recording end and the mixing end, thereby obtaining the corresponding echo delay and echo strength for the current frame, and acquiring the filter weight for the current frame. Once the filter parameters for the current frame are acquired, the adaptive filter is configured, thereby improving the efficiency of the adaptive filter echo retrieval and reducing the echo cancellation waiting time.

[0113] In one embodiment, when the current frame is the first frame, the distance change and filter weight of the current frame are obtained, including: obtaining a preset initial distance change, and using the preset initial distance change as the distance change of the current frame; obtaining a preset initial filter weight, and using the preset initial filter weight as the filter weight of the current frame.

[0114] The initial distance change can be understood as being directly considered to be 1 when the current frame is the first frame, that is, when there is no previous frame.

[0115] In an exemplary embodiment, when the current frame is the first frame, the recording end 102 obtains a preset initial distance change 1, and uses the preset initial distance change 1 as the distance change of the current frame, obtains a preset initial filter weight w(0), and uses the preset initial filter weight w(0) as the filter weight of the current frame.

[0116] According to the above embodiment, in the first frame, by obtaining the preset initial distance change and initial filter weight, and using them as the distance change and filter weight for the current frame, setting reasonable initial values in the first frame helps mitigate the impact of environmental changes and noise, thereby improving the cancellation effect of echo cancellation and enhancing the accuracy and efficiency of the mixing process.

[0117] In one embodiment, when the current frame is not the first frame, the distance change and filter weight of the current frame are obtained, including: obtaining the signal strength of the communication connection between the recording end and the mixing end of the current frame, and obtaining the distance change of the current frame based on the signal strength; obtaining a preset initial step factor, the filter weight of the previous frame, the reference echo signal and the error signal; the previous frame is the previous frame of the current frame; based on the distance change and the initial step factor of the current frame, the step factor of the current frame is obtained; according to the step factor of the current frame, the filter weight of the previous frame, the reference echo signal and the error signal, the filter weight of the current frame is obtained.

[0118] The signal strength is an indication of the strength of the received signal at the recording end, typically expressed in dBm. A greater signal strength (closer to 0) indicates a stronger signal. The initial step size factor represents information retrieval, optimization, or a search process, referring to the distance or step size moved incrementally in the search space or parameter space. The reference echo signal is the echo signal obtained by updating the original echo signal based on the change in distance. The error signal is the error audio signal caused by delayed microphone recording.

[0119] For example, when the current frame is not the first frame, the recording end 102 obtains the signal strength of the communication connection between the recording end and the mixing end of the current frame, and calculates the distance change of the current frame using the signal strength of the first frame and the signal strength of the current frame. , and obtain the preset initial step factor μ(0), the filter weight w(n) of the previous frame of the current frame, the reference echo signal and error signal e(n), based on the distance change of the current frame and the initial step size factor μ(0), the step size factor μ(n+1) of the current frame is obtained, and the filter weight w(n) of the previous frame is obtained according to the step size factor μ(n+1) of the current frame, the filter weight w(n) of the previous frame, and the reference echo signal And the error signal e(n), get the filter weight w(n+1) of the current frame. The specific calculation process is shown below:

[0120]

[0121]

[0122]

[0123] Among them, n+1 is the current frame, n is the previous frame of the current frame, is a constant parameter, and △d is the distance change of the previous frame.

[0124] Based on the aforementioned implementation, the signal strength of the communication connection between the recording end and the mixing end in the current frame is obtained, and the distance change of the current frame is calculated based on the signal strength. A preset initial step size factor is obtained, and the step size factor corresponding to the current frame is calculated based on the distance change and the initial step size factor. The filter weight of the current frame is then calculated based on the filter weight of the previous frame, the reference echo signal, and the error signal, and the step size factor of the current frame. The displacement change between terminals is calculated based on the point-to-point connection signal strength, and the filter parameters of the current frame are calculated using the relevant parameters of the previous frame of the current frame. This improves the efficiency of adaptive filter echo retrieval and reduces echo cancellation wait time.

[0125] In one embodiment, after obtaining the estimated echo signal corresponding to the current frame, it also includes: obtaining a delayed recorded audio signal of the current frame collected by a microphone after the echo delay; and using the difference between the delayed recorded audio signal and the estimated echo signal as the error signal of the current frame.

[0126] The delayed recorded audio signal can be understood as the audio signal collected by the microphone during the echo delay period after the current frame.

[0127] In an exemplary embodiment, after obtaining the estimated echo signal Estimated_Echo_Signal corresponding to the current frame, the recording end obtains the delayed recorded audio signal d(n+Delay) of the current frame collected by the microphone after the echo delay Delay, and uses the difference between the delayed recorded audio signal and the estimated echo signal as the error signal of the current frame - (mic-est_echo)=d(n+Delay)-Estimated_Echo_Signal.

[0128] According to the above embodiment, after obtaining the estimated echo signal corresponding to the current frame, it is also necessary to obtain the delayed recorded audio signal collected after the echo delay of the current frame, and use the difference between the delayed recorded audio signal and the estimated echo signal as the error signal of the current frame, which lays a data foundation for updating the filter weights of the adaptive filter in the next frame. At the same time, the filter weights are updated according to the error signal, thereby improving the matching degree between the filter weights and each frame.

[0129] In one embodiment, an original echo signal and a recorded audio signal are input into an adaptive filter configured using filter parameters of a current frame to obtain an estimated echo signal corresponding to the current frame, including: inputting the original echo signal and the recorded audio signal into the adaptive filter, updating the original echo signal using a distance change of the current frame to obtain a reference echo signal of the current frame; inputting the reference echo signal and the recorded audio signal into a feature extraction unit of the adaptive filter to obtain corresponding acoustic features; and inputting the acoustic features into an intelligent control unit of the adaptive filter to obtain an estimated echo signal corresponding to the current frame.

[0130] Among them, acoustic features can be understood as digital descriptions extracted from audio signals to describe their acoustic properties, which may include Mel-Frequency Cepstral Coefficients (MFCC), short-time energy, fundamental frequency, resonance peaks, etc.; the intelligent control unit can be understood as a functional module for echo prediction, and it also has the task of automatically updating the filter parameters of each frame. Generally, the intelligent control unit outputs the estimated echo signal only after the filter parameters are updated.

[0131] Exemplarily, the recording end 102 inputs the original echo signal and the recorded audio signal into the adaptive filter, and updates the original echo signal using the distance change of the current frame: , get the reference echo signal of the current frame The reference echo signal and the recorded audio signal are input into the feature extraction unit of the adaptive filter for feature extraction to obtain acoustic features, and the acoustic features are input into the intelligent control unit of the adaptive filter for prediction to obtain the estimated echo signal corresponding to the current frame.

[0132] Based on the aforementioned implementation, the original echo signal is updated using the distance change of the current frame to obtain a reference echo signal for the current frame. Feature extraction is performed based on the reference echo signal and the recorded audio signal, and the extracted features are used to predict an estimated echo signal. This lays the data foundation for subsequent echo cancellation of the recorded audio signal using the estimated echo signal. Furthermore, calculating the estimated echo on the recording end reduces computational costs on the mixing end, thereby improving the efficiency of the mixing process.

[0133] In one embodiment, the mixing weight includes a human voice weight and an echo weight. Before performing echo cancellation on the recorded audio signal using the preset mixing weight and the estimated echo signal, the method further includes: using the mixing weight sent by the mixing end as the preset mixing weight; and performing echo cancellation on the recorded audio signal using the preset mixing weight and the estimated echo signal, including: performing echo cancellation on the recorded audio signal using the echo weight, the human voice weight and the estimated echo signal.

[0134] In an exemplary embodiment, the recording end 102 uses the mixing weight sent by the mixing end 104 as the pre-set mixing weight, that is, to achieve data synchronization of the mixing parameters of both ends, wherein the mixing weight includes the human voice weight and the echo weight, and the echo weight, the human voice weight and the estimated echo signal are used to perform echo cancellation and limit adjustment on the recorded audio signal.

[0135] Assuming the human voice weight is a, the echo weight is c, and the limit amplitude is d, the specific process of echo cancellation and limit adjustment is as follows:

[0136] Enhanced_Clean_Speech=d*[a*MIC_AUDIO-(1-c)*Est_Echo_Signal]

[0137] According to the above embodiment, the recording end 102 obtains the mixing weights after data synchronization with the mixing end. Based on the weights assigned to the recorded audio signal and the echo signal, the estimated echo signal is used to perform echo cancellation on the recorded audio signal, and subsequently, to perform clipping adjustment. By having the recording end handle the echo cancellation and clipping tasks, the mixing end significantly reduces the number of mixing steps required, thereby increasing the speed of the mixing process.

[0138] In an exemplary embodiment, Figure 3 As shown, a microphone mixing method is provided, which is applied to Figure 1 The mixing terminal 104 in FIG. 1 is taken as an example to illustrate the method, which includes the following steps S301 and S302.

[0139] Step S301: Send the original echo signal corresponding to the current frame to the recording end; the recording end is communicatively connected to the mixing end; the recording end is used to obtain the recorded audio signal of the current frame collected by the microphone, receive the original echo signal, input the original echo signal and the recorded audio signal into an adaptive filter configured using the filter parameters of the current frame, and obtain an estimated echo signal corresponding to the current frame; use the preset mixing weights and the estimated echo signal to echo cancel the recorded audio signal; perform audio encoding on the recorded audio signal after echo cancellation, and return it.

[0140] Exemplarily, when the current frame is the first frame, the mixing end 104 plays the original mixed audio source and sends the original echo signal corresponding to the current frame to the recording end 102 to which it is in communication; when the current frame is not the first frame, the mixing end 104 plays the mixed audio of the previous frame and sends the original echo signal corresponding to the current frame to the recording end 102. The recording end 102 obtains the recorded audio signal of the current frame collected by the microphone, receives the original echo signal, inputs the original echo signal and the recorded audio signal into an adaptive filter configured using the filter parameters of the current frame for prediction, obtains an estimated echo signal corresponding to the current frame, uses pre-set mixing weights to perform echo cancellation and limiter adjustment on the recorded audio signal, obtains the recorded audio signal after echo cancellation, and finally performs audio encoding on the recorded audio signal after echo cancellation and returns it to the mixing end 104.

[0141] Based on the aforementioned implementation, while the mixing end is playing the mixed audio source, it also needs to send a corresponding original echo signal to the recording end. The recording end then performs echo cancellation on the recorded audio signal based on the received original echo signal and the collected recorded audio signal. By sending an original echo signal to the recording end for each frame, data synchronization between the two ends is ensured, thereby ensuring consistency in the mixing processing of the mixed audio signal and improving the processing quality of the mixing process.

[0142] Step S302: receiving the encoded recorded audio signal, performing audio decoding on the encoded recorded audio signal, and performing mixing processing on the decoded recorded audio signal.

[0143] In an exemplary embodiment, the mixing end 104 receives the encoded recorded audio signal, decodes the encoded recorded audio signal, reconverts the digital signal into an analog signal, and mixes the decoded recorded audio signal according to a preset mixing weight.

[0144] According to the above embodiment, the mixing end decodes the received encoded recorded audio signal and converts the digital signal back into a digital signal, which is conducive to the subsequent mixing of the analog signal. At the same time, according to the preset mixing weight, the decoded recorded audio signal is mixed without the need for echo cancellation, which effectively improves the processing speed of audio mixing and avoids the problem of audio and video being out of sync.

[0145] In one embodiment, mixing processing is performed on the decoded recorded audio signal, including: obtaining background sound and pre-set mixing weights; and mixing processing is performed on the decoded recorded audio signal using the background sound weight and the background sound included in the mixing weights.

[0146] The mixing process can be understood as the process of mixing with background sound and performing dynamic range compression, equalizer, reverberation and other operations on the mixed audio.

[0147] Exemplarily, the mixing end 104 obtains the background sound BG_MUSIC and the pre-set mixing weights, which include the human voice weight a, the background sound weight b and the echo weight c. The background sound weight b and the background sound BG_MUSIC included in the mixing weights are used to mix the decoded recorded audio signal - SPK_AUDIO=Enhanced_Clean_Speech+b*BG_MUSIC, and perform dynamic range compression, equalizer, and reverberation, and output it to the speaker for playback.

[0148] Based on the aforementioned method, by obtaining background sound and pre-set mixing weights, the background sound is mixed into the decoded recorded audio signal using the background sound weight and background sound included in the mixing weights. Dynamic range compression, equalizer, and reverberation are then performed. By adding background sound, the audio is given a more spatial and natural feel, creating a more realistic auditory scene. The mixing weights are used to precisely control the ratio of background sound to the recording, achieving personalized audio adjustments to meet different needs. The combination of dynamic range compression, equalizer, and reverberation helps balance the various frequency bands of the audio, reducing excessive or insufficient sound pressure differences, and improving overall clarity and naturalness.

[0149] In one embodiment, after mixing the decoded recorded audio signal, the method includes: sending a mixing weight to a recording end; and the recording end is configured to receive the mixing weight and use the mixing weight as a preset mixing weight.

[0150] In an exemplary embodiment, after mixing the decoded recorded audio signal, the mixing end 104 sends the mixing weight to the recording end 102. The recording end 102 receives the mixing weight and uses the mixing weight as the pre-set mixing weight for the echo cancellation of the recorded audio signal of the next frame.

[0151] According to the aforementioned embodiment, after the decoded recorded audio signal is mixed, the mixing end sends the mixing weights to the recording end, and the corresponding mixing parameters are synchronized for each frame, which is conducive to ensuring data synchronization between the two ends, thereby ensuring data synchronization of the mixing processing and improving the mixing quality of the audio mixing.

[0152] In one embodiment, Figure 4 As shown, a microphone mixing system is provided, the system includes a recording end and a mixing end, and the recording end and the mixing end are communicatively connected;

[0153] The mixing end is used to send the original echo signal corresponding to the current frame to the recording end;

[0154] The recording end is used to obtain the recorded audio signal of the current frame collected by the microphone and receive the original echo signal, input the original echo signal and the recorded audio signal into the adaptive filter configured with the filter parameters of the current frame, obtain the estimated echo signal corresponding to the current frame, and use the preset mixing weights and the estimated echo signal to perform echo cancellation on the recorded audio signal; perform audio encoding on the recorded audio signal after echo cancellation, and send the encoded recorded audio signal to the mixing end;

[0155] The mixing end is used to receive the encoded recorded audio signal, perform audio decoding on the encoded recorded audio signal, and perform mixing processing on the decoded recorded audio signal.

[0156] Exemplarily, a microphone mixing system includes a recording end and a mixing end, the recording end is communicatively connected to the mixing end, the mixing end sends the original echo signal corresponding to the current frame to the recording end, the recording end receives the original echo signal, and obtains the recorded audio signal of the current frame collected by the microphone, inputs the original echo signal and the recorded audio signal into an adaptive filter configured using the filter parameters of the current frame, obtains an estimated echo signal corresponding to the current frame, uses the human voice weight and echo weight included in the pre-set mixing weight, and the estimated echo signal, to perform echo cancellation and limit adjustment on the recorded audio signal to obtain the recorded audio signal after echo cancellation, audio encodes the recorded audio signal after echo cancellation, and sends the encoded recorded audio signal to the mixing end, the mixing end receives the encoded recorded audio signal, audio decodes the encoded recorded audio signal, and performs mixing processing on the decoded recorded audio signal. The recording end takes on part of the mixing processing work - echo cancellation and limiter adjustment, and at the same time the mixing end also participates in the parameter construction and update process of this process, ensuring the synchronization of parameters on both ends. Secondly, having the recording end take on part of the work also greatly reduces the data processing volume of the mixing end, effectively improving the mixing processing efficiency of the mixing and avoiding the problem of audio and video asynchrony caused by audio delay.

[0157] See the comparison diagram of the original method for microphone mixing and the present application. Figure 5 In the original method, the recording end only performs recording operations, and echo cancellation and mixing are performed on the mixing end; this application transfers some of the echo cancellation and mixing work that can be coordinated to the recording end; this application uses the signal changes of the point-to-point connection to infer the changes in the recording distance, which is used to update the parameters of the adaptive filter (the following data are all specific examples, and do not limit this application to be implemented only in this case).

[0158] 1. Wi-Fi Direct: A Wi-Fi technology that allows devices to connect directly to each other without going through a wireless router (AP). It is commonly used for point-to-point transmission, such as direct communication between a mobile phone and a printer, or a mobile phone and a TV.

[0159] 2. BT (Bluetooth): A short-range wireless communication technology, commonly used for data exchange between devices such as headphones, keyboards, and bracelets.

[0160] 3. RSSI (Received Signal Strength Indicator): Indicates the received signal strength of the wireless signal, usually measured in dBm. A larger RSSI (closer to 0) indicates a stronger signal.

[0161] 4. Mixing: The process of combining multiple audio signals into one. For example, combining microphone, system sound, and background music.

[0162] 5. Limiter: A dynamic processing tool used to limit the maximum output amplitude of an audio signal to prevent overload and distortion. Commonly used in recording, live streaming, and other scenarios.

[0163] 6. Echo Reference: The reference signal used in Acoustic Echo Cancellation (AEC), usually the audio signal from the playback end, is used to infer the echo component received by the microphone.

[0164] 7. Inferred Echo Signal: In the echo cancellation algorithm, the echo component contained in the microphone is inferred by analyzing the echo reference signal and the microphone signal.

[0165] 8. Echo Strength: refers to the strength of the inferred echo portion of the microphone signal, usually used to adjust the processing strength of the AEC algorithm.

[0166] 9. Echo Delay: The delay time it takes for the playback signal to reach the microphone, usually measured in milliseconds. It affects the accuracy of the echo cancellation algorithm.

[0167] 10. Adaptive Filter: A filter that can dynamically adjust its own parameters according to the input signal. It is often used in scenarios such as noise suppression, echo cancellation, and channel estimation.

[0168] The system architecture of this application includes a recording terminal and a mixing terminal. Figure 6 .

[0169] Mixer:

[0170] (1) Mixing engine: weighted additive mixing processing, dynamic range compression, equalization and reverberation;

[0171] (2) Audio decoder: responsible for decoding the transmitted human voice audio;

[0172] (3) Transmission module: responsible for transmitting mixing parameters and echo signals to the recording end, and receiving the human voice audio transmitted by the recording end.

[0173] (4) Audio output: Play mixed audio.

[0174] (5) WiFi / BT: Establish a point-to-point connection.

[0175] Recording end:

[0176] (1) Audio decoder: encodes the processed audio stream into an audio format;

[0177] (2) Mixing pre-processing module: responsible for audio signal echo cancellation, weighting adjustment and limiting;

[0178] (3) Distance estimation module: estimates the distance change based on the RSSI of the WiFi / BT module.

[0179] (4) Transmission module: responsible for receiving mixing parameters and transmitting echo signals to the recording end, and transmitting human voice audio to the mixing end.

[0180] (5) Adaptive filter: responsible for estimating the echo signal and labeling the human voice and echo signal.

[0181] (6) Collaborative control module: responsible for adaptive filter adjustment and recording parameter adjustment.

[0182] (7) WiFi / BT: Establish a point-to-point connection.

[0183] (8) Microphone: responsible for recording sound streams, including human voices and echo sounds.

[0184] The implementation process of this application includes four steps: initialization, recording and pre-processing, mixing, and post-processing. Figure 7 :

[0185] Step 1: Initialization:

[0186] The mixing end and the recording end are connected via WiFi or Bluetooth point-to-point;

[0187] The mixing end synchronizes parameters (sampling rate, format, reverberation weighting coefficient, and limiting index) with the recording end;

[0188] Configure the microphone on the recording end (sampling rate, recording format);

[0189] Get the RSSI of the point-to-point connection and initialize the distance R0 according to the logarithmic path loss model.

[0190] Step 2: Recording and pre-processing (recording end):

[0191] According to the initialization distance d0, the displacement △d=1, and the adaptive filter is configured;

[0192] Call the microphone to record:

[0193] Adaptive filters are used to mark human voices and echoes;

[0194] According to the reverberation weighting coefficient, the weighted amplitude of the human voice and echo are adjusted and limited respectively;

[0195] The adjusted sound stream is encoded and transmitted to the mixing end.

[0196] Step 3: Mixing process (mixing end):

[0197] Receive and decode the sound stream;

[0198] Adjust the background sound volume (normalize) and mix the sound stream with the background sound stream into a mixed audio stream;

[0199] Sound post-processing, including dynamic range compression, equalizer, reverb, etc.;

[0200] The mixed audio stream signal is output, one copy is given to the speaker for playback, and the other copy is transmitted to the adaptive filter at the recording end to generate the estimated echo signal.

[0201] Step 4: Post-mixing processing (recording end):

[0202] Re-obtain the RSSI value and calculate the distance change;

[0203] Configure adaptive filters based on distance changes;

[0204] According to the mixed audio signal input from the mixing end, an estimated echo signal is generated through an adaptive filter for echo cancellation of the next frame of audio recording signal.

[0205] Model Description:

[0206] The following are the neural models of the distance estimation model and distance-based adaptive filter mentioned in the implementation process:

[0207] 1) Distance estimation model: calculate △d.

[0208] Initialize distance d0, signal strength RSSI, and n path attenuation coefficient:

[0209]

[0210]

[0211] Adaptive echo cancellation model: DDE-Net (Distance-aware Deep Echo Canceller), which uses a DNN to control the adaptive filter (LMS (Least Mean Squares) or NLMS (Normalized Least Mean Squares)). For a schematic diagram of the model, see Figure 8 .

[0212] Echo delay is positively correlated with distance:

[0213] Delay=h*(△d-1)+Delay0

[0214] Where Delay0 is the initial echo delay and h is a constant factor.

[0215] Adaptive filter weights, step factor adjustment:

[0216] w filter weight:

[0217] e(n) error signal (mic-est_echo) = d(n+Delay)-Estimated_Echo_Signal, where d(n) is the signal received by the microphone in time window n.

[0218] μ step size factor (the closer the distance, the higher the filter weight and the more sensitive it is), is a constant factor.

[0219] x(n): The signal output by the speaker in time window n, the reference signal vector. The farther the distance, the smaller the echo reference becomes.

[0220]

[0221]

[0222]

[0223] Echo intensity is negatively correlated with distance:

[0224]

[0225] The human voice can be echo-cancelled by weighting the mix:

[0226] CleanSpeech=a*MIC_AUDIO-(1-c)*Estimated_Echo_Signal

[0227] a is the weight corresponding to MIC_AUDIO, c is the weight corresponding to Estimated_Echo_Signal, and CleanSpeech is the recorded audio signal after echo cancellation.

[0228] One embodiment of the present application uses a mobile phone + smart TV multi-terminal collaborative karaoke service mixing optimization, the flow chart is as follows Figure 9 shown.

[0229] initialization:

[0230] 1. The user opens the karaoke app.

[0231] 2. Check the mixing parameters of the smart TV, such as 48KHz sampling rate, PCM encoding format, mixing weight (voice / background / echo = 1:0.6:0.1), and limit to 32000.

[0232] 3. Smart TV and mobile phone are connected point-to-point (WiFi direct or Bluetooth).

[0233] 4. Synchronize the mixing parameters on the mobile phone, configure the microphone module with a 48KHz sampling rate and PCM encoding format, and the microphone starts recording.

[0234] Start the first loop:

[0235] 5. Check the point-to-point connection RSSI on the mobile phone = -40dBm.

[0236] 6. The mobile phone configures an adaptive filter, imports the parameters of the adaptive filter, initializes △d=1, delay(0), filter weight w(0) and search step size μ(0).

[0237] 7. The smart TV plays the audio source and simultaneously transmits the echo stream SPK_AUDIO to the mobile phone's adaptive filter to generate an echo stream signal x(0).

[0238] 8. Adaptive filter: Find the estimated echo signal Esti_Echo_Signal based on the echo stream x(n) and calculate the error e(0). e(0) is used to update the adaptive filter weights for the next time.

[0239] 9. The user sings synchronously and the microphone records MIC_AUDIO.

[0240] 10. The phone eliminates the recorded stream and the estimated echo signal based on the mixing weight, and adjusts the limit to Enhanced_Clean_Speech = 1*(1*MIC_AUDIO-0.9*Est_Echo_Signal).

[0241] 11. Encode the mixed audio stream into PCM (Pulse-Code Modulation) and transmit it to the smart TV.

[0242] 12. The smart TV decodes the audio and mixes it with the background stream SPK_AUDIO=Enhanced_Clean_Speech+0.6*BG_MUSIC, performs dynamic range compression, equalizer, and reverberation, and outputs it to the speaker for playback.

[0243] Enter the nth loop until the karaoke playback ends (n≥2):

[0244] 13. Check the point-to-point connection of the mobile phone and find RSSI = -48dBm. Calculate △d = 1.84 (loss n = 3).

[0245] 14. The mobile phone configures an adaptive filter, imports the parameters of the adaptive filter, and calculates delay(n), filter weight w(n) and search step size μ(n).

[0246] 15. The smart TV plays the audio source and simultaneously transmits the echo stream SPK_AUDIO to the mobile phone's adaptive filter to generate an echo stream signal x(n).

[0247] 16. Adaptive filter: Update the echo stream reference signal based on △d according to the echo stream x(n) , find the estimated echo signal Est_Echo_Signal and calculate the error e(n), which is used to update the adaptive filter weights for the next time.

[0248] 17. The user sings synchronously and the microphone records MIC_AUDIO.

[0249] 18. The phone eliminates the recorded stream and the estimated echo signal based on the mixing weight, and performs a limit adjustment Enhanced_Clean_Speech = 1*(1*MIC_AUDIO-0.9*Est_Echo_Signal).

[0250] 19. Encode the mixed audio stream into PCM (n) and transmit it to the smart TV.

[0251] 20. The smart TV decodes the audio and mixes it with the background stream SPK_AUDIO=Enhanced_Clean_Speech+0.6*BG_MUSIC, performs dynamic range compression, equalizer, and reverberation, and outputs it to the speaker for playback.

[0252] Compared with the existing technology, the technical solution proposed in this application improves the efficiency of echo cancellation and mixing and reduces performance requirements:

[0253] 1. Improved echo cancellation efficiency: The displacement change between terminals is calculated based on the signal strength of the point-to-point connection, thereby improving the efficiency of the adaptive filter echo retrieval and reducing the echo cancellation waiting time;

[0254] 2. Improved mixing efficiency: While canceling echoes, the microphone recording stream is adjusted for mixing weight and clipped in advance, reducing mixing processing time.

[0255] 3. Reduce the performance requirements of the mixing end: Moving the processing of echo cancellation and adaptive filtering DNN to the recording end (such as a smartphone with an AI chip) can reduce the CPU / GPU performance requirements of the mixing end (such as a smart TV).

[0256] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0257] Based on the same inventive concept, embodiments of the present application also provide a microphone mixing device for implementing the aforementioned microphone mixing method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more microphone mixing device embodiments provided below can be found in the above-described limitations of the microphone mixing method and will not be further elaborated here.

[0258] In an exemplary embodiment, Figure 10 As shown, a microphone mixing device 1000 is provided, which is applied to a recording end and includes: a first signal acquisition module 1010, a second signal acquisition module 1020, an echo cancellation module 1030 and a first signal sending module 1040, wherein:

[0259] The first signal acquisition module 1010 is used to obtain the recorded audio signal of the current frame collected by the microphone and receive the original echo signal corresponding to the current frame sent by the mixing end; the mixing end is in communication with the recording end;

[0260] The second signal acquisition module 1020 is configured to input the original echo signal and the recorded audio signal into an adaptive filter configured using the filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame;

[0261] The echo cancellation module 1030 is used to perform echo cancellation on the recorded audio signal using the preset mixing weights and the estimated echo signal;

[0262] The first signal sending module 1040 is used to audio encode the recorded audio signal after echo cancellation and send the encoded recorded audio signal to the mixing end; the mixing end is used to audio decode the encoded recorded audio signal and mix the decoded recorded audio signal.

[0263] In one embodiment, based on the above-mentioned microphone mixing device, the first signal acquisition module acquires the recorded audio signal of the current frame collected by the microphone, and receives the original echo signal corresponding to the current frame sent by the mixing end connected to the first signal acquisition module, and transmits the recorded audio signal and the original echo signal to the second signal acquisition module. The second signal acquisition module inputs the original echo signal and the recorded audio signal into the adaptive filter configured with the filter parameters of the current frame to obtain the estimated echo signal corresponding to the current frame, and transmits the estimated echo signal and the recorded audio signal to the echo cancellation module. The echo cancellation module uses the preset mixing weight and the estimated echo signal to perform echo cancellation on the recorded audio signal, and then sends the recorded audio signal after echo cancellation to the first sending module. The first signal sending module audio encodes the recorded audio signal after echo cancellation and sends the encoded recorded audio signal to the mixing end. The mixing end is used to audio decode the encoded recorded audio signal and perform mixing processing on the decoded recorded audio signal. The recording end is responsible for echo cancellation, performing echo cancellation for the current frame simultaneously with the recording audio. The original echo signal sent by the mixing end also participates in the echo cancellation process, ensuring consistent mixing. The recording and mixing ends share the mixing process, which speeds up mixing efficiency while ensuring mixing quality and avoiding audio delays that can cause audio and video asynchrony.

[0264] In one embodiment, the second signal acquisition module 1020 is further used to obtain the distance change and filter weight of the current frame, as well as the preset initial echo delay and initial echo intensity; obtain the echo delay of the current frame based on the distance change and initial echo delay of the current frame; obtain the echo intensity of the current frame based on the distance change and initial echo intensity of the current frame; use the filter weight, echo delay and echo intensity of the current frame as filter parameters of the current frame, and configure the adaptive filter using the filter parameters of the current frame.

[0265] In one embodiment, when the current frame is the first frame, the signal second acquisition module 1020 is further used to obtain a preset initial distance change, and use the preset initial distance change as the distance change of the current frame; obtain a preset initial filter weight, and use the preset initial filter weight as the filter weight of the current frame.

[0266] In one embodiment, when the current frame is not the first frame, the second signal acquisition module 1020 is further used to obtain the signal strength of the communication connection between the recording end and the mixing end of the current frame, and obtain the distance change of the current frame based on the signal strength; obtain a preset initial step factor, the filter weight of the previous frame, the reference echo signal and the error signal; the previous frame is the previous frame of the current frame; based on the distance change and the initial step factor of the current frame, obtain the step factor of the current frame; according to the step factor of the current frame, the filter weight of the previous frame, the reference echo signal and the error signal, obtain the filter weight of the current frame.

[0267] In one embodiment, after obtaining the estimated echo signal corresponding to the current frame, the microphone mixing device 1000 further includes an error signal acquisition module for obtaining a delayed recorded audio signal of the current frame collected by the microphone after the echo delay; and using the difference between the delayed recorded audio signal and the estimated echo signal as the error signal of the current frame.

[0268] In one embodiment, the second signal acquisition module 1020 is further used to input the original echo signal and the recorded audio signal into the adaptive filter, update the original echo signal using the distance change of the current frame, and obtain a reference echo signal of the current frame; input the reference echo signal and the recorded audio signal into the feature extraction unit of the adaptive filter to obtain corresponding acoustic features; and input the acoustic features into the intelligent control unit of the adaptive filter to obtain an estimated echo signal corresponding to the current frame.

[0269] In one embodiment, the mixing weight includes a vocal weight and an echo weight. The echo cancellation module 1030 is further configured to use the mixing weight sent by the mixing end as a pre-set mixing weight, and to perform echo cancellation on the recorded audio signal using the echo weight, the vocal weight, and the estimated echo signal.

[0270] In an exemplary embodiment, Figure 11 As shown, a microphone mixing device 1100 is provided, which is applied to a mixing end and includes: a second signal sending module 1110 and a signal receiving module 1120, wherein:

[0271] The second signal sending module 1110 is configured to send the original echo signal corresponding to the current frame to the recording end; the recording end is communicatively connected to the mixing end; the recording end is configured to obtain the recorded audio signal of the current frame collected by the microphone, receive the original echo signal, input the original echo signal and the recorded audio signal into an adaptive filter configured using the filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame; perform echo cancellation on the recorded audio signal using pre-set mixing weights and the estimated echo signal; perform audio encoding on the echo-cancelled recorded audio signal, and return the result;

[0272] The signal receiving module 1120 is configured to receive the encoded recorded audio signal, perform audio decoding on the encoded recorded audio signal, and perform mixing processing on the decoded recorded audio signal.

[0273] In one embodiment, the signal receiving module 1120 is further configured to obtain background sound and preset mixing weights; and perform mixing processing on the decoded recorded audio signal using the background sound weight and the background sound included in the mixing weights.

[0274] In one embodiment, after mixing the decoded recorded audio signal, the microphone mixing device 1100 further includes a mixing weight sending module for sending the mixing weight to the recording end; the recording end is configured to receive the mixing weight and use the mixing weight as a pre-set mixing weight.

[0275] Each module in the microphone mixing device described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a communication device in hardware form, or can be stored in a memory in the communication device in software form, so that the processor can call and execute the corresponding operations of each module.

[0276] In an exemplary embodiment, a communication device is provided. The communication device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 12As shown. The communication device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the communication device is used to provide computing and control capabilities. The memory of the communication device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the communication device is used to exchange information between the processor and external devices. The communication interface of the communication device is used to communicate with external terminals via wired or wireless means. The wireless means can be implemented via Wi-Fi, a mobile cellular network, near field communication (NFC), or other technologies. When executed by the processor, the computer program implements a microphone mixing method. The display unit of the communication device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the communication device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the communication device casing, or an external keyboard, touchpad or mouse.

[0277] Those skilled in the art will understand that Figure 12 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present application, and does not constitute a limitation on the communication device to which the scheme of the present application is applied. The specific communication device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0278] In an exemplary embodiment, a communication device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the microphone mixing method of the above embodiment when executing the computer program.

[0279] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the microphone mixing method of the above embodiment is implemented.

[0280] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the microphone mixing method of the above embodiment is implemented.

[0281] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0282] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0283] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0284] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A microphone mixing method, characterized in that: Applied to the recording end, the method includes: Acquiring a recorded audio signal of a current frame collected by a microphone, and receiving an original echo signal corresponding to the current frame sent by a mixing end; the mixing end is communicatively connected to the recording end; Inputting the original echo signal and the recorded audio signal into an adaptive filter configured using the filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame; performing echo cancellation on the recorded audio signal using a preset mixing weight and the estimated echo signal; The recorded audio signal after echo cancellation is audio-encoded and sent to the mixing end; the mixing end is used to perform audio decoding on the encoded recorded audio signal and perform mixing processing on the decoded recorded audio signal.

2. The method according to claim 1, characterized in that Before inputting the original echo signal and the recorded audio signal into the adaptive filter configured using the filter parameters of the current frame, the method further includes: Obtaining the distance change and filter weight of the current frame, as well as the preset initial echo delay and initial echo intensity; Obtaining an echo delay of the current frame according to the distance variation of the current frame and the initial echo delay; Obtaining the echo intensity of the current frame based on the distance change of the current frame and the initial echo intensity; The filter weight, echo delay, and echo strength of the current frame are used as filter parameters of the current frame, and the adaptive filter is configured using the filter parameters of the current frame.

3. The method according to claim 2, characterized in that In a case where the current frame is the first frame, obtaining the distance change and the filter weight of the current frame includes: Obtaining a preset initial distance variation, and using the preset initial distance variation as the distance variation of the current frame; Acquire a preset initial filter weight, and use the preset initial filter weight as the filter weight of the current frame.

4. The method according to claim 2, characterized in that In a case where the current frame is not the first frame, obtaining the distance change and the filter weight of the current frame includes: Obtaining a signal strength of a communication connection between the recording end and the mixing end in the current frame, and obtaining a distance change in the current frame according to the signal strength; Obtaining a preset initial step factor, a filter weight of a previous frame, a reference echo signal, and an error signal; the previous frame is a frame previous to the current frame; Obtaining a step size factor for the current frame based on the distance change of the current frame and the initial step size factor; The filter weight of the current frame is obtained according to the step factor of the current frame, the filter weight of the previous frame, a reference echo signal, and an error signal.

5. The method according to claim 4, characterized in that After obtaining the estimated echo signal corresponding to the current frame, the method further includes: Acquire a delayed recorded audio signal of the current frame collected by the microphone after the echo delay; The difference between the delayed recorded audio signal and the estimated echo signal is used as the error signal of the current frame.

6. The method according to claim 1, characterized in that Inputting the original echo signal and the recorded audio signal into an adaptive filter configured using filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame includes: Inputting the original echo signal and the recorded audio signal into the adaptive filter, and updating the original echo signal using the distance variation of the current frame to obtain a reference echo signal of the current frame; Inputting the reference echo signal and the recorded audio signal into the feature extraction unit of the adaptive filter to obtain corresponding acoustic features; The acoustic features are input into the intelligent control unit of the adaptive filter to obtain the estimated echo signal corresponding to the current frame.

7. The method according to any one of claims 1 to 6, characterized in that The mixing weight includes a human voice weight and an echo weight. Before performing echo cancellation on the recorded audio signal using the preset mixing weight and the estimated echo signal, the method further includes: Using the mixing weight sent by the mixing end as the preset mixing weight; The method of performing echo cancellation on the recorded audio signal by using the preset mixing weight and the estimated echo signal includes: Echo cancellation is performed on the recorded audio signal using the echo weight, the human voice weight and the estimated echo signal.

8. A microphone mixing method, characterized in that: Applied to a mixing end, the method includes: The recording end is communicatively connected to the mixing end; the recording end is configured to obtain a recorded audio signal of the current frame collected by a microphone, receive the original echo signal, input the original echo signal and the recorded audio signal into an adaptive filter configured using filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame; perform echo cancellation on the recorded audio signal using pre-set mixing weights and the estimated echo signal; perform audio encoding on the echo-cancelled recorded audio signal, and return the result; The encoded recorded audio signal is received, audio decoding is performed on the encoded recorded audio signal, and mixing processing is performed on the decoded recorded audio signal.

9. The method according to claim 8, characterized in that The mixing process of the decoded recorded audio signal includes: Get background sounds and preset mix weights; The decoded recorded audio signal is mixed using the background sound weight included in the mixing weight and the background sound.

10. The method according to any one of claims 8 to 9, characterized in that: After the mixing process is performed on the decoded recorded audio signal, the method includes: The mixing weight is sent to the recording end; the recording end is used to receive the mixing weight and use the mixing weight as a preset mixing weight.

11. A microphone mixing system, characterized in that: The system includes a recording end and a mixing end, wherein the recording end and the mixing end are communicatively connected; The mixing end is used to send the original echo signal corresponding to the current frame to the recording end; The recording end is configured to obtain a recorded audio signal of a current frame collected by a microphone, and receive the original echo signal, input the original echo signal and the recorded audio signal into an adaptive filter configured using filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame, perform echo cancellation on the recorded audio signal using preset mixing weights and the estimated echo signal; perform audio encoding on the recorded audio signal after echo cancellation, and send the encoded recorded audio signal to the mixing end; The mixing end is used to receive the encoded recorded audio signal, perform audio decoding on the encoded recorded audio signal, and perform mixing processing on the decoded recorded audio signal.

12. A microphone mixing device, characterized in that: Applied to the recording end, the device includes: a first signal acquisition module, configured to acquire a recorded audio signal of a current frame collected by a microphone, and receive an original echo signal corresponding to the current frame sent by a mixing end; the mixing end is communicatively connected to the recording end; a second signal acquisition module, configured to input the original echo signal and the recorded audio signal into an adaptive filter configured using the filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame; an echo cancellation module, configured to perform echo cancellation on the recorded audio signal using a preset mixing weight and the estimated echo signal; The first signal sending module is used to audio encode the recorded audio signal after echo cancellation and send the encoded recorded audio signal to the mixing end; the mixing end is used to audio decode the encoded recorded audio signal and perform mixing processing on the decoded recorded audio signal.

13. A microphone mixing device, characterized in that: Applied to a mixing end, the device comprises: A second signal sending module is configured to send an original echo signal corresponding to a current frame to a recording end; the recording end is communicatively connected to the mixing end; the recording end is configured to obtain a recorded audio signal of the current frame collected by a microphone, receive the original echo signal, input the original echo signal and the recorded audio signal into an adaptive filter configured using filter parameters of the current frame to obtain an estimated echo signal corresponding to the current frame; perform echo cancellation on the recorded audio signal using pre-set mixing weights and the estimated echo signal; perform audio encoding on the echo-cancelled recorded audio signal, and return the result; The signal receiving module is used to receive the encoded recorded audio signal, perform audio decoding on the encoded recorded audio signal, and perform mixing processing on the decoded recorded audio signal.

14. A communication device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.