Echo cancellation method and apparatus, conference system, electronic device, and storage medium

By using a multi-source signal fusion network and echo cancellation technology, the echo cancellation problem caused by different speaker delays in multi-person conferences is solved, realizing automatic management of terminal audio signals and efficient echo cancellation, thus improving conference quality and stability.

CN115665602BActive Publication Date: 2026-04-17IFLYTEK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2022-10-12
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies cannot effectively eliminate echoes caused by the different time delays of multiple speakers in multi-person conference scenarios, and therefore cannot achieve effective echo cancellation.

Method used

By acquiring reference signals and microphone signals from each terminal, feature extraction and fusion are performed using a multi-source signal fusion network to estimate echo signal features. Echo cancellation is then performed based on these echo signal features. Finally, pickup control is implemented by combining terminal status and voice detection results to achieve automatic management.

Benefits of technology

It enables echo cancellation for multiple terminals in multi-person conference scenarios, improving conference quality and stability, avoiding the inconvenience of manual control, and enhancing the automated management of audio signal acquisition and playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115665602B_ABST
    Figure CN115665602B_ABST
Patent Text Reader

Abstract

The application provides an echo cancellation method and device, a conference system, electronic equipment and a storage medium, wherein the method comprises: obtaining reference signals of each terminal and a microphone signal of any terminal in the terminals; performing feature extraction on the reference signals of each terminal and the microphone signal of the terminal, respectively, and determining an echo signal feature based on the reference signal features of each terminal and the microphone signal feature of the terminal obtained through the feature extraction; and performing echo cancellation on the microphone signal of the terminal based on the echo signal feature to obtain an echo cancellation signal of the terminal, which overcomes the defect that the traditional echo cancellation method cannot perform echo cancellation for a multi-person conference scene, and simultaneously realizes automatic management of terminal audio signal acquisition and playing, avoids the inconvenience of manual control, and improves the stability of the conference process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information processing technology, and in particular to an echo cancellation method, apparatus, conferencing system, electronic device, and storage medium. Background Technology

[0002] With the development of information technology, smart devices are being used more and more widely in various fields. Echo cancellation, as an indispensable part of smart device interaction, has always been a research hotspot for technicians in related fields.

[0003] Echo cancellation addresses speaker-microphone coupling by eliminating or removing the far-end audio signal picked up by the microphone and output by the speaker, thus preventing the far-end audio signal from being echoed back to the far end. Common echo cancellation methods utilize adaptive filters, which adaptively update the transfer function between the speaker and microphone using an algorithm.

[0004] However, in multi-person conference scenarios, there are often multiple speakers on. Due to differences in distance, network, hardware, etc., the latency of the audio signals played by each speaker is also different. This makes traditional echo cancellation methods ineffective for echo cancellation because they cannot estimate multiple latencys at the same time, and therefore cannot perform echo cancellation for this scenario. Summary of the Invention

[0005] This invention provides an echo cancellation method, apparatus, conferencing system, electronic device, and storage medium to address the shortcomings of existing technologies that cannot simultaneously estimate multiple delays, thus failing to effectively cancel echoes in multi-person conferencing scenarios.

[0006] This invention provides an echo cancellation method, comprising:

[0007] Acquire the reference signal of each terminal, and the microphone signal of any one of the terminals;

[0008] Feature extraction is performed on the reference signals of each terminal and the microphone signal of any terminal respectively, and the echo signal features are determined based on the feature extraction of the reference signal features of each terminal and the microphone signal features of any terminal.

[0009] Based on the echo signal characteristics, echo cancellation is performed on the microphone signal of any terminal to obtain the echo cancellation signal of any terminal.

[0010] According to an echo cancellation method provided by the present invention, determining the echo signal features based on the reference signal features of each terminal obtained by feature extraction and the microphone signal features of any terminal includes:

[0011] The reference signal features of each terminal are fused to obtain reference signal fusion features;

[0012] Based on the reference signal fusion characteristics and the microphone signal characteristics of any of the terminals, the echo signal characteristics are determined.

[0013] An echo cancellation method provided by the present invention further includes:

[0014] Based on the status of each terminal and / or the voice detection results of each terminal, determine the participation status of each participant corresponding to each terminal;

[0015] Based on the participation status of each participant corresponding to each terminal, the sound pickup control is performed on each terminal;

[0016] The state can be either handheld or placed, and the meeting state can be any one of discussion, presentation, or listening.

[0017] According to an echo cancellation method provided by the present invention, determining the participation status of each participant corresponding to each terminal based on the status of each terminal and / or the voice detection results of each terminal includes:

[0018] If all terminals are in the placement state, then the participation state of each participant corresponding to each terminal is determined to be in the discussion state.

[0019] If any terminal is in a handheld state and the voice detection result of any terminal indicates that any terminal has detected sound, then the participant corresponding to any terminal is determined to be in a speaking state, and the other participants are determined to be in a listening state.

[0020] According to an echo cancellation method provided by the present invention, the step of controlling the sound pickup of each terminal based on the participation status of each participant corresponding to each terminal includes:

[0021] When the participant status of any terminal is in the presentation state, obtain the echo cancellation signal of the terminal.

[0022] Based on the voiceprint features of each participant, speech separation is performed on the echo cancellation signal of any terminal to obtain the speech separation signal of any terminal. The voiceprint features of each participant are obtained by voiceprint extraction based on the historical echo cancellation signals of each terminal.

[0023] According to an echo cancellation method provided by the present invention, the step of controlling the sound pickup of each terminal based on the participation status of each participant corresponding to each terminal includes:

[0024] When all participants are in discussion mode, the microphone signals of each terminal are acquired.

[0025] The microphone signals of each terminal are audio aligned, and the target terminal is determined based on the energy of each terminal corresponding to the audio aligned microphone signals. The target terminal is then used as the microphone terminal for the current speaker among the participants in the discussion state.

[0026] The present invention also provides an echo cancellation device, comprising:

[0027] The signal acquisition unit is used to acquire the reference signal of each terminal and the microphone signal of any one of the terminals.

[0028] The echo determination unit is used to extract features from the reference signals of each terminal and the microphone signal of any terminal, and to determine the echo signal features based on the features of the reference signals of each terminal and the microphone signal features of any terminal obtained from the feature extraction.

[0029] An echo cancellation unit is used to perform echo cancellation on the microphone signal of any terminal based on the echo signal characteristics, so as to obtain an echo cancellation signal of any terminal.

[0030] The present invention also provides a conference system, including terminals and an echo cancellation device;

[0031] The echo cancellation device is used to determine echo signal characteristics based on the reference signal characteristics of the reference signals of each terminal and the microphone signal characteristics of the microphone signal of any terminal, and to perform echo cancellation on the microphone signal of any terminal based on the echo signal characteristics to obtain the echo cancellation signal of any terminal.

[0032] According to a conference system provided by the present invention, the echo cancellation device is further configured to determine the participation status of each participant corresponding to each terminal based on the status of each terminal and / or the voice detection results of each terminal, and to perform sound pickup control on each terminal based on the participation status of each participant corresponding to each terminal;

[0033] The state can be either handheld or placed, and the meeting state can be any one of discussion, presentation, or listening.

[0034] According to a conference system provided by the present invention, the echo cancellation device is specifically used to acquire the echo cancellation signal of any terminal when the participant's participation status is in the presentation state, and to perform speech separation on the echo cancellation signal of any terminal based on the voiceprint characteristics of each participant to obtain the speech separation signal of any terminal.

[0035] The voiceprint features of each participant are obtained by voiceprint extraction based on the historical echo cancellation signals of each terminal.

[0036] According to a conference system provided by the present invention, the echo cancellation device is specifically used to acquire the microphone signals of each terminal when all participants are in the discussion state, perform audio alignment on the microphone signals of each terminal, determine the target terminal based on the energy of each terminal corresponding to the audio-aligned microphone signals, and use the target terminal as the pickup terminal of the current speaker among the participants in the discussion state.

[0037] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the echo cancellation method as described above.

[0038] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the echo cancellation method as described above.

[0039] The echo cancellation method, apparatus, conferencing system, electronic device, and storage medium provided by this invention can simultaneously acquire reference signals from each terminal, and use the reference signals from each terminal and the microphone signal of any terminal to perform time delay estimation to obtain the channel propagation parameters of the reference signals of each terminal, thereby obtaining echo signal characteristics. Based on these echo signal characteristics, echo cancellation is performed on the microphone signal of the terminal to obtain the echo-cancelled signal of the terminal. This overcomes the shortcomings of traditional echo cancellation methods that cannot perform echo cancellation in multi-person conferencing scenarios. At the same time, it realizes automatic management of terminal audio signal acquisition and playback, avoids the inconvenience of manual control, and improves the stability of the conferencing process. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0041] Figure 1 This is a flowchart illustrating the echo cancellation method provided by the present invention;

[0042] Figure 2 This is a schematic diagram of the process for determining echo signal characteristics provided by the present invention;

[0043] Figure 3 This is a general framework diagram of the echo cancellation process provided by the present invention;

[0044] Figure 4 This is a flowchart illustrating the sound pickup control process provided by the present invention;

[0045] Figure 5 This is a schematic diagram of the speech separation process in the sound pickup control provided by the present invention;

[0046] Figure 6 This is a general framework diagram of the speech separation process provided by the present invention;

[0047] Figure 7 This is a schematic diagram of the microphone selection process during the sound pickup control provided by the present invention;

[0048] Figure 8 This is a schematic diagram of the echo cancellation device provided by the present invention;

[0049] Figure 9 This is a schematic diagram of the conference system provided by the present invention;

[0050] Figure 10 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0052] Multi-person meetings, a frequent activity in daily office work, help people solve problems quickly and efficiently. During multi-person meetings, the acquisition, processing, and playback of audio signals are key technologies. The entire process involves acquiring audio signals from participants in different locations, processing noise in the audio, and eliminating echoes between devices.

[0053] Currently, the main audio acquisition and processing solutions for multi-person conferencing scenarios are as follows:

[0054] Firstly, there's the gooseneck microphone solution. This solution is mostly used in large standard conferences and has poor versatility. It involves deploying multiple gooseneck microphones to capture audio signals and supports dynamic switching. The audio signals captured by the gooseneck microphones are centrally connected to a sound card for processing. Existing gooseneck microphone deployment solutions are commonly found in large professional conference rooms.

[0055] Secondly, the single-terminal solution is mostly used in conference discussion scenarios where participants share a single microphone array to join the meeting. However, in offline multi-person conference scenarios, this solution will result in poor sound quality at a distance.

[0056] Thirdly, the multi-distributed terminal solution: For the convenience of work collaboration, participants will use their own devices (terminals) to access the meeting. This results in multiple terminals in the same meeting room simultaneously collecting and playing audio signals. In this case, in order to avoid echo interference, participants need to manually control the terminals, that is, manually turn the microphone and speaker on and off, which is very inconvenient.

[0057] Furthermore, in multi-person conference scenarios, there are often multiple speakers that are turned on. The audio signals played by each speaker will propagate to the microphone after transmission and spatial reflection. Due to differences in distance, network, hardware, etc., the time delay of the audio signals played by each speaker is also different. Since current echo cancellation methods cannot simultaneously estimate the time delay of audio signals played by multiple sound sources, they cannot effectively cancel echoes. In other words, current echo cancellation methods cannot be used for echo cancellation in multi-person conference scenarios.

[0058] To address this issue, this invention provides an echo cancellation method, aiming to propose a multi-source echo cancellation method applicable to multi-person conference scenarios. This method simultaneously acquires reference signals and estimates time delays from multiple terminals, and then performs echo cancellation based on these estimates. This overcomes the limitation of traditional echo cancellation methods that cannot be applied to multi-person conference scenarios. Furthermore, it achieves automatic management of terminal audio signal acquisition and playback, avoiding the inconvenience of manual control. Figure 1 This is a flowchart illustrating the echo cancellation method provided by the present invention, as shown below. Figure 1 As shown, the execution subject of this method is a conference system, and the method includes:

[0059] Step 110: Obtain the reference signal of each terminal and the microphone signal of any terminal among them.

[0060] Here, the terminal refers to the device used by the participants to access the meeting, which can be a smartphone, tablet, etc. This embodiment of the invention does not specifically limit this. For the convenience of work collaboration, participants can use their own devices, i.e., terminals, to access the meeting.

[0061] The microphone signal is the audio signal picked up by the microphone, which can be captured by the microphone of any of the terminals connected to the conference. However, in multi-person conference scenarios, the microphone signal contains multiple echo signals. Therefore, multiple reference signals are needed for echo cancellation. These reference signals can be understood as the source signals that need to be eliminated through echo cancellation; they are also audio signals, specifically the speaker signals from each terminal connected to the conference. In other words, the conference system collects the speaker signals from each terminal as their reference signals. Taking hands-free calling on a mobile phone as an example, the microphone signal is the audio signal picked up by the phone's microphone, and the reference signal is the audio signal output by the phone's speaker.

[0062] Step 120: Extract features from the reference signal and microphone signal of each terminal respectively, and determine the echo signal features based on the extracted features of the reference signal and microphone signal of each terminal.

[0063] Specifically, in this embodiment of the invention, after acquiring the reference signals of each terminal and the microphone signal of any terminal, the conferencing system can estimate multiple channel propagation parameters to achieve echo cancellation for multiple sound sources. The specific process includes the following steps:

[0064] In conference systems, audio signals output from the terminal speakers often experience echoes as they travel through multiple feedback loops to the microphone. For example, the audio signal of the first participant might be output through the speaker of the second participant's terminal, reflect spatially, and then re-enter the microphone of that terminal. This time, the audio signal input to the microphone contains not only the second participant's audio signal but also the echo signal from the first participant. Consequently, the audio signal output from the speaker of the first participant's terminal contains both of these signals. In other words, the first participant hears the second participant's voice superimposed with their own, severely impacting conference quality. In such cases, echo cancellation is necessary. Echo cancellation utilizes techniques to estimate the magnitude of the echo signal and subtract this estimate from the microphone signal to cancel it out.

[0065] Because there are multiple echo signals in the microphone signal in a multi-person conference scenario, and there is an acoustic transmission process between the reference signal of each terminal and the actual echo signal, traditional adaptive echo cancellation methods need to estimate this transmission path parameter in real time. However, due to differences in distance, network, hardware, etc., the time delay of each reference signal is also different, making it difficult to estimate the time delay for multiple sound sources at the same time, and thus unable to effectively cancel echoes.

[0066] Therefore, in this embodiment of the invention, considering that the microphone signal is composed of multiple echo signals formed by the propagation of reference signals from each terminal through the channel, and the target signal superimposed, where the target signal can be understood as the audio signal of the participant corresponding to the terminal, multi-source fusion processing can be performed on the microphone signal and the reference signal of each terminal to estimate the channel propagation parameters of the reference signal of each terminal through multi-source fusion processing, thereby obtaining the echo signal in the microphone signal. Specifically, firstly, feature extraction can be performed on the reference signal and the microphone signal of each terminal to obtain the reference signal features of each terminal and the microphone signal features of the terminal; then, the echo signal features can be solved based on the reference signal features and the microphone signal features of each terminal.

[0067] Here, the feature extraction process for the reference signal and microphone signal, as well as the process for determining the echo signal features, can be implemented through a multi-source signal fusion network. That is, the reference signal and microphone signal of each terminal can be input into the multi-source signal fusion network. The multi-source signal fusion network extracts features from the input reference signal and microphone signal of each terminal respectively. Based on the feature-extracted reference signal features of each terminal, and using the microphone signal features of the terminal, the echo signal input from the microphone of each terminal after the reference signal of each terminal has been propagated through the channel is predicted. In other words, based on the reference signal features and microphone signal features of each terminal, the echo component in the microphone signal is predicted, and finally the echo signal features of the terminal output by the multi-source signal fusion network are obtained.

[0068] It is worth noting that before inputting the reference signals and microphone signals of each terminal into the multi-source signal fusion network, the multi-source signal fusion network can be pre-trained. The training process of the multi-source signal fusion network includes the following steps: First, a large number of sample microphone signals, sample reference signals of multiple terminals, and sample echo signal features are collected. Then, the initial multi-source signal fusion network is trained based on the sample reference signals of each terminal, the sample microphone signals of any terminal, and the sample echo signal features, thereby obtaining a multi-source signal fusion network with echo signal prediction capabilities.

[0069] In this embodiment of the invention, training the multi-source signal fusion network enables it to learn the implicit relationship between the features of the sample reference signal, the sample microphone signal, and the sample echo signal. This implicit relationship allows for the implicit calculation of the time delay of multiple sound sources and the effective prediction of the echo component in the microphone signal.

[0070] Furthermore, by simultaneously acquiring reference signals from each terminal and microphone signals from any terminal, and using these as inputs to the multi-source signal fusion network, the multi-source signal fusion network can predict the echo signal of the microphone that arrives at the terminal after propagation through the channel, based on the microphone signals and reference signals, during the multi-source fusion process.

[0071] Step 130: Based on the echo signal characteristics, perform echo cancellation on the microphone signal of the terminal to obtain the echo cancellation signal of the terminal.

[0072] In this embodiment of the invention, after multi-source fusion processing to obtain echo signal characteristics, the conference system can use these echo signal characteristics to cancel the echo of the microphone signal of the terminal, thereby canceling out multiple echo signals present in the microphone signal and ensuring conference quality.

[0073] Specifically, in step 130, the conference system can use the echo signal features as a basis to eliminate multiple echo signals present in the microphone signal of the terminal, thereby obtaining an echo cancellation signal. The specific process can be to input the echo signal features and the microphone signal of the terminal into a deep neural network. Here, the deep neural network can be understood as a pre-trained echo cancellation network for echo cancellation, so that the echo cancellation network can predict the echo components in the microphone signal of the terminal based on the echo signal features, and mask them, ultimately obtaining the audio mask of the microphone signal of the terminal output by the echo cancellation network.

[0074] Then, this audio mask can be used to eliminate the echo component in the microphone signal of the terminal. That is, the microphone signal of the terminal can be multiplied by the audio mask output by the echo cancellation network. In this way, multiple echo signals present in the microphone signal of the terminal can be canceled. It can also be understood that the echo signal in the microphone signal of the terminal can be set to zero, and the audio signal of the corresponding participant of the terminal, i.e., the target signal, can be preserved, thus obtaining the echo cancellation signal.

[0075] It should be noted that before inputting the echo signal features and microphone signals into the echo cancellation network, the echo cancellation network can be pre-trained. It is important to note that this network is trained based on the criterion of obtaining the optimal echo cancellation signal, that is, with the goal of minimizing the echo component in the echo cancellation signal. The training process includes the following steps: First, a large number of sample microphone signals, sample reference signals from multiple terminals, and sample echo cancellation signals are collected. Based on the sample reference signals of each terminal and the sample microphone signal of any terminal, the sample echo signal features are determined. Then, the initial network is trained based on the sample echo signal features, sample microphone signals, and sample echo cancellation signals, thereby obtaining the trained echo cancellation network.

[0076] Here, the initial network used for training can be built on top of Long Short-Term Memory (LSTM), Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), etc., and this embodiment of the invention does not specifically limit it.

[0077] The echo cancellation method provided by this invention can simultaneously acquire reference signals from each terminal, and use the reference signals from each terminal and the microphone signal of any terminal to perform time delay estimation to obtain the channel propagation parameters of the reference signals of each terminal, thereby obtaining echo signal characteristics. Based on these echo signal characteristics, echo cancellation is performed on the microphone signal of the terminal to obtain the echo-cancelled signal of the terminal. This overcomes the defect of traditional echo cancellation methods that cannot perform echo cancellation in multi-person conference scenarios. At the same time, it realizes automatic management of terminal audio signal acquisition and playback, avoids the inconvenience of manual control, and improves the stability of the conference process.

[0078] Based on the above embodiments, Figure 2 This is a schematic diagram of the process for determining echo signal characteristics provided by the present invention, as shown below. Figure 2 As shown, in step 120, based on the reference signal features of each terminal obtained from feature extraction and the microphone signal features of the terminal, the echo signal features are determined, including:

[0079] Step 210: Fuse the reference signal features of each terminal to obtain the reference signal fused features;

[0080] Step 220: Determine the echo signal characteristics based on the reference signal fusion characteristics and the microphone signal characteristics of the terminal.

[0081] Specifically, in step 120, after extracting features from the reference signal and the microphone signal of each terminal to obtain the reference signal features and the microphone signal features of each terminal, the process of determining the echo signal features based on the reference signal features and the microphone signal features of each terminal may include the following steps:

[0082] Step 210: First, the reference signal features of each terminal obtained by feature extraction can be fused to obtain reference signal fusion features. Specifically, feature fusion can be performed based on the reference signal features of each terminal to obtain reference signal fusion features.

[0083] Here, the fusion method for the reference signal features of each terminal can be splicing, addition, weighted fusion, etc., and the embodiments of the present invention do not specifically limit this. The fusion of the reference signal features of each terminal helps the subsequent echo prediction process, that is, it enables the multi-source signal fusion network to more accurately predict the echo components in the microphone signal, thereby facilitating the echo cancellation process for the microphone signal.

[0084] Step 220: Subsequently, the echo signal characteristics can be determined based on the reference signal fusion characteristics obtained by feature fusion and the microphone signal characteristics of the terminal. Specifically, the reference signal fusion characteristics obtained by feature fusion of the reference signal characteristics of each terminal can be used as a benchmark. The channel propagation parameters of the reference signals of each terminal can be estimated using the microphone signal characteristics of the terminal, thereby solving the echo signal characteristics of multiple echo signals in the microphone signal. In other words, based on the microphone signal characteristics of the terminal, the echo signal input from the microphone of the terminal after the reference signal of each terminal propagates through the channel is predicted according to the reference signal fusion characteristics, that is, the echo component in the microphone signal is predicted to obtain the echo signal characteristics.

[0085] Based on the above embodiments, Figure 3 This is a general framework diagram of the echo cancellation process provided by the present invention, as shown below. Figure 3 As shown, firstly, reference signals from each terminal and the microphone signal from any one of the terminals need to be obtained. Then, the reference signals from each terminal and the microphone signal from any one of the terminals can be input into a multi-source signal fusion network to obtain the echo signal features output by the multi-source signal fusion network. Subsequently, the microphone signal and echo signal features of that terminal can be input into an echo cancellation network to obtain the audio mask of the microphone signal output by the echo cancellation network. The echo cancellation network here consists of CNN layers (convolutional layers) and Bi-LSTM layers (bidirectional long short-term memory). The echo cancellation network consists of a convolutional layer (CLS) and a fully connected layer (FC). Specifically, the microphone signal of the terminal is input into the convolutional layer of the echo cancellation network to obtain the microphone signal features output by the convolutional layer. The microphone signal features output by the convolutional layer, along with the echo signal features output by the multi-source signal fusion network, are input into the bidirectional long short-time memory (LSTM) layer of the echo cancellation network to solve for the audio mask of the microphone signal. The audio mask of the microphone signal is then output through the fully connected layer. After that, the audio mask of the microphone signal output by the fully connected layer is multiplied with the microphone signal of the terminal to obtain the echo cancellation signal.

[0086] In this embodiment of the invention, reference signal acquisition, time delay estimation, and echo cancellation can be performed simultaneously for multiple sound sources, overcoming the shortcomings of traditional schemes that cannot simultaneously estimate the time delay of reference signals from multiple sound sources, thus failing to effectively cancel echoes.

[0087] Based on the above embodiments, Figure 4This is a flowchart illustrating the sound pickup control process provided by the present invention, as shown below. Figure 4 As shown, the method also includes:

[0088] Step 410: Based on the status of each terminal and / or the voice detection results of each terminal, determine the participation status of each participant corresponding to each terminal.

[0089] Step 420: Based on the participation status of each participant corresponding to each terminal, perform sound pickup control on each terminal; the status is either handheld or placed, and the participation status is any one of discussion, presentation, or listening.

[0090] Specifically, in addition to echo cancellation of the microphone signals from each terminal, the conference system can also determine the participation status of each participant based on the status of each terminal, and control the audio pickup of each terminal accordingly. The specific process includes the following steps:

[0091] Due to the portability of terminals and the randomness of speaking during multi-person conferences, the participation status of participants can be adjusted by changes in terminal status and / or the voice detection results of terminals. Specifically, step 110 can be executed first to determine the participation status of each participant corresponding to each terminal based on the status of each terminal and / or the voice detection results of each terminal. Here, the terminal status is either a placed state or a held state, i.e., whether the terminal is placed on the table or held in the hand of the corresponding participant. The participation status corresponding to the terminal status can be any one of discussion state, presentation state, and listening state, while the voice detection results are used to indicate whether the corresponding terminal has detected sound.

[0092] Then, step 420 is executed. Based on the participation status of the participants corresponding to each terminal, the sound pickup control of each terminal can be performed. Specifically, the microphone movement during the sound pickup process of each terminal can be adjusted based on the participation status of the participants corresponding to each terminal, and / or, the echo-cancelled signal picked up by each terminal after echo cancellation can be separated into speech. That is, the microphone movement selection during the sound pickup process can be controlled based on the participation status of the participants corresponding to each terminal, and / or, the speech separation of the echo-cancelled signal obtained after echo cancellation can be controlled.

[0093] In this embodiment of the invention, the portability and flexibility of the terminal, as well as the randomness of speaking during the meeting, are fully utilized to adjust the meeting mode or the meeting status of each participant, realizing dynamic switching between different meeting modes or participation statuses, improving the convenience of communication during the meeting and optimizing the meeting quality.

[0094] Based on the above embodiments, step 410 includes:

[0095] With all terminals in the "place" state, the participation state of each participant corresponding to each terminal is determined to be "discussion".

[0096] If any terminal is in a handheld state and the voice detection result of that terminal indicates that the terminal has detected sound, then the participant corresponding to that terminal is determined to be in a speaking state, and the participants of other participants are determined to be in a listening state.

[0097] Specifically, in step 410, the process of determining the participation status of each participant corresponding to each terminal based on the status of each terminal and / or the voice detection results of each terminal can be divided into the following two cases:

[0098] Firstly, when all terminals are in a placed state, meaning that each participant places their corresponding terminal on the table during the meeting, the participation status of each participant can be set to discussion state. In other words, the meeting mode at this time is a free discussion mode, and all participants are in a free discussion state. In this meeting mode, each participant speaks randomly. In other words, each participant may speak at any time in the free discussion mode. Therefore, the microphone can be controlled to select the direction during the sound pickup process.

[0099] Secondly, in addition to the discussion state, this embodiment of the invention also considers the flexibility of the terminal and determines a speaking state and a corresponding listening state. That is, when any terminal is in a handheld state and the voice detection result of the terminal shows that the terminal has detected sound, in other words, when the terminal is held in the hand of the corresponding participant and the participant is speaking properly, it can be determined that the participant is in a speaking state, that is, the participant's participation state is a speaking state, while the participation states of other participants are in a listening state.

[0100] It should be noted that when any participant is giving a presentation, only the microphone of that participant's terminal is turned on, while the microphones of other terminals are turned off. This is to ensure the participant's focus during the presentation and to avoid interference from microphone signals input from other terminals.

[0101] Based on the above embodiments, Figure 5 This is a schematic diagram of the speech separation process in the sound pickup control provided by the present invention, as shown below. Figure 5 As shown, step 420 includes:

[0102] Step 421-A: When the participant status of any terminal is in presentation mode, obtain the echo cancellation signal of that terminal.

[0103] Step 422-A: Based on the voiceprint features of each participant, perform speech separation on the echo cancellation signal of the terminal to obtain the speech separation signal of the terminal. The voiceprint features of each participant are obtained by voiceprint extraction based on the historical echo cancellation signals of each terminal.

[0104] Specifically, in step 420, during the process of controlling the sound pickup of each terminal according to the participation status of each participant, the process of performing speech separation on the echo-cancelled signal obtained after echo cancellation can include the following steps:

[0105] Step 421-A: When any participant is in a speaking state, that is, when the participant corresponding to any terminal is in a speaking state and the participants corresponding to other terminals are in a listening state, the echo cancellation signal of that terminal can be obtained; the echo cancellation signal here is the echo cancellation signal obtained after the microphone signal picked up by the terminal has undergone the echo cancellation process described above. The echo cancellation process has been described in detail above and will not be repeated here.

[0106] Step 422-A: After obtaining the echo cancellation signal of the terminal, before performing speech separation on the echo cancellation signal of the terminal, it is necessary to determine the voiceprint features of each participant. These voiceprint features can be determined based on the historical echo cancellation signals of each terminal. That is, voiceprint extraction can be performed on the historical echo cancellation signals of each terminal to extract the voiceprint features of each participant. Then, based on the voiceprint features of each participant, speech separation can be performed on the echo cancellation signal of the terminal to obtain the speech separation signal.

[0107] Here, speech separation for echo cancellation signals can be achieved through a speech separation model. That is, the voiceprint features of each participant and the echo cancellation signal of the terminal can be input into the speech separation model. The speech separation model performs speech separation on the echo cancellation signal of the terminal based on the input voiceprint features of each participant, and finally obtains the audio mask of the echo cancellation signal output by the speech separation model. Multiplying this audio mask by the echo cancellation signal yields the speech separation signal.

[0108] In this embodiment of the invention, voice separation is performed on the echo cancellation signal of the terminal, which can separate the audio signal of the participants corresponding to the terminal from the echo cancellation signal. This ensures that the audio signal is not interfered with by background noise and the audio signals of other participants when played through a speaker, thus guaranteeing focus during the speech and clarity of the audio signal during the speech.

[0109] Based on the above embodiments, Figure 6 This is a general framework diagram of the speech separation process provided by the present invention, as shown below. Figure 6As shown, when the participant status of any terminal is in the presentation state, the first step is to obtain the echo cancellation signal of that terminal. Then, the voiceprint features of each participant and the echo cancellation signal of that terminal can be input into the speech separation model to obtain the audio mask of the echo cancellation signal output by the speech separation model. This speech separation model consists of a CNN layer (convolutional layer), a Bi-LSTM layer (bidirectional long short-term memory layer), and a fully connected layer (FC layer). Specifically, the echo cancellation signal of the terminal can be input into the convolutional layer of the speech separation model to obtain the echo cancellation signal features output by the convolutional layer. The echo cancellation signal features output by the convolutional layer, along with the voiceprint features of each participant, can then be input into the bidirectional long short-term memory layer of the speech separation model to solve for the audio mask of the echo cancellation signal. This mask is then output through the fully connected layer. Afterward, the audio mask of the echo cancellation signal output by the fully connected layer can be multiplied with the echo cancellation signal of the terminal to obtain the speech separation signal.

[0110] The voiceprint features of each participant are determined based on the historical echo cancellation signals of each terminal. In other words, the voiceprint can be extracted based on the historical echo cancellation signals of each terminal to obtain the voiceprint features of each participant corresponding to each terminal. The historical echo cancellation signal can be the longest echo cancellation signal picked up by the corresponding terminal in the past after echo cancellation, so as to ensure the accuracy of the extracted voiceprint features.

[0111] In this embodiment of the invention, under the guidance of the voiceprint characteristics of each participant, voice separation is performed on the echo cancellation signal. This allows the audio signal of the participant corresponding to the terminal to be extracted and retained, thereby ensuring that the participant's presentation is not interfered with by background noise or the audio signals of other participants. In other words, when the meeting mode is presentation mode, the interference of background noise and other voices can be suppressed to ensure focus during the presentation. Furthermore, the speaker only plays the extracted audio signal of the participant, which ensures the clarity of the audio signal in the presentation mode, thereby improving the meeting quality.

[0112] Based on the above embodiments, Figure 7 This is a schematic diagram of the microphone selection process during the sound pickup control provided by the present invention, as shown below. Figure 7 As shown, step 420 includes:

[0113] Step 421-B: With all participants in the discussion state, obtain the microphone signals from each terminal.

[0114] Step 422-B: Perform audio alignment on the microphone signals of each terminal, and determine the target terminal based on the energy of each terminal corresponding to the audio-aligned microphone signals, and use the target terminal as the microphone terminal for the current speaker among the participants in the discussion state.

[0115] Specifically, in step 420, during the process of controlling the sound pickup of each terminal based on the participation status of each participant, the process of controlling the microphone's movement selection during sound pickup may further include the following steps:

[0116] Step 421-B: When all participants are in discussion mode, that is, when the participation mode of each participant corresponding to each terminal is discussion mode, the microphone signal of each terminal can be obtained; the microphone signal here is the audio signal picked up by the microphone on each terminal connected to the conference.

[0117] Step 422-B: Since multiple microphones are often used to pick up sound during multi-person conferences, and the positions of different microphones are different (in other words, the distance between different microphones and the speakers is different), there is a transmission delay in the microphone signals picked up by the microphones. In this embodiment of the invention, audio alignment can be performed on the microphone signals of each terminal. That is, an autocorrelation function can be used to align the audio signals of the microphone signals of each terminal to eliminate the transmission delay. This is a preprocessing process, which is a prerequisite for the next step (energy level). That is, only after audio alignment is performed to ensure that the microphone signals are synchronized in time can energy judgment be performed. In other words, this process provides crucial assistance for the subsequent selection of microphone movement.

[0118] Then, the energy of each terminal corresponding to each microphone signal after audio alignment can be determined, and the target terminal can be determined based on the energy of each terminal, and the target terminal can be used as the sound pickup terminal of the current speaker. Specifically, the terminal with the highest microphone energy can be determined as the target terminal based on the energy of the microphones in each terminal, and then the target terminal can be used as the sound pickup terminal of the current speaker among the participants in the discussion state, that is, the target terminal can be used as the terminal to pick up the audio signal of the current speaker.

[0119] It is worth mentioning that the target terminal selected by energy here must be the terminal closest to the current speaker. This can ensure the sound pickup quality during the sound pickup process to the maximum extent. In addition, while using the target terminal as the current speaker's sound pickup terminal, the microphones of other terminals will also be turned off. This ensures that the terminal closest to the current speaker is always the sound pickup terminal, and different speakers can be assigned to different sound pickup terminals, realizing the dynamic selection of sound pickup terminals or microphones.

[0120] Unlike traditional meeting scenarios where all microphones are turned on and used for sound pickup after each terminal connects, making it impossible to select microphone direction and guarantee the sound quality during the meeting, this invention integrates all terminals connected to the meeting into a single unit and proposes a microphone direction selection scheme. This scheme can automatically select the microphone with the highest gain and turn off the other microphones, thereby achieving automated management of terminals after they join the meeting.

[0121] Based on the above embodiments, the overall process of the echo cancellation method provided by the present invention includes the following steps: The conference system first publishes a conference APP (Application), through which functions such as remote conference creation and conference access can be realized.

[0122] After each terminal joins the conference, the conference system first needs to acquire the reference signal of each terminal, as well as the microphone signal of any terminal. Then, it can extract features from the reference signal and microphone signal of each terminal to obtain the reference signal features and microphone signal features of each terminal. The reference signal features of each terminal can be fused to obtain the reference signal fused features. Based on the reference signal fused features and the microphone signal features of the terminal, the echo signal features are determined. Subsequently, based on the echo signal features, echo cancellation can be performed on the microphone signal of the terminal to obtain the echo cancellation signal of the terminal.

[0123] In addition, the conferencing system can determine the participation status of each participant corresponding to each terminal based on the status of each terminal and / or the voice detection results of each terminal; and control the sound pickup of each terminal based on the participation status of each participant corresponding to each terminal; here, the status is either handheld or placed, and the participation status is any one of discussion, presentation, or listening.

[0124] Furthermore, the process of determining the participation status of each participant corresponding to each terminal based on the status of each terminal and / or the voice detection results of each terminal may specifically include: when all terminals are in a placed state, determining that the participation status of each participant corresponding to each terminal is in a discussion state; when any terminal is in a handheld state and the voice detection result of that terminal indicates that the terminal has detected sound, determining that the participation status of the participant corresponding to that terminal is in a presentation state, and the participation status of other participants is in a listening state.

[0125] The process of controlling the sound pickup of each terminal based on the participation status of each participant can include the following steps: When the participation status of a participant corresponding to any terminal is in a presentation state, acquire the echo cancellation signal of that terminal; based on the voiceprint features of each participant, perform speech separation on the echo cancellation signal of that terminal to obtain the speech separation signal of that terminal; the voiceprint features of each participant are extracted based on the historical echo cancellation signals of each terminal. Correspondingly, when the participation status of each participant is in a discussion state, acquire the microphone signal of each terminal; perform audio alignment on the microphone signals of each terminal, and based on the energy of each terminal corresponding to the audio-aligned microphone signals, determine the target terminal, and use the target terminal as the sound pickup terminal for the current speaker among the participants in the discussion state.

[0126] In this embodiment of the invention, the conference system can access multiple terminals through a conference APP. Each terminal accessing the conference is a microphone terminal for the corresponding participant. At the same time, the conference system can adjust the participant status based on the terminal status and / or voice detection results. Furthermore, when a participant corresponding to any terminal is in a presentation state, the system can also perform voice separation on the echo cancellation signal of that terminal to suppress background noise and other voice interference during the presentation state, which greatly improves the conference system's ability to handle complex scenarios.

[0127] Furthermore, in the multi-distributed terminal solution, this embodiment of the invention can fully utilize the networking and linkage of each terminal to meet the needs of intelligent interaction among the terminals, leverage the personalized characteristics of the terminals, enhance the microphone's pickup gain, and realize the microphone's dynamic selection.

[0128] The method provided in this invention can simultaneously acquire reference signals from each terminal, and use the reference signals from each terminal and the microphone signal of any terminal to perform time delay estimation to obtain the channel propagation parameters of the reference signals of each terminal, thereby obtaining echo signal characteristics. Based on these echo signal characteristics, echo cancellation is performed on the microphone signal of the terminal to obtain the echo cancellation signal of the terminal. This overcomes the defect of traditional echo cancellation methods that cannot perform echo cancellation in multi-person conference scenarios. At the same time, it realizes automatic management of terminal audio signal acquisition and playback, avoids the inconvenience of manual control, and improves the stability of the conference process.

[0129] The echo cancellation device provided by the present invention is described below. The echo cancellation device described below can be referred to in correspondence with the echo cancellation method described above.

[0130] Figure 8 This is a schematic diagram of the echo cancellation device provided by the present invention, as shown below. Figure 8 As shown, the device includes:

[0131] The signal acquisition unit 810 is used to acquire the reference signal of each terminal and the microphone signal of any one of the terminals.

[0132] The echo determination unit 820 is used to extract features from the reference signal and the microphone signal of each terminal, and to determine the echo signal features based on the extracted features of the reference signal and the microphone signal of each terminal.

[0133] The echo cancellation unit 830 is used to cancel the echo of the microphone signal of the terminal based on the echo signal characteristics to obtain the echo cancellation signal of the terminal.

[0134] The echo cancellation device provided by this invention can simultaneously acquire reference signals from each terminal, and use the reference signals from each terminal and the microphone signal of any terminal to perform time delay estimation to obtain the channel propagation parameters of the reference signals of each terminal, thereby obtaining echo signal characteristics. Based on these echo signal characteristics, echo cancellation is performed on the microphone signal of the terminal to obtain the echo-cancelled signal of the terminal. This overcomes the shortcomings of traditional echo cancellation methods that cannot perform echo cancellation in multi-person conference scenarios. At the same time, it realizes automatic management of terminal audio signal acquisition and playback, avoids the inconvenience of manual control, and improves the stability of the conference process.

[0135] Based on the above embodiments, the echo determination unit 820 is used for:

[0136] The reference signal features of each terminal are fused to obtain reference signal fusion features;

[0137] Based on the reference signal fusion characteristics and the microphone signal characteristics of the terminal, the echo signal characteristics are determined.

[0138] Based on the above embodiments, the device further includes a sound pickup control unit, used for:

[0139] Based on the status of each terminal and / or the voice detection results of each terminal, determine the participation status of each participant corresponding to each terminal;

[0140] Based on the participation status of each participant corresponding to each terminal, the sound pickup control is performed on each terminal;

[0141] The state can be either handheld or placed, and the meeting state can be any one of discussion, presentation, or listening.

[0142] Based on the above embodiments, the pickup control unit is used for:

[0143] If all terminals are in the placement state, then the participation state of each participant corresponding to each terminal is determined to be in the discussion state.

[0144] If any terminal is in a handheld state and the voice detection result of that terminal indicates that the terminal has detected sound, then the participant corresponding to that terminal is determined to be in a speaking state, and the participants of other participants are determined to be in a listening state.

[0145] Based on the above embodiments, the pickup control unit is used for:

[0146] If the participant status of any terminal is in the presentation state, obtain the echo cancellation signal of that terminal.

[0147] Based on the voiceprint features of each participant, speech separation is performed on the echo cancellation signal of the terminal to obtain the speech separation signal of the terminal. The voiceprint features of each participant are obtained by voiceprint extraction based on the historical echo cancellation signals of each terminal.

[0148] Based on the above embodiments, the pickup control unit is used for:

[0149] When all participants are in discussion mode, the microphone signals of each terminal are acquired.

[0150] The microphone signals of each terminal are audio aligned, and the target terminal is determined based on the energy of each terminal corresponding to the audio aligned microphone signals. The target terminal is then used as the microphone terminal for the current speaker among the participants in the discussion state.

[0151] Figure 9 This is a schematic diagram of the conference system provided by the present invention, as shown below. Figure 9 As shown, the conference system includes terminals 910 and an echo cancellation device 800;

[0152] The echo cancellation device 800 is used to determine the echo signal characteristics based on the reference signal characteristics of the reference signal of each terminal 910 and the microphone signal characteristics of the microphone signal of any terminal among the terminals 910, and to perform echo cancellation on the microphone signal of the terminal based on the echo signal characteristics to obtain the echo cancellation signal of the terminal.

[0153] Specifically, in this embodiment of the invention, the conference system includes terminals 910 and an echo cancellation device 800. Here, the terminal refers to the device used by the participants to access the conference, which can be a smartphone, tablet, etc. For the convenience of work collaboration, participants can connect their own devices, i.e., terminals, to the conference. The microphone signal is the audio signal picked up by the microphone. Since there are multiple echo signals in the microphone signal in a multi-person conference scenario, multiple reference signals are required in the process of echo cancellation of the microphone signal. The reference signals are also audio signals, which can be understood as the source signals that need to be canceled by echo cancellation, and they are the speaker signals of each terminal accessing the conference.

[0154] In conference systems, audio signals output from the terminal speakers often experience echoes as they travel through multiple feedback loops to the microphone. For example, the audio signal of the first participant might be output through the speaker of the second participant's terminal, reflect spatially, and then re-enter the microphone of that terminal. This time, the audio signal input to the microphone contains not only the second participant's audio signal but also the echo signal from the first participant. Consequently, the audio signal output from the speaker of the first participant's terminal contains both of these signals. In other words, the first participant hears the second participant's voice superimposed with their own, severely impacting conference quality. In such cases, echo cancellation is necessary. Echo cancellation utilizes techniques to estimate the magnitude of the echo signal and subtract this estimate from the microphone signal to cancel it out.

[0155] Because there are multiple echo signals in the microphone signal in a multi-person conference scenario, and there is an acoustic transmission process between the reference signal of each terminal and the actual echo signal, traditional adaptive echo cancellation methods need to estimate this transmission path parameter in real time. However, due to differences in distance, network, hardware, etc., the time delay of each reference signal is also different, making it difficult to estimate the time delay for multiple sound sources at the same time, and thus unable to effectively cancel echoes.

[0156] Therefore, in this embodiment of the invention, the echo cancellation device 800 first acquires the reference signals of each terminal 910 (terminal 1, terminal 2, terminal 3, etc.) in the conference system, as well as the microphone signal of any terminal among the terminals 910. Then, considering that the microphone signal is formed by multiple echo signals generated by the propagation of the reference signals of each terminal 910 through the channel, and the target signal is superimposed, the target signal here can be understood as the audio signal of the participant corresponding to the terminal. Therefore, the echo cancellation device 800 can perform multi-source fusion processing on the microphone signal and the reference signal of each terminal 910 to estimate the channel propagation parameters of the reference signal of each terminal 910, thereby obtaining the echo signal in the microphone signal. That is, the echo cancellation device 800 can perform feature extraction on the reference signal and the microphone signal of each terminal 910 respectively to obtain the reference signal features and the microphone signal features of each terminal 910. Then, the echo signal features can be solved based on the reference signal features and the microphone signal features of each terminal 910.

[0157] After this, the echo cancellation device 800 can use the echo signal characteristics to cancel the echo in the microphone signal of the terminal, thereby canceling multiple echo signals in the microphone signal and ensuring the meeting quality. Specifically, the echo cancellation device 800 can use the echo signal characteristics as a reference to predict the echo components in the microphone signal of the terminal and output a mask to obtain the audio mask of the microphone signal of the terminal. Then, it can use this audio mask to cancel the echo components in the microphone signal of the terminal, that is, multiply the microphone signal of the terminal by the audio mask. In this way, multiple echo signals in the microphone signal of the terminal can be canceled. In other words, the echo signals in the microphone signal of the terminal can be set to zero, and the audio signals of the participants corresponding to the terminal, i.e., the target signals, can be retained, and finally the echo-canceling signal is obtained.

[0158] The conference system provided by this invention includes various terminals and an echo cancellation device. The echo cancellation device can use the reference signal of each terminal and the microphone signal of any terminal to perform time delay estimation, thereby obtaining the channel propagation parameters of the reference signal of each terminal, thus obtaining the echo signal characteristics. Echo cancellation can be performed based on these echo signal characteristics to obtain the echo cancellation signal of the terminal. This overcomes the shortcomings of traditional solutions that cannot perform echo cancellation in multi-person conference scenarios. At the same time, it realizes the automatic management of terminal audio signal acquisition and playback, avoids the inconvenience of manual control, and improves the stability of the conference system.

[0159] Based on the above embodiments, the echo cancellation device 80 is also used to determine the participation status of each participant corresponding to each terminal 910 based on the status of each terminal 910 and / or the voice detection results of each terminal 910, and to perform sound pickup control on each terminal 910 based on the participation status of each participant corresponding to each terminal 910.

[0160] The status can be either held or placed, and the meeting status can be any one of discussion, presentation, or listening.

[0161] Specifically, in addition to canceling the echo of the microphone signals from each terminal 910, the echo cancellation device 800 in the conference system can also be used to control the sound pickup of each terminal 910. Specifically, due to the portability of the terminals and the randomness of speaking during multi-person conferences, the echo cancellation device 800 can adjust the participation status of participants based on changes in the terminal status and / or the voice detection results of the terminals. That is, the echo cancellation device 800 can adjust the participation status of each terminal 910 based on its status and / or the voice pickup results of each terminal 910. Based on the sound detection results, the participation status of each participant corresponding to each terminal 910 is determined. Based on the participation status of the participants corresponding to each terminal 910, sound pickup control can be performed on each terminal 910. The sound pickup control here is actually the adjustment of the microphone movement during the sound pickup process of each terminal 910, and / or the speech separation of the echo cancellation signal of each terminal 910. That is, based on the participation status of the participants corresponding to each terminal 910, the microphone movement selection during the sound pickup process can be controlled, and / or the speech separation of the echo cancellation signal obtained after echo cancellation can be controlled.

[0162] Here, the terminal's state is either placed or held, meaning whether the terminal is placed on a table or held in the hand of the corresponding participant. The participant state corresponding to the terminal's state can be any one of discussion, presentation, or listening. The voice detection result is used to indicate whether the corresponding terminal has detected sound.

[0163] In this embodiment of the invention, the echo cancellation device in the conference system makes full use of the portability and flexibility of the terminal, as well as the randomness of the speeches during the conference, to adjust the conference mode or the conference status of each participant, thereby realizing the dynamic switching between different conference modes or participant statuses in the conference system, improving the flexibility and ease of use of the conference system.

[0164] Based on the above embodiments, the echo cancellation device 800 is specifically used to acquire the echo cancellation signal of the terminal when the participant status of any terminal is in the presentation state, and to perform speech separation on the echo cancellation signal of the terminal based on the voiceprint characteristics of each participant to obtain the speech separation signal of the terminal.

[0165] The voiceprint features of each participant were extracted based on the historical echo cancellation signals of each terminal 910.

[0166] Specifically, when the echo cancellation device 800 in the conference system performs sound pickup control on each terminal 910, the voice separation process for the echo cancellation signal can be as follows: when any participant (the participant corresponding to any terminal) is in a speaking state and other participants are in a listening state, the echo cancellation device 800 can acquire the echo cancellation signal of that terminal and perform voice separation on the echo cancellation signal of that terminal to obtain a voice separation signal. The voice separation process here is based on the voiceprint characteristics of each participant, and the voiceprint characteristics of each participant are determined based on the historical echo cancellation signals of each terminal 910. In short, the echo cancellation device 800 can perform voice separation on the echo cancellation signal of the terminal based on the voiceprint characteristics of each participant to obtain a voice separation signal.

[0167] The echo cancellation signal here is the echo cancellation signal obtained after the microphone signal picked up by the terminal has undergone echo cancellation by the echo cancellation device 800. The echo cancellation process has been described in detail above and will not be repeated here.

[0168] In this embodiment of the invention, the voice separation function of the echo cancellation device can separate the audio signal of the participant corresponding to the terminal from the echo cancellation signal, so that when the audio signal is played through the speaker, it is not disturbed by background noise and the audio signals of other participants, thus ensuring the focus of the speech process and the clarity of the audio signal during the speech.

[0169] Based on the above embodiments, the echo cancellation device 800 is specifically used to acquire the microphone signals of each terminal 910 when all participants are in the discussion state, perform audio alignment on the microphone signals of each terminal 910, determine the target terminal based on the energy of each terminal 910 corresponding to the audio-aligned microphone signals, and use the target terminal as the sound pickup terminal of the current speaker among the participants in the discussion state.

[0170] Specifically, when the echo cancellation device 800 in the conference system controls the sound pickup of each terminal 910, it selects the microphone movement during the sound pickup process. Specifically, since multiple microphones are often used for sound pickup in multi-person conferences, and different microphones are in different positions (i.e., different microphones are at different distances from the speakers), the microphone signals picked up by the microphones experience transmission delay. Therefore, the echo cancellation device 800 in this embodiment can acquire the microphone signals of each terminal 910 while all participants are in a discussion state, and perform audio alignment on the microphone signals of each terminal 910. This is done by using an autocorrelation function to eliminate transmission delay. This is a preprocessing step, a prerequisite for the next step (determining energy level). Only after audio alignment, ensuring that the microphone signals are synchronized in time, can energy determination be performed. In other words, this process provides crucial assistance for subsequent microphone movement selection.

[0171] Then, the echo cancellation device 800 can determine the energy of each terminal 910 corresponding to each microphone signal after audio alignment, and determine the target terminal based on the energy of each terminal 910, and use the target terminal as the pickup terminal of the current speaker. That is, the terminal with the largest microphone energy can be selected from each terminal 910 as the target terminal, and then the target terminal is used as the pickup terminal of the current speaker among the participants in the discussion state, that is, the target terminal is used as the terminal to pick up the audio signal of the current speaker.

[0172] It is worth mentioning that the target terminal selected by energy here must be the terminal closest to the current speaker. This can ensure the sound pickup quality during the sound pickup process to the maximum extent. In addition, while using the target terminal as the current speaker's sound pickup terminal, the microphones of other terminals will also be turned off. This ensures that the terminal closest to the current speaker is always the sound pickup terminal, and different speakers can be assigned to different sound pickup terminals, realizing the dynamic selection of sound pickup terminals or microphones.

[0173] Unlike traditional meeting scenarios where all microphones are turned on and picking up sound after each terminal connects, making it impossible to select microphones and guarantee the sound quality during the meeting, this invention integrates all terminals 910 connected to the meeting into a network for coordinated operation. It proposes a microphone selection scheme that can automatically select the microphone with the highest gain and turn off the others, thereby achieving automated management of the meeting system.

[0174] Based on the above embodiments, the conference system provided by the present invention includes various terminals 910 and an echo cancellation device 800.

[0175] The echo cancellation device 800 is used to determine the echo signal characteristics based on the reference signal characteristics of the reference signal of each terminal 910 and the microphone signal characteristics of the microphone signal of any terminal among the terminals 910, and to perform echo cancellation on the microphone signal of the terminal based on the echo signal characteristics to obtain the echo cancellation signal of the terminal.

[0176] The echo cancellation device 800 is also used to determine the participation status of each participant corresponding to each terminal 910 based on the status of each terminal 910 and / or the voice detection results of each terminal 910, and to perform sound pickup control on each terminal 910 based on the participation status of each participant corresponding to each terminal 910; the status is either handheld or placed, and the participation status is any one of discussion, presentation, or listening.

[0177] The echo cancellation device 800 is specifically used to acquire the echo cancellation signal of any terminal when the participant's participation status is in the presentation state, and to perform speech separation on the echo cancellation signal of the terminal based on the voiceprint features of each participant to obtain the speech separation signal of the terminal; the voiceprint features of each participant are obtained by voiceprint extraction based on the historical echo cancellation signals of each terminal 910.

[0178] The echo cancellation device 800 is specifically used to acquire the microphone signals of each terminal 910 when all participants are in the discussion state, perform audio alignment on the microphone signals of each terminal 910, and determine the target terminal based on the energy of each terminal 910 corresponding to the audio-aligned microphone signals, and use the target terminal as the sound pickup terminal of the current speaker among the participants in the discussion state.

[0179] In this embodiment of the invention, each terminal can access the conference system through a conference APP (which can realize functions such as remote conference creation and conference access). Each terminal accessing the conference is the corresponding audio pickup terminal for the participant. At the same time, the conference system can adjust the participant status based on the terminal status and / or voice detection results. Furthermore, when the participant corresponding to any terminal is in a presentation state, the echo cancellation signal of that terminal can be separated into voices to suppress background noise and other voice interference during the presentation state, which greatly improves the conference system's ability to handle complex scenarios.

[0180] Furthermore, in the multi-distributed terminal solution, this embodiment of the invention can fully utilize the networking and linkage of each terminal to meet the needs of intelligent interaction among the terminals, leverage the personalized characteristics of the terminals, enhance the microphone's pickup gain, and realize the microphone's dynamic selection.

[0181] Figure 10 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 10 As shown, the electronic device may include a processor 1010, a communications interface 1020, a memory 1030, and a communication bus 1040, wherein the processor 1010, the communications interface 1020, and the memory 1030 communicate with each other via the communication bus 1040. The processor 1010 can call logic instructions in the memory 1030 to execute an echo cancellation method, which includes: acquiring reference signals of each terminal and a microphone signal of any one of the terminals; performing feature extraction on the reference signals and the microphone signal of each terminal respectively, and determining echo signal features based on the feature extraction of the reference signal features and the microphone signal features of each terminal; and performing echo cancellation on the microphone signal of the terminal based on the echo signal features to obtain the echo cancellation signal of the terminal.

[0182] Furthermore, the logical instructions in the aforementioned memory 1030 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0183] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is able to execute the echo cancellation method provided by the above methods, the method comprising: acquiring reference signals of each terminal and a microphone signal of any one of the terminals; performing feature extraction on the reference signals of each terminal and the microphone signal of the terminal respectively, and determining echo signal features based on the feature extraction of the reference signal features of each terminal and the microphone signal features of the terminal; and performing echo cancellation on the microphone signal of the terminal based on the echo signal features to obtain an echo cancellation signal of the terminal.

[0184] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the echo cancellation method provided by the above methods. The method includes: acquiring reference signals of each terminal and a microphone signal of any one of the terminals; performing feature extraction on the reference signals of each terminal and the microphone signal of the terminal respectively, and determining echo signal features based on the feature extraction of the reference signal features of each terminal and the microphone signal features of the terminal; and performing echo cancellation on the microphone signal of the terminal based on the echo signal features to obtain an echo cancellation signal of the terminal.

[0185] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0186] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An echo cancellation method, characterized by, Multi-source echo cancellation is applied to multi-person conference scenarios where multiple terminals access the conference, including: Acquire reference signals from each terminal accessing the conference, as well as the microphone signal from any one of the terminals; the reference signals from each terminal are the speaker signals from each terminal. Feature extraction is performed on the reference signals of each terminal and the microphone signal of any terminal respectively, and the echo signal features are determined based on the feature extraction features of the reference signals of each terminal and the microphone signal features of any terminal; the echo signal features characterize the echo components in the microphone signal. Based on the echo signal characteristics, echo cancellation is performed on the microphone signal of any terminal to obtain the echo cancellation signal of any terminal.

2. The echo cancellation method of claim 1, wherein, The determination of echo signal features based on the reference signal features of each terminal obtained from feature extraction and the microphone signal features of any terminal includes: The reference signal features of each terminal are fused to obtain reference signal fusion features; Based on the reference signal fusion characteristics and the microphone signal characteristics of any of the terminals, the echo signal characteristics are determined.

3. The echo cancellation method of claim 1, wherein, Also includes: Based on the status of each terminal and / or the voice detection results of each terminal, determine the participation status of each participant corresponding to each terminal; Based on the participation status of each participant corresponding to each terminal, the sound pickup control is performed on each terminal; The state can be either handheld or placed, and the meeting state can be any one of discussion, presentation, or listening.

4. The echo cancellation method of claim 3, wherein, The determination of the participation status of each participant corresponding to each terminal based on the status of each terminal and / or the voice detection results of each terminal includes: If all terminals are in the placement state, then the participation state of each participant corresponding to each terminal is determined to be in the discussion state. If any terminal is in a handheld state and the voice detection result of any terminal indicates that any terminal has detected sound, then the participant corresponding to any terminal is determined to be in a speaking state, and the other participants are determined to be in a listening state.

5. The echo cancellation method of claim 3, wherein, The method of controlling the sound pickup of each terminal based on the participation status of each participant corresponding to each terminal includes: When the participant status of any terminal is in the presentation state, obtain the echo cancellation signal of the terminal. Based on the voiceprint features of each participant, speech separation is performed on the echo cancellation signal of any terminal to obtain the speech separation signal of any terminal. The voiceprint features of each participant are obtained by voiceprint extraction based on the historical echo cancellation signals of each terminal.

6. The echo cancellation method of claim 3, wherein, The method of controlling the sound pickup of each terminal based on the participation status of each participant corresponding to each terminal includes: When all participants are in discussion mode, the microphone signals of each terminal are acquired. The microphone signals of each terminal are audio aligned, and the target terminal is determined based on the energy of each terminal corresponding to the audio aligned microphone signals. The target terminal is then used as the microphone terminal for the current speaker among the participants in the discussion state.

7. An echo cancellation device, characterized by Multi-source echo cancellation is applied to multi-person conference scenarios where multiple terminals access the conference, including: The signal acquisition unit is used to acquire the reference signal of each terminal accessing the conference, as well as the microphone signal of any one of the terminals; the reference signal of each terminal is the speaker signal of each terminal. An echo determination unit is used to extract features from the reference signals of each terminal and the microphone signal of any terminal, and to determine echo signal features based on the extracted features of the reference signals of each terminal and the microphone signal features of any terminal; the echo signal features characterize the echo components in the microphone signal. An echo cancellation unit is used to perform echo cancellation on the microphone signal of any terminal based on the echo signal characteristics, so as to obtain an echo cancellation signal of any terminal.

8. A conferencing system, characterized by Multi-source echo cancellation is applied to multi-person conference scenarios where multiple terminals access the conference. The system includes each terminal accessing the conference and an echo cancellation device. The echo cancellation device is used to determine echo signal characteristics based on the reference signal characteristics of the reference signals of each terminal and the microphone signal characteristics of the microphone signal of any one of the terminals, and to perform echo cancellation on the microphone signal of any one terminal based on the echo signal characteristics to obtain the echo cancellation signal of any one terminal; the echo signal characteristics characterize the echo component in the microphone signal; the reference signal of each terminal is the speaker signal of each terminal.

9. The conference system according to claim 8, characterized in that, The echo cancellation device is also used to determine the participation status of each participant corresponding to each terminal based on the status of each terminal and / or the voice detection results of each terminal, and to perform sound pickup control on each terminal based on the participation status of each participant corresponding to each terminal. The state can be either handheld or placed, and the meeting state can be any one of discussion, presentation, or listening.

10. The conference system according to claim 9, characterized in that, The echo cancellation device is specifically used to acquire the echo cancellation signal of any terminal when the participant's participation status is in the presentation state, and to perform speech separation on the echo cancellation signal of any terminal based on the voiceprint characteristics of each participant, so as to obtain the speech separation signal of any terminal. The voiceprint features of each participant are obtained by voiceprint extraction based on the historical echo cancellation signals of each terminal.

11. The conference system according to claim 9, characterized in that, The echo cancellation device is specifically used to acquire the microphone signals of each terminal when all participants are in discussion mode, perform audio alignment on the microphone signals of each terminal, determine the target terminal based on the energy of each terminal corresponding to the audio-aligned microphone signals, and use the target terminal as the pickup terminal of the current speaker among the participants in the discussion mode.

12. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the echo cancellation method as described in any one of claims 1 to 6.

13. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the echo cancellation method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice signal processing method and device, electronic equipment and readable storage medium

    CN113823304A