The algorithm selects the model's training method, echo cancellation method, device, and equipment.

CN116665690BActive Publication Date: 2026-08-14JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0007]本发明实施例提供了一种算法选取模型的训练方法、回声消除方法、装置及设备,以解决不同的回声消除算法所适用的语音场景不同的问题,提高回声消除算法与复杂语音场景的适配性,进而提高回声消除后的语音质量

Benefits of technology

[0037]本发明实施例的技术方案,通过采用回声消除算法集中的至少两个预设回声消除算法,基于近端扬声器播放的训练远端声音数据,对近端麦克风采集的训练近端声音数据分别进行回声消除处理得到训练消除数据集,并将训练远端声音数据和训练消除数据集输入到未训练完成的初始算法选取模型中,得到输出的预测算法选取结果,基于预测算法选取结果和标准算法选取结果,对初始算法选取模型的模型参数进行调整得到训练完成的目标算法选取模型,本发明实施例训练得到的算法选取模型能够预测出最适合当前语音场景的回声消除算法,解决了不同的回声消除算法所适用的语音场景不同的问题,保证了回声消除算法与当前语音场景的适配性,进而提高了回声消除后的语音质量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116665690B_ABST
    Figure CN116665690B_ABST
Patent Text Reader

Abstract

This invention discloses a training method for an algorithm selection model, an echo cancellation method, an apparatus, and a device. The method includes: acquiring training near-end sound data collected by a near-end microphone and training far-end sound data played by a near-end speaker; employing at least two preset echo cancellation algorithms from a set of echo cancellation algorithms, and performing echo cancellation processing on the training near-end sound data based on the training far-end sound data to obtain a training cancellation dataset; inputting the training far-end sound data and the training cancellation dataset into an untrained initial algorithm selection model to obtain an output predicted algorithm selection result; and adjusting the model parameters of the initial algorithm selection model based on the predicted algorithm selection result and the standard algorithm selection result to obtain a trained target algorithm selection model. The algorithm selection model in this embodiment can predict the most suitable echo cancellation algorithm for the current speech scene, ensuring the adaptability of the echo cancellation algorithm to the current speech scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a training method, echo cancellation method, apparatus, and equipment for an algorithm selection model. Background Technology

[0002] In recent years, offline communication methods such as meetings and face-to-face exchanges have increasingly been replaced by online communication methods such as real-time video conferencing and voice chat. As a result, users have more stringent requirements for the sound quality of audio and video.

[0003] Echo is a significant factor affecting the audio and video communication experience. It occurs when a speaker's voice data, after passing through a remote microphone, a nearby speaker, and other devices, and undergoing propagation and reflection in the spatial environment, is transmitted back to the remote speaker, making the speaker hear themselves again. Currently, many algorithms have been developed for echo cancellation, such as adaptive filtering algorithms and deep learning models.

[0004] In the process of realizing this invention, at least the following technical problems were found in the prior art:

[0005] Adaptive filtering algorithms perform well in eliminating noise in far-end single-talk scenarios, but their performance is poor in far-end and near-end dual-talk scenarios or scenarios with nonlinear distortion (such as high noise, high latency, or poor speaker quality). Deep learning models, on the other hand, exhibit the opposite performance in eliminating noise.

[0006] Therefore, different echo cancellation algorithms are applicable to different speech scenarios, and no echo cancellation algorithm that can be applied to complex speech scenarios has been developed yet. Summary of the Invention

[0007] This invention provides a training method for an algorithm selection model, an echo cancellation method, an apparatus, and a device to address the problem that different echo cancellation algorithms are applicable to different speech scenarios, improve the adaptability of echo cancellation algorithms to complex speech scenarios, and thus improve the speech quality after echo cancellation.

[0008] According to an embodiment of the present invention, a training method for selecting a model by an algorithm is provided, the method comprising:

[0009] Acquire training near-end sound data collected by the near-end microphone and training far-end sound data played by the near-end speaker;

[0010] Using at least two preset echo cancellation algorithms from the set of echo cancellation algorithms, and based on the training far-end sound data, perform echo cancellation processing on the training near-end sound data to obtain a training cancellation dataset;

[0011] The training far-end sound data and the training elimination dataset are input into the untrained initial algorithm selection model to obtain the output prediction algorithm selection result.

[0012] Based on the selection results of the prediction algorithm and the standard algorithm, the model parameters of the initial algorithm selection model are adjusted to obtain the trained target algorithm selection model.

[0013] An echo cancellation method is provided according to an embodiment of the present invention, the method comprising:

[0014] Acquire near-end sound data to be tested collected by a near-end microphone and far-end sound data to be tested played by a near-end speaker; wherein, the near-end sound data to be tested includes first near-end sound data to be tested and second near-end sound data to be tested, and the far-end sound data to be tested includes first far-end sound data to be tested and second far-end sound data to be tested.

[0015] Using at least two preset echo cancellation algorithms from the set of echo cancellation algorithms, based on the first test far-end sound data, the first test near-end sound data is subjected to echo cancellation processing to obtain the first test cancellation dataset.

[0016] Based on the first far-end sound data to be tested, the first cancellation dataset to be tested, and the pre-trained target algorithm selection model, the target echo cancellation algorithm is determined.

[0017] Using the target echo cancellation algorithm, based on the second far-end sound data to be tested, echo cancellation processing is performed on the second near-end sound data to be tested to obtain the second canceled data;

[0018] Based on the target echo cancellation algorithm, the second echo cancellation data to be tested, and the first echo cancellation dataset to be tested, the echo cancellation data corresponding to the near-end sound data to be tested is determined;

[0019] The target algorithm selection model is obtained by using the training method of the algorithm selection model described in any embodiment of the present invention.

[0020] According to another embodiment of the present invention, a training apparatus for selecting an algorithmic model is provided, the apparatus comprising:

[0021] The training sound data acquisition module is used to acquire training near-end sound data collected by the near-end microphone and training far-end sound data played by the near-end speaker;

[0022] The training cancellation dataset determination module is used to perform echo cancellation processing on the training near-end sound data based on the training far-end sound data using at least two preset echo cancellation algorithms from the echo cancellation algorithm set to obtain the training cancellation dataset.

[0023] The prediction algorithm selection result output module is used to input the training far-end sound data and the training elimination dataset into the untrained initial algorithm selection model to obtain the output prediction algorithm selection result.

[0024] The target algorithm selection model determination module is used to adjust the model parameters of the initial algorithm selection model based on the prediction algorithm selection result and the standard algorithm selection result to obtain the trained target algorithm selection model.

[0025] According to another embodiment of the present invention, an echo cancellation device is provided, the device comprising:

[0026] The test sound data acquisition module is used to acquire test near-end sound data collected by the near-end microphone and test far-end sound data played by the near-end speaker; wherein, the test near-end sound data includes first test near-end sound data and second test near-end sound data, and the test far-end sound data includes first test far-end sound data and second test far-end sound data.

[0027] The first echo cancellation dataset determination module is used to perform echo cancellation processing on the first near-end sound data based on the first far-end sound data to be tested, using at least two preset echo cancellation algorithms from the echo cancellation algorithm set to obtain the first echo cancellation dataset.

[0028] The target echo cancellation algorithm determination module is used to determine the target echo cancellation algorithm based on the first far-end sound data to be tested, the first cancellation dataset to be tested, and the pre-trained target algorithm selection model.

[0029] The second elimination data determination module is used to perform echo cancellation processing on the second near-end sound data based on the target echo cancellation algorithm to obtain the second elimination data.

[0030] The echo cancellation data determination module is used to determine the echo cancellation data corresponding to the near-end sound data to be tested based on the target echo cancellation algorithm, the second cancellation data to be tested, and the first cancellation dataset to be tested;

[0031] The target algorithm selection model is obtained by using the training method of the algorithm selection model described in any embodiment of the present invention.

[0032] According to another embodiment of the present invention, an electronic device is provided, the electronic device comprising:

[0033] At least one processor; and

[0034] A memory communicatively connected to the at least one processor; wherein,

[0035] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the training method for the algorithm selection model according to any embodiment of the present invention, and / or to perform the echo cancellation method according to any embodiment of the present invention.

[0036] According to another embodiment of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute, implement the training method of the algorithm selection model described in any embodiment of the present invention, and / or be able to execute the echo cancellation method described in any embodiment of the present invention.

[0037] The technical solution of this invention employs at least two preset echo cancellation algorithms from a set of echo cancellation algorithms. Based on training far-end sound data played by a near-end speaker, echo cancellation processing is performed on training near-end sound data collected by a near-end microphone to obtain a training cancellation dataset. The training far-end sound data and the training cancellation dataset are then input into an untrained initial algorithm selection model to obtain the output predicted algorithm selection result. Based on the predicted algorithm selection result and the standard algorithm selection result, the model parameters of the initial algorithm selection model are adjusted to obtain a trained target algorithm selection model. The algorithm selection model trained in this invention can predict the most suitable echo cancellation algorithm for the current speech scene, solving the problem that different echo cancellation algorithms are applicable to different speech scenes, ensuring the adaptability of the echo cancellation algorithm to the current speech scene, and thus improving the speech quality after echo cancellation.

[0038] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 A flowchart illustrating a training method for selecting a model according to an embodiment of the present invention;

[0041] Figure 2 This is a schematic diagram of an audio / video conferencing voice scenario provided in one embodiment of the present invention;

[0042] Figure 3 This is a network architecture diagram of an initial algorithm selection model provided in one embodiment of the present invention;

[0043] Figure 4 A flowchart illustrating an echo cancellation method provided in one embodiment of the present invention;

[0044] Figure 5 A flowchart illustrating another echo cancellation method provided in an embodiment of the present invention;

[0045] Figure 6 A flowchart illustrating a specific example of an echo cancellation method provided in an embodiment of the present invention;

[0046] Figure 7 This is a network architecture diagram of an echo cancellation model provided in one embodiment of the present invention;

[0047] Figure 8 A schematic diagram of the structure of a training device for an algorithm selection model provided in an embodiment of the present invention;

[0048] Figure 9 This is a schematic diagram of the structure of an echo cancellation device provided in one embodiment of the present invention;

[0049] Figure 10 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present invention. Detailed Implementation

[0050] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0051] It should be noted that the terms "first," "second," "initial," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0052] Figure 1 This is a flowchart illustrating a training method for an algorithm selection model according to an embodiment of the present invention. This embodiment is applicable to training a network model for selecting an echo cancellation algorithm. The method can be executed by a training device for the algorithm selection model, which can be implemented in hardware and / or software and can be configured in a terminal device. Figure 1 As shown, the method includes:

[0053] S110. Acquire training near-end sound data collected by the near-end microphone and training far-end sound data played by the near-end speaker.

[0054] For example, the embodiments of the present invention can be applied to voice scenarios such as audio and video conferencing, voice games, and smart speaker control. Specifically, audio and video conferencing voice scenarios include a near-end speaker, a near-end microphone, a far-end speaker, and a far-end microphone, while voice scenarios such as voice games and smart speaker control only include a near-end speaker and a near-end microphone. In voice scenarios such as voice games and smart speaker control, the near-end microphone may simultaneously receive the user's input voice signal as well as sound signals such as teammates' voices and background music played by the near-end speaker, thus also requiring echo cancellation.

[0055] Although the concept of "far end" does not exist in voice scenarios such as voice games and smart speaker control, in order to unify the description and distinguish it from near end sound data, in this embodiment, the sound data played by the near end speaker in different voice scenarios is defined as far end sound data.

[0056] Figure 2This is a schematic diagram of an audio / video conferencing voice scenario provided in an embodiment of the present invention. Specifically, the voice scenario includes a near-end speaker, a near-end microphone, a far-end speaker, and a far-end microphone. Here, "near-end" refers to the spatial environment where the main speaker is located, and "far-end" refers to the spatial environment where other participants in the call are located. Specifically, n represents the sampling point index after sampling the audio / video stream, lpb(n) represents the far-end sound data played by the near-end speaker, lpb_e(n) represents the sound data captured by the near-end microphone after environmental reflection and spatial propagation, s(n) represents the main speaker's voice data, and v(n) represents the noise data in the near-end spatial environment. It is easy to understand that the near-end sound data mic(n) collected by the near-end microphone = s(n) + lpb_e(n) + v(n). The purpose of echo cancellation is to remove lpb_e(n) and v(n) from the near-end sound data mic(n) to obtain echo-cancelled data S(n).

[0057] by Figure 2 For example, training near-end sound data corresponds to Figure 2 In the context of mic(n), the training term corresponds to the remote audio data. Figure 2 lpb(n) in the system equipment can be obtained through the sampling path.

[0058] Specifically, the training near-end sound data and the training far-end sound data correspond to the same sound time period. For example, the training near-end sound data and the training far-end sound data can use open-source sound data or actually collected sound data.

[0059] S120. Using at least two preset echo cancellation algorithms from the echo cancellation algorithm set, and based on the training far-end sound data, perform echo cancellation processing on the training near-end sound data to obtain the training cancellation dataset.

[0060] Specifically, the echo cancellation algorithm set includes at least two preset echo cancellation algorithms. For example, the preset echo cancellation algorithms included in the echo cancellation algorithm set include, but are not limited to, filter-based echo cancellation algorithms, transform-based echo cancellation algorithms, adaptive filtering algorithms, and echo cancellation models, etc. The number and type of algorithms in the echo cancellation algorithm set are not limited here.

[0061] In one embodiment, the principle of the filter-based echo cancellation algorithm is to input training near-end sound data into a filter to obtain output training cancellation data, wherein the cutoff frequency of the filter can be set based on the delay time of the training far-end sound data.

[0062] In one embodiment, the principle of the transformation-based echo cancellation algorithm is to use the forward and inverse transformation algorithm to reverse the training far-end sound data, and then cancel the reversed training far-end sound data with the training near-end sound data to obtain the training cancellation data.

[0063] In one embodiment, the adaptive filtering algorithm can be classified into four modules: time delay estimation, linear filtering, double-talk detection, and residual echo suppression. The algorithm principle is to estimate an approximate echo path by adjusting the weight vector of the filter to approximate the real echo path, thereby obtaining a simulated training far-end sound data. The simulated training far-end sound data is then removed from the training near-end sound data to obtain the training canceled data.

[0064] In one embodiment, the input data of the echo cancellation model are training far-end sound data and training near-end sound data, and the output data is training cancellation data. The network architecture and training method of the echo cancellation model are not limited here.

[0065] Specifically, the training cancellation dataset includes training cancellation data obtained using at least two preset echo cancellation algorithms for training near-end sound data. For example, the training cancellation dataset can be {S|S1(n),S2(n),…,S}. m} represents the training elimination dataset, where m represents the total number of training elimination data in the training elimination dataset.

[0066] S130. Input the training far-end sound data and the training cancellation dataset into the untrained initial algorithm selection model to obtain the output prediction algorithm selection result.

[0067] In one specific embodiment, the present invention does not limit the network architecture of the initial algorithm selection model.

[0068] In another specific embodiment, the initial algorithm selection model includes a feature extraction module, a concatenation module, a self-attention module, a temporal pooling module, and an output layer. The feature extraction module extracts features from at least two training elimination data points in the training far-end speech data and the training elimination dataset, respectively, to obtain training far-end speech features and at least two training elimination features. The concatenation module concatenates the training far-end speech features and each training elimination feature output by the feature extraction module to obtain concatenated speech features. The self-attention module determines updated speech features based on the concatenated speech features output by the concatenation module. The temporal pooling module compresses the updated speech features output by the self-attention module to obtain compressed speech features. The output layer outputs the prediction algorithm selection result based on the compressed speech features output by the temporal pooling module.

[0069] For example, the feature extraction module may employ a convolutional neural network (CNN), and the network architecture of the feature extraction module is not limited here.

[0070] For example, the output layer includes a fully connected layer and a softmax function; the network architecture of the output layer is not limited here.

[0071] Specifically, the self-attention module enables each temporal speech feature in the concatenated speech features to update itself using surrounding speech features, resulting in more accurate updated speech features compared to the concatenated speech features.

[0072] Figure 3 This is a network architecture diagram of an initial algorithm selection model provided in one embodiment of the present invention. Specifically, the feature extraction module in the initial algorithm selection model uses a CNN network, and the output layer includes a fully connected layer and a softmax function. The input data of the initial algorithm selection model includes training far-end sound data lpb(n) and at least two training cancellation data (S1(n), S2(n), ..., S...) from the training cancellation dataset. m (n)), the output data are the prediction results p1, p2, ..., p1 corresponding to each preset echo cancellation algorithm. m .

[0073] For example, the prediction result can be a predicted selection probability, a predicted non-selection probability, or whether to select. Specifically, the sum of all prediction results is 1. For instance, assuming the echo cancellation algorithm set includes Algorithm 1, Algorithm 2, and Algorithm 3, the predicted algorithm selection result can be [0.6 0.2 0.2], or it can be [1 0 0], where "1" indicates selection and "0" indicates non-selection. There is no limitation on the predicted algorithm selection result here.

[0074] S140. Based on the selection results of the prediction algorithm and the selection results of the standard algorithm, the model parameters of the initial algorithm selection model are adjusted to obtain the target algorithm selection model after training.

[0075] Based on the above embodiments, specifically, the method further includes: using a preset evaluation algorithm, determining the algorithm evaluation results corresponding to each preset echo cancellation algorithm based on at least one training cancellation data in the training cancellation dataset; and determining the standard algorithm selection result corresponding to the initial algorithm selection model based on the evaluation results of each algorithm.

[0076] For example, the preset evaluation algorithms include, but are not limited to, at least one of the following: PESQ (Perceptual evaluation of speech quality), ERLE (echo return loss enhancement), CER (character error rate), and STOI (Short-Time Objective Intelligibility). The preset evaluation algorithm is not limited here.

[0077] The PESQ algorithm is an objective, full-reference speech quality assessment method that provides a subjective MOS (Mean Offset of Speech) prediction value for objective speech quality assessment. The calculation process includes preprocessing, time alignment, perceptual filtering, masking effects, and more. Specifically, the PESQ algorithm's evaluation score ranges from -0.5 to 4.5; a higher score indicates a better echo cancellation effect from the preset echo cancellation algorithm.

[0078] The ERLE algorithm is commonly used to evaluate echo cancellation effectiveness in single-talk scenarios within near-space environments. For example, the ERLE algorithm satisfies the following formula:

[0079]

[0080] Among them, when the training data is subsequently used for speech recognition, the CER algorithm is an important indicator for measuring the effect of echo cancellation. The CER algorithm is used to measure the similarity between two sequences.

[0081] Among them, the STOI algorithm is an important indicator for measuring speech intelligibility. The STOI algorithm evaluation results score ranges from 0 to 1. The higher the score, the better the echo cancellation effect of the preset echo cancellation algorithm.

[0082] For example, the standard algorithm selection result can be used to characterize the echo cancellation effect of the preset echo cancellation algorithm on the training near-end sound data. For instance, when the standard algorithm selection result is the standard selection probability corresponding to each preset echo cancellation algorithm, the higher the standard selection probability, the better the echo cancellation effect of the preset echo cancellation algorithm.

[0083] Based on the above embodiments, specifically, the model parameters of the initial algorithm selection model are adjusted based on the prediction algorithm selection results and the standard algorithm selection results to obtain the trained target algorithm selection model. This includes: constructing a loss function based on the prediction algorithm selection results and the standard algorithm selection results; adjusting the model parameters of the initial algorithm selection model based on the loss function to obtain the adjusted initial algorithm selection model; based on the adjusted initial algorithm selection model, returning to the step of inputting the training far-end sound data and the training cancellation dataset into the untrained initial algorithm selection model to obtain the output prediction algorithm selection results; until the loss function converges, the adjusted initial algorithm selection model is used as the trained target algorithm selection model.

[0084] For example, the loss function types include, but are not limited to, squared loss function, logarithmic loss function, exponential loss function, mean squared error loss function, logistic regression loss function, Huber loss function, cross-entropy loss function, and Kullback-Leibler divergence loss function, etc. The specific type of loss function is not limited here.

[0085] In one specific embodiment, the loss function is a cross-loss function. For example, the cross-loss function L satisfies the formula:

[0086]

[0087] Where y represents the selection result of the standard algorithm. This indicates the result selected by the prediction algorithm.

[0088] The technical solution of this embodiment employs at least two preset echo cancellation algorithms from a set of echo cancellation algorithms. Based on training far-end sound data played by a near-end speaker, echo cancellation processing is performed on training near-end sound data collected by a near-end microphone to obtain a training cancellation dataset. The training far-end sound data and the training cancellation dataset are then input into an untrained initial algorithm selection model to obtain the output predicted algorithm selection result. Based on the predicted algorithm selection result and the standard algorithm selection result, the model parameters of the initial algorithm selection model are adjusted to obtain a trained target algorithm selection model. The algorithm selection model trained in this embodiment can predict the most suitable echo cancellation algorithm for the current speech scene, solving the problem that different echo cancellation algorithms are applicable to different speech scenes, ensuring the adaptability of the echo cancellation algorithm to the current speech scene, and thus improving the speech quality after echo cancellation.

[0089] Figure 4This is a flowchart illustrating an echo cancellation method according to an embodiment of the present invention. This embodiment is applicable to the situation of canceling echo signals in sound data collected in a speech scene. The method can be executed by an echo cancellation device, which can be implemented in hardware and / or software and can be configured in a terminal device. Figure 4 As shown, the method includes:

[0090] S210: Acquire the near-end sound data to be tested collected by the near-end microphone and the far-end sound data to be tested played by the near-end speaker.

[0091] For example, the embodiments of the present invention can be applied to voice scenarios such as audio and video conferencing, voice games, and smart speaker control. Specifically, the voice scenario in audio and video conferencing includes a near-end speaker, a near-end microphone, a far-end speaker, and a far-end microphone, while the voice scenarios in voice games and smart speaker control only include a near-end speaker and a near-end microphone. In voice scenarios such as voice games and smart speaker control, the near-end microphone may simultaneously receive the user's input voice signal as well as sound signals such as teammates' voices and background music played by the near-end speaker, thus also requiring echo cancellation.

[0092] In this embodiment, the near-end sound data to be tested includes first near-end sound data to be tested and second near-end sound data to be tested, and the far-end sound data to be tested includes first far-end sound data to be tested and second far-end sound data to be tested.

[0093] Specifically, the near-end sound data to be tested is divided into a first near-end sound data to be tested and a second near-end sound data to be tested, and the far-end sound data to be tested is divided into a first far-end sound data to be tested and a second far-end sound data to be tested. The first near-end sound data to be tested and the second near-end sound data to be tested do not overlap, and the first far-end sound data to be tested and the second far-end sound data to be tested do not overlap.

[0094] S220. Using at least two preset echo cancellation algorithms from the echo cancellation algorithm set, based on the first far-end sound data to be tested, perform echo cancellation processing on the first near-end sound data to be tested to obtain the first canceled dataset.

[0095] Specifically, the echo cancellation algorithm set includes at least two preset echo cancellation algorithms. For example, the preset echo cancellation algorithms included in the echo cancellation algorithm set include, but are not limited to, filter-based echo cancellation algorithms, transform-based echo cancellation algorithms, adaptive filtering algorithms, and echo cancellation models, etc. The number and type of algorithms in the echo cancellation algorithm set are not limited here.

[0096] Specifically, the first dataset to be canceled includes first canceled data obtained by using at least two preset echo cancellation algorithms for the first near-end sound data to be tested. For example, the first dataset to be canceled can be {S|S1(n),S2(n),…,S m} represents the first elimination dataset to be tested, where m represents the total number of training elimination data in the first elimination dataset.

[0097] S230. Based on the first far-end sound data to be tested, the first cancellation dataset to be tested, and the pre-trained target algorithm selection model, determine the target echo cancellation algorithm.

[0098] Specifically, the first far-end sound data to be tested and the first echo cancellation dataset to be tested are input into a pre-trained target algorithm selection model to obtain the target algorithm selection result. Based on the target algorithm selection result, the target echo cancellation algorithm is determined.

[0099] In this embodiment, the target algorithm selection model is obtained by using the training method of the algorithm selection model provided in the above embodiments of the present invention, which will not be described in detail here.

[0100] For example, the target algorithm selection result can be the target selection probability, target non-selection probability, or whether to select for each preset echo cancellation algorithm. Specifically, when the prediction algorithm selection result is the target selection probability corresponding to each preset echo cancellation algorithm, the preset echo cancellation algorithm with the highest target selection probability is selected as the target echo cancellation algorithm, or the preset echo cancellation algorithm with a target selection probability exceeding a preset probability threshold is selected as the target echo cancellation algorithm, such as a preset probability threshold of 0.85; when the prediction algorithm selection result is the target non-selection probability corresponding to each preset echo cancellation algorithm, the preset echo cancellation algorithm with the lowest target non-selection probability is selected as the target echo cancellation algorithm, or the preset echo cancellation algorithm with a target selection probability not exceeding a preset probability threshold is selected as the target echo cancellation algorithm, such as a preset probability threshold of 0.25; when the target algorithm selection result is whether to select for each preset echo cancellation algorithm, the selected preset echo cancellation algorithm is selected as the target echo cancellation algorithm.

[0101] S240. Using a target echo cancellation algorithm, based on the second far-end sound data to be tested, echo cancellation processing is performed on the second near-end sound data to be tested to obtain the second canceled data.

[0102] S250. Based on the target echo cancellation algorithm, the second echo cancellation data to be tested, and the first echo cancellation dataset to be tested, determine the echo cancellation data corresponding to the near-end sound data to be tested.

[0103] Specifically, the first elimination data obtained by using the target echo cancellation algorithm in the first elimination dataset is used as the target elimination data. Based on the target elimination data and the second elimination data, the echo cancellation data corresponding to the near-end sound data to be tested is generated.

[0104] The technical solution of this embodiment employs at least two preset echo cancellation algorithms from a set of echo cancellation algorithms to perform echo cancellation processing on the first near-end sound data to be tested, thereby obtaining a first canceled data set. A pre-trained target algorithm selection model is used to predict the target echo cancellation algorithm. Then, based on the second far-end sound data to be tested, the target echo cancellation algorithm is used to perform echo cancellation processing on the second near-end sound data to be tested, thereby obtaining the second canceled data set. This embodiment of the invention uses an algorithm selection model to determine the most suitable echo cancellation algorithm for the current speech scene, solving the problem that different echo cancellation algorithms are applicable to different speech scenes, ensuring the adaptability of the echo cancellation algorithm to the current speech scene, and thus improving the speech quality after echo cancellation.

[0105] Figure 5 This is a flowchart of another echo cancellation method provided in one embodiment of the present invention. The embodiment further refines the "acquiring the near-end sound data to be tested collected by the near-end microphone and the far-end sound data to be tested played by the near-end speaker" in the above embodiment, such as... Figure 5 As shown, the method includes:

[0106] S310. In response to detecting the elimination trigger command, a first test sound period and a second test sound period are generated based on the current trigger time, the first preset voice duration, and the second preset voice duration.

[0107] In this embodiment, the first preset speech duration is shorter than the second preset speech duration. This setting improves the efficiency of the subsequent target echo cancellation algorithm determination process and enhances the overall efficiency of the echo cancellation method.

[0108] For example, assuming the current trigger time is 8:00, the first preset voice duration and the second preset voice duration are 1 minute and 59 minutes respectively, then the first test sound period and the second test sound period are 8:00-8:01 and 8:01-9:00 respectively.

[0109] Based on the above embodiments, specifically, the method further includes: in response to detecting a voice start command, generating a cancellation trigger command when voice data is present in the sound data collected by the near-end microphone; and / or, generating a cancellation trigger command when the current acquisition time exceeds the second sound period to be tested, and returning to execute the step of generating a first sound period to be tested and a second sound period to be tested based on the current trigger time, a first preset voice duration, and a second preset voice duration; and / or, in response to detecting a silence period, generating a cancellation trigger command when the silence duration corresponding to the silence period exceeds a silence duration threshold, and returning to execute the step of generating a first sound period to be tested and a second sound period to be tested based on the current trigger time, a first preset voice duration, and a second preset voice duration.

[0110] In one embodiment, when a voice call begins, voice recognition is performed on the sound data collected by the near-end microphone. If the recognition result contains voice data, a cancellation trigger command is generated. The voice data represents socially significant sound signals emitted by human vocal organs.

[0111] In one embodiment, in a voice scenario, the near-end microphone collects the near-end sound data to be tested in real time. If the current collection time exceeds the second test sound time period, a cancellation trigger instruction is generated and the process returns to execute S310.

[0112] The advantage of this setup is that, in actual voice scenarios, the likelihood of changes in the voice environment for both parties is very small. For a period of time, the voice environment will remain unchanged, whether it's a single-talk scenario, a two-talk scenario, or a non-linear distortion scenario. Therefore, the same echo cancellation algorithm can be used for echo cancellation processing within the second preset voice duration. However, as the voice duration increases, the likelihood of changes in the voice environment increases. This embodiment of the invention limits the second preset voice duration to achieve the purpose of reselecting the echo cancellation algorithm at regular intervals. This allows the echo cancellation method to adapt to complex voice scenarios with changing voice environments, improving the flexibility of the echo cancellation process and further ensuring the voice quality after echo cancellation.

[0113] In one embodiment, for example, the silence duration threshold can be 5 minutes or 10 minutes.

[0114] The advantage of this setting is that the longer the silence duration, the greater the possibility of changes in the speech environment. By detecting the silence duration, this embodiment of the invention enables the echo cancellation method to adapt to complex speech scenarios with changing speech scenes, improving the flexibility of the echo cancellation process and further ensuring the speech quality after echo cancellation.

[0115] S320: Acquire the first and second near-end sound data to be tested, which are collected by the near-end microphone and correspond to the first and second test sound time periods, respectively.

[0116] S330: Obtain the first and second remote sound data to be tested, which are played by the near-end player and correspond to the first and second test sound periods, respectively.

[0117] S340. Using at least two preset echo cancellation algorithms from the echo cancellation algorithm set, based on the first far-end sound data to be tested, perform echo cancellation processing on the first near-end sound data to be tested to obtain the first canceled dataset.

[0118] S350. Based on the first test far-end sound data, the first test cancellation dataset, and the pre-trained target algorithm selection model, determine the target echo cancellation algorithm.

[0119] S360. Using a target echo cancellation algorithm, based on the second far-end sound data to be tested, echo cancellation processing is performed on the second near-end sound data to be tested to obtain the second canceled data.

[0120] S370. Based on the target echo cancellation algorithm, the second echo cancellation data to be tested, and the first echo cancellation dataset to be tested, determine the echo cancellation data corresponding to the near-end sound data to be tested.

[0121] Figure 6 This is a flowchart illustrating a specific example of an echo cancellation method provided in an embodiment of the present invention. Specifically, after a call begins, it is detected in real time whether the near-end microphone has collected sound data. If so, a first test sound period t1 and a second test sound period t2 are generated based on the current trigger time, a first preset voice duration, and a second preset voice duration. At least two preset echo cancellation algorithms from the echo cancellation algorithm set are used to perform echo cancellation processing on the near-end audio of t1, resulting in at least two first test cancellation data (S1(n), S2(n), ..., S...). m (n)). Based on the far-end audio of time period t1, each first-to-be-tested echo cancellation data, and the target algorithm selection model, the target echo cancellation algorithm is determined. It is determined whether the current acquisition time is within time period t2. If not, the process returns to the step of generating the first-to-be-tested sound time period t1 and the second-to-be-tested sound time period t2 based on the current trigger time, the first preset speech duration, and the second preset speech duration. If so, the target echo cancellation algorithm is used to perform echo cancellation processing on the near-end audio of t2, obtaining the second-to-be-tested echo cancellation data x(n). Correspondingly, the echo cancellation data S(n) corresponding to time period t1+t2 is x(n)+S x (n), where S x(n) represents the first elimination data to be tested obtained by using the target echo cancellation algorithm in the first elimination dataset.

[0122] The technical solution of this embodiment, in response to the detection of a cancellation trigger command, generates a first test sound period and a second test sound period based on the current trigger time, a first preset voice duration, and a second preset voice duration. It acquires first test near-end sound data and second test near-end sound data corresponding to the first test sound period and the second test sound period, respectively, collected by the near-end microphone. It also acquires first test far-end sound data and second test far-end sound data corresponding to the first test sound period and the second test sound period, respectively, played by the near-end player. This solves the problem of the singular selection time of the echo cancellation algorithm, improves the flexibility of the selection time of the echo cancellation algorithm, further improves the flexibility of the echo cancellation process, and thus further ensures the voice quality after echo cancellation.

[0123] Based on the above embodiments, specifically, the echo cancellation algorithm set includes an echo cancellation model. The echo cancellation model includes an input module, an intra-block module, a conversion module, an extra-block module, and an output module. The input module is used to determine a first speech feature to be tested based on the first far-end sound data to be tested and the first near-end sound data to be tested. The intra-block module is used to perform frequency domain feature processing on the first speech feature to be tested output by the input module and output a first frequency domain speech feature. The conversion module is used to convert the first frequency domain speech feature output by the intra-block module to obtain a first converted speech feature. The extra-block module is used to perform time domain feature processing on the first converted speech feature output by the conversion module to obtain a first time domain speech feature. The output module is used to output the first canceled data to be tested based on the first time domain speech feature output by the extra-block module and the first near-end sound data to be tested.

[0124] Specifically, the echo cancellation model uses a dual-path network architecture.

[0125] In one specific embodiment, the input module includes a preprocessing layer, a concatenation layer, and a convolutional network connected in series. The preprocessing layer is used to perform short-time Fourier transform (STFT) on the input first test far-end sound data and the first test near-end sound data respectively to obtain the first test far-end speech spectrum and the first test near-end speech spectrum. The concatenation layer is used to concatenate the first test far-end speech spectrum and the first test near-end speech spectrum output by the preprocessing layer to obtain the first test concatenated speech spectrum. The convolutional network is used to extract features from the first test concatenated speech spectrum output by the concatenation layer and output the first test speech features.

[0126] In one specific embodiment, the in-block module includes a self-attention network, and the out-of-block module includes a Long Short-Term Memory (LSTM) network.

[0127] In one specific embodiment, the output module includes a series of deconvolutional networks, a fusion unit, and an inverse short-time Fourier transform layer. The deconvolutional network is used to perform deconvolution processing on the first temporal speech features output by the out-of-block module to obtain first deconvolutional speech features. The fusion unit is used to fuse the first near-end sound data to be tested and the first deconvolutional speech features output by the deconvolutional network to obtain first fused speech features to be tested. The inverse short-time Fourier transform layer is used to perform inverse short-time Fourier transform (iSTFT) on the first fused speech features to be tested output by the fusion unit to output the first canceled data to be tested.

[0128] For example, the fusion operation of the fusion unit can be bitwise multiplication.

[0129] Figure 7 This is a network architecture diagram of an echo cancellation model provided in one embodiment of the present invention. Specifically, the input data of the echo cancellation model are first far-end sound data to be tested and first near-end sound data to be tested, and the output data is the first canceled data to be tested. In the self-attention network used in the block modules, each frequency domain point (Q) within the block is evaluated in relation to other frequency domain points (K) and represented by weights. These weights are then used to weight and sum the corresponding frequency domain values ​​(V) of other frequency domain points to obtain the vector representation of that frequency domain point based on the self-attention mechanism. The light-colored frequency domain points in "Q" represent the process of obtaining the light-colored frequency domain points in "V".

[0130] The advantage of this design is that traditional echo cancellation models also use Long Short-Term Memory (LSTM) networks within their intra-block modules. This means that the intra-block modules can only represent the frequency domain speech features at a specific moment, without any temporal sequence requirements. Therefore, LSM networks can only learn limited feature information, resulting in limited modeling capabilities. In contrast, the echo cancellation model in this embodiment uses a self-attention network with a self-attention mechanism within its intra-block modules. Compared to LSM networks, these modules can not only learn the temporal sequence information between frequency domain points but also enrich the feature representation of the current frequency domain point based on the evaluated closeness of the relationship between the current frequency domain point and other frequency domain points, thus enhancing modeling capabilities.

[0131] Figure 8 This is a schematic diagram of a training device for selecting an algorithmic model, provided in one embodiment of the present invention. Figure 8As shown, the device includes: a training sound data acquisition module 410, a training elimination dataset determination module 420, a prediction algorithm selection result output module 430, and a target algorithm selection model determination module 440.

[0132] Among them, the training sound data acquisition module 410 is used to acquire training near-end sound data collected by the near-end microphone and training far-end sound data played by the near-end speaker;

[0133] The training cancellation dataset determination module 420 is used to perform echo cancellation processing on the training near-end sound data based on the training far-end sound data by employing at least two preset echo cancellation algorithms from the echo cancellation algorithm set to obtain the training cancellation dataset.

[0134] The prediction algorithm selection result output module 430 is used to input the training far-end sound data and the training cancellation dataset into the untrained initial algorithm selection model to obtain the output prediction algorithm selection result.

[0135] The target algorithm selection model determination module 440 is used to adjust the model parameters of the initial algorithm selection model based on the prediction algorithm selection results and the standard algorithm selection results to obtain the trained target algorithm selection model.

[0136] The technical solution of this embodiment employs at least two preset echo cancellation algorithms from a set of echo cancellation algorithms. Based on training far-end sound data played by a near-end speaker, echo cancellation processing is performed on training near-end sound data collected by a near-end microphone to obtain a training cancellation dataset. The training far-end sound data and the training cancellation dataset are then input into an untrained initial algorithm selection model to obtain the output predicted algorithm selection result. Based on the predicted algorithm selection result and the standard algorithm selection result, the model parameters of the initial algorithm selection model are adjusted to obtain a trained target algorithm selection model. The algorithm selection model trained in this embodiment can predict the most suitable echo cancellation algorithm for the current speech scene, solving the problem that different echo cancellation algorithms are applicable to different speech scenes, ensuring the adaptability of the echo cancellation algorithm to the current speech scene, and thus improving the speech quality after echo cancellation.

[0137] Based on the above embodiments, specifically, the initial algorithm selection model includes a feature extraction module, a concatenation module, a self-attention module, a temporal pooling module, and an output layer. The feature extraction module extracts features from at least two training elimination data points in the training far-end speech data and the training elimination dataset, respectively, to obtain training far-end speech features and at least two training elimination features. The concatenation module concatenates the training far-end speech features and each training elimination feature output by the feature extraction module to obtain concatenated speech features. The self-attention module determines updated speech features based on the concatenated speech features output by the concatenation module. The temporal pooling module compresses the updated speech features output by the self-attention module to obtain compressed speech features. The output layer outputs the prediction algorithm selection result based on the compressed speech features output by the temporal pooling module.

[0138] Based on the above embodiments, specifically, the device further includes:

[0139] The standard algorithm selection result determination module is used to determine the algorithm evaluation results corresponding to each preset echo cancellation algorithm based on at least one training cancellation data in the training cancellation dataset, using a preset evaluation algorithm.

[0140] Based on the evaluation results of each algorithm, the standard algorithm selection result corresponding to the initial algorithm selection model is determined.

[0141] Based on the above embodiments, specifically, the target algorithm selection model determination module 440 is used for:

[0142] A loss function is constructed based on the selection results of the prediction algorithm and the standard algorithm.

[0143] The model parameters of the initial algorithm selection model are adjusted based on the loss function to obtain the adjusted initial algorithm selection model;

[0144] Based on the adjusted initial algorithm selection model, return to the step of inputting the training far-end sound data and the training cancellation dataset into the untrained initial algorithm selection model to obtain the output prediction algorithm selection result;

[0145] Once the loss function converges, the adjusted initial algorithm selection model is used as the target algorithm selection model after training is complete.

[0146] The training device for the algorithm selection model provided in the embodiments of the present invention can execute the training method for the algorithm selection model provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0147] Figure 9 This is a schematic diagram of the structure of an echo cancellation device according to an embodiment of the present invention. Figure 9As shown, the device includes: a test sound data acquisition module 510, a first test elimination dataset determination module 520, a target echo cancellation algorithm determination module 530, a second test elimination data determination module 540, and an echo cancellation data determination module 550.

[0148] The test sound data acquisition module 510 is used to acquire the test near-end sound data collected by the near-end microphone and the test far-end sound data played by the near-end speaker; wherein the test near-end sound data includes first test near-end sound data and second test near-end sound data, and the test far-end sound data includes first test far-end sound data and second test far-end sound data.

[0149] The first echo cancellation dataset determination module 520 is used to perform echo cancellation processing on the first near-end sound data based on the first far-end sound data to be tested, using at least two preset echo cancellation algorithms from the echo cancellation algorithm set to obtain the first echo cancellation dataset.

[0150] The target echo cancellation algorithm determination module 530 is used to determine the target echo cancellation algorithm based on the first far-end sound data to be tested, the first cancellation dataset to be tested, and the pre-trained target algorithm selection model.

[0151] The second elimination data determination module 540 is used to perform echo cancellation processing on the second near-end sound data based on the second far-end sound data to obtain the second elimination data.

[0152] The echo cancellation data determination module 550 is used to determine the echo cancellation data corresponding to the near-end sound data to be tested based on the target echo cancellation algorithm, the second cancellation data to be tested and the first cancellation dataset to be tested.

[0153] The target algorithm selection model is obtained by using the training method of the algorithm selection model provided in the above embodiments of the present invention.

[0154] The technical solution of this embodiment employs at least two preset echo cancellation algorithms from a set of echo cancellation algorithms to perform echo cancellation processing on the first near-end sound data to be tested, thereby obtaining a first canceled data set. A pre-trained target algorithm selection model is used to predict the target echo cancellation algorithm. Then, based on the second far-end sound data to be tested, the target echo cancellation algorithm is used to perform echo cancellation processing on the second near-end sound data to be tested, thereby obtaining the second canceled data set. This embodiment of the invention uses an algorithm selection model to determine the most suitable echo cancellation algorithm for the current speech scene, solving the problem that different echo cancellation algorithms are applicable to different speech scenes, ensuring the adaptability of the echo cancellation algorithm to the current speech scene, and thus improving the speech quality after echo cancellation.

[0155] Based on the above embodiments, specifically, the sound data acquisition module 510 is used for:

[0156] In response to the detection of a cancellation trigger command, a first test sound period and a second test sound period are generated based on the current trigger time, a first preset voice duration, and a second preset voice duration; wherein, the first preset voice duration is less than the second preset voice duration;

[0157] Acquire the first and second near-end sound data to be tested, which are collected by the near-end microphone and correspond to the first and second test sound time periods, respectively.

[0158] Obtain the first and second remote sound data to be tested, which are played by the near-end player and correspond to the first and second test sound periods, respectively.

[0159] Based on the above embodiments, specifically, the device further includes:

[0160] The cancellation trigger command generation module is used to generate a cancellation trigger command in response to the detection of a voice start command, provided that voice data is present in the sound data acquired by the near-end microphone; and / or,

[0161] If the current acquisition time exceeds the second test sound period, generate a trigger cancellation command and return to execute the steps of generating the first test sound period and the second test sound period based on the current trigger time, the first preset speech duration, and the second preset speech duration; and / or,

[0162] In response to the detection of a silent period, if the silence duration corresponding to the silent period exceeds the silence duration threshold, a trigger cancellation instruction is generated, and the process returns to execute the steps of generating a first test sound period and a second test sound period based on the current trigger time, a first preset voice duration, and a second preset voice duration.

[0163] Based on the above embodiments, specifically, the echo cancellation algorithm set includes an echo cancellation model. The echo cancellation model includes an input module, an intra-block module, a conversion module, an extra-block module, and an output module. The input module is used to determine a first speech feature to be tested based on the first far-end sound data to be tested and the first near-end sound data to be tested. The intra-block module is used to perform frequency domain feature processing on the first speech feature to be tested output by the input module and output a first frequency domain speech feature. The conversion module is used to convert the first frequency domain speech feature output by the intra-block module to obtain a first converted speech feature. The extra-block module is used to perform time domain feature processing on the first converted speech feature output by the conversion module to obtain a first time domain speech feature. The output module is used to output the first canceled data to be tested based on the first time domain speech feature output by the extra-block module and the first near-end sound data to be tested.

[0164] Based on the above embodiments, specifically, the in-block module includes a self-attention network, and the out-of-block module includes a long short-term memory network.

[0165] Based on the above embodiments, specifically, the input module includes a preprocessing layer, a concatenation layer, and a convolutional network in series. The preprocessing layer is used to perform short-time Fourier transform on the first test far-end sound data and the first test near-end sound data respectively to obtain the first test far-end speech spectrum and the first test near-end speech spectrum. The concatenation layer is used to concatenate the first test far-end speech spectrum and the first test near-end speech spectrum output by the preprocessing layer to obtain the first test concatenated speech spectrum. The convolutional network is used to extract features from the first test concatenated speech spectrum output by the concatenation layer and output the first test speech features.

[0166] Based on the above embodiments, specifically, the output module includes a series of deconvolutional networks, a fusion unit, and a short-time Fourier transform layer. The deconvolutional network is used to perform deconvolution processing on the first temporal speech features output by the out-of-block module to obtain the first deconvolutional speech features. The fusion unit is used to perform fusion processing on the first near-end sound data to be tested and the first deconvolutional speech features output by the deconvolutional network to obtain the first fused speech features to be tested. The short-time Fourier transform layer is used to perform short-time Fourier transform on the first fused speech features to be tested output by the fusion unit to output the first canceled data to be tested.

[0167] The echo cancellation device provided in the embodiments of the present invention can execute the echo cancellation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.

[0168] Figure 10 This is a schematic diagram of an electronic device provided according to one embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0169] like Figure 10As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0170] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0171] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the algorithm selection model training method provided in the above embodiments, and / or, echo cancellation method.

[0172] In some embodiments, the algorithm selection model training method and / or echo cancellation method provided in the above embodiments can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded into and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the algorithm selection model training method and / or echo cancellation method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to execute the algorithm selection model training method and / or echo cancellation method by any other suitable means (e.g., by means of firmware).

[0173] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0174] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0175] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0176] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0177] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0178] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0179] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0180] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A training method for selecting a model using an algorithm, characterized in that, include: Acquire training near-end sound data collected by the near-end microphone and training far-end sound data played by the near-end speaker; Using at least two preset echo cancellation algorithms from the set of echo cancellation algorithms, and based on the training far-end sound data, perform echo cancellation processing on the training near-end sound data to obtain a training cancellation dataset; The training far-end sound data and the training elimination dataset are input into the untrained initial algorithm selection model to obtain the output prediction algorithm selection result. Based on the selection results of the prediction algorithm and the standard algorithm, the model parameters of the initial algorithm selection model are adjusted to obtain the trained target algorithm selection model.

2. The method according to claim 1, characterized in that, The initial algorithm selection model includes a feature extraction module, a concatenation module, a self-attention module, a temporal pooling module, and an output layer. The feature extraction module extracts features from at least two training elimination data points in the training far-end speech data and the training elimination dataset to obtain training far-end speech features and at least two training elimination features. The concatenation module concatenates the training far-end speech features output by the feature extraction module and each of the training elimination features to obtain concatenated speech features. The self-attention module determines updated speech features based on the concatenated speech features output by the concatenation module. The temporal pooling module compresses the updated speech features output by the self-attention module to obtain compressed speech features. The output layer outputs the prediction algorithm selection result based on the compressed speech features output by the temporal pooling module.

3. The method according to claim 1, characterized in that, The method further includes: Using a preset evaluation algorithm, the algorithm evaluation results corresponding to each preset echo cancellation algorithm are determined based on at least one training cancellation data in the training cancellation dataset. Based on the evaluation results of each algorithm, the standard algorithm selection result corresponding to the initial algorithm selection model is determined.

4. The method according to any one of claims 1-3, characterized in that, The step of adjusting the model parameters of the initial algorithm selection model based on the selection results of the prediction algorithm and the standard algorithm to obtain the trained target algorithm selection model includes: Based on the selection results of the prediction algorithm and the standard algorithm, a loss function is constructed; The model parameters of the initial algorithm selection model are adjusted based on the loss function to obtain the adjusted initial algorithm selection model; Based on the adjusted initial algorithm selection model, the process returns to the step of inputting the training far-end sound data and the training cancellation dataset into the untrained initial algorithm selection model to obtain the output prediction algorithm selection result. When the loss function converges, the adjusted initial algorithm selection model is used as the target algorithm selection model after training is completed.

5. An echo cancellation method, characterized in that, include: Acquire near-end sound data to be tested collected by a near-end microphone and far-end sound data to be tested played by a near-end speaker; wherein, the near-end sound data to be tested includes first near-end sound data to be tested and second near-end sound data to be tested, and the far-end sound data to be tested includes first far-end sound data to be tested and second far-end sound data to be tested. Using at least two preset echo cancellation algorithms from the set of echo cancellation algorithms, based on the first test far-end sound data, the first test near-end sound data is subjected to echo cancellation processing to obtain the first test cancellation dataset. Based on the first far-end sound data to be tested, the first cancellation dataset to be tested, and the pre-trained target algorithm selection model, the target echo cancellation algorithm is determined. Using the target echo cancellation algorithm, based on the second far-end sound data to be tested, echo cancellation processing is performed on the second near-end sound data to be tested to obtain the second canceled data; Based on the target echo cancellation algorithm, the second echo cancellation data to be tested, and the first echo cancellation dataset to be tested, the echo cancellation data corresponding to the near-end sound data to be tested is determined; The target algorithm selection model is obtained by training the algorithm selection model according to any one of claims 1-4.

6. The method according to claim 5, characterized in that, The acquisition of the near-end sound data to be tested collected by the near-end microphone and the far-end sound data to be tested played by the near-end speaker includes: In response to the detection of a cancellation trigger command, a first test sound period and a second test sound period are generated based on the current trigger time, a first preset voice duration, and a second preset voice duration; wherein, the first preset voice duration is less than the second preset voice duration; Acquire the first and second near-end sound data to be tested, which are collected by the near-end microphone and correspond to the first and second test sound time periods, respectively. Obtain the first and second remote sound data to be tested, which are played by the near-end player and correspond to the first and second test sound periods, respectively.

7. The method according to claim 6, characterized in that, The method further includes: In response to the detection of a voice start command, if voice data is present in the sound data acquired by the near-end microphone, a cancellation trigger command is generated; and / or, If the current acquisition time exceeds the second test sound period, a trigger cancellation instruction is generated, and the process returns to execute the steps of generating the first test sound period and the second test sound period based on the current trigger time, the first preset speech duration, and the second preset speech duration; and / or, In response to the detection of a silent period, if the silence duration corresponding to the silent period exceeds the silence duration threshold, a trigger cancellation instruction is generated, and the process returns to the execution of the steps of generating a first test sound period and a second test sound period based on the current trigger time, a first preset voice duration, and a second preset voice duration.

8. The method according to claim 5, characterized in that, The echo cancellation algorithm set includes an echo cancellation model, which comprises an input module, an intra-block module, a conversion module, an extra-block module, and an output module. The input module determines a first speech feature to be tested based on the first far-end sound data and the first near-end sound data to be tested. The intra-block module performs frequency domain feature processing on the first speech feature output by the input module to output a first frequency domain speech feature. The conversion module converts the first frequency domain speech feature output by the intra-block module to obtain a first converted speech feature. The extra-block module performs time domain feature processing on the first converted speech feature output by the conversion module to obtain a first time domain speech feature. The output module outputs first canceled data based on the first time domain speech feature output by the extra-block module and the first near-end sound data to be tested.

9. The method according to claim 8, characterized in that, The in-block module includes a self-attention network, and the out-of-block module includes a long short-term memory network.

10. The method according to claim 8, characterized in that, The input module includes a preprocessing layer, a concatenation layer, and a convolutional network connected in series. The preprocessing layer performs short-time Fourier transform on the input first test far-end sound data and the first test near-end sound data to obtain the first test far-end speech spectrum and the first test near-end speech spectrum. The concatenation layer concatenates the first test far-end speech spectrum and the first test near-end speech spectrum output by the preprocessing layer to obtain the first test concatenated speech spectrum. The convolutional network extracts features from the first test concatenated speech spectrum output by the concatenation layer to output the first test speech features.

11. The method according to claim 10, characterized in that, The output module includes a series of deconvolutional networks, a fusion unit, and a short-time Fourier transform layer. The deconvolutional network is used to perform deconvolution processing on the first temporal speech features output by the out-of-block module to obtain first deconvolutional speech features. The fusion unit is used to fuse the first near-end sound data to be tested and the first deconvolutional speech features output by the deconvolutional network to obtain first fused speech features to be tested. The short-time Fourier transform layer is used to perform short-time Fourier transform on the first fused speech features to be tested output by the fusion unit to output first canceled data to be tested.

12. A training device for selecting an algorithmic model, characterized in that, include: The training sound data acquisition module is used to acquire training near-end sound data collected by the near-end microphone and training far-end sound data played by the near-end speaker; The training cancellation dataset determination module is used to perform echo cancellation processing on the training near-end sound data based on the training far-end sound data using at least two preset echo cancellation algorithms from the echo cancellation algorithm set to obtain the training cancellation dataset. The prediction algorithm selection result output module is used to input the training far-end sound data and the training elimination dataset into the untrained initial algorithm selection model to obtain the output prediction algorithm selection result. The target algorithm selection model determination module is used to adjust the model parameters of the initial algorithm selection model based on the prediction algorithm selection result and the standard algorithm selection result to obtain the trained target algorithm selection model.

13. An echo cancellation device, characterized in that, include: The test sound data acquisition module is used to acquire test near-end sound data collected by the near-end microphone and test far-end sound data played by the near-end speaker; wherein, the test near-end sound data includes first test near-end sound data and second test near-end sound data, and the test far-end sound data includes first test far-end sound data and second test far-end sound data. The first echo cancellation dataset determination module is used to perform echo cancellation processing on the first near-end sound data based on the first far-end sound data to be tested, using at least two preset echo cancellation algorithms from the echo cancellation algorithm set to obtain the first echo cancellation dataset. The target echo cancellation algorithm determination module is used to determine the target echo cancellation algorithm based on the first far-end sound data to be tested, the first cancellation dataset to be tested, and the pre-trained target algorithm selection model. The second elimination data determination module is used to perform echo cancellation processing on the second near-end sound data based on the target echo cancellation algorithm to obtain the second elimination data. The echo cancellation data determination module is used to determine the echo cancellation data corresponding to the near-end sound data to be tested based on the target echo cancellation algorithm, the second cancellation data to be tested, and the first cancellation dataset to be tested; The target algorithm selection model is obtained by training the algorithm selection model according to any one of claims 1-4.

14. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, which enables the at least one processor to perform the training method for the algorithm selection model according to any one of claims 1-4, and / or to perform the echo cancellation method according to any one of claims 5-11.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the training method for the algorithm selection model according to any one of claims 1-4, and / or are capable of executing the echo cancellation method according to any one of claims 5-11.

Citation Information

Patent Citations

  • Voice processing method and device, electronic equipment and storage medium

    CN115294997A

  • Classified noise reduction method and device, equipment and storage medium

    CN115938378A