Signal processing method, model training method, apparatus, device, and storage medium

By combining linear filtering and nonlinear filtering models, the problem of nonlinear echo components in real-time voice communication cannot be completely eliminated is solved, and efficient echo removal is achieved on the terminal device, avoiding excessive suppression of human voices and improving the quality of voice communication.

WO2025200881A1PCT designated stage Publication Date: 2025-10-02BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Application Number
PCT/CN2025/078263
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-28
Filing Date
2025-02-20
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

In real-time voice communication, existing technologies cannot effectively eliminate nonlinear echo components, resulting in excessive suppression of human voice components, affecting the quality of voice communication, especially in complex environments such as multi-person chorus scenes.

Method used

By obtaining a mixed sound signal and a reference sound signal, a linear filter is used to filter out the linear echo component, and a pre-trained nonlinear filter model is used to filter out the nonlinear echo component. The filtering amount of the nonlinear echo component is controlled to avoid excessive suppression of the human voice. A lightweight neural network model is used to implement it on the terminal device.

Benefits of technology

It effectively removes nonlinear echoes and improves voice call quality, especially in complex environments such as multi-person chorus, ensuring the restoration of human voice components and improving communication effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025078263_02102025_PF_FP_ABST
    Figure CN2025078263_02102025_PF_FP_ABST
Patent Text Reader

Abstract

A signal processing method, a model training method, an apparatus, a device, and a storage medium. The signal processing method comprises: acquiring a mixed sound signal and a reference sound signal (S101); on the basis of the reference sound signal, filtering a linear echo component in the mixed sound signal to obtain a first filtered signal (S102); and obtaining an output signal by means of a pre-trained nonlinear filtering model and the first filtered signal, and sending the output signal to a remote terminal device, wherein for the output signal, a nonlinear echo component of a target energy level is filtered out relative to the first filtered signal, the target energy level is less than a preset energy level, and the preset energy level is determined on the basis of the energy level of a human voice component in the first filtered signal (S103).A nonlinear component in a sample filtered signal is filtered out on the basis of a nonlinear filtering model, and the amount of the nonlinear component filtered out is controlled, avoiding over-suppression of the human voice component, and improving the voice communication quality in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Signal processing method, model training method, device, equipment and storage medium

[0001] This application claims priority to Chinese Patent Application No. 202410371423.8 filed on March 28, 2024, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field

[0002] The embodiments of the present disclosure relate to a signal processing method, a model training method, an apparatus, a device, and a storage medium. Background Art

[0003] In real-time voice communication scenarios, the sound played through the speaker of one device in external speaker mode is easily re-collected by the microphone, causing echo interference and affecting the quality of the sound signal received by the other device.

[0004] In related technologies, linear filtering of the sound signal collected by the microphone is performed using a filter, which can eliminate some linear echo components in the sound signal. However, the nonlinear echo components in the sound signal can only be tentatively suppressed through a pre-trained processing model, and cannot be completely eliminated.

[0005] However, in the process of suppressing nonlinear echo components, there is a problem that the human voice component in the sound signal is over-suppressed, which affects the quality of voice communication in complex environments. Summary of the Invention

[0006] Embodiments of the present disclosure provide a signal processing method, a model training method, an apparatus, a device, and a storage medium to overcome the problem of excessive suppression of human voice components in sound signals.

[0007] In a first aspect, an embodiment of the present disclosure provides a signal processing method, including:

[0008] Acquire a mixed sound signal and a reference sound signal, wherein the reference sound signal includes a sound signal sent by a remote terminal device to a near-end terminal device, and the mixed sound signal includes an ambient sound signal collected during the process of the near-end terminal device playing the reference sound signal; based on the reference sound signal, filter the linear echo component in the mixed sound signal to obtain a first filtered signal; obtain an output signal through a pre-trained nonlinear filtering model and the first filtered signal, and send the output signal to the remote terminal device; wherein, the output signal filters out the nonlinear echo component of a target energy level relative to the first filtered signal, and the target energy level is less than a preset energy level, and the preset energy level is determined based on the energy level of the human voice component in the first filtered signal.

[0009] In a second aspect, an embodiment of the present disclosure provides a model training method, including:

[0010] Acquire sample data, the sample data including a sample mixed sound signal and a corresponding sample reference sound signal, the sample reference sound signal representing a sound signal sent from a far-end terminal device to a near-end terminal device, and the sample mixed sound signal representing an ambient sound signal collected during the process of the near-end terminal device playing the sample reference sound signal; based on the sample reference sound signal, filter the linear echo component in the sample mixed sound signal to obtain a sample filtered signal; use the sample filtered signal and the corresponding sample pure speech signal to train an initial filtering model until the initial filtering model reaches a preset convergence condition to obtain a nonlinear filtering model, wherein the nonlinear filtering model is used to filter out the nonlinear echo component of a target energy level in the sound signal, the target energy level is less than a preset energy level, and the preset energy level is determined based on the energy level of the human voice component in the sound signal.

[0011] In a third aspect, an embodiment of the present disclosure provides a signal processing device, including:

[0012] a collection unit, configured to obtain a mixed sound signal and a reference sound signal, wherein the reference sound signal comprises a sound signal sent by a remote terminal device to a near-end terminal device, and the mixed sound signal comprises an ambient sound signal collected during the process of the near-end terminal device playing the reference sound signal;

[0013] a linear filtering unit, configured to filter the linear echo component in the mixed sound signal based on the reference sound signal to obtain a first filtered signal;

[0014] a nonlinear filtering unit, configured to obtain an output signal through a pre-trained nonlinear filtering model and the first filtered signal, and send the output signal to the remote terminal device;

[0015] The output signal filters out nonlinear echo components of a target energy level relative to the first filtered signal, and the target energy level is less than a preset energy level, which is determined based on the energy level of the human voice component in the first filtered signal.

[0016] In a fourth aspect, an embodiment of the present disclosure provides a model training device, comprising:

[0017] a sample acquisition unit, configured to acquire sample data, the sample data including a sample mixed sound signal and a corresponding sample reference sound signal, the sample reference sound signal representing a sound signal sent by a remote terminal device to a near-end terminal device, and the sample mixed sound signal representing an ambient sound signal collected during the process of the near-end terminal device playing the sample reference sound signal;

[0018] a sample processing unit, configured to filter the linear echo component in the sample mixed sound signal based on the sample reference sound signal to obtain a sample filtered signal;

[0019] A training unit is used to train an initial filtering model using the sample filtered signal and the corresponding sample pure speech signal until the initial filtering model reaches a preset convergence condition, thereby obtaining a nonlinear filtering model, wherein the nonlinear filtering model is used to filter out nonlinear echo components of a target energy level in the signal, the target energy level being less than a preset energy level, and the preset energy level being determined based on the energy level of the human voice component in the signal.

[0020] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;

[0021] The memory stores computer-executable instructions;

[0022] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the signal processing method described in the first aspect and various possible designs of the first aspect, or executes the model training method described in the second aspect and various possible designs of the second aspect.

[0023] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the signal processing method described in the first aspect and various possible designs of the first aspect is implemented, or the model training method described in the second aspect and various possible designs of the second aspect is implemented.

[0024] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, comprising a computer program, which, when executed by a processor, implements the signal processing method described in the first aspect and various possible designs of the first aspect, or implements the model training method described in the second aspect and various possible designs of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, a brief introduction to the drawings required for the embodiments will be given below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0026] FIG1 is a diagram of an application scenario of a signal processing method provided by an embodiment of the present disclosure;

[0027] FIG2 is a flow chart of a signal processing method according to an embodiment of the present disclosure;

[0028] FIG3 is a flowchart of a specific implementation method of step S102 in the embodiment shown in FIG2 ;

[0029] FIG4 is a flowchart of a specific implementation of step S103 in the embodiment shown in FIG2 ;

[0030] FIG5 is a second flow chart of a signal processing method provided by an embodiment of the present disclosure;

[0031] FIG6 is a schematic diagram of a process for generating a model input signal according to an embodiment of the present disclosure;

[0032] FIG7 is a flowchart of a specific implementation of step S207A;

[0033] FIG8 is a schematic diagram of a process for correcting an amplitude mask spectrum according to an embodiment of the present disclosure;

[0034] FIG9 is a flow chart of a model training method according to an embodiment of the present disclosure;

[0035] FIG10 is a flowchart of a specific implementation of step S303 in the embodiment shown in FIG9 ;

[0036] FIG11 is a schematic diagram of a model parameter space provided by an embodiment of the present disclosure;

[0037] FIG12 is a schematic structural diagram of an initial filtering model provided by an embodiment of the present disclosure;

[0038] FIG13 is a flowchart of a specific implementation method of step S3032 in the embodiment shown in FIG10 ;

[0039] FIG14 is a structural block diagram of a signal processing device provided by an embodiment of the present disclosure;

[0040] FIG15 is a structural block diagram of a model training device provided by an embodiment of the present disclosure;

[0041] FIG16 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure; and

[0042] FIG17 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0044] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0045] The following explains the application scenarios of the embodiments of the present disclosure:

[0046] Figure 1 is an application scenario diagram of the signal processing method provided by the embodiment of the present disclosure. The signal processing method provided by the embodiment of the present disclosure can be applied to an application (APP) with a voice communication function, such as a live broadcast application, a game application, an instant messaging application, etc. More specifically, it can be applied to an application scenario of a multi-person song chorus. The execution subject of this embodiment can be a terminal device that runs the above-mentioned application with a voice communication function, or a server that deploys the service end corresponding to the above-mentioned application, or other electronic devices that perform similar functions.

[0047] Among them, in some embodiments, the terminal device or server can implement the signal processing method provided by the embodiment of the present disclosure by running various computer executable instructions or computer programs. For example, computer executable instructions can be program-level commands, machine instructions or software instructions. The computer program can be a native program or software module in the operating system; it can be a local application, that is, a program that needs to be installed in the operating system to run, or it can be a small program embedded in any APP, that is, a program that runs based on a browser environment. In summary, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form, and the specific implementation form can be configured as needed. Further, in some embodiments, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud storage, cloud communications, cloud databases, cloud computing, cloud functions, network services, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, wherein the cloud service can be an interactive processing service for terminal devices to call.

[0048] Referring to FIG1 , taking a terminal device as an example, in a scenario where at least two terminal devices (terminal device A and terminal device B are shown in the figure) are conducting a voice call, the near-end terminal device A is the executor of the signal processing method provided in the embodiment of the present disclosure. After receiving the far-end voice signal sent by the far-end terminal device B, terminal device A plays it through the speaker. At the same time, terminal device A collects the sound emitted by the near-end user through the microphone, generates a near-end voice signal, and returns it to terminal device B, thereby completing the voice call process between the far-end terminal device B and the near-end terminal device A. During this process, the far-end voice signal emitted by the speaker on the terminal device A side will form an echo after being reflected by the environment and re-collected by the microphone. As a result, the near-end voice signal sent to terminal device B is mixed with the echo component, resulting in a decrease in the quality of voice communication. The signal processing method provided in this embodiment can be applied to the case where terminal device A collects the sound emitted by the near-end user through the microphone and processes the sound signal mixed with the echo to generate a near-end voice signal with a suppressed echo level.

[0049] For example, nonlinear echo components in sound signals are typically tentatively suppressed using pre-trained processing models, but they cannot be completely eliminated. In practice, due to factors such as difficulty in achieving precise model convergence and impurities in the training samples, the processing model can over-suppress the human voice component while suppressing the echo component. This can cause the human voice to become muffled and blurred, impacting voice communication quality. This impact is particularly severe in scenarios involving multiple voices, such as singing a chorus.

[0050] The embodiments of the present disclosure provide a signal processing method to solve the above problems.

[0051] Referring to FIG2 , FIG2 is a flow chart of a signal processing method according to an embodiment of the present disclosure. The method of this embodiment can be applied in a terminal device, and the signal processing method includes:

[0052] Step S101: Acquire a mixed sound signal and a reference sound signal.

[0053] Step S102: filtering the linear echo component in the mixed sound signal based on the reference sound signal to obtain a first filtered signal.

[0054] For example, referring to the application scenario diagram shown in FIG1 , the terminal device is, for example, a smartphone. After running a voice call application, for example, the terminal device (near-end terminal device) obtains the sound signal sent by the terminal device (far-end terminal device) on the opposite side through the application (client) via the server (server), namely, the reference sound signal. That is, the reference sound signal includes the sound signal sent by the far-end terminal device to the near-end terminal device; the terminal device will subsequently filter out the echo component based on the reference sound information and play the reference sound signal sent by the terminal device on the opposite side through the speaker. On the other hand, the terminal device (near-end terminal device) collects the ambient sound through the built-in microphone and obtains a sound signal containing the echo generated by playing the reference sound signal, namely, a mixed sound signal. In one possible case, the mixed sound signal includes the echo signal generated by the terminal device playing the reference sound signal, the sound signal generated by the terminal device playing background music, accompaniment, dialogue and other sound content, and may also include other noise information such as wind in the environment. The above-mentioned collection of various sound signals in the environment that can be collected by the sound collection unit of the terminal device is the mixed sound signal. That is, the mixed sound signal includes the ambient sound signal collected during the playback of the reference sound signal by the near-end terminal device. Specifically, the mixed sound signal is the sound signal collected by the terminal device within a preset time period after the mixed sound signal is played, that is, the two are correlated; the mixed sound signal and the reference sound signal are both time domain signals of a certain length, which can be the same or slightly different, for example, both have 1024 points. Furthermore, based on the specific sampling rate, the corresponding time length varies accordingly. In this embodiment, there is no limit on the signal length of the mixed sound signal and the reference sound signal, and they can be configured according to needs and device performance.

[0055] After obtaining the mixed sound signal and the reference sound signal, since the echo is generated by playing the reference sound through the speaker, the reference sound signal is equivalent to the source signal of the echo component. The terminal device filters the mixed sound signal based on the signal spectrum characteristics of the reference sound signal to filter out the linear echo component. In one possible implementation, as shown in Figure 3, the specific implementation of step S102 includes:

[0056] Step S1021: performing frequency domain transformation on the mixed sound signal and the reference sound signal to obtain a first time-frequency domain complex spectrum corresponding to the mixed sound signal and a second time-frequency domain complex spectrum corresponding to the reference sound signal.

[0057] Step S1022: Align the second time-frequency domain complex spectrum based on the first time-frequency domain complex spectrum to generate a third time-frequency domain complex spectrum.

[0058] Step S1023: Using the third time-frequency domain complex spectrum as a reference, a linear filter module is used to filter the first time-frequency domain complex spectrum to obtain a first filtered signal.

[0059] Exemplarily, first, the terminal device performs a short-time Fourier transform (STFT) on the mixed sound signal and the reference sound signal to obtain the time-frequency domain complex signals corresponding to the two, namely the first time-frequency domain complex spectrum and the second time-frequency domain complex spectrum. Among them, the short-time Fourier transform is a common mathematical change. On the basis of the Fourier transform, it adds a translation window to calculate the frequency complex spectrum of the time domain signal corresponding to each translation window to obtain the frequency characteristics corresponding to different local time periods. The complex spectrum is the output result after the time domain sound signal is Fourier transformed. The more detailed calculation process will not be repeated this time.

[0060] Afterwards, due to the influence of the sound propagation speed and path, there is a certain delay between the echo component in the mixed sound signal and the reference sound signal. Therefore, after obtaining the first time-frequency domain complex spectrum and the second time-frequency domain complex spectrum, the second time-frequency domain complex spectrum, that is, the reference sound signal, is calibrated and aligned using the phase information contained in the complex spectrum to obtain a third time-frequency domain complex spectrum. Afterwards, using the frequency characteristics represented by the third time-frequency domain complex spectrum as a reference, the first time-frequency domain complex spectrum is filtered using a linear filter module to filter the echo frequency represented by the third time-frequency domain complex spectrum to obtain a first filtered signal from which the linear echo component has been filtered out, wherein the first filtered signal can also be a complex spectrum based on a complex form, and its data size is the same as the data size of the first time-frequency domain complex spectrum, the second time-frequency domain complex spectrum, and the third time-frequency domain complex spectrum. In the above steps, the step of aligning the second time-frequency domain complex spectrum based on the first time-frequency domain complex spectrum and the step of filtering the first time-frequency domain complex spectrum using the linear filter module can be implemented by a preset functional module, and the specific implementation method will not be repeated here.

[0061] In the steps of this embodiment, the mixed sound signal and the reference sound signal are processed separately to obtain the corresponding time-frequency domain complex spectra of the two, and then phase alignment and linear filtering are performed based on the time-frequency domain complex spectra to obtain a first filtered signal from which the linear echo component is filtered out. Thereafter, subsequent nonlinear echo component filtering is performed based on the first filtered signal, which can further improve the effect of nonlinear echo filtering and thereby improve the voice call quality.

[0062] Step S103: Obtain an output signal through a pre-trained nonlinear filtering model and a first filtering signal, and send the output signal to a remote terminal device, wherein the output signal filters out the nonlinear echo component of a target energy level relative to the first filtering signal, and the target energy level is less than a preset energy level, and the preset energy level is determined based on the energy level of the human voice component in the first filtering signal.

[0063] Exemplarily, after obtaining the first filtered signal, the terminal device uses a preset nonlinear filtering model to further filter out the nonlinear echo components in the first filtered signal to obtain an output signal, and then sends the output signal with the nonlinear echo components filtered out to the remote terminal device. At this time, the sound signal received by the remote terminal device does not contain the echo component, thereby achieving the purpose of removing the echo. In one possible implementation, the first filtered signal can be directly input into the nonlinear filtering model for processing, and the output signal is generated by the nonlinear filtering model. In another possible implementation, the first filtered signal can be preprocessed first to generate a model input signal that matches the nonlinear filtering model, and then input into the nonlinear filtering model for processing. Exemplarily, as shown in FIG4 , the specific implementation of step S103 includes:

[0064] Step S1031: Obtain the first amplitude spectrum, the second amplitude spectrum, and the third amplitude spectrum corresponding to the first time-frequency domain complex spectrum, the third time-frequency domain complex number, and the first filtered signal, respectively.

[0065] Step S1032: stack the first amplitude spectrum, the second amplitude spectrum, and the third amplitude spectrum, and perform feature extraction on each channel based on the target characteristic frequency point to obtain a model input signal.

[0066] Step S1033: Input the model input signal into the nonlinear filtering model to obtain an output signal.

[0067] For example, referring to the first time-frequency domain complex spectrum and the third time-frequency domain complex number generated in the steps of the embodiment shown in FIG3 , in the steps of this embodiment, the obtained first filtered signal (complex spectrum) is combined with the first time-frequency domain complex spectrum and the third time-frequency domain complex number to perform channel stacking to obtain multi-channel stacked data. For example, the data size of the first filtered signal, the first time-frequency domain complex spectrum, and the third time-frequency domain complex number are all T×F, where T is the time series length and F is the frequency series length. After channel stacking the first filtered signal, the first time-frequency domain complex spectrum, and the third time-frequency domain complex number, the resulting multi-channel stacked data has a data size of T×F×3. Feature extraction is then performed on this multi-channel stacked data, for example, by extracting frequency values ​​at multiple target characteristic frequency points to obtain a model input signal. The data size of this model input signal is then T×F'×3, where F' is the length of the frequency feature, i.e., the number of target characteristic frequency points. For example, if the number of target characteristic frequency points is 176, the data size of the model input signal is T×176×3.

[0068] Thereafter, the model input signal is input into the nonlinear filter model for processing. For example, the nonlinear filter model filters out energy of the nonlinear echo component, and the output signal is obtained by processing the model input signal.

[0069] The nonlinear filtering model is a pre-trained, lightweight neural network model that can run on devices with low computing resources, such as mobile terminals. It not only filters out nonlinear echo components in the signal, but also has the ability to control the amount of nonlinear echo components removed from the signal.

[0070] In one possible implementation, the aforementioned capabilities of the nonlinear filtering model are realized through processing logic configured within the nonlinear filtering model. For example, the nonlinear filtering model first determines a preset energy level based on the vocal component in the first filtered signal, and then filters out the nonlinear echo component in the first filtered signal based on the preset energy level, thereby ensuring that the nonlinear echo component filtered out of the first filtered signal is less than the preset energy level, thereby resolving the problem of over-suppression of the vocal component. In one possible embodiment, the implementation of step S1033 includes:

[0071] Step S1033A: Process the model input signal through the nonlinear filtering model to obtain a first energy level of the human voice component in the first filtered signal.

[0072] Step S1033B: Determine the corresponding preset energy level through the first energy level.

[0073] Step S1033C: filtering the first filtered signal using a nonlinear filtering model with a preset energy level as a parameter to obtain an output signal.

[0074] In the steps of this embodiment, the first energy level of the human voice component is first determined through the processing logic configured in the nonlinear filtering model, and then a preset energy level is determined based on the first energy level. For example, a preset energy level is determined based on a multiple of the first energy level, a power of the first energy level, or a shift of the first energy level. Then, based on the preset energy level, the nonlinear echo component is filtered out, so that the actual filtering amount of the nonlinear echo component (target energy level) is less than the preset energy level, thereby achieving the purpose of controlling the filtering amount of the nonlinear component and avoiding excessive suppression of the human voice.

[0075] In another possible implementation, the aforementioned capabilities of the nonlinear filtering model are achieved through training. Specifically, compared to conventional echo filtering models, the nonlinear filtering model provided in this embodiment can, for a model input signal, output an output signal with a relatively high residual nonlinear echo component and a relatively low vocal component. Because the nonlinear filtering model can output an output signal with these characteristics, communications based on this output signal can reduce echo interference while also minimizing vocal suppression. This significantly enhances call quality in scenarios such as multi-person speech and chorus.

[0076] In this embodiment, a mixed sound signal and a reference sound signal are obtained; a linear echo component in the mixed sound signal is filtered based on the reference sound signal to obtain a first filtered signal; an output signal is obtained using a pre-trained nonlinear filtering model and the first filtered signal, and the output signal is transmitted to a remote terminal device; wherein the output signal filters out a target energy level of nonlinear echo components relative to the first filtered signal, the target energy level being less than a preset energy level, the preset energy level being determined based on the energy level of the human voice component in the first filtered signal. A sample filtered signal is obtained by linearly filtering sample data for echo components; the nonlinear components in the sample filtered signal are then filtered based on the nonlinear filtering model; and the ability of the nonlinear filtering model to control the amount of nonlinear components filtered is utilized to ensure that the amount of nonlinear echo components filtered in the output signal is less than a preset energy level determined based on the human voice component, thereby avoiding excessive suppression of the human voice component and prioritizing the restoration of the human voice component, thereby improving the quality of voice communication in complex environments.

[0077] Referring to FIG5 , FIG5 is a second flow chart of the signal processing method provided by the embodiment of the present disclosure. Based on the embodiment shown in FIG2 , this embodiment further refines step S103 , and the signal processing method includes:

[0078] Step S201: Acquire a mixed sound signal and a reference sound signal.

[0079] Step S202: performing frequency domain transformation on the mixed sound signal and the reference sound signal to obtain a first time-frequency domain complex spectrum corresponding to the mixed sound signal and a second time-frequency domain complex spectrum corresponding to the reference sound signal.

[0080] Step S203: aligning the second time-frequency domain complex spectrum based on the first time-frequency domain complex spectrum to generate a third time-frequency domain complex spectrum.

[0081] Step S204: using the third time-frequency domain complex spectrum as a reference, and using a linear filter module to filter the first time-frequency domain complex spectrum to obtain a first filtered signal.

[0082] Step S205: Obtain the first amplitude spectrum, the second amplitude spectrum, and the third amplitude spectrum corresponding to the first time-frequency domain complex spectrum, the third time-frequency domain complex number, and the first filtered signal, respectively.

[0083] Step S206: stack the first amplitude spectrum, the second amplitude spectrum, and the third amplitude spectrum by channel, and perform feature extraction on each channel based on the target characteristic frequency point to obtain a model input signal.

[0084] Figure 6 is a schematic diagram of a process for generating a model input signal provided by an embodiment of the present disclosure. The above process is introduced below in conjunction with Figure 6. As shown in Figure 6, first, the terminal device collects the ambient sound in the near-end real environment through a microphone, that is, a mixed sound signal (shown as y in the figure), and at the same time obtains the real-time communication voice sent by the remote terminal device received through the application, that is, a reference sound signal (shown as x in the figure). After that, the mixed sound signal and the reference sound signal are converted into a first time-frequency domain complex spectrum (shown as Y in the figure) and a second time-frequency domain complex spectrum (shown as X in the figure), and then the first time-frequency domain complex spectrum and the second time-frequency domain complex spectrum are first input into a delay compensation module, and the second time-frequency domain complex spectrum corresponding to the reference sound signal is subjected to delay compensation to generate a third time-frequency domain complex spectrum (shown as X' in the figure). After that, the third time-frequency domain complex spectrum and the first time-frequency domain complex spectrum are input into a linear filter to filter out the linear echo component to generate a first filtered signal (shown as E in the figure). After that, the first time-frequency domain complex spectrum, the third time-frequency domain complex number and the first filtered signal are respectively converted into the corresponding first amplitude spectrum (shown as |Y| in the figure), the second amplitude spectrum (shown as |X'| in the figure) and the third amplitude spectrum (shown as |E| in the figure), and after channel stacking, multi-channel stacked data is generated (shown as [|X'|, |Y|, |E|] in the figure). Then, F' target feature frequency points are extracted from the multi-channel stacked data to generate the model input signal stack([|X'|, |Y|, |E|])∈R 3×T×F′ . Where T is the number of time domain points.

[0085] Afterwards, the model input signal is input into the nonlinear filter model for processing to generate the output signal. The specific implementation of each step in the above process has been described in detail in the embodiment shown in FIG2 and will not be repeated here.

[0086] Optionally, before step S206, the method further includes (not shown in the figure):

[0087] Step S205A: Obtain the corresponding target frequency number according to the pre-configured echo cancellation parameters.

[0088] Step S205B: determining a target characteristic frequency point based on the characteristic frequency points of the target frequency point quantity in the low frequency interval, wherein the frequency value corresponding to the upper limit of the low frequency interval is less than a first frequency value, and the first frequency value is determined based on the human voice component in the mixed sound signal.

[0089] For example, in one possible implementation, the target characteristic frequencies can be stored in a pre-set fixed frequency sequence. The terminal device uses this fixed frequency sequence to obtain fixed target characteristic frequencies. For example, the fixed frequency sequence contains 3000 characteristic frequencies, each separated by 2 Hz. In another possible implementation, the number of target frequencies is dynamically determined. As shown in the above steps, the terminal device first determines the target frequency number based on pre-configured echo cancellation parameters. Generally speaking, a higher number of target frequencies results in better echo removal, but correspondingly, greater computational resource overhead. A corresponding target frequency number, such as 176, is determined based on the pre-configured echo cancellation parameters. A low-frequency interval is then determined based on the range of the vocal component in the mixed sound signal. This low-frequency interval can be dynamically determined based on analysis of the mixed sound signal. For example, a first frequency value is determined by analyzing the interval containing the vocal component in the mixed sound signal. A frequency interval, i.e., the low-frequency interval, is then determined using this first frequency value as the upper limit. Of course, the low-frequency interval can also be a fixed value, such as the 0-8 kHz range. Afterwards, based on the above target frequency number and low-frequency range, 176 characteristic frequency points, namely the target characteristic frequency points, are evenly extracted.

[0090] Step S207: Processing the model input signal through the nonlinear filtering model to obtain an amplitude mask spectrum for the first filtered signal, where the amplitude mask spectrum is used to adjust the frequency amplitude of the target characteristic frequency point in the first filtered signal.

[0091] Exemplarily, after obtaining a model input signal, the model input signal is input into a nonlinear filtering model for processing, and the nonlinear filtering model outputs an amplitude mask spectrum of the first filtered signal. The amplitude mask spectrum includes amplitude masks corresponding to the number of points of the first filtered signal (i.e., the number of target frequency points). Each amplitude mask is used to adjust the frequency amplitude of the corresponding characteristic frequency point in the first filtered signal. In subsequent steps, by applying the amplitude mask spectrum to the corresponding first filtered signal, the sound pressure suppression of a specific frequency in the first filtered signal is achieved, thereby achieving filtering of the nonlinear echo component in the first filtered signal.

[0092] Optionally, after step S207, the method further includes:

[0093] Step S207A: performing smoothing correction on the amplitude mask spectrum within the target frequency range to obtain an optimized amplitude mask spectrum.

[0094] Exemplarily, after obtaining the amplitude mask spectrum, in order to further provide a nonlinear echo suppression effect, the amplitude mask spectrum within the target frequency range can be further smoothed and corrected. For example, the mask average value of the m amplitude masks before and after the nth amplitude mask in the target frequency range is calculated, and the mask average value is used as the value of the nth amplitude mask, thereby achieving smooth correction. The target frequency range can be a frequency range preset based on the spectral distribution characteristics of the human voice component, wherein the nth amplitude mask referred to above can be determined based on indicators such as the mask value and kurtosis value of the amplitude mask. Exemplarily, the target frequency range includes a first frequency range. As shown in FIG7 , the specific implementation method of step S207A includes:

[0095] Step S207A-1: Obtain the first N maximum values ​​of the amplitude mask in the first frequency interval.

[0096] Step S207A-2: Obtain an amplitude mask mean value based on the first N maximum amplitude mask values.

[0097] Step S207A-3: determining a characteristic frequency point to be corrected according to the amplitude mask mean, wherein the amplitude mask of the characteristic frequency point to be corrected is greater than the amplitude mask mean.

[0098] Step S207A-4: performing weighted calculation on the amplitude mask of the characteristic frequency point to be corrected based on the amplitude mask mean to obtain an optimized amplitude mask spectrum corresponding to the first frequency interval, wherein the optimized amplitude mask spectrum of the characteristic frequency point to be corrected is smaller than the corresponding amplitude mask spectrum.

[0099] FIG8 is a schematic diagram of a process for correcting an amplitude mask spectrum provided by an embodiment of the present disclosure. The above process is described below in conjunction with FIG8. As shown in FIG8, exemplarily, the first frequency interval is [2k, 6k] Hz (i.e., 2000 to 6000 Hz, the same below). The terminal device first obtains the characteristic frequency points corresponding to the first N maximum values ​​of the amplitude mask in the first frequency interval, such as points P1, P2, and P3, and then calculates the average value of the amplitude mask of points P1, P2, and P3, such as Value_1. Afterwards, on the amplitude mask spectrum, the abnormal points whose amplitude mask is greater than the amplitude mask mean are determined, that is, the characteristic frequency points to be corrected, such as points P4 and P5, whose corresponding amplitude masks are Value_2 and Value_3, respectively. Finally, the amplitude masks of the above characteristic frequency points to be corrected are weightedly calculated based on the amplitude mask mean to obtain the optimized amplitude masks corresponding to the characteristic frequency points to be corrected, such as Value_2' and Value_3', respectively. Then, the optimized amplitude mask is used to replace the original amplitude mask of the characteristic frequency point to be corrected, and the optimized amplitude mask spectrum is obtained.

[0100] Regarding the nonlinear filtering model proposed in the embodiment of the present disclosure, due to its low computational complexity and limited modeling capabilities, it is prone to occasional slight echo leakage (similar to musical noise, which is perceived as occasional "sizzling", mainly occurring in chorus scenes, when the echo and near-end human voice spectral structures are similar or even overlapped). Although the slight echo residue in the chorus scene can be masked by appropriate mixing strategies, it cannot be completely eliminated. To this end, in the embodiment of the present disclosure, the amplitude mask output by the model is further optimized to achieve correction of abnormal amplitude masks and avoid the problem of slight echo leakage. In specific implementation, the inventors found in practice that the nonlinear filtering model proposed in this embodiment produces echo leakage similar to musical noise mainly because certain frequency points in the 2k-6k Hz frequency band have large abnormal amplitude mask values. Therefore, in the embodiment of the present disclosure, the amplitude mask spectrum is smoothed to reduce the abnormal values, thereby achieving further suppression of slight echo residues.

[0101] Optionally, the target frequency interval further includes a second frequency interval and a third frequency interval, the lower limit of the second frequency interval is greater than or equal to the upper limit of the first frequency interval; the upper limit of the third frequency interval is less than or equal to the lower limit of the first frequency interval, and further includes:

[0102] Step S207A-5: Obtain the first J maximum amplitude mask values ​​in the second frequency interval and the first K maximum amplitude mask values ​​in the third frequency interval.

[0103] Step S207A-6: If the average of the first J amplitude mask maxima is greater than the first mask threshold and the average of the K amplitude mask maxima is less than the second mask threshold, then the optimized amplitude mask spectrum corresponding to the second frequency interval is obtained based on the amplitude mask average corresponding to the first frequency interval.

[0104] Exemplarily, in order to further reduce echo leakage, in the embodiment of the present disclosure, the abnormal points in the high-frequency interval and the low-frequency interval of the first frequency interval are further located and processed. Specifically, as shown in the above steps, first obtain the first J amplitude mask maxima and the corresponding characteristic frequency points in the first frequency interval, for example, obtain the characteristic frequency points corresponding to the first 20 amplitude mask maxima in the [6k, 8k] frequency interval, and calculate the corresponding amplitude mask mean, for example, M_high; then obtain the characteristic frequency points corresponding to the first 10 amplitude mask maxima in the [0, 2k] frequency interval, and calculate the corresponding amplitude mask mean, for example, M_low. Afterwards, if M_high<0.3 and M_low is greater than 0.3, obtain the amplitude mask mean corresponding to the first frequency interval, which is the amplitude mask mean corresponding to the first frequency interval, for example, the average of the first 20 amplitude mask maxima in the [4k, 6k] frequency interval, for example, M_mid. Afterwards, the amplitude mask corresponding to the second frequency interval is weighted based on M_mid, for example, M_high is updated by the following formula:

[0105] M_high=M_high*0.5+M_mid*0.5

[0106] Finally, based on the amplitude mask M_high, the outliers in the second frequency interval (i.e., the frequency points whose amplitude mask is greater than M_high) are processed. For example, the amplitude mask of the outliers is set to M_high, or the amplitude mask of the outliers is weighted based on M_high. The specific implementation method is similar to the above-mentioned process of processing the outliers in the first frequency interval, and will not be repeated here.

[0107] In the steps of this embodiment, by extending the processing of abnormal points in the high frequency range, echo leakage can be further reduced and call quality can be improved. Wherein, N, J, and K are all positive integers.

[0108] Step S208: using the optimized amplitude mask spectrum to modify the first filtered signal to obtain a second filtered signal.

[0109] Step S209: performing time domain transformation on the second filtered signal to generate an output signal, and sending the output signal to a remote terminal device.

[0110] For example, after obtaining the optimized amplitude mask spectrum, each optimized amplitude mask in the optimized amplitude mask spectrum is used to adjust the amplitude of the corresponding characteristic frequency point in the first filtered signal to obtain a second filtered signal. Subsequently, an inverse Fourier transform is performed on the second filtered signal to restore it to a time domain signal, thereby generating an output signal. The output signal is then transmitted to the remote terminal device, completing the voice communication process.

[0111] Of course, it is understood that, in this embodiment, since step S207A is an optional step, in another possible implementation, after executing step S207, step S208 can be directly executed to generate the second filtered signal, that is, the first filtered signal is modified using the amplitude mask spectrum to obtain the second filtered signal. The specific implementation process is similar and will not be repeated here.

[0112] In this embodiment, the implementation of steps S201 to 2066 has been introduced in the embodiment shown in FIG. 2 of the present disclosure (steps S101 to S102 ), and the specific implementation methods are the same and will not be described again here.

[0113] The key to the signal processing method provided in the above embodiment lies in the use of a special processing model that can quantitatively filter out nonlinear echo components in the signal, thereby ensuring that the nonlinear echo components in the generated output signal have a higher energy level than those in the output signal output by a general echo filtering model. This is the nonlinear filtering model proposed in this embodiment, which solves the problem of excessive suppression of human voice components during echo removal in two-way voice communication applications. The nonlinear filtering model achieves this capability due to its special model structure and feature training method. The following describes the training method for this nonlinear filtering model.

[0114] FIG9 is a flow chart of a model training method provided by an embodiment of the present disclosure, which is used to implement the training of the nonlinear filtering model used in the above embodiments. The execution subject of the method of this embodiment can be a terminal device or a server, and an electronic device with similar functions can be obtained. The execution subject of this embodiment can be the same as or different from the execution subject of the method embodiments shown in FIG2-FIG8. For example, as shown in FIG9, the method includes:

[0115] Step S301: Acquire sample data, where the sample data includes a sample mixed sound signal and a corresponding sample reference sound signal.

[0116] Step S302: Based on the sample reference sound signal, filter the linear echo component in the sample mixed sound signal to obtain a sample filtered signal.

[0117] Step S303: Use the sample filter signal and the corresponding sample pure speech signal to train the initial filter model until the initial filter model reaches the preset convergence condition, thereby obtaining a nonlinear filter model, wherein the nonlinear filter model is used to filter out the nonlinear echo component of the target energy level in the sound signal, the target energy level is less than the preset energy level, and the preset energy level is determined based on the energy level of the human voice component in the sound signal.

[0118] Exemplarily, the sample reference sound signal represents the sound signal sent by the far-end terminal device to the near-end terminal device, and the sample mixed sound signal represents the ambient sound signal collected during the near-end terminal device playing the sample reference sound signal. Specifically, for example, the sample mixed sound signal refers to the ambient sound signal actually collected by the microphone, or the simulated sound signal obtained by simulation. The sample mixed sound signal contains an echo component, and the sample reference sound signal is the source signal for the echo component. The specific implementation methods of the sample mixed sound signal and the corresponding sample reference sound signal can refer to the mixed sound signal and reference sound signal introduced in the embodiment shown in Figure 2, which will not be repeated here. Afterwards, based on the sample reference sound signal, a linear filter is used to filter the linear echo component in the sample mixed sound signal to obtain a sample filtered signal, wherein the sample filtered signal is equivalent to the first filtered signal in the above method embodiment. The process of filtering the linear echo component in the sample mixed sound signal can refer to the process of generating the first filtered signal in the previous method embodiment, which will not be repeated here.

[0119] Afterwards, after obtaining the sample filtered signal, the model input information is constructed using the sample filtered signal, the aligned sample reference sound signal (complex spectrum) and the sample mixed sound signal (complex spectrum). Afterwards, the sample clean speech signal corresponding to the sample mixed sound signal in the sample data is obtained. The sample clean speech signal is an ideal sound signal without echo components, corresponding to the output signal sent by the near-end terminal device to the far-end terminal device. The model input information is used as the model input, and the sample clean speech signal is used as the "correct output" of the model to train the initial filtering model until it converges to the nonlinear filtering model used in the above method embodiment. The trained nonlinear filtering model can filter out the nonlinear echo component of the target energy level in the sound signal, and the target energy level is less than the preset energy level. The preset energy level is determined based on the energy level of the human voice component in the sound signal input to the nonlinear filtering model.

[0120] Furthermore, in a possible implementation, as shown in FIG10 , a specific implementation of step S303 includes:

[0121] Step S3031: Obtaining a combined loss function, wherein the combined loss function is composed of a first loss function and a second loss function, wherein the first loss function is used to guide the initial filtering model to filter out nonlinear echo components in the model input signal, and the second loss function is used to generate a penalty value for filtering results exceeding a preset energy level when the initial filtering model filters out the nonlinear echo components in the model input signal, wherein the preset energy level is determined based on the energy level of the human voice component in the sample mixed sound signal;

[0122] Step S3032: Based on the combined loss function, the initial filter model is trained using the sample filtered signal and the corresponding sample clean speech signal until the initial filter model converges to a nonlinear filter model.

[0123] Exemplarily, in this embodiment, the training process of the initial filter model is implemented using a combined loss function, which is composed of a first loss function and a second loss function. The first loss function is used to guide the initial filter model to filter out nonlinear echo components in the model input signal. This loss function functions similarly to the damage function used in the training process of the echo cancellation model. Through this loss function, the initial filter model can gradually converge to a model capable of filtering out nonlinear echo components. The second loss function is an auxiliary loss function used to generate a penalty value when the filter output of the model in each round filters out excessive nonlinear echo components during the training process of the initial filter model, thereby controlling the convergence direction of the initial filter model. The combined loss function composed of the first and second loss functions can guide the initial filter model to converge to a model that meets the performance requirements for filtering out nonlinear echo components in the signal while controlling the model's convergence path, thereby causing the initial filter model to converge along a specific path.

[0124] Afterwards, the sample filter signal is input into the initial filter model to be trained to obtain the model output. The sample clean speech signal and the combined loss function are then used to calculate the residual value and reversely update the parameters of the initial filter model until the initial filter model converges.

[0125] The following is a more detailed description of the role of the combined loss function used in the above embodiment:

[0126] FIG11 is a schematic diagram of a model parameter space provided by an embodiment of the present disclosure. As shown in FIG11 , the two axes of the model parameter space are the w1 axis and the w2 axis, respectively representing the two model parameters of the control model. After the model parameters are randomly initialized, the initial filter model to be trained is obtained. At this time, the initial filter model (the model parameters) is located at the P0 point in the parameter space. Afterwards, after processing the training samples, calculating the residual values, updating the gradients, and updating the model parameters, the initial filter model (the model parameters) will gradually move closer to the optimal solution position P1. Under ideal conditions, the model will eventually converge to P1, thereby generating an ideal filter model that can completely filter out nonlinear echoes from the signal. At the same time, since the direction of model convergence is random, the model can converge along the path L1 in the figure, or it can converge along the path L2 in the figure. Exemplarily, when the initial filtering model converges along path L1 in the figure and converges to within the range of Z1, the model can now filter out fewer nonlinear echo components from the signal, thereby achieving better restoration of the human voice component (collected by the near-end terminal device), but at the same time there may be a slight leakage echo characteristic. Afterwards, the model continues to converge to point P1 until it reaches the convergence limit, such as point P3 in the figure; and when the initial filtering model converges along path L2 in the figure and converges to within the range of Z2, the model can now filter out more nonlinear echo components from the signal, thereby achieving stronger echo component suppression, but at the same time may cause excessive suppression of the human voice component. Afterwards, the model continues to converge to point P1 until it reaches the convergence limit, such as point P4 in the figure.

[0127] However, in actual applications, due to objective reasons (training sample quality, training cost), the model cannot converge to the ideal point P1, but can only be trained to the convergence limits P3 and P4. In this case, due to the random convergence path of the model, the model can converge via path L2 and finally converge to the P4 position. The converged model obtained in this case will filter out more nonlinear echo components, and at the same time lead to the problem of excessive suppression of human voice components, that is, the nonlinear filtering model required in this embodiment cannot be obtained.

[0128] In the embodiment of the present disclosure, the initial filtering model is trained by a combined loss function composed of two loss functions (a first loss function and a second loss function), which enables the model to converge toward ideal parameters (controlled by the first loss function) while converging along the path L1 shown in, for example, FIG11 (controlled by the second loss function), thereby enabling the trained function to have the characteristic of tending to filter out fewer echo components while retaining more human voice components.

[0129] Furthermore, FIG12 is a schematic diagram of the structure of an initial filtering model provided by an embodiment of the present disclosure. The following describes the process of training the initial filtering model in conjunction with the schematic diagram of the structure of the initial filtering model shown in FIG12. Exemplarily, as shown in FIG12, the initial filtering model includes an encoding module, a feature extraction module, a mask prediction module, and an auxiliary mask prediction module, wherein the encoding module, the feature extraction module, and the mask prediction module are sequentially connected in series, the feature extraction module includes a plurality of feature extraction layers connected in series, and the auxiliary mask prediction module is connected to a feature extraction layer located in the middle layer of the feature extraction module.

[0130] The encoding module encodes the model input signal to produce an encoded signal. Specifically, the encoding module consists of several two-dimensional convolutional layers, which compress the frequency dimension of the input features (model input signal) while amplifying the channel dimension, mapping the features to a higher-dimensional latent space. The feature extraction module progressively extracts the nonlinear echo components from the encoded signal through feature extraction layers. Specifically, the feature extraction module consists of four layers of gated recurrent units (GRUs) with consistent dimensions. These modules model long-range dependencies and global spectral patterns, gradually extracting features that are useful for distinguishing echoes from near-end speech signals. The mask prediction module is used to predict the precise amplitude mask spectrum based on the first spectral feature. The precise amplitude mask spectrum is used to adjust the frequency amplitude of the target characteristic frequency point in the model input signal to filter out the nonlinear echo components in the model input signal. The first spectral feature is the spectral feature extracted by the terminal feature extraction layer of the feature extraction module. Specifically, the mask estimation module consists of a layer of two-dimensional convolution, a layer of fully connected layer, and a layer of Sigmoid nonlinear activation function. The mask prediction module is a module that further predicts the amplitude mask based on the output of the feature extraction module to achieve echo filtering. The amplitude mask it outputs is the output of the initial filtering model. The auxiliary mask prediction module is used to predict the rough amplitude mask spectrum based on the second spectral feature. The rough amplitude mask spectrum is used to calculate the penalty value of the filtering result. The second spectral feature is the characteristic frequency extracted by the non-terminal feature extraction layer of the feature extraction module.

[0131] Correspondingly, since the nonlinear filtering model used in the above embodiment is obtained after training the initial filtering model, the two are identical or similar in model structure. Therefore, in a possible implementation method, the nonlinear filtering model also has the above model structure, that is, the nonlinear filtering model includes an encoding module, a feature extraction module, a mask prediction module and an auxiliary mask prediction module, wherein the encoding module is used to encode the model input signal to obtain an encoded signal; the feature extraction module includes multiple feature extraction layers connected in series, which are used to gradually extract the nonlinear echo components in the encoded signal based on the feature extraction layer; the mask prediction module is used to predict the precise amplitude mask spectrum according to the first spectrum feature, and the precise amplitude mask spectrum is used to adjust the frequency amplitude of the target characteristic frequency point in the model input signal to filter out the nonlinear echo components in the model input signal, and the first spectrum feature is the spectrum feature extracted by the terminal feature extraction layer of the feature extraction module; the auxiliary mask prediction module is used to predict the rough amplitude mask spectrum according to the second spectrum feature, and the rough amplitude mask spectrum is used to calculate the penalty value of the filtering result, and the second spectrum feature is the characteristic frequency extracted by the non-terminal feature extraction layer of the feature extraction module.

[0132] Furthermore, based on the structure of the initial filtering model, as shown in FIG13 , the specific implementation steps of step S3032 include:

[0133] Step S3032-1: Process the sample filtered signal to obtain a model input signal based on a multi-channel amplitude spectrum.

[0134] Step S3032-2: Input the model input signal into the initial filtering model, and process it through the encoding module and the feature extraction module respectively to obtain the first frequency feature and the second frequency feature.

[0135] Step S3032-3: Process the first frequency feature through the mask prediction module to obtain a precise amplitude mask spectrum, and process the second frequency feature through the auxiliary mask prediction module to obtain a rough amplitude mask spectrum.

[0136] For example, first, the sample filter signal is processed to generate a model input signal based on a multi-channel amplitude spectrum that can be received by the initial filter model. The specific implementation method and generation method of the model input signal have been introduced in detail in the previous embodiment and will not be repeated here. Afterwards, the model input signal is input into the initial filter model and processed in sequence by the encoding module and the feature extraction module. The last feature extraction layer of the feature extraction module outputs the first frequency feature, and the second feature extraction layer of the feature extraction module outputs the second frequency feature. The first frequency feature and the second frequency feature can be represented by a multi-dimensional matrix, which will not be repeated here.

[0137] Afterwards, the precise amplitude mask spectrum and the rough amplitude mask spectrum are calculated by the mask prediction module and the auxiliary mask prediction module respectively. The precise amplitude mask spectrum is used to guide the model to converge to the optimal parameter position, and the rough amplitude mask spectrum is used to control the convergence path of the guided model.

[0138] Step S3032-4: applying the precise amplitude mask spectrum and the rough amplitude mask spectrum to the sample filtered signals respectively to obtain corresponding first sample filtered signals and second sample filtered signals;

[0139] Exemplarily, thereafter, the precise amplitude mask spectrum and the rough amplitude mask spectrum are respectively applied to the sample filtered signal. Exemplarily, in order to achieve better spectrum recovery, in this embodiment, the precise amplitude mask spectrum and the rough amplitude mask spectrum are applied to the sample filtered signal in the form of deep filtering, thereby recovering the amplitude of the signal and generating a first sample filtered signal and a second sample filtered signal. The first sample filtered signal and the second sample filtered signal achieve suppression of energy levels at different frequency points.

[0140] Step S3032-5: Based on the complex spectrum of the sample clean speech signal, using the first loss function and the second loss function respectively, calculate a first error value corresponding to the first sample filtered signal and a second error value corresponding to the second sample filtered signal, and generate a comprehensive error value based on the first error value and the second error value;

[0141] Step S3032-6: Update the model parameters of the initial filtering model based on the comprehensive error value, and update to the next set of sample filtering signals.

[0142] Then, exemplarily, the complex spectrum of the sample pure speech signal is obtained, and the first loss function is used to obtain the first error value corresponding to the first sample time domain signal; the second loss function is used to obtain the second error value corresponding to the second sample time domain signal, and then the first error value and the second error value are added to obtain a comprehensive error value, and then reverse transmission is performed based on the comprehensive error value to update the parameters of the initial filtering model.

[0143] Furthermore, illustratively, the second loss function is implemented as shown in formulas (1)-(2): baseLoss=sL1(|R1 r |,|R2 r |)+sL1(|R1 i |,|R2 i |)+sL1(|R1|,|R2|) (2)

[0144] The implementation of the combined loss function is shown in formula (3): Loss final=(1-β)*Loss2+β*Loss1 (3)

[0145] Here, Loss1 is the first loss function, Loss2 is the second loss function, sL1 is the smooth L1 loss function, i.e., smoothL1Loss. R1 represents the model prediction result, i.e., the first sample filtered signal, and R2 represents the true result, i.e., the complex spectrum of the sample clean speech signal. The subscripts r and i represent the real and imaginary parts, respectively. || represents the modulo operation. α is the dynamic penalty coefficient, such as 1.1 and 2.2. β is the weight coefficient used to control the ratio of the two loss functions, such as 0.3, 0.6, 1.3, etc.

[0146] Furthermore, the second loss function, which is equivalent to the loss function of the first stage, uses a dynamic penalty on the frequency points of overvoltage by calculating local features, so that the model tends to have residual echoes; and the first loss function, which is equivalent to the loss function of the second stage, allows the model to continue to converge on the basis of the previous stage and achieve less echo residuals.

[0147] Afterwards, the next set of sample filtered signals is loaded, and the process returns to step S3032-1, repeating the above process until a preset convergence condition is met. Convergence conditions may include, for example, the number of cycles reaching an upper limit, or the comprehensive error value being less than an error threshold. The determination step of whether the preset convergence condition is met can be performed before step S3032-1 or after step S3032-7, and this is not limited here.

[0148] The model training method provided in this embodiment sets an auxiliary mask prediction module in the model, obtains a rough amplitude mask spectrum based on the auxiliary mask prediction module, and then uses the rough amplitude mask spectrum to obtain a second loss function. The combined loss function formed by the second loss function guides the model to converge to a convergence path that tends to retain echo residue. In actual application, the model can converge to a parameter position that retains more echo components, thereby having a strong special effect of retaining human voice components and avoiding excessive suppression of human voices. It is particularly suitable for voice communication scenarios with choral accompaniment (linear echo), such as multi-person chorus.

[0149] Optionally, before executing step S201, the method further includes (not shown in the figure):

[0150] Step S200A: Acquire original sample data, where the original sample data includes at least two training samples.

[0151] Step S200B: Processing the original sample data based on the pre-trained echo filter model to obtain pre-processing information corresponding to each training sample, where the pre-processing information is a sound signal with the echo component in the training sample filtered out;

[0152] Step S200C: determining at least one interference sample according to the energy amplitude of the preprocessing information, and removing the interference sample from the original sample data to obtain sample data, wherein the energy amplitude of the interference sample is greater than a preset test threshold.

[0153] This embodiment also provides a data pre-cleaning step. Specifically, first, after the original sample data is processed using a pre-trained echo filter model, the original sample data is processed to obtain pre-processed information output by the echo filter model. The original sample data includes at least two training samples, each of which may include, for example, the sample mixed sound signal, the sample reference sound signal, and the sample pure speech signal described in the previous embodiment. The echo filter model is a model that filters out echo components from sound signals. The echo filter model can be trained based on the model training method provided in this embodiment. Subsequently, after the training samples are processed using the echo filter model, the pre-processed information obtained is the echo filtering result. If the energy amplitude of the echo filtering result is too large (greater than a test threshold), it indicates that the sample data originally contained sound signals collected by the near-end device. Since this part of the sound signal is not filtered out by the model, the energy amplitude of the output pre-processed information is large. In this type of sample, since the near-end sound component is mixed in, the near-end sound component acts as interference and is therefore not suitable for training the model. In this embodiment, each training sample is preprocessed through a threshold model to obtain corresponding preprocessing information, and then based on the preprocessing information, it is judged whether the training sample is a sample that is not suitable for model training of a nonlinear filtering model, and unsuitable "dirty data" is eliminated, thereby improving the quality of the training sample and improving the model training effect.

[0154] Corresponding to the signal processing methods of the above embodiments, FIG14 is a block diagram of the structure of a signal processing device provided by an embodiment of the present disclosure. The methods described in the above embodiments can be performed by this signal processing device, which can be implemented in software and / or hardware and integrated into an electronic device with certain data processing functions. Electronic devices may include, but are not limited to, mobile terminals with big data processing capabilities, as well as fixed terminals with big data processing capabilities, such as desktop computers and supercomputers.

[0155] For ease of illustration, only the parts related to the embodiment of the present disclosure are shown. Referring to FIG14 , the signal processing device 4 includes:

[0156] A collection unit 41 is configured to obtain a mixed sound signal and a reference sound signal, wherein the reference sound signal includes a sound signal sent by a remote terminal device to a local terminal device, and the mixed sound signal includes an ambient sound signal collected during the process of the local terminal device playing the reference sound signal;

[0157] a linear filtering unit 42 for filtering the linear echo component in the mixed sound signal based on the reference sound signal to obtain a first filtered signal;

[0158] The nonlinear filtering unit 43 is used to obtain an output signal through a pre-trained nonlinear filtering model and a first filtered signal, and send the output signal to a remote terminal device; wherein the output signal filters out the nonlinear echo component of a target energy level relative to the first filtered signal, and the target energy level is less than a preset energy level, and the preset energy level is determined based on the energy level of the human voice component in the first filtered signal.

[0159] According to one or more embodiments of the present disclosure, the linear filtering unit 42 is specifically used to: perform frequency domain transformation on the mixed sound signal and the reference sound signal to obtain a first time-frequency domain complex spectrum corresponding to the mixed sound signal and a second time-frequency domain complex spectrum corresponding to the reference sound signal; align the second time-frequency domain complex spectrum based on the first time-frequency domain complex spectrum to generate a third time-frequency domain complex spectrum; use the third time-frequency domain complex spectrum as a reference amount, and use the linear filter module to filter the first time-frequency domain complex spectrum to obtain a first filtered signal.

[0160] According to one or more embodiments of the present disclosure, when the nonlinear filtering unit 43 obtains an output signal through a pre-trained nonlinear filtering model and a first filtered signal, it is specifically used to: obtain the first time-frequency domain complex spectrum, the first amplitude spectrum, the second amplitude spectrum and the third amplitude spectrum corresponding to the third time-frequency domain complex number and the first filtered signal respectively; stack the first amplitude spectrum, the second amplitude spectrum and the third amplitude, and perform feature extraction on each channel based on the target characteristic frequency point to obtain a model input signal; input the model input signal into the nonlinear filtering model to obtain an output signal.

[0161] According to one or more embodiments of the present disclosure, the nonlinear filtering unit 43 is further used to: obtain the corresponding target frequency point number according to the preconfigured echo cancellation parameters; determine the target characteristic frequency point according to the characteristic frequency point of the target frequency point number in the low-frequency interval, wherein the frequency value corresponding to the upper limit of the low-frequency interval is less than the first frequency value, and the first frequency value is determined based on the human voice component in the mixed sound signal.

[0162] According to one or more embodiments of the present disclosure, the nonlinear filtering unit 43 is specifically used to: process the first filtered signal to obtain a model input signal based on a multi-channel amplitude spectrum; process the model input signal through a nonlinear filtering model to obtain an amplitude mask spectrum for the first filtered signal, and the amplitude mask spectrum is used to adjust the frequency amplitude of the target characteristic frequency point in the first filtered signal; use the amplitude mask spectrum to correct the first filtered signal to obtain a second filtered signal; perform time domain transformation on the second filtered signal to generate an output signal.

[0163] According to one or more embodiments of the present disclosure, after processing the model input signal through the nonlinear filtering model to obtain the amplitude mask spectrum for the first filtered signal, the nonlinear filtering unit 43 is further used to: perform smoothing correction on the amplitude mask spectrum within the target frequency range to obtain an optimized amplitude mask spectrum; when the nonlinear filtering unit 43 uses the amplitude mask spectrum to correct the first filtered signal to obtain the second filtered signal, it is specifically used to: use the optimized amplitude mask spectrum to correct the first filtered signal to obtain the second filtered signal.

[0164] According to one or more embodiments of the present disclosure, the target frequency interval includes a first frequency interval; when the nonlinear filtering unit 43 performs smoothing correction on the amplitude mask spectrum within the target frequency interval to obtain an optimized amplitude mask spectrum, it is specifically used to: obtain the first N amplitude mask maxima within the first frequency interval; obtain the amplitude mask mean based on the first N amplitude mask maxima; determine the characteristic frequency point to be corrected based on the amplitude mask mean, and the amplitude mask of the characteristic frequency point to be corrected is greater than the amplitude mask mean; perform weighted calculation on the amplitude mask of the characteristic frequency point to be corrected based on the amplitude mask mean to obtain the optimized amplitude mask spectrum corresponding to the first frequency interval, wherein the optimized amplitude mask spectrum of the characteristic frequency point to be corrected is smaller than the corresponding amplitude mask spectrum.

[0165] According to one or more embodiments of the present disclosure, the target frequency interval also includes a second frequency interval and a third frequency interval, the lower limit of the second frequency interval is greater than or equal to the upper limit of the first frequency interval; the upper limit of the third frequency interval is less than or equal to the lower limit of the first frequency interval; the nonlinear filtering unit 43 is also used to: obtain the first J amplitude mask maxima in the second frequency interval and the first K amplitude mask maxima in the third frequency interval; if the average of the first J amplitude mask maxima is greater than the first mask threshold and the average of the K amplitude mask maxima is less than the second mask threshold, then according to the amplitude mask mean corresponding to the first frequency interval, the optimized amplitude mask spectrum corresponding to the second frequency interval is obtained.

[0166] The acquisition unit 41, linear filtering unit 42 and nonlinear filtering unit 43 are connected in sequence. The signal processing device 3 provided in this embodiment can implement the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, which will not be repeated here.

[0167] Corresponding to the model training method of the above embodiment, Figure 15 is a structural block diagram of the model training device provided by the embodiment of the present disclosure. The method introduced in the above embodiment can be executed by the model training device, which can be implemented by software and / or hardware, and the device can be integrated into an electronic device with certain data processing functions. Among them, the electronic device may include but is not limited to mobile terminals with big data processing capabilities, as well as fixed terminals with big data processing capabilities such as desktop computers and supercomputers.

[0168] For ease of illustration, only the parts related to the embodiment of the present disclosure are shown. Referring to FIG15 , the model training device 5 includes:

[0169] The sample acquisition unit 51 is configured to acquire sample data, wherein the sample data includes a sample mixed sound signal and a corresponding sample reference sound signal, wherein the sample reference sound signal represents a sound signal sent by the remote terminal device to the local terminal device, and the sample mixed sound signal represents an ambient sound signal collected during the process of the local terminal device playing the sample reference sound signal;

[0170] The sample processing unit 52 is configured to filter the linear echo component in the sample mixed sound signal based on the sample reference sound signal to obtain a sample filtered signal;

[0171] The training unit 53 is used to train the initial filtering model using the sample filtered signal and the corresponding sample pure speech signal until the initial filtering model reaches a preset convergence condition, thereby obtaining a nonlinear filtering model, wherein the nonlinear filtering model is used to filter out the nonlinear echo component of the target energy level in the signal, the target energy level is less than the preset energy level, and the preset energy level is determined based on the energy level of the human voice component in the signal.

[0172] According to one or more embodiments of the present disclosure, the training unit 53 is specifically used to: obtain a combined loss function, wherein the combined loss function is composed of a first loss function and a second loss function, the first loss function is used to guide the initial filtering model to filter out the nonlinear echo components in the model input signal, and the second loss function is used to generate a penalty value for the filtering results that exceed a preset energy level when the initial filtering model filters out the nonlinear echo components in the model input signal, and the preset energy level is determined based on the energy level of the human voice component in the sample mixed sound signal; based on the combined loss function, the initial filtering model is trained using the sample filtered signal and the corresponding sample pure speech signal until the initial filtering model converges to a nonlinear filtering model.

[0173] According to one or more embodiments of the present disclosure, the training unit 53 trains the initial filtering model based on the combined loss function using the sample filter signal and the corresponding sample clean speech signal until the initial filtering model converges to a nonlinear filtering model, and is specifically used to: loop through the following steps until the convergence condition is reached: process the sample filter signal to obtain a model input signal based on the multi-channel amplitude spectrum; input the model input signal into the initial filtering model, and process it through the encoding module and the feature extraction module respectively to obtain a first frequency feature and a second frequency feature; process the first frequency feature through the mask prediction module to obtain an accurate amplitude mask spectrum, and process the second frequency feature through the auxiliary mask prediction module to obtain a rough amplitude mask spectrum; apply the accurate amplitude mask spectrum and the rough amplitude mask spectrum to the sample filter signal respectively to obtain the corresponding first sample filter signal and the second sample filter signal; calculate a first error value corresponding to the first sample filter signal and a second error value corresponding to the second sample filter signal based on the complex spectrum of the sample clean speech signal using the first loss function and the second loss function respectively, and generate a comprehensive error value based on the first error value and the second error value; update the model parameters of the initial filtering model based on the comprehensive error value, and update to the next group of sample filter signals.

[0174] According to one or more embodiments of the present disclosure, the initial filtering model includes an encoding module, a feature extraction module, a mask prediction module and an auxiliary mask prediction module, wherein the encoding module is used to encode the model input signal to obtain an encoded signal; the feature extraction module includes multiple feature extraction layers connected in series, which are used to gradually extract the nonlinear echo components in the encoded signal based on the feature extraction layers; the mask prediction module is used to predict an accurate amplitude mask spectrum based on a first spectral feature, and the accurate amplitude mask spectrum is used to adjust the frequency amplitude of the target characteristic frequency point in the model input signal to filter out the nonlinear echo components in the model input signal, and the first spectral feature is a spectral feature extracted by the terminal feature extraction layer of the feature extraction module; the auxiliary mask prediction module is used to predict a rough amplitude mask spectrum based on a second spectral feature, and the rough amplitude mask spectrum is used to calculate the penalty value of the filtering result, and the second spectral feature is a characteristic frequency extracted by the non-terminal feature extraction layer of the feature extraction module.

[0175] According to one or more embodiments of the present disclosure, before acquiring sample data, the sample acquisition unit 51 is further used to: acquire original sample data, which includes at least two training samples; process the original sample data based on a pre-trained echo filter model to obtain preprocessing information corresponding to each training sample, where the preprocessing information is a sound signal with the echo component in the training sample filtered out; determine at least one interference sample based on the energy amplitude of the preprocessing information, and remove the interference sample from the original sample data to obtain sample data, wherein the energy amplitude of the interference sample is greater than a preset test threshold.

[0176] FIG16 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. As shown in FIG16 , the electronic device 6 includes:

[0177] A processor 61, and a memory 62 communicatively connected to the processor 61;

[0178] Memory 62 stores computer-executable instructions;

[0179] The processor 61 executes the computer-executable instructions stored in the memory 62 to implement the signal processing method in any of the embodiments shown in Figures 2 to 8, or to implement the model training method in any of the embodiments shown in Figures 9 to 13.

[0180] Optionally, the processor 61 and the memory 62 are connected via a bus 63 .

[0181] The relevant explanations can be understood by referring to the relevant descriptions and effects corresponding to the steps in the embodiments corresponding to Figures 2 to 13, and no further details will be given here.

[0182] An embodiment of the present disclosure provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, they are used to implement the signal processing method provided in any one of the embodiments corresponding to Figures 2 to 8 of the present disclosure, or to implement the model training method provided in any one of the embodiments corresponding to Figures 9 to 13.

[0183] An embodiment of the present disclosure provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the signal processing method provided in any one of the embodiments corresponding to Figures 2 to 8 of the present disclosure, or is used to implement the model training method provided in any one of the embodiments corresponding to Figures 9 to 13.

[0184] In order to implement the above embodiment, the embodiment of the present disclosure further provides an electronic device.

[0185] Referring to FIG17 , a schematic diagram of the structure of an electronic device 900 suitable for implementing an embodiment of the present disclosure is shown. The electronic device 900 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG17 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.

[0186] As shown in FIG17 , the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0187] Typically, the following devices can be connected to the I / O interface 905: input devices 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 908 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 909. The communication device 909 can allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although FIG17 shows an electronic device 900 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0188] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0189] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0190] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0191] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.

[0192] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).

[0193] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0194] The units or modules involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit or module does not, in some cases, limit the unit itself.

[0195] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0196] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0197] In a first aspect, according to one or more embodiments of the present disclosure, a signal processing method is provided, including:

[0198] Acquire a mixed sound signal and a reference sound signal, wherein the reference sound signal includes a sound signal sent by a remote terminal device to a near-end terminal device, and the mixed sound signal includes an ambient sound signal collected during the process of the near-end terminal device playing the reference sound signal; based on the reference sound signal, filter the linear echo component in the mixed sound signal to obtain a first filtered signal; obtain an output signal through a pre-trained nonlinear filtering model and the first filtered signal, and send the output signal to the remote terminal device; wherein, the output signal filters out the nonlinear echo component of a target energy level relative to the first filtered signal, and the target energy level is less than a preset energy level, and the preset energy level is determined based on the energy level of the human voice component in the first filtered signal.

[0199] According to one or more embodiments of the present disclosure, filtering the linear echo component in the mixed sound signal based on the reference sound signal to obtain a first filtered signal includes: performing frequency domain transform on the mixed sound signal and the reference sound signal to obtain a first time-frequency domain complex spectrum corresponding to the mixed sound signal and a second time-frequency domain complex spectrum corresponding to the reference sound signal; aligning the second time-frequency domain complex spectrum based on the first time-frequency domain complex spectrum to generate a third time-frequency domain complex spectrum; and filtering the first time-frequency domain complex spectrum using a linear filter module with the third time-frequency domain complex spectrum processed as a reference to obtain the first filtered signal.

[0200] According to one or more embodiments of the present disclosure, the output signal is obtained by using a pre-trained nonlinear filtering model and the first filtered signal, including: obtaining the first amplitude spectrum, the second amplitude spectrum and the third amplitude spectrum corresponding to the first time-frequency domain complex spectrum, the third time-frequency domain complex number and the first filtered signal respectively; channel stacking the first amplitude spectrum, the second amplitude spectrum and the third amplitude, and performing feature extraction on each channel based on the target characteristic frequency point to obtain a model input signal; inputting the model input signal into the nonlinear filtering model to obtain the output signal.

[0201] According to one or more embodiments of the present disclosure, the method further includes: obtaining a corresponding target frequency number based on preconfigured echo cancellation parameters; determining the target characteristic frequency point based on the characteristic frequency points of the target frequency number in the low-frequency interval, wherein the frequency value corresponding to the upper limit of the low-frequency interval is less than a first frequency value, and the first frequency value is determined based on the human voice component in the mixed sound signal.

[0202] According to one or more embodiments of the present disclosure, the output signal is obtained by using a pre-trained nonlinear filtering model and the first filtered signal, including: processing the first filtered signal to obtain a model input signal based on a multi-channel amplitude spectrum; processing the model input signal through the nonlinear filtering model to obtain an amplitude mask spectrum for the first filtered signal, and the amplitude mask spectrum is used to adjust the frequency amplitude of the target characteristic frequency point in the first filtered signal; using the amplitude mask spectrum to correct the first filtered signal to obtain a second filtered signal; and performing time domain transformation on the second filtered signal to generate the output signal.

[0203] According to one or more embodiments of the present disclosure, after processing the model input signal through the nonlinear filtering model to obtain the amplitude mask spectrum for the first filtered signal, it also includes: performing smoothing correction on the amplitude mask spectrum within the target frequency range to obtain an optimized amplitude mask spectrum; correcting the first filtered signal using the amplitude mask spectrum to obtain the second filtered signal includes: correcting the first filtered signal using the optimized amplitude mask spectrum to obtain the second filtered signal.

[0204] According to one or more embodiments of the present disclosure, the target frequency interval includes a first frequency interval; the amplitude mask spectrum is smoothed and corrected within the target frequency interval to obtain an optimized amplitude mask spectrum, including: obtaining the first N amplitude mask maximum values ​​within the first frequency interval; obtaining the amplitude mask mean based on the first N amplitude mask maximum values; determining the characteristic frequency point to be corrected based on the amplitude mask mean, the amplitude mask of the characteristic frequency point to be corrected is greater than the amplitude mask mean; performing weighted calculation on the amplitude mask of the characteristic frequency point to be corrected based on the amplitude mask mean to obtain the optimized amplitude mask spectrum corresponding to the first frequency interval, wherein the optimized amplitude mask spectrum of the characteristic frequency point to be corrected is smaller than the corresponding amplitude mask spectrum.

[0205] According to one or more embodiments of the present disclosure, the target frequency interval also includes a second frequency interval and a third frequency interval, the lower limit of the second frequency interval is greater than or equal to the upper limit of the first frequency interval; the upper limit of the third frequency interval is less than or equal to the lower limit of the first frequency interval; and it also includes: obtaining the first J amplitude mask maximum values ​​in the second frequency interval and the first K amplitude mask maximum values ​​in the third frequency interval; if the average of the first J amplitude mask maximum values ​​is greater than the first mask threshold and the average of the K amplitude mask maximum values ​​is less than the second mask threshold, then according to the amplitude mask mean corresponding to the first frequency interval, the optimized amplitude mask spectrum corresponding to the second frequency interval is obtained.

[0206] In a second aspect, according to one or more embodiments of the present disclosure, a model training method is provided, comprising:

[0207] Acquire sample data, the sample data including a sample mixed sound signal and a corresponding sample reference sound signal, the sample reference sound signal representing a sound signal sent from a far-end terminal device to a near-end terminal device, and the sample mixed sound signal representing an ambient sound signal collected during the process of the near-end terminal device playing the sample reference sound signal; based on the sample reference sound signal, filter the linear echo component in the sample mixed sound signal to obtain a sample filtered signal; use the sample filtered signal and the corresponding sample pure speech signal to train an initial filtering model until the initial filtering model reaches a preset convergence condition to obtain a nonlinear filtering model, wherein the nonlinear filtering model is used to filter out the nonlinear echo component of a target energy level in the sound signal, the target energy level is less than a preset energy level, and the preset energy level is determined based on the energy level of the human voice component in the sound signal.

[0208] According to one or more embodiments of the present disclosure, the sample filter signal and the corresponding sample clean speech signal are used to train the initial filter model until the initial filter model reaches a preset convergence condition to obtain a nonlinear filter model, including: obtaining a combined loss function, wherein the combined loss function is composed of a first loss function and a second loss function, the first loss function is used to guide the initial filter model to filter out the nonlinear echo component in the model input signal, and the second loss function is used to generate a penalty for the filtering result exceeding a preset energy level when the initial filter model filters out the nonlinear echo component in the model input signal, and the preset energy level is determined based on the energy level of the human voice component in the sample mixed sound signal; based on the combined loss function, the sample filter signal and the corresponding sample clean speech signal are used to train the initial filter model until the initial filter model converges to the nonlinear filter model.

[0209] According to one or more embodiments of the present disclosure, the initial filtering model is trained based on the combined loss function using the sample filtering signal and the corresponding sample clean speech signal until the initial filtering model converges to the nonlinear filtering model, including: looping the following steps until the convergence condition is reached: processing the sample filtering signal to obtain a model input signal based on a multi-channel amplitude spectrum; inputting the model input signal into the initial filtering model, and processing it through an encoding module and a feature extraction module respectively to obtain a first frequency feature and a second frequency feature; processing the first frequency feature through a mask prediction module to obtain an accurate amplitude mask spectrum, and The code prediction module processes the second frequency feature to obtain a rough amplitude mask spectrum; the precise amplitude mask spectrum and the rough amplitude mask spectrum are respectively applied to the sample filtered signal to obtain corresponding first sample filtered signals and second sample filtered signals; based on the complex spectrum of the sample clean speech signal, the first loss function and the second loss function are respectively used to calculate the first error value corresponding to the first sample filtered signal and the second error value corresponding to the second sample filtered signal, and generate a comprehensive error value based on the first error value and the second error value; based on the comprehensive error value, the model parameters of the initial filtering model are updated, and updated to the next group of sample filtered signals.

[0210] According to one or more embodiments of the present disclosure, the initial filtering model includes an encoding module, a feature extraction module, a mask prediction module and an auxiliary mask prediction module, wherein the encoding module is used to encode the model input signal to obtain an encoded signal; the feature extraction module includes multiple feature extraction layers connected in series, which are used to gradually extract the nonlinear echo components in the encoded signal based on the feature extraction layers; the mask prediction module is used to predict an accurate amplitude mask spectrum based on a first spectral feature, and the accurate amplitude mask spectrum is used to adjust the frequency amplitude of the target characteristic frequency point in the model input signal to filter out the nonlinear echo components in the model input signal, and the first spectral feature is the spectral feature extracted by the terminal feature extraction layer of the feature extraction module; the auxiliary mask prediction module is used to predict a rough amplitude mask spectrum based on a second spectral feature, and the rough amplitude mask spectrum is used to calculate the penalty value of the filtering result, and the second spectral feature is the characteristic frequency extracted by the non-terminal feature extraction layer of the feature extraction module.

[0211] According to one or more embodiments of the present disclosure, before obtaining sample data, it also includes: obtaining original sample data, wherein the original sample data includes at least two training samples; processing the original sample data based on a pre-trained echo filtering model to obtain preprocessing information corresponding to each training sample, wherein the preprocessing information is a sound signal obtained by filtering out the echo component in the training sample; determining at least one interference sample based on the energy amplitude of the preprocessing information, and removing the interference sample from the original sample data to obtain the sample data, wherein the energy amplitude of the interference sample is greater than a preset test threshold.

[0212] In a third aspect, according to one or more embodiments of the present disclosure, a signal processing device is provided, including:

[0213] a collection unit, configured to obtain a mixed sound signal and a reference sound signal, wherein the reference sound signal comprises a sound signal sent by a remote terminal device to a near-end terminal device, and the mixed sound signal comprises an ambient sound signal collected during the process of the near-end terminal device playing the reference sound signal;

[0214] a linear filtering unit, configured to filter the linear echo component in the mixed sound signal based on the reference sound signal to obtain a first filtered signal;

[0215] A nonlinear filtering unit is used to obtain an output signal through a pre-trained nonlinear filtering model and the first filtered signal, and send the output signal to the remote terminal device; wherein, the output signal filters out the nonlinear echo component of the target energy level relative to the first filtered signal, and the target energy level is less than a preset energy level, and the preset energy level is determined based on the energy level of the human voice component in the first filtered signal.

[0216] According to one or more embodiments of the present disclosure, the linear filtering unit is specifically used to: perform frequency domain transformation on the mixed sound signal and the reference sound signal to obtain a first time-frequency domain complex spectrum corresponding to the mixed sound signal and a second time-frequency domain complex spectrum corresponding to the reference sound signal; align the second time-frequency domain complex spectrum based on the first time-frequency domain complex spectrum to generate a third time-frequency domain complex spectrum; and use the third time-frequency domain complex spectrum as a reference quantity to filter the first time-frequency domain complex spectrum using a linear filter module to obtain the first filtered signal.

[0217] According to one or more embodiments of the present disclosure, when the nonlinear filtering unit obtains the output signal through the pre-trained nonlinear filtering model and the first filtered signal, it is specifically used to: obtain the first amplitude spectrum, the second amplitude spectrum and the third amplitude spectrum corresponding to the first time-frequency domain complex spectrum, the third time-frequency domain complex number and the first filtered signal respectively; stack the first amplitude spectrum, the second amplitude spectrum and the third amplitude, and perform feature extraction on each channel based on the target characteristic frequency point to obtain a model input signal; input the model input signal into the nonlinear filtering model to obtain the output signal.

[0218] According to one or more embodiments of the present disclosure, the nonlinear filtering unit is further used to: obtain the corresponding target frequency point number based on the preconfigured echo cancellation parameters; determine the target characteristic frequency point based on the characteristic frequency points of the target frequency point number in the low-frequency interval, wherein the frequency value corresponding to the upper limit of the low-frequency interval is less than the first frequency value, and the first frequency value is determined based on the human voice component in the mixed sound signal.

[0219] According to one or more embodiments of the present disclosure, the nonlinear filtering unit is specifically used to: process the first filtered signal to obtain a model input signal based on a multi-channel amplitude spectrum; process the model input signal through the nonlinear filtering model to obtain an amplitude mask spectrum for the first filtered signal, and the amplitude mask spectrum is used to adjust the frequency amplitude of the target characteristic frequency point in the first filtered signal; use the amplitude mask spectrum to correct the first filtered signal to obtain a second filtered signal; perform time domain transformation on the second filtered signal to generate the output signal.

[0220] According to one or more embodiments of the present disclosure, after the model input signal is processed by the nonlinear filtering model to obtain the amplitude mask spectrum for the first filtered signal, the nonlinear filtering unit is further used to: perform smoothing correction on the amplitude mask spectrum within the target frequency range to obtain an optimized amplitude mask spectrum; when the nonlinear filtering unit uses the amplitude mask spectrum to correct the first filtered signal to obtain the second filtered signal, the nonlinear filtering unit is specifically used to: use the optimized amplitude mask spectrum to correct the first filtered signal to obtain the second filtered signal.

[0221] According to one or more embodiments of the present disclosure, the target frequency interval includes a first frequency interval; when the nonlinear filtering unit performs smoothing correction on the amplitude mask spectrum within the target frequency interval to obtain an optimized amplitude mask spectrum, it is specifically used to: obtain the first N amplitude mask maxima in the first frequency interval; obtain the amplitude mask mean based on the first N amplitude mask maxima; determine the characteristic frequency point to be corrected based on the amplitude mask mean, and the amplitude mask of the characteristic frequency point to be corrected is greater than the amplitude mask mean; perform weighted calculation on the amplitude mask of the characteristic frequency point to be corrected based on the amplitude mask mean to obtain the optimized amplitude mask spectrum corresponding to the first frequency interval, wherein the optimized amplitude mask spectrum of the characteristic frequency point to be corrected is smaller than the corresponding amplitude mask spectrum.

[0222] According to one or more embodiments of the present disclosure, the target frequency interval also includes a second frequency interval and a third frequency interval, the lower limit of the second frequency interval is greater than or equal to the upper limit of the first frequency interval; the upper limit of the third frequency interval is less than or equal to the lower limit of the first frequency interval; the nonlinear filtering unit is also used to: obtain the first J amplitude mask maxima in the second frequency interval and the first K amplitude mask maxima in the third frequency interval; if the average of the first J amplitude mask maxima is greater than the first mask threshold and the average of the K amplitude mask maxima is less than the second mask threshold, then according to the amplitude mask mean corresponding to the first frequency interval, the optimized amplitude mask spectrum corresponding to the second frequency interval is obtained.

[0223] In a fourth aspect, according to one or more embodiments of the present disclosure, a model training device is provided, comprising:

[0224] a sample acquisition unit, configured to acquire sample data, the sample data including a sample mixed sound signal and a corresponding sample reference sound signal, the sample reference sound signal representing a sound signal sent by a remote terminal device to a near-end terminal device, and the sample mixed sound signal representing an ambient sound signal collected during the process of the near-end terminal device playing the sample reference sound signal;

[0225] a sample processing unit, configured to filter the linear echo component in the sample mixed sound signal based on the sample reference sound signal to obtain a sample filtered signal;

[0226] A training unit is used to train an initial filtering model using the sample filtered signal and the corresponding sample pure speech signal until the initial filtering model reaches a preset convergence condition, thereby obtaining a nonlinear filtering model, wherein the nonlinear filtering model is used to filter out nonlinear echo components of a target energy level in the signal, the target energy level being less than a preset energy level, and the preset energy level being determined based on the energy level of the human voice component in the signal.

[0227] According to one or more embodiments of the present disclosure, the training unit is specifically used to: obtain a combined loss function, wherein the combined loss function is composed of a first loss function and a second loss function, the first loss function is used to guide the initial filtering model to filter out nonlinear echo components in the model input signal, and the second loss function is used to generate a penalty for filtering results that exceed a preset energy level when the initial filtering model filters out nonlinear echo components in the model input signal, and the preset energy level is determined based on the energy level of the human voice component in the sample mixed sound signal; based on the combined loss function, the initial filtering model is trained using the sample filtered signal and the corresponding sample pure speech signal until the initial filtering model converges to the nonlinear filtering model.

[0228] According to one or more embodiments of the present disclosure, the training unit trains the initial filtering model based on the combined loss function and using the sample filtering signal and the corresponding sample clean speech signal until the initial filtering model converges to the nonlinear filtering model, and is specifically used to: loop through the following steps until the convergence condition is reached: process the sample filtering signal to obtain a model input signal based on a multi-channel amplitude spectrum; input the model input signal into the initial filtering model, and process it through the encoding module and the feature extraction module respectively to obtain a first frequency feature and a second frequency feature; process the first frequency feature through the mask prediction module to obtain an accurate amplitude mask spectrum, and pass The second frequency feature is processed by an auxiliary mask prediction module to obtain a rough amplitude mask spectrum; the precise amplitude mask spectrum and the rough amplitude mask spectrum are respectively applied to the sample filtered signal to obtain corresponding first sample filtered signals and second sample filtered signals; based on the complex spectrum of the sample pure speech signal, the first loss function and the second loss function are respectively used to calculate the first error value corresponding to the first sample filtered signal and the second error value corresponding to the second sample filtered signal, and a comprehensive error value is generated according to the first error value and the second error value; the model parameters of the initial filtering model are updated based on the comprehensive error value, and updated to the next group of sample filtered signals.

[0229] According to one or more embodiments of the present disclosure, the initial filtering model includes an encoding module, a feature extraction module, a mask prediction module and an auxiliary mask prediction module, wherein the encoding module is used to encode the model input signal to obtain an encoded signal; the feature extraction module includes multiple feature extraction layers connected in series, which are used to gradually extract the nonlinear echo components in the encoded signal based on the feature extraction layers; the mask prediction module is used to predict an accurate amplitude mask spectrum based on a first spectral feature, and the accurate amplitude mask spectrum is used to adjust the frequency amplitude of the target characteristic frequency point in the model input signal to filter out the nonlinear echo components in the model input signal, and the first spectral feature is the spectral feature extracted by the terminal feature extraction layer of the feature extraction module; the auxiliary mask prediction module is used to predict a rough amplitude mask spectrum based on a second spectral feature, and the rough amplitude mask spectrum is used to calculate the penalty value of the filtering result, and the second spectral feature is the characteristic frequency extracted by the non-terminal feature extraction layer of the feature extraction module.

[0230] According to one or more embodiments of the present disclosure, before acquiring sample data, the sample acquisition unit is further used to: acquire original sample data, wherein the original sample data includes at least two training samples; process the original sample data based on a pre-trained echo filter model to obtain preprocessing information corresponding to each of the training samples, wherein the preprocessing information is a sound signal obtained by filtering out the echo component in the training sample; determine at least one interference sample based on the energy amplitude of the preprocessing information, and remove the interference sample from the original sample data to obtain the sample data, wherein the energy amplitude of the interference sample is greater than a preset test threshold.

[0231] In a fifth aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: at least one processor and a memory;

[0232] The memory stores computer-executable instructions;

[0233] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the signal processing method described in the first aspect and various possible designs of the first aspect, or executes the model training method described in the second aspect and various possible designs of the second aspect.

[0234] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the signal processing method described in the first aspect and various possible designs of the first aspect is implemented, or the model training method described in the second aspect and various possible designs of the second aspect is implemented.

[0235] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the signal processing method described in the first aspect and various possible designs of the first aspect, or implements the model training method described in the second aspect and various possible designs of the second aspect.

[0236] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0237] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0238] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A signal processing method, comprising: Acquire a mixed sound signal and a reference sound signal, wherein the reference sound signal includes a sound signal sent by a remote terminal device to a local terminal device, and the mixed sound signal includes an ambient sound signal collected during the process of the local terminal device playing the reference sound signal; filtering the linear echo component in the mixed sound signal based on the reference sound signal to obtain a first filtered signal; Obtaining an output signal through a pre-trained nonlinear filtering model and the first filtered signal, and sending the output signal to the remote terminal device; The output signal filters out nonlinear echo components of a target energy level relative to the first filtered signal, and the target energy level is less than a preset energy level, which is determined based on the energy level of the human voice component in the first filtered signal.

2. The method according to claim 1, wherein The filtering of the linear echo component in the mixed sound signal based on the reference sound signal to obtain a first filtered signal includes: Performing frequency domain transformation on the mixed sound signal and the reference sound signal to obtain a first time-frequency domain complex spectrum corresponding to the mixed sound signal and a second time-frequency domain complex spectrum corresponding to the reference sound signal; Aligning the second time-frequency domain complex spectrum based on the first time-frequency domain complex spectrum to generate a third time-frequency domain complex spectrum; The first time-frequency domain complex spectrum is filtered using a linear filter module with the third time-frequency domain complex spectrum processed as a reference to obtain the first filtered signal.

3. The method according to claim 2, wherein: The method of obtaining an output signal by using a pre-trained nonlinear filtering model and the first filtering signal includes: Obtain a first amplitude spectrum, a second amplitude spectrum, and a third amplitude spectrum corresponding to the first time-frequency domain complex spectrum, the third time-frequency domain complex number, and the first filtered signal, respectively; Channel stacking is performed on the first amplitude spectrum, the second amplitude spectrum, and the third amplitude spectrum, and feature extraction is performed on each channel based on a target characteristic frequency point to obtain a model input signal; The model input signal is input into the nonlinear filtering model to obtain the output signal.

4. The method according to claim 3, further comprising: According to the pre-configured echo cancellation parameters, the corresponding target frequency points are obtained; The target characteristic frequency point is determined based on the characteristic frequency points of the target frequency point quantity in the low-frequency interval, wherein a frequency value corresponding to an upper limit of the low-frequency interval is less than a first frequency value, and the first frequency value is determined based on a human voice component in the mixed sound signal.

5. The method according to claim 1, wherein The method of obtaining an output signal by using a pre-trained nonlinear filtering model and the first filtering signal includes: processing the first filtered signal to obtain a model input signal based on a multi-channel amplitude spectrum; Processing the model input signal through the nonlinear filtering model to obtain an amplitude mask spectrum for the first filtered signal, wherein the amplitude mask spectrum is used to adjust the frequency amplitude of a target characteristic frequency point in the first filtered signal; modifying the first filtered signal using the amplitude mask spectrum to obtain a second filtered signal; Performing a time domain transform on the second filtered signal to generate the output signal.

6. The method according to claim 5, wherein: After processing the model input signal through the nonlinear filtering model to obtain an amplitude mask spectrum for the first filtered signal, the method further includes: For the amplitude mask spectrum, smoothing correction is performed within the target frequency range to obtain an optimized amplitude mask spectrum; The modifying the first filtered signal by using the amplitude mask spectrum to obtain a second filtered signal includes: The first filtered signal is modified using the optimized amplitude mask spectrum to obtain a second filtered signal.

7. The method according to claim 6, wherein: The target frequency interval includes a first frequency interval; and performing smoothing correction on the amplitude mask spectrum within the target frequency interval to obtain an optimized amplitude mask spectrum includes: Get the first N maximum values ​​of the amplitude mask in the first frequency interval; Obtaining an amplitude mask mean value according to the first N amplitude mask maximum values; Determining a characteristic frequency point to be corrected according to the amplitude mask mean, wherein the amplitude mask of the characteristic frequency point to be corrected is greater than the amplitude mask mean; The amplitude mask of the characteristic frequency point to be corrected is weightedly calculated based on the amplitude mask mean to obtain an optimized amplitude mask spectrum corresponding to the first frequency interval, wherein the optimized amplitude mask spectrum of the characteristic frequency point to be corrected is smaller than the corresponding amplitude mask spectrum.

8. The method according to claim 7, wherein: The target frequency interval further includes a second frequency interval and a third frequency interval, and the lower limit of the second frequency interval is greater than or equal to the upper limit of the first frequency interval; The upper limit of the third frequency interval is less than or equal to the lower limit of the first frequency interval; The method further comprises: Obtaining the first J maximum amplitude mask values ​​in the second frequency interval and the first K maximum amplitude mask values ​​in the third frequency interval; If the average of the first J amplitude mask maxima is greater than the first mask threshold and the average of the K amplitude mask maxima is less than the second mask threshold, then the optimized amplitude mask spectrum corresponding to the second frequency interval is obtained according to the amplitude mask average corresponding to the first frequency interval.

9. A model training method comprising: Acquire sample data, where the sample data includes a sample mixed sound signal and a corresponding sample reference sound signal, where the sample reference sound signal represents a sound signal sent by a remote terminal device to a local terminal device, and the sample mixed sound signal represents an ambient sound signal collected during the process of the local terminal device playing the sample reference sound signal; Based on the sample reference sound signal, filtering the linear echo component in the sample mixed sound signal to obtain a sample filtered signal; The initial filtering model is trained using the sample filtering signal and the corresponding sample pure speech signal until the initial filtering model reaches a preset convergence condition to obtain a nonlinear filtering model, wherein the nonlinear filtering model is used to filter out the nonlinear echo component of the target energy level in the sound signal, and the target energy level is less than the preset energy level, and the preset energy level is determined based on the energy level of the human voice component in the sound signal.

10. The method according to claim 9, wherein: The initial filter model is trained using the sample filter signal and the corresponding sample clean speech signal until the initial filter model reaches a preset convergence condition, thereby obtaining a nonlinear filter model, including: Obtaining a combined loss function, wherein the combined loss function is composed of a first loss function and a second loss function, the first loss function being used to guide the initial filtering model to filter out nonlinear echo components in the model input signal, and the second loss function being used to generate a penalty value for a filtering result exceeding a preset energy level when the initial filtering model filters out the nonlinear echo components in the model input signal, the preset energy level being determined based on an energy level of a human voice component in the sample mixed sound signal; Based on the combined loss function, the initial filter model is trained using the sample filtered signal and the corresponding sample clean speech signal until the initial filter model converges to the nonlinear filter model.

11. The method according to claim 10, wherein: The method of training the initial filter model based on the combined loss function and using the sample filter signal and the corresponding sample clean speech signal until the initial filter model converges to the nonlinear filter model includes: The following steps are executed repeatedly until the convergence condition is reached: Processing the sample filtered signal to obtain a model input signal based on a multi-channel amplitude spectrum; Inputting the model input signal into the initial filtering model, and processing it through the encoding module and the feature extraction module respectively to obtain a first frequency feature and a second frequency feature; Processing the first frequency feature through a mask prediction module to obtain a precise amplitude mask spectrum, and processing the second frequency feature through an auxiliary mask prediction module to obtain a rough amplitude mask spectrum; Applying the precise amplitude mask spectrum and the rough amplitude mask spectrum to the sample filtered signal respectively to obtain a corresponding first sample filtered signal and a second sample filtered signal; Based on the complex spectrum of the sample clean speech signal, using the first loss function and the second loss function respectively, calculate a first error value corresponding to the first sample filtered signal and a second error value corresponding to the second sample filtered signal, and generate a comprehensive error value according to the first error value and the second error value; The model parameters of the initial filtering model are updated based on the comprehensive error value and updated to the next set of sample filtering signals.

12. The method according to any one of claims 9 to 11, wherein: The initial filtering model includes an encoding module, a feature extraction module, a mask prediction module and an auxiliary mask prediction module, wherein: The encoding module is used to encode the model input signal to obtain an encoded signal; The feature extraction module includes a plurality of feature extraction layers connected in series, and is used for gradually extracting nonlinear echo components in the coded signal based on the feature extraction layers; The mask prediction module is configured to predict an accurate amplitude mask spectrum based on a first spectrum feature, wherein the accurate amplitude mask spectrum is configured to adjust the frequency amplitude of a target characteristic frequency point in the model input signal to filter out nonlinear echo components in the model input signal, wherein the first spectrum feature is a spectrum feature extracted by a terminal feature extraction layer of the feature extraction module; The auxiliary mask prediction module is used to predict a rough amplitude mask spectrum based on a second spectrum feature, and the rough amplitude mask spectrum is used to calculate a penalty value of a filtering result. The second spectrum feature is a characteristic frequency extracted by a non-terminal feature extraction layer of the feature extraction module.

13. The method according to any one of claims 9 to 12, wherein: Before acquiring the sample data, the method further includes: Acquire original sample data, where the original sample data includes at least two training samples; Based on the pre-trained echo filter model, the original sample data is processed to obtain pre-processing information corresponding to each training sample, wherein the pre-processing information is a sound signal after the echo component in the training sample is filtered out; At least one interference sample is determined according to the energy amplitude of the preprocessing information, and the interference sample is removed from the original sample data to obtain the sample data, wherein the energy amplitude of the interference sample is greater than a preset test threshold.

14. A signal processing device comprising: a collection unit configured to obtain a mixed sound signal and a reference sound signal, wherein the reference sound signal includes a sound signal sent by a remote terminal device to a near-end terminal device, and the mixed sound signal includes an ambient sound signal collected during the process of the near-end terminal device playing the reference sound signal; a linear filtering unit configured to filter the linear echo component in the mixed sound signal based on the reference sound signal to obtain a first filtered signal; a nonlinear filtering unit configured to obtain an output signal through a pre-trained nonlinear filtering model and the first filtered signal, and send the output signal to the remote terminal device; The output signal filters out nonlinear echo components of a target energy level relative to the first filtered signal, and the target energy level is less than a preset energy level, which is determined based on the energy level of the human voice component in the first filtered signal.

15. A model training device comprising: a sample acquisition unit configured to acquire sample data, the sample data including a sample mixed sound signal and a corresponding sample reference sound signal, the sample reference sound signal representing a sound signal sent by a remote terminal device to a near-end terminal device, and the sample mixed sound signal representing an ambient sound signal collected during the process of the near-end terminal device playing the sample reference sound signal; a sample processing unit configured to filter the linear echo component in the sample mixed sound signal based on the sample reference sound signal to obtain a sample filtered signal; The training unit is configured to use the sample filtered signal and the corresponding sample pure speech signal to train the initial filtering model until the initial filtering model reaches a preset convergence condition, thereby obtaining a nonlinear filtering model, wherein the nonlinear filtering model is used to filter out the nonlinear echo component of the target energy level in the signal, the target energy level is less than a preset energy level, and the preset energy level is determined based on the energy level of the human voice component in the signal.

16. An electronic device comprising: processor and memory, wherein The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the signal processing method according to any one of claims 1 to 13.

17. A computer-readable storage medium storing computer-executable instructions, wherein: When a processor executes the computer-executable instructions, the signal processing method according to any one of claims 1 to 13 is implemented.

18. A computer program product comprising a computer program, wherein When the computer program is executed by a processor, the signal processing method according to any one of claims 1 to 13 is implemented.

Citation Information

Patent Citations

  • Signal processing method and device, model training method and device, equipment and storage medium

    CN120727022A

  • Voice call echo processing method and device, equipment and storage medium

    CN112217948A

  • Method and device for eliminating echo signal, computing equipment and storage medium

    CN113763977A

  • Echo cancellation method and device, electronic equipment and storage medium

    CN115602184A

  • Audio processing method and device, storage medium and electronic equipment

    CN117612548A

Cited By

  • Acoustic model training method and device

    CN121687118A