Audio signal processing method and apparatus, storage medium, and electronic device

By using the target neural network model to obtain phase estimation information in audio signal processing, and extracting the target audio signal in combination with spectrum features, the problems of large calculation volume and low processing efficiency in the prior art are solved, and more efficient and accurate audio signal processing is achieved.

WO2025130212A1PCT designated stage expired Publication Date: 2025-06-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
PCT/CN2024/120318
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-09-23
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The prior art has a large amount of calculation when processing noise-free audio signals, resulting in low processing efficiency and fails to effectively solve the efficiency problem of audio signal processing methods.

Method used

By acquiring the first spectral characteristics of the noisy audio signal, the target neural network model is input to obtain reference phase estimation information, and the second spectral characteristics are determined in combination with the phase sine value and the cosine value, thereby extracting the target audio signal.

Benefits of technology

This improves the accuracy and efficiency of audio signal processing, and avoids the increase in the computational volume of using complex neural network models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024120318_26062025_PF_FP_ABST
    Figure CN2024120318_26062025_PF_FP_ABST
Patent Text Reader

Abstract

An audio signal processing method, performed by an electronic device, the method comprising: obtaining a first spectral feature of a noisy audio signal, wherein the noisy audio signal comprises a reference audio signal and a target audio signal to be extracted, and the reference audio signal and the target audio signal are each a noise signal or a clean speech signal in the noisy audio signal (S202); inputting the first spectral feature into a target neural network model to obtain reference phase estimation information matching the noisy audio signal, wherein the target neural network model is obtained by performing pretraining on the basis of sample noisy audio signals and sample target audio signals, the reference phase estimation information is used to indicate phase differences between the target audio signal and the noisy audio signal across multiple frequency components, and the sample target audio signals are either sample noise signals or sample clean speech signals in the sample noisy audio signals (S204); on the basis of a first phase sine value and a first phase cosine value indicated by the first spectral feature, and a reference phase sine value and a reference phase cosine value indicated by the reference phase estimation information, determining a second phase sine value and a second phase cosine value used to determine a second spectral feature (S206); and on the basis of the second phase sine value and the second phase cosine value, determining the second spectral feature, and, on the basis of the second spectral feature, determining the target audio signal (S208).
Need to check novelty before this filing date? Find Prior Art

Description

Audio signal processing method and device, storage medium and electronic device

[0001] Related applications

[0002] This application claims priority to Chinese patent application number 2023117580234, filed on December 19, 2023, entitled “Audio Signal Processing Method and Device, Storage Medium and Electronic Device,” the entire text of which is hereby incorporated by reference. Technical Field

[0003] The present application relates to the field of computers, and in particular to a method and device for processing audio signals, a storage medium, and an electronic device. Background Art

[0004] With the continuous innovation of audio signal processing technology, speech enhancement and noise reduction technology has flourished. Currently, speech enhancement technology is widely used in scenarios such as calls, video conferencing, smart speakers, and voice recognition front-ends, bringing great benefits to people's production and life.

[0005] To improve the accuracy of analysis results, related techniques employ complex neural networks to model and analyze noisy audio signals. While this process can yield good results, it incurs a significant computational burden, hindering widespread adoption of these speech analysis methods. In other words, these audio signal processing methods suffer from low processing efficiency.

[0006] To address the above-mentioned problems, no effective solutions have been proposed so far.

[0007] Summary of the Invention

[0008] Embodiments of the present application provide a method and apparatus for processing an audio signal, a storage medium, and an electronic device.

[0009] According to one aspect of an embodiment of the present application, a method for processing an audio signal is provided, which is performed by an electronic device and includes:

[0010] Obtaining a first spectral feature of a noisy audio signal, wherein the noisy audio signal includes a reference audio signal and a target audio signal to be extracted; the reference audio signal and the target audio signal are respectively one of a noise signal and a clean speech signal in the noisy audio signal;

[0011] Inputting the first spectral feature into a target neural network model to obtain reference phase estimation information that matches the noisy audio signal, wherein the target neural network model is pre-trained based on sample noisy audio signals and sample target audio signals, and the reference phase estimation information is used to indicate the phase difference between the target audio signal and the noisy audio signal in multiple frequency components; the sample target audio signal is a sample noise signal or a sample clean speech signal in the sample noisy audio signal;

[0012] Determining a second phase sine value and a second phase cosine value for determining a second spectrum feature based on a first phase sine value and a first phase cosine value indicated by the first spectrum feature, and a reference phase sine value and a reference phase cosine value indicated by the reference phase estimation information; and

[0013] The second frequency spectrum feature is determined according to the second phase sine value and the second phase cosine value, and the target audio signal is determined according to the second frequency spectrum feature.

[0014] According to another aspect of an embodiment of the present application, a method for training an audio signal processing model is provided, which is performed by an electronic device and includes:

[0015] Obtaining a first spectral feature of a sample noisy audio signal and a training label matching the sample noisy audio signal, wherein the sample noisy audio signal includes a sample reference audio signal and a sample target audio signal, and the training label includes a phase difference label, the phase difference label being determined based on a cosine value and a sine value of a first sample vector indicated by the first spectral feature of the sample noisy audio signal, and a cosine value and a sine value of a second sample vector indicated by a second spectral feature of the sample target audio signal; the sample reference audio signal and the sample target audio signal are respectively one of a sample noise signal or a sample clean speech signal in the sample noisy audio signal;

[0016] Inputting the first spectral feature into an audio signal processing model to be trained to obtain reference phase estimation information matching the sample noisy audio signal;

[0017] Training the audio signal processing model according to the reference phase estimation information and the training loss determined by the training label; and

[0018] When the training loss satisfies a target convergence condition, the trained audio signal processing model is determined as a target neural network model.

[0019] According to another aspect of the embodiments of the present application, there is further provided an audio signal processing device, including:

[0020] an acquiring unit, configured to acquire a first spectral feature of a noisy audio signal, wherein the noisy audio signal comprises a reference audio signal and a target audio signal to be extracted; the reference audio signal and the target audio signal are respectively one of a noise signal and a clean speech signal in the noisy audio signal;

[0021] a processing unit, configured to input the first spectral feature into a target neural network model to obtain matching information with the noisy audio signal and reference phase estimation information, wherein the target neural network model is pre-trained based on sample noisy audio signals and sample target audio signals, and the reference phase estimation information is used to indicate a phase difference between the target audio signal and the noisy audio signal in multiple frequency components; the sample target audio signal is a sample noise signal or a sample clean speech signal in the sample noisy audio signal;

[0022] a first determining unit, configured to determine a second phase sine value and a second phase cosine value for determining a second spectral feature based on a first phase sine value and a first phase cosine value indicated by the first spectral feature, and a reference phase sine value and a reference phase cosine value indicated by the reference phase estimation information; and

[0023] The second determining unit is configured to determine the second frequency spectrum feature according to the second phase sine value and the second phase cosine value, and determine the target audio signal according to the second frequency spectrum feature.

[0024] According to another aspect of the embodiments of the present application, there is further provided an audio signal processing device, comprising:

[0025] an acquiring unit, configured to acquire a first spectral feature of a noisy audio signal, wherein the noisy audio signal comprises a reference audio signal and a target audio signal to be extracted; the reference audio signal and the target audio signal are respectively one of a noise signal and a clean speech signal in the noisy audio signal;

[0026] a processing unit, configured to input the first spectral feature into a target neural network model to obtain matching information with the noisy audio signal and reference phase estimation information, wherein the target neural network model is pre-trained based on sample noisy audio signals and sample target audio signals, and the reference phase estimation information is used to indicate a phase difference between the target audio signal and the noisy audio signal in multiple frequency components; the sample target audio signal is a sample noise signal or a sample clean speech signal in the sample noisy audio signal;

[0027] a first determining unit, configured to determine a second phase sine value and a second phase cosine value for determining a second spectral feature based on a first phase sine value and a first phase cosine value indicated by the first spectral feature, and a reference phase sine value and a reference phase cosine value indicated by the reference phase estimation information; and

[0028] The second determining unit is configured to determine the second frequency spectrum feature according to the second phase sine value and the second phase cosine value, and determine the target audio signal according to the second frequency spectrum feature.

[0029] According to another aspect of the embodiments of the present application, a computer-readable storage medium is further provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned audio signal processing method or audio signal processing model training method when running.

[0030] According to another aspect of the embodiments of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the above-described audio signal processing method or audio signal processing model training method.

[0031] According to another aspect of an embodiment of the present application, an electronic device is also provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned audio signal processing method or audio signal processing model training method through the computer program.

[0032] The details of one or more embodiments of the present application are set forth in the following drawings and description. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.

[0034] FIG1 is a schematic diagram of an application environment of an optional audio signal processing method according to an embodiment of the present application;

[0035] FIG2 is a flowchart of an optional method for processing an audio signal according to an embodiment of the present application;

[0036] FIG3 is a schematic diagram of an optional method for processing an audio signal according to an embodiment of the present application;

[0037] FIG4 is a schematic diagram of another optional method for processing an audio signal according to an embodiment of the present application;

[0038] FIG5 is a schematic diagram of another optional method for processing an audio signal according to an embodiment of the present application;

[0039] FIG6 is a flowchart of an optional method for training an audio signal processing model according to an embodiment of the present application;

[0040] FIG7 is a schematic diagram of another optional method for processing an audio signal according to an embodiment of the present application;

[0041] FIG8 is a schematic diagram of another optional method for processing an audio signal according to an embodiment of the present application;

[0042] FIG9 is a schematic diagram of another optional method for processing an audio signal according to an embodiment of the present application;

[0043] FIG10 is a schematic structural diagram of an optional audio signal processing device according to an embodiment of the present application;

[0044] FIG11 is a schematic structural diagram of an optional electronic device according to an embodiment of the present application;

[0045] FIG12 is a schematic structural diagram of an optional audio signal processing model training device according to an embodiment of the present application;

[0046] FIG13 is a schematic structural diagram of another optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0047] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0048] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0049] According to one aspect of an embodiment of the present application, a method for processing an audio signal is provided. As an optional implementation, the audio signal processing method can be, but is not limited to, applied in the environment shown in FIG1 . As shown in FIG1 , a terminal device 102 includes a memory 104 for storing various data generated during the operation of the terminal device 102, a processor 106 for processing and calculating the aforementioned data, and a display 108. The terminal device 102 can exchange data with a server 112 via a network 110. The server 112 is connected to a database 114, which is used to store various data.

[0050] Optionally, the terminal device 102 may run an application for collecting audio signals, such as a recording application. The recording application may collect audio signals in the environment, send the collected audio signals to the server 112 for noise reduction processing, and obtain the audio signals processed by the server 112.

[0051] Furthermore, the specific application process of the above method in the environment shown in FIG1 is shown in the following steps:

[0052] In steps S102 and S104, terminal device 102 collects a noisy audio signal to be processed. The noisy audio signal includes a reference audio signal and a target audio signal to be extracted. The reference audio signal and the target audio signal are either a noise signal or a clean speech signal in the noisy audio signal. Terminal device 102 transmits the noisy audio signal to server 112 via network 110.

[0053] Then, the server 112 executes steps S106-S112 to obtain a first spectral feature of the noisy audio signal; inputs the first spectral feature into the target neural network model to obtain reference phase estimation information that matches the noisy audio signal, wherein the target neural network model is pre-trained based on the sample noisy audio signal and the sample target audio signal, and the reference phase estimation information is used to indicate the phase difference between the target audio signal and the noisy audio signal on multiple frequency components; the sample target audio signal is a sample noise signal or a sample pure speech signal in the sample noisy audio signal; based on the first spectral feature and the reference phase estimation information, the second spectral feature of the target audio signal is determined; and the target audio signal is determined based on the second spectral feature. Wherein, determining the second spectral feature of the target audio signal based on the first spectral feature and the reference phase estimation information includes: determining the second phase sine value and the second phase cosine value for determining the second spectral feature based on the first phase sine value and the first phase cosine value indicated by the first spectral feature, and the reference phase sine value and the reference phase cosine value indicated by the reference phase estimation information; determining the second spectral feature based on the second phase sine value and the second phase cosine value, and determining the target audio signal based on the second spectral feature;

[0054] Then the server 112 executes step S114 to send the target audio signal to the terminal device 102 via the network 110 .

[0055] In an embodiment of the present application, a first spectral feature of a noisy audio signal is first obtained, wherein the noisy audio signal includes a reference audio signal and a target audio signal to be extracted; the reference audio signal and the target audio signal are respectively one of a noise signal or a pure speech signal in the noisy audio signal; the first spectral feature is input into a target neural network model to obtain reference amplitude estimation information and reference phase estimation information that match the noisy audio signal, wherein the target neural network model is pre-trained based on sample noisy audio signals and sample target audio signals, and the reference amplitude estimation information is used to indicate the difference between the target audio signal and the noisy audio signal in multiple frequency distributions. The reference phase estimation information is used to indicate the phase difference between the target audio signal and the noisy audio signal in the multiple frequency components; according to the first spectral feature, the reference amplitude estimation information and the reference phase estimation information, the second spectral feature of the target audio signal is determined; according to the second spectral feature, the target audio signal is determined, thereby realizing the prediction of the estimated values ​​for characterizing the phase difference and amplitude difference between the audio signals through the target neural network model, and then determining the processed target audio signal according to the predicted reference amplitude estimation information and reference phase estimation information and the spectral feature of the original audio signal.

[0056] According to the above-mentioned embodiment of the present application, a technical solution is proposed that can simultaneously model and analyze the amplitude information and phase information of a noisy audio signal, thereby parsing the audio signal interfered by noise through comprehensive processing of the phase information, thereby improving the accuracy of audio signal processing; at the same time, the target neural network is used to obtain phase information estimation, avoiding the use of a complex neural network model, reducing the amount of calculation, and improving the processing efficiency of audio signal processing.

[0057] Optionally, in this embodiment, the above-mentioned terminal device can be a terminal device configured with a target client, which can include but is not limited to at least one of the following: a mobile phone (such as an Android phone, an iOS phone, etc.), a laptop computer, a tablet computer, a PDA, an MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV, etc. The target client can be a video client, an instant messaging client, a browser client, an education client, etc. The above-mentioned network can include but is not limited to: a wired network, a wireless network, wherein the wired network includes: a local area network, a metropolitan area network and a wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that realize wireless communication. The above-mentioned server can be a single server, or it can be a server cluster composed of multiple servers, or a cloud server. The above is only an example, and this embodiment does not impose any limitation on this.

[0058] Optionally, as an optional implementation, as shown in FIG2 , the audio signal processing method includes:

[0059] S202 , obtaining a first spectrum feature of a noisy audio signal, wherein the noisy audio signal includes a reference audio signal and a target audio signal to be extracted; the reference audio signal and the target audio signal are respectively one of a noise signal and a clean speech signal in the noisy audio signal.

[0060] S204: Input the first spectral feature into the target neural network model to obtain reference phase estimation information that matches the noisy audio signal, wherein the target neural network model is pre-trained based on the sample noisy audio signal and the sample target audio signal, and the reference phase estimation information is used to indicate the phase difference between the target audio signal and the noisy audio signal in multiple frequency components; the sample target audio signal is the sample noise signal or the sample clean speech signal in the sample noisy audio signal.

[0061] When the reference audio signal is a noise signal in a noisy audio signal, the target audio signal is a clean speech signal in the noisy audio signal, and the sample target audio signal is a sample clean speech signal in the sample noisy audio signal. When the reference audio signal is a clean speech signal in a noisy audio signal, the target audio signal is a noise signal in the noisy audio signal, and the sample target audio signal is a sample noise signal in the sample noisy audio signal.

[0062] S206, determining a second phase sine value and a second phase cosine value for determining a second spectrum feature based on the first phase sine value and the first phase cosine value indicated by the first spectrum feature, and the reference phase sine value and the reference phase cosine value indicated by the reference phase estimation information.

[0063] S208 : Determine a second frequency spectrum feature according to the second phase sine value and the second phase cosine value, and determine a target audio signal according to the second frequency spectrum feature.

[0064] In this embodiment, a technical solution is proposed for modeling and analyzing the phase information of a noisy audio signal, thereby parsing the audio signal interfered with by noise through comprehensive processing of the phase information, thereby improving the accuracy of audio signal processing; at the same time, a target neural network is used to obtain phase information estimation, avoiding the use of a complex neural network model, reducing the amount of calculation, and improving the processing efficiency of audio signal processing.

[0065] It should be noted that the above-mentioned audio signal processing method of the present application can be applied to scenarios including but not limited to audio noise reduction, or scenarios of extracting noise signals based on noisy audio signals. In the audio noise reduction scenario, the above-mentioned target audio signal is the audio signal after noise reduction processing, such as speech audio, instrumental audio, etc. Correspondingly, in this scenario, the above-mentioned reference audio signal is the noise signal in the above-mentioned noisy audio signal; in the scenario of extracting noise signals, the above-mentioned target audio signal is noise audio, and correspondingly, in this scenario, the above-mentioned reference audio signal is the pure audio signal in the above-mentioned noisy audio signal, such as speech audio, instrumental audio, etc. It can be understood that in the scenario of extracting noise signals, the extracted noise audio can be used to further analyze the noise signal, or to synthesize noisy sample audio to train the noise reduction network model.

[0066] Specifically, in the scenario of extracting noise signals, the above-mentioned noisy audio signal is an audio signal collected in the target environment, and the above-mentioned target audio signal is the ambient noise in the target environment. The target ambient noise in the target environment can be extracted through the above-mentioned implementation method, and the obtained target ambient noise is synthesized with different pure audio signals (such as speech signals, instrumental audio signals) to obtain a noisy audio signal, and the synthesized noisy audio signal is used to train the noise reduction network to obtain a neural network model for processing ambient noise in a specific environment.

[0067] In the audio noise reduction scenario, it can be specifically used in the audio signal noise reduction processing scenarios of audio recording, voice calls, video calls, video conferences, cameras, smart home appliances, etc. In the case where the above-mentioned audio signal processing method is applied to the audio signal noise reduction processing scenario of audio recording, the above-mentioned noisy audio signal can be an audio signal collected from the audio recording environment, which carries the target audio signal to be extracted, that is, a pure audio signal (such as a voice signal, an instrumental audio signal light) and environmental noise, and then the noise reduction processing is performed through the above-mentioned implementation method to obtain an audio signal that is not interfered by noise; in the case where the above-mentioned audio signal processing method is applied to the audio signal noise reduction processing scenario of voice calls, the above-mentioned method can be used to perform noise reduction processing on the audio signal collected during the voice call to obtain a voice signal with noise removed. In the case where the above-mentioned audio signal processing method is applied to the audio signal noise reduction processing scenario of a camera, the above-mentioned method can be used to perform noise reduction processing on the audio signal collected by the camera to obtain a voice signal with noise removed. When the above-mentioned audio signal processing method is applied to the audio signal noise reduction processing scenario of smart home appliances, the above-mentioned method can be used to perform noise reduction processing on the audio signal collected by the smart home appliance to obtain a noise-removed voice signal.

[0068] Furthermore, when the above-mentioned audio signal processing method is used for speech noise reduction, the noisy audio signal in the above-mentioned step S202 can be used, but is not limited to, to indicate the original speech signal collected by the terminal device, where the speech signal includes a noise signal and a speech signal to be extracted (target audio signal). For example, assuming that the above-mentioned audio signal processing method is applied to an audio signal noise reduction processing scenario for a voice call, and the current user object A is talking to user object B through terminal device a, then the above-mentioned target audio signal can be the audio signal emitted by the environment where user object A is located, which is collected by terminal device a, and the audio signal includes the noise signal existing in the environment where user object A is located and the speech signal emitted by user object A.

[0069] Optionally, the method of obtaining the first spectrum feature of the noisy audio signal in step S202 may be to perform frequency domain conversion processing on the noisy audio signal to obtain the noisy frequency domain feature corresponding to the noisy audio signal. It is understandable that in the above-mentioned noisy audio signal represented by x n In the case of , the corresponding noisy frequency domain feature obtained by frequency domain conversion of the above-mentioned noisy frequency signal can be expressed as X k (plural). It is understandable that X k The real part and the imaginary part of can each correspond to an array. The complex number composed of each corresponding value in the real array and the imaginary array can be used to represent the frequency domain signal at a specific frequency, which can be further expressed by X k Characterize the combination of frequency domain signals in multiple different frequency domains to obtain the above-mentioned first spectrum feature X k .

[0070] After obtaining the first spectral feature of the noisy audio signal, as in steps S204 and S206, the trained target neural network model can be used to obtain corresponding reference phase estimation information based on the first spectral feature, and the audio signal can be extracted based on the reference phase estimation information.

[0071] In an optional embodiment, the above-mentioned determination of the second spectral feature based on the second phase sine value and the second phase cosine value determines the second spectral feature of the target audio signal based on the first spectral feature, reference amplitude estimation information and reference phase estimation information, including: obtaining reference amplitude estimation information output by the target neural network model based on the first spectral feature, wherein the reference amplitude estimation information is used to indicate the amplitude difference between the target audio signal and the noisy audio signal in multiple frequency components; determining the second audio amplitude that matches the second spectral feature based on the first audio amplitude indicated by the first spectral feature and the reference amplitude estimation information; determining the second spectral feature based on the second audio amplitude, the second phase sine value and the second phase cosine value.

[0072] In the above embodiment, the second spectrum feature may be extracted using the reference amplitude estimation information and reference phase estimation information output by the target neural network model.

[0073] The above-mentioned target neural network model may include but is not limited to a CRN (Counterfactual Recurrent Network) network, a convolutional neural network (CNN) or a recurrent neural network (RNN), which can be used to extract corresponding reference amplitude estimation information and reference phase estimation information based on the first spectrum feature.

[0074] It should be noted that the reference amplitude estimation information can be used to characterize the amplitude difference between the audio signal to be extracted and the noisy audio signal in the frequency domain, and the reference phase estimation information can be used to characterize the phase difference between the audio signal to be extracted and the noisy audio signal in the frequency domain.

[0075] In an optional embodiment, the amplitude difference may be the absolute value of the amplitude difference between the first spectral feature and the second spectral feature of the target audio signal, that is, the reference amplitude estimation information is the difference between the amplitude value of the first spectral feature and the amplitude value of the second spectral feature.

[0076] In another optional embodiment, the amplitude difference may also be the amplitude ratio of the first spectral feature to the second spectral feature of the target audio signal, that is, the reference phase difference is the ratio between the amplitude value of the first spectral feature and the amplitude value of the second spectral feature.

[0077] It can be understood that the amplitude value of the above-mentioned first spectral feature can be specifically expressed as a first amplitude vector obtained by combining the amplitude values ​​of multiple frequency components, and the amplitude value of the above-mentioned second spectral feature can be specifically expressed as a second amplitude vector obtained by combining the amplitude values ​​of multiple frequency components. Correspondingly, when the amplitude difference is the absolute value of the amplitude difference between the first spectral feature and the second spectral feature of the target audio signal, the above-mentioned reference amplitude estimation information can be expressed as an amplitude vector corresponding to the vector difference between the above-mentioned multiple first amplitude vectors and the second amplitude vector; when the amplitude difference is the amplitude ratio of the first spectral feature to the second spectral feature of the target audio signal, the above-mentioned reference amplitude estimation information can be expressed as an amplitude ratio vector composed of the element ratios between the corresponding elements in the above-mentioned multiple first amplitude vectors and the second amplitude vector.

[0078] In an optional implementation, the phase difference may be a sine value and a cosine value corresponding to a difference between a signal phase represented by the first spectral feature and a signal phase represented by the second spectral feature of the target audio signal.

[0079] For example, when the phase difference is the phase difference between the signals represented by the first spectrum feature and the second spectrum feature, the difference can be represented by the following sine and cosine values:

[0080] in, is the first spectrum phase characterized by the above first spectrum feature, The second spectrum phase represented by the above second spectrum feature. is the cosine value used to characterize the phase difference, is the sine value used to characterize the phase difference. It can be understood that the above and are two vectors of the same length, each element in the vector is used to represent the phase corresponding to a frequency component. Correspondingly, the above is the cosine value vector, is a vector of sine values.

[0081] In another optional embodiment, when the phase difference is the phase difference between the signals represented by the second spectrum feature and the first spectrum feature, the difference can be represented by the following sine and cosine values:

[0082] After obtaining the reference amplitude estimation information and reference phase estimation information through the above steps, the second spectrum feature can be obtained based on the reference amplitude estimation information, reference phase estimation information and the first spectrum feature, and then the corresponding target audio signal can be obtained based on the second spectrum feature.

[0083] Optionally, in the above step S204, determining the second phase sine value and the second phase cosine value for determining the second spectral feature based on the first phase sine value and the first phase cosine value indicated by the first spectral feature, and the reference phase sine value and the reference phase cosine value indicated by the reference phase estimation information includes: traversing the N frequency components indicated by the first spectral feature, obtaining the first phase sine value and the first phase cosine value corresponding to the currently traversed frequency component, when the reference amplitude estimation information is used to indicate the phase difference between the target audio signal and the noisy audio signal on N frequency components. , where N is an integer greater than 1; obtaining a reference phase sine value and a reference phase cosine value corresponding to the currently traversed frequency component from the reference phase estimation information; obtaining a first product of the first phase sine value and the reference phase cosine value, and a second product of the first phase cosine value and the reference phase sine value, and determining the sum of the first product and the second product as the second phase sine value; obtaining a third product of the first phase cosine value and the reference phase cosine value, and a fourth product of the first phase sine value and the reference phase sine value, and determining the second phase cosine value according to the difference between the third product and the fourth product.

[0084] Specifically, when the reference amplitude estimation information is used to indicate the phase difference between the target audio signal and the noisy audio signal at N frequency components, the obtained reference phase sine value is And the reference phase cosine value is In the case of , correspondingly, the second phase sine value and the second phase cosine value can be obtained as follows:

[0085] in, is the second phase cosine value, is the second phase sine value. It can be understood that in the above formula as well as They are all vector representations of the same latitude.

[0086] In an embodiment of the present application, a first spectral feature of a noisy audio signal is first obtained, wherein the noisy audio signal includes a reference audio signal and a target audio signal to be extracted; the reference audio signal and the target audio signal are respectively one of a noise signal or a pure speech signal in the noisy audio signal; the first spectral feature is input into a target neural network model to obtain reference amplitude estimation information and reference phase estimation information that match the noisy audio signal, wherein the target neural network model is pre-trained based on sample noisy audio signals and sample target audio signals, and the reference amplitude estimation information is used to indicate the difference between the target audio signal and the noisy audio signal in multiple frequency distributions. The reference phase estimation information is used to indicate the phase difference between the target audio signal and the noisy audio signal in the multiple frequency components; according to the first spectral feature, the reference amplitude estimation information and the reference phase estimation information, the second spectral feature of the target audio signal is determined; according to the second spectral feature, the target audio signal is determined, thereby realizing the prediction of the estimated values ​​for characterizing the phase difference and amplitude difference between the audio signals through the target neural network model, and then determining the processed target audio signal according to the predicted reference amplitude estimation information and reference phase estimation information and the spectral feature of the original audio signal.

[0087] According to the above-mentioned embodiment of the present application, a technical solution is proposed that can simultaneously model and analyze the amplitude information and phase information of a noisy audio signal, thereby parsing the audio signal interfered with by noise through comprehensive processing of the amplitude information and phase information, thereby improving the accuracy of audio signal processing; at the same time, a target neural network is used to obtain amplitude information and phase information estimation, avoiding the use of a complex neural network model, avoiding an increase in calculation amount, and improving the processing efficiency of audio signal processing.

[0088] Optionally, the above-mentioned determination of the second audio amplitude for determining the second spectral feature based on the first audio amplitude indicated by the first spectral feature and the reference amplitude estimation information includes: when the reference amplitude estimation information is used to indicate the reference amplitude ratio of the target audio signal to the noisy audio signal on N frequency components, obtaining the first amplitude values ​​on the N frequency components indicated by the first spectral feature and the N reference amplitude ratios indicated by the reference amplitude estimation information, wherein N is an integer greater than 1; and determining the product value of the N first amplitude values ​​and their respective corresponding reference amplitude ratios as the second amplitude value on the N frequency components matching the second spectral feature.

[0089] Specifically, when the reference amplitude estimation information obtained by the target neural network model is used to represent the reference amplitude ratio of the target audio signal to the noisy audio signal on N frequency components, the second audio amplitude can be obtained as follows:

[0090] Assume that the reference amplitude estimation information output by the above target neural network model is The first audio amplitude of the noisy frequency signal is |X k |, the second audio amplitude corresponding to the second spectral feature of the target audio signal is:

[0091] It can be understood that |X k |、 are all vectors of the same dimension.

[0092] In an optional embodiment, the above-mentioned determination of the second spectral characteristics based on the second audio amplitude, the second phase sine value and the second phase cosine value includes: determining the real part representation of the second spectral characteristics based on the second audio amplitude and the second phase cosine value; determining the imaginary part representation of the second spectral characteristics based on the second audio amplitude and the second phase sine value; and determining the second spectral characteristics based on the real part representation and the imaginary part representation.

[0093] In obtaining the above as well as In the case of , the second spectrum feature can be expressed as follows:

[0094] in, is the real part representation of the second spectrum feature, It is the imaginary part representation of the second spectrum feature.

[0095] Through the above-mentioned implementation mode of the present application, a technical solution is proposed that can simultaneously model and analyze the amplitude information and phase information of a noisy audio signal, thereby parsing the audio signal interfered with by noise through the comprehensive processing of the amplitude information and the phase information, thereby improving the accuracy of audio signal processing; at the same time, the target neural network is used to obtain amplitude information and phase information estimation, avoiding the use of a complex neural network model, avoiding the increase in calculation amount, and improving the processing efficiency of audio signal processing.

[0096] In an optional embodiment, the above-mentioned inputting the first spectral feature into the target neural network model to obtain reference amplitude estimation information and reference phase estimation information that match the noisy frequency signal includes: in the target neural network model, encoding the first spectral feature through the encoding network to obtain a first encoding result; analyzing the first encoding result through a recurrent neural network constructed based on a gated recurrent unit to obtain a first intermediate result carrying timing information; inputting the first encoding result and the first intermediate result into the decoding network in the target neural network model to obtain reference amplitude estimation information and reference phase estimation information that match the noisy frequency signal, wherein the sub-network in the decoding network is obtained based on the adjustment of the sub-network in the encoding network.

[0097] Optionally, in an embodiment of the present application, the target neural network model obtained through training may, but is not limited to, adopt an encoder-decoder interactive structure.

[0098] Optionally, the above-mentioned encoding processing of the first spectral features through the encoding network includes: encoding the first spectral features in sequence through M encoding sub-networks with a connection relationship in the encoding network to obtain M first encoding results, wherein each encoding sub-network includes: a convolution layer, a normalization layer and an activation layer, and when the first spectral features corresponding to each frame are convolutionally processed in the convolution layer, the first spectral features corresponding to the adjacent previous frame will be referred to, and M is a natural number greater than or equal to 2.

[0099] Optionally, the above-mentioned analysis of the first encoding result by a recurrent neural network constructed based on a gated recurrent unit to obtain a first intermediate result carrying timing information includes: inputting the first encoding result output by the Mth encoding sub-network into the recurrent neural network to obtain a first intermediate result carrying timing information.

[0100] Optionally, the above-mentioned inputting the first encoding result and the first intermediate result into the decoding network in the target neural network model to obtain reference amplitude estimation information and reference phase estimation information matching the noisy frequency signal includes: inputting the first intermediate result and the first encoding result output by the i-th encoding subnetwork into the M-i+1-th decoding subnetwork, wherein the decoding network includes M decoding subnetworks with a connection relationship, each decoding subnetwork includes: a transposed convolution layer, a normalization layer and an activation layer associated with the convolution layer, i is an integer greater than or equal to 1 and less than or equal to M, and a jump connection is set between the i-th encoding subnetwork and the M-i+1-th decoding subnetwork; obtaining the first decoding result output by the M-th encoding subnetwork; and determining the first decoding result as the reference amplitude estimation information and reference phase estimation information matching the noisy frequency signal.

[0101] The model structure of the target neural network model is described below in conjunction with Figure 3. As shown in Figure 3, the target neural network model includes an encoding network 301, a recurrent neural network RNN ​​302, and a decoding network 303. The encoding network 301 includes multiple encoding sub-networks, and the decoding network 303 includes multiple decoding sub-networks that are jump-connected to the multiple encoding sub-networks. The input of each decoding sub-network is the output result of the corresponding encoding sub-network and the output result of RNN 302.

[0102] More specifically, each encoding sub-network in the encoding network 301 in Figure 3 can be an EncConv2d module, and its network structure can be shown in Figure 4, consisting of a convolution layer 402 (i.e., two-dimensional convolution (Conv2d)), a normalization layer 404 (i.e., normalization (BatchNorm)), and an activation layer 406 (i.e., activation function (PReLU)). Among them, the convolution kernel size (Kernel Size) of each layer of EncConv2d is (5, 2), which means that the frequency domain field of view is 5 and the time domain field of view is 2. That is, the analysis and processing of the signal features of each frame will refer to the signal of the previous frame, which can be regarded as a streaming convolution structure, ensuring the causality of the network. The stride of the convolution can be, but is not limited to, set to (2, 1), that is, the frequency domain stride of the convolution is 2 and the time domain stride is 1. This can reduce the number of frequency domain features of the signal by half layer by layer, while the dimension of the time domain features remains unchanged, which not only maintains the time domain continuity of the information but also reduces the amount of computation.

[0103] Specifically, the recurrent neural network constructed based on the gated recurrent unit can be used, but is not limited to, to indicate the extraction module. Specifically, the extraction module can be a recurrent neural network (RNN) composed of a stack of gated recurrent units (GRUs), which is used to extract the timing information from the encoder output.

[0104] In addition, each decoding subnetwork in the decoding network 303 (i.e., the Decoder module) in Figure 3 can be a subnetwork corresponding to the encoding subnetwork. Specifically, each decoding subnetwork in the decoding network 303 can be used to restore the frequency domain feature number with the first frequency domain feature. The Decoder module can be, but is not limited to, composed of a stack of DecTConv2d modules. The DecTConv2d structure is highly similar to EncConv2d, and includes: a transposed convolution layer (i.e., a transposed convolution network (ConvTranspose2d)) corresponding to the convolution layer (i.e., two-dimensional convolution (Conv2d)) in EncConv2d, a normalization layer (i.e., normalization (BatchNorm)), and an activation layer (i.e., activation function (PReLU)). The number of DecTConv2d layers included in the Decoder is the same as the number of EncConv2d layers included in the Encoder. The parameters of each layer of DecTConv2d are also the same as the parameters of the EncConv2 layer of the corresponding layer. In addition, the output of each layer of Encoder can be used as the influencing parameter of the corresponding layer in the Decoder module by using skip connections, thereby realizing layer-by-layer restoration of signal feature dimensions.

[0105] The following is a detailed description of the process of processing the first frequency domain feature of the noisy frequency signal by the target neural network model with reference to FIG3. The encoder (encoding network 301) receives the short-time Fourier transform representation of the noisy frequency signal from the signal preprocessing module. Then, EncConv2d extracts high-dimensional features layer by layer, and the corresponding output is given to DecTConv2d (decoding network 303) through skip connections. RNNs (RNN302) accept the output features from the last layer of EncConv2d, perform temporal information extraction and analysis, and give the input to the Decoder. The Decoder accepts the output from RNNs and Encoder, and increases the dimension layer by layer through transposed convolution. The number of channels of the last layer of DecTConv2d is 3, and finally three outputs are obtained, namely, the amplitude spectrum mask estimation Phase spectrum mask estimation cosine value Sine value

[0106] After obtaining the mask estimate of the speech signal, the STFT of the enhanced speech can be obtained by modulating the complex spectrum of the original noisy signal short-time Fourier transform. The expression is as follows:

[0107] The enhanced speech amplitude spectrum estimate is determined as follows:

[0108] The phase cosine estimation information is determined as follows:

[0109] The phase sine estimation information is determined as follows:

[0110] The enhanced speech STFT estimation information is determined as follows:

[0111] After obtaining the short-time Fourier transform estimation information of the pure speech, the inverse short-time Fourier transform (iSTFT) corresponding to the STFT is finally performed to obtain the time domain waveform signal of the enhanced speech.

[0112] Through the above-mentioned implementation of the present application, a speech enhancement solution is provided for simultaneously modeling and analyzing the amplitude spectrum and phase spectrum of a noisy frequency signal. Without almost increasing the network computational complexity, the model is enabled to learn pure human voice phase information, thereby significantly improving the noise reduction effect.

[0113] In an optional embodiment, the above-mentioned obtaining of the first spectral characteristics of the noisy audio signal includes: obtaining the noisy audio signal; resampling the noisy audio signal according to a target sampling rate to obtain a first reference audio signal; performing time-domain framing and windowing processing on the first reference audio signal according to a framing parameter to obtain multiple first audio sub-signals; and performing discrete Fourier transform on the multiple first audio sub-signals to obtain first spectral characteristics corresponding to the multiple first audio sub-signals.

[0114] The above embodiment provides a method for obtaining the first spectrum feature. First, the noisy frequency signal x n Resampling is performed to resample all audio data of all sampling rates to 48kHz. After the resampling operation is completed, the long audio signal is then subjected to time-domain frame segmentation and windowing. For example, but not limited to, the target audio signal can be segmented into multiple frames of short signals of fixed length, with a single frame consisting of 1024 sampling points (i.e., a single frame length of 1024) and a frame shift of 512 (i.e., the overlap length between each two adjacent frames is 512). A Hamming window is then used to modulate each frame of the target audio signal to prevent spectral leakage.

[0115] It should be noted that the windowing processing method used for windowing the target audio signal is not limited to the Hamming window, and other methods such as rectangular window, Hanning window, etc. can also be used. This is not limited in this embodiment.

[0116] After performing frame windowing processing on the noisy frequency signal, the short-time Fourier transform (STFT) method can also be used to obtain the above-mentioned first spectral characteristics. Specifically, the above-mentioned short-time Fourier transform (STFT) is a mathematical transformation related to the Fourier transform, which is used to determine the frequency and phase of the sine wave in the local area of ​​the time-varying signal. Its core logic is to select a time-frequency localized window function, assuming that the analysis window function g(t) is stable (pseudo-stationary) within a short time interval, and move the window function so that f(t) and g(t) are stationary signals within different finite time widths, thereby calculating the power spectrum at different moments. Pseudo-stationary means that the fluctuation is within a preset fluctuation range.

[0117] It should be noted that the amplitude spectrum described above is a curve plotting signal amplitude against frequency (angular frequency). In the frequency domain description of a signal, frequency is used as the independent variable, and the amplitudes of the signal's individual frequency components are used as the dependent variable. This frequency function is called the amplitude spectrum, which characterizes the distribution of the signal's amplitude over frequency. For frequency domain descriptions of random signals, the power spectrum is often used, which characterizes the distribution of the signal's energy over frequency.

[0118] Correspondingly, the above-mentioned determination of the target audio signal based on the second spectral characteristics includes: obtaining multiple second spectral characteristics determined based on the first spectral characteristics corresponding to the multiple first audio sub-signals; performing an inverse short-time Fourier transform on the multiple second spectral characteristics to obtain multiple second audio sub-signals; and determining the target audio signal based on the splicing results of the multiple second audio sub-signals.

[0119] It is understood that in the above embodiment, after obtaining the second spectrum feature of the target audio signal, an inverse short-time Fourier transform (iSTFT) corresponding to the STFT can be performed on it to obtain the time domain waveform signal of the target audio signal.

[0120] It is understood that the first spectral features corresponding to each first audio sub-signal can be used to obtain reference amplitude estimation information and reference phase estimation information corresponding to each first spectral feature through the target neural network model, and then multiple second spectral features can be determined based on each first spectral feature and the corresponding reference amplitude estimation information and reference phase estimation information. After performing an inverse short-time Fourier transform on each second spectral feature, multiple second audio sub-signals can be obtained. The number of second audio sub-signals is related to the processing parameters used in the time-domain framing and windowing process of the noisy audio signal. The processing parameters include, but are not limited to, a single frame length parameter and a frame shift parameter. Based on these parameters, the multiple second audio sub-signals can be spliced ​​together to obtain the target audio signal.

[0121] In an optional embodiment, after determining the second spectral characteristics of the target audio signal based on the first spectral characteristics, reference amplitude estimation information and reference phase estimation information, the method further includes: determining a reference spectral characteristic of the reference audio signal based on the first spectral characteristics and the second spectral characteristics; and performing an inverse short-time Fourier transform on the reference spectral characteristics to obtain a reference audio signal.

[0122] It can be understood that in this embodiment, after extracting the target audio signal from the noisy audio signal through the above embodiment, the reference spectrum characteristics of the reference audio signal can be further determined based on the first spectrum characteristics and the second spectrum characteristics, thereby realizing the extraction of the reference spectrum characteristics, and then performing an inverse short-time Fourier transform based on the reference spectrum characteristics to obtain the reference audio signal.

[0123] When the second audio feature is an audio signal obtained after noise reduction, the reference audio signal may be a noise signal; when the second audio feature is an extracted noise signal, the reference audio signal may be an audio signal obtained after noise reduction.

[0124] A complete audio processing process of the present application is described below with reference to Figure 5. In this embodiment, the noisy audio signal is a noisy audio signal, the target audio signal is a clean speech signal to be extracted, and the reference audio signal is a noise signal.

[0125] As shown in FIG5 , the specific implementation system framework is mainly divided into three modules, namely, an audio signal pre-processing and feature extraction module 501 , a neural network model inference module 502 and a post-processing speech generation module 503 .

[0126] The pre-processing and feature extraction module 501 first processes the noisy frequency signal x nResampling is performed to resample all audio data of different sampling rates to 48kHz. After the resampling operation is completed, the long audio signal is subjected to time domain framing and windowing. According to the single frame length of 1024 and the frame shift of 512 (overlap 512), the original audio signal is divided into multiple frames of short signals with fixed lengths, and each frame signal is modulated using a Hamming window to prevent spectrum leakage. After the framing and windowing operation is completed, the modulated signal is subjected to a discrete Fourier transform (DFT) operation to extract the frequency domain features and obtain the noisy frequency signal x. n Frequency domain representation of X k The combination of audio signal framing and windowing with the discrete Fourier transform (DFT) is also known as the short-time Fourier transform (STFT).

[0127] In the neural network model inference module 502, this embodiment adopts the Encoder-Decoder framework. The Encoder part is mainly composed of the EncConv2d structure with a two-dimensional convolution (Conv2d) as the kernel. The kernel size of each layer of EncConv2d is (5, 2), which means that the frequency domain field of view is 5 and the time domain field of view is 2. The analysis and processing of each frame signal feature will refer to the previous frame signal. The convolution stride is (2, 1), which can reduce the number of signal frequency domain features by half layer by layer, while the number of time domain frames remains unchanged, which plays a role in reducing dimensionality and reducing the amount of calculation. The Decoder part is mainly composed of DecTConv2d with a transposed two-dimensional convolution (ConvTranspose2d) as the kernel. The parameters of DecTConv2d of each layer are the same as those of the corresponding EncConv2d, which realizes the restoration of the signal dimension. Between the Encoder and the Decoder, the present invention uses a recurrent neural network module RNNs composed of stacked GRUs (Gated Recurrent Units). The main function of RNNs is to extract and analyze the inter-frame temporal information of the audio signal. Therefore, the workflow of the deep learning network module is as follows: the Encoder receives the short-time Fourier transform representation of the noisy audio signal from the signal preprocessing module. Then, EncConv2d extracts high-dimensional features layer by layer, and the corresponding output is given to DecTConv2d through skip connections. RNNs accept the output features from their last layer of EncConv2d, perform temporal information extraction and analysis, and give the input to the Decoder. The Decoder accepts the output from RNNs and Encoder, and increases the dimension layer by layer through transposed convolution. The number of channels in the last layer of DecTConv2d is 3, and finally three outputs are obtained, namely, the amplitude spectrum mask estimate, Phase spectrum mask estimation cosine value Sine value

[0128] After obtaining the mask estimate of the speech signal, the post-processing speech generation module 503 modulates the short-time Fourier transform complex spectrum of the original noisy signal to obtain the STFT of the enhanced speech, which is expressed as follows:

[0129] The enhanced speech amplitude spectrum estimate is determined as follows:

[0130] The phase cosine estimation information is determined as follows:

[0131] The phase sine estimation information is determined as follows:

[0132] The enhanced speech STFT estimation information is determined as follows:

[0133] After obtaining the short-time Fourier transform estimation information of the pure speech, the inverse short-time Fourier transform (iSTFT) corresponding to the STFT is finally performed to obtain the time domain waveform signal of the enhanced speech.

[0134] Through the above-mentioned implementation of the present application, a speech enhancement solution is provided for simultaneously modeling and analyzing the amplitude spectrum and phase spectrum of a noisy frequency signal. Without almost increasing the network computational complexity, the model is enabled to learn pure human voice phase information, thereby significantly improving the noise reduction effect.

[0135] In an optional embodiment, the present application further provides a method for training an audio signal processing model, as shown in FIG6 , comprising the following steps:

[0136] S602. Obtain a first spectral feature of a sample noisy frequency signal and a training label matching the sample noisy frequency signal, wherein the sample noisy frequency signal includes a sample reference audio signal and a sample target audio signal, and the training label includes a phase difference label, which is determined based on a cosine value and a sine value of a first sample vector indicated by the first spectral feature of the sample noisy frequency signal, and a cosine value and a sine value of a second sample vector indicated by a second spectral feature of the sample target audio signal; the sample reference audio signal and the sample target audio signal are respectively one of a sample noise signal or a sample clean speech signal in the sample noisy frequency signal.

[0137] Wherein, when the reference audio signal is a noise signal in a noisy audio signal, the target audio signal is a clean speech signal in the noisy audio signal, the sample reference audio signal is a sample noise signal in the sample noisy audio signal, and the sample target audio signal is a sample clean speech signal in the sample noisy audio signal. When the reference audio signal is a clean speech signal in a noisy audio signal, the target audio signal is a noise signal in the noisy audio signal, the sample reference audio signal is a sample clean speech signal in the sample noisy audio signal, and the sample target audio signal is a sample noise signal in the sample noisy audio signal.

[0138] S604: Input the first spectrum feature into the audio signal processing model to be trained to obtain reference phase estimation information that matches the sample noisy audio signal.

[0139] S606 : Train the audio signal processing model according to the reference phase estimation information and the training loss determined by the training label.

[0140] S608: When the training loss satisfies the target convergence condition, the trained audio signal processing model is determined as the target neural network model.

[0141] In an optional embodiment, the above-mentioned training of the audio signal processing model based on the training loss determined according to the reference phase estimation information and the training label includes: inputting the first spectral feature into the audio signal processing model to be trained to obtain reference amplitude estimation information matching the sample noisy frequency signal, wherein the amplitude difference label is determined based on the amplitude difference between the sample noisy frequency signal and the sample target audio signal in multiple frequency components; and training the audio signal processing model based on the training loss determined according to the reference amplitude estimation information, the reference phase estimation information and the training label.

[0142] In the above-mentioned implementation manner of the present application, the corresponding sample noisy audio signal and the sample target audio signal can be determined according to the actual use of the above-mentioned target neural network model.

[0143] When the target neural network model is used to perform speech noise reduction operations, the sample noisy audio signal may be a noisy sample speech signal, the sample target audio signal may be a pure sample speech signal, and the reference sample speech signal may be a noise signal.

[0144] When the target neural network model is used to perform a noise extraction operation, the sample noisy audio signal may be a noisy sample speech signal, the sample target audio signal may be a noise signal, and the reference sample speech signal may be a pure speech signal.

[0145] Taking the above target neural network model for speech noise reduction as an example, in the above training process, a clean speech data set (sample target audio signal) and a noise data set (reference noise signal) can be used to mix and generate a noisy audio signal (sample noisy audio signal), and by controlling the noise mixing ratio to simulate the signal-to-noise ratio of the noisy audio signal in different noise environments, the supervised learning method is used to train the model. Assume that the clean speech signal is s n , the noise signal is d n , the corresponding short-time Fourier transforms are S k and D k , then there is a noisy frequency signal x n is: x n =s n +d n

[0146] Correspondingly, the representation after short-time Fourier transform is as follows: X k =S k +D k

[0147] Among them, X k is x n Frequency domain representation, S k For s n Frequency domain representation of D k d n Frequency domain representation of .

[0148] In an optional embodiment, the above-mentioned acquisition of the first spectral feature of the sample noisy frequency signal and the training label matching the sample noisy frequency signal includes: acquiring a sample target audio signal and a sample reference audio signal, and obtaining the sample noisy frequency signal based on mixing the sample target audio signal and the sample reference audio signal; respectively acquiring the first spectral feature of the sample noisy frequency signal and the second spectral feature of the sample target audio signal; and determining the amplitude difference label and the phase difference label based on the first spectral feature and the second spectral feature.

[0149] In an optional embodiment, the above-mentioned determination of the amplitude difference label and the phase difference label based on the first spectral feature and the second spectral feature includes: obtaining the first sample amplitude value of the sample noisy frequency signal on N frequency components according to the first spectral feature, and obtaining the second sample amplitude value of the sample target audio signal on N frequency components according to the second spectral feature, where N is an integer greater than 1; determining the amplitude difference label according to the ratio between the N second sample amplitude values ​​and their respective corresponding first sample amplitude values; obtaining the first sample vector cosine value and the first sample vector sine value of the sample noisy frequency signal on N frequency components according to the first spectral feature, and obtaining the second sample vector cosine value and the second sample vector sine value of the sample target audio signal on N frequency components according to the second spectral feature; determining the phase difference label according to the N first sample vector cosine values, the first sample vector sine values ​​and the respective corresponding second sample vector cosine values, the second sample vector sine values.

[0150] The ideal amplitude mask |m for determining the amplitude difference label can be determined in the above way. k |, which can be specifically expressed as:

[0151] The ideal phase cosine mask for determining the phase difference tag can be determined in the above way. Specifically, it can be expressed as:

[0152] The ideal phase sinusoidal mask for determining the phase difference tag can be determined in the above way. Specifically, it can be expressed as:

[0153] It is understandable that after the phase difference labels and the amplitude difference labels are determined through the above embodiments of the present application, the audio signal processing model shown in FIG3 can be trained based on the above phase difference labels and amplitude difference labels.

[0154] In an optional embodiment, the above-mentioned inputting the first spectral feature into the audio signal processing model to be trained to obtain reference amplitude estimation information and reference phase estimation information that match the sample noisy frequency signal includes: in the audio signal processing model, encoding the first spectral feature through an encoding network to obtain a first encoding result; analyzing the first encoding result through a recurrent neural network constructed based on a gated recurrent unit to obtain a first intermediate result carrying timing information; inputting the first encoding result and the first intermediate result into a decoding network in the audio signal processing model to obtain reference amplitude estimation information and reference phase estimation information that match the sample noisy frequency signal, wherein the sub-network in the decoding network is obtained based on the adjustment of the sub-network in the encoding network.

[0155] It can be understood that during the training process, the audio signal processing model shown in Figure 3 is used based on the first spectral feature to obtain reference amplitude estimation information and reference phase estimation information that match the sample noisy frequency signal, and then the training loss is determined in combination with the training label that matches the sample noisy frequency signal.

[0156] Specifically, the above training loss can be obtained as follows:

[0157] in, For characterization The loss in the time domain may be any one of the following: mean square error loss function (MSE Loss, referred to as MSE), error loss function (Mean Absolute Error, referred to as MAE), scale invariant signal-to-noise ratio (Scale invariant Signal-to-Noise Ratio, referred to as SI-SNR), etc. for The frequency domain loss can be either the mean square error (MSE) loss function or the mean absolute error (MAE) loss function. λ is the loss weight distribution coefficient, which is determined based on the experimental training results.

[0158] Furthermore, 1000 sets of test data with a signal-to-noise ratio range of [-10, 30] dB and a step of 2 dB were used to obtain the test results of this embodiment. Among them, the speech perceptual quality parameter PESQ, the scale-invariant signal-to-noise ratio parameter SI-SNR, and the simulated subjective audio quality perception parameter DNSMOS were selected as effect evaluation indicators to determine the test result value. Specifically, Figure 7 is the PESQ indicator test result, Figure 8 is the SI-SNR indicator test result, and Figure 9 is the MOS_OVL indicator test result. It can be seen that the target neural network model determined by the above-mentioned embodiment of the present application has the ability to learn pure human voice phase information. Compared with the existing technical solutions, the present invention obviously has a higher effect upper limit, and higher learning efficiency and output efficiency.

[0159] Through the above-mentioned implementation mode of the present application, the first spectral feature of the noisy audio signal is first obtained, wherein the noisy audio signal includes a reference audio signal and a target audio signal to be extracted; the reference audio signal and the target audio signal are respectively one of the noise signal or the pure speech signal in the noisy audio signal; the above-mentioned first spectral feature is input into the target neural network model to obtain reference amplitude estimation information and reference phase estimation information matching the above-mentioned noisy audio signal, wherein the above-mentioned target neural network model is pre-trained based on sample noisy audio signals and sample target audio signals, and the above-mentioned reference amplitude estimation information is used to indicate the difference between the above-mentioned target audio signal and the above-mentioned noisy audio signal in multiple frequencies. The reference phase estimation information is used to indicate the phase difference between the target audio signal and the noisy audio signal on the multiple frequency components; the second spectrum feature of the target audio signal is determined according to the first spectrum feature, the reference amplitude estimation information and the reference phase estimation information; the target audio signal is determined according to the second spectrum feature, thereby realizing the prediction of the estimated values ​​for characterizing the phase difference and amplitude difference between the audio signals through the target neural network model, and then determining the processed target audio signal according to the predicted reference amplitude estimation information and reference phase estimation information and the spectrum feature of the original audio signal.

[0160] According to the above-mentioned embodiment of the present application, a technical solution is proposed that can simultaneously model and analyze the amplitude information and phase information of a noisy audio signal, thereby parsing the audio signal interfered with by noise through comprehensive processing of the amplitude information and phase information, thereby improving the accuracy of audio signal processing; at the same time, a target neural network is used to obtain amplitude information and phase information estimation, avoiding the use of a complex neural network model, avoiding an increase in calculation amount, and improving the processing efficiency of audio signal processing.

[0161] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0162] According to another aspect of the embodiments of the present application, there is also provided an audio signal processing device for implementing the above-mentioned audio signal processing method. As shown in FIG10 , the device includes:

[0163] The acquisition unit 1002 is configured to acquire a first spectral feature of a noisy audio signal, wherein the noisy audio signal includes a reference audio signal and a target audio signal to be extracted; the reference audio signal and the target audio signal are respectively one of a noise signal and a clean speech signal in the noisy audio signal.

[0164] The processing unit 1004 is used to input the above-mentioned first spectral feature into the target neural network model to obtain reference phase estimation information matching the above-mentioned noisy audio signal, wherein the above-mentioned target neural network model is pre-trained based on the sample noisy audio signal and the sample target audio signal, and the above-mentioned reference phase estimation information is used to indicate the phase difference between the above-mentioned target audio signal and the above-mentioned noisy audio signal in the above-mentioned multiple frequency components.

[0165] The first determination unit 1006 is used to determine the second phase sine value and the second phase cosine value for determining the second spectral feature based on the first phase sine value and the first phase cosine value indicated by the first spectral feature, and the reference phase sine value and the reference phase cosine value indicated by the reference phase estimation information.

[0166] The second determining unit 1008 is configured to determine a second frequency spectrum feature according to the second phase sine value and the second phase cosine value, and determine a target audio signal according to the second frequency spectrum feature.

[0167] Optionally, the first determination unit 1006 includes: an acquisition module for acquiring reference amplitude estimation information output by the target neural network model based on the first spectral feature, wherein the reference amplitude estimation information is used to indicate the amplitude difference between the target audio signal and the noisy audio signal in multiple frequency components; a first determination module for determining a second audio amplitude matching the second spectral feature based on the first audio amplitude indicated by the first spectral feature and the reference amplitude estimation information; and a second determination module for determining the second spectral feature based on the second audio amplitude, the second phase sine value and the second phase cosine value.

[0168] Optionally, the above-mentioned third determination module is used to: when the above-mentioned reference amplitude estimation information is used to indicate the reference amplitude ratio of the above-mentioned target audio signal to the above-mentioned noisy audio signal on the above-mentioned N frequency components, obtain the first amplitude values ​​on the above-mentioned N frequency components indicated by the above-mentioned first spectral characteristics, and the N reference amplitude ratios indicated by the above-mentioned reference amplitude estimation information, wherein the above-mentioned N is an integer greater than 1; and determine the product value of the above-mentioned N first amplitude values ​​and their respective corresponding reference amplitude ratios as the second amplitude value on the above-mentioned N frequency components matching the above-mentioned second spectral characteristics.

[0169] Optionally, the second determination module is configured to: when the reference amplitude estimation information is used to indicate the phase difference between the target audio signal and the noisy audio signal on the N frequency components, traverse the N frequency components indicated by the first spectral feature, and obtain the first phase sine value and the first phase cosine value corresponding to the currently traversed frequency component, wherein N is an integer greater than 1; obtain the reference phase sine value and the reference phase cosine value corresponding to the currently traversed frequency component from the reference phase estimation information; obtain a first product of the first phase sine value and the reference phase cosine value, and a second product of the first phase cosine value and the reference phase sine value, and determine the sum of the first product and the second product as the second phase sine value; obtain a third product of the first phase cosine value and the reference phase cosine value, and a fourth product of the first phase sine value and the reference phase sine value, and determine the second phase cosine value based on the difference between the third product and the fourth product.

[0170] Optionally, the third determination module is configured to: determine a real part representation of the second spectral feature based on the second audio amplitude and the second phase cosine value; determine an imaginary part representation of the second spectral feature based on the second audio amplitude and the second phase sine value; and determine the second spectral feature based on the real part representation and the imaginary part representation.

[0171] Optionally, the above-mentioned processing unit 1004 includes: a first processing module, used to encode the above-mentioned first spectral features through the encoding network in the above-mentioned target neural network model to obtain a first encoding result; a second processing module, used to analyze the above-mentioned first encoding result through a recurrent neural network constructed based on a gated recurrent unit to obtain a first intermediate result carrying timing information; a third processing module, used to input the above-mentioned first encoding result and the above-mentioned first intermediate result into the decoding network in the above-mentioned target neural network model to obtain reference amplitude estimation information and reference phase estimation information matching the above-mentioned noisy frequency signal, wherein the sub-network in the above-mentioned decoding network is obtained based on the adjustment of the sub-network in the above-mentioned encoding network.

[0172] Optionally, the above-mentioned first processing module is used to: sequentially encode the above-mentioned first spectral features through M encoding sub-networks with a connection relationship in the above-mentioned encoding network to obtain M above-mentioned first encoding results, wherein each of the above-mentioned encoding sub-networks respectively includes: a convolution layer, a normalization layer and an activation layer, and when the above-mentioned first spectral features corresponding to each frame are convolutionally processed in the above-mentioned convolution layer, the above-mentioned first spectral features corresponding to the adjacent previous frame will be referred to, and M is a natural number greater than or equal to 2; the above-mentioned second processing module is used to: analyze the above-mentioned first encoding results through the recurrent neural network constructed based on the gated recurrent unit to obtain the first intermediate result carrying time series information, including: inputting the above-mentioned first encoding result output by the Mth above-mentioned encoding sub-network into the above-mentioned recurrent neural network to obtain the above-mentioned first intermediate result carrying time series information; the above-mentioned third processing module is used to: The first encoding result and the above-mentioned first intermediate result are input into the decoding network in the above-mentioned target neural network model to obtain the reference amplitude estimation information and reference phase estimation information matching the above-mentioned noisy frequency signal, including: inputting the above-mentioned first intermediate result and the above-mentioned first encoding result output by the i-th above-mentioned encoding subnetwork into the M-i+1-th decoding subnetwork, wherein the above-mentioned decoding network includes M decoding subnetworks with a connection relationship, and each of the above-mentioned decoding subnetworks respectively includes: a transposed convolution layer, a normalization layer and an activation layer associated with the above-mentioned convolution layer, the above-mentioned i is an integer greater than or equal to 1 and less than or equal to the above-mentioned M, and a jump connection is set between the i-th above-mentioned encoding subnetwork and the M-i+1-th above-mentioned decoding subnetwork; obtaining the first decoding result output by the M-th above-mentioned encoding subnetwork; and determining the above-mentioned first decoding result as the reference amplitude estimation information and reference phase estimation information matching the above-mentioned noisy frequency signal.

[0173] Optionally, the acquisition unit 1002 is configured to: acquire the noisy audio signal; resample the noisy audio signal according to a target sampling rate to obtain a first reference audio signal; perform time-domain framing and windowing processing on the first reference audio signal according to a framing parameter to obtain a plurality of first audio sub-signals; and perform discrete Fourier transform on each of the plurality of first audio sub-signals to obtain the first spectral features corresponding to each of the plurality of first audio sub-signals.

[0174] Optionally, the acquisition unit 1002 is configured to: acquire a plurality of second spectral features determined based on the first spectral features corresponding to the plurality of first audio sub-signals; perform inverse short-time Fourier transform on the plurality of second spectral features to obtain a plurality of second audio sub-signals; and determine the target audio signal based on a concatenation result of the plurality of second audio sub-signals.

[0175] Optionally, the audio signal processing device is further configured to: determine a reference spectrum feature of the reference audio signal based on the first spectrum feature and the second spectrum feature; and perform an inverse short-time Fourier transform on the reference spectrum feature to obtain the reference audio signal.

[0176] For a specific embodiment, please refer to the example shown in the above-mentioned audio signal processing method, and this embodiment will not be described in detail here.

[0177] According to another aspect of the embodiments of the present application, an electronic device for implementing the above-mentioned audio signal processing method is also provided. This embodiment is illustrated using the electronic device as a terminal as an example. As shown in Figure 11, the electronic device includes a memory 1102 and a processor 1104. The memory 1102 stores a computer program, and the processor 1104 is configured to execute the steps of any of the above-mentioned method embodiments through the computer program. Optionally, in this embodiment, the above-mentioned electronic device can be located in at least one network device among multiple network devices of a computer network.

[0178] Alternatively, those skilled in the art will appreciate that the structure shown in FIG11 is merely illustrative, and the electronic device may also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal device. FIG11 does not limit the structure of the electronic device. For example, the electronic device may include more or fewer components (such as a network interface, etc.) than those shown in FIG11, or may have a configuration different from that shown in FIG11.

[0179] Memory 1102 can be used to store software programs and modules, such as program instructions / modules corresponding to the audio signal processing method and apparatus in the embodiments of the present application. Processor 1104 executes the software programs and modules stored in memory 1102 to perform various functional applications and data processing, thereby implementing the aforementioned audio signal processing method. Memory 1102 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, memory 1102 may further include memory remotely located from processor 1104, which can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. Memory 1102 can specifically, but is not limited to, be used to store information such as the target audio signal. As an example, as shown in FIG11 , memory 1102 may, but is not limited to, include the acquisition unit 1002, processing unit 1104, first determination unit 1006, and second determination unit 1008 of the aforementioned audio signal processing apparatus. In addition, other module units in the above-mentioned audio signal processing device may also be included but not limited to, which will not be described in detail in this example.

[0180] Optionally, the transmission device 1106 is configured to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one embodiment, the transmission device 1106 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 1106 is a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0181] In addition, the electronic device further includes: a display 1108 and a connection bus 1110 for connecting various module components in the electronic device.

[0182] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes via network communication. The nodes may form a point-to-point network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the point-to-point network.

[0183] According to one aspect of the present application, a computer program product is provided, comprising a computer program / instructions containing program code for executing the above-described method. In such an embodiment, the computer program can be downloaded and installed from a network via a communication component and / or installed from a removable medium. When the computer program is executed by a central processing unit, the various functions provided in the embodiments of the present application are performed.

[0184] According to one aspect of the present application, a computer-readable storage medium is provided. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device performs the above-mentioned audio signal processing method.

[0185] According to another aspect of the embodiments of the present application, a training device for an audio signal processing model for implementing the above-mentioned training method for an audio signal processing model is also provided. As shown in FIG12 , the device includes:

[0186] An acquisition unit 1202 is configured to acquire a first spectral feature of a sample noisy frequency signal and a training label matching the sample noisy frequency signal, wherein the sample noisy frequency signal includes a sample reference audio signal and a sample target audio signal; the training label includes a phase difference label, and the phase difference label is determined based on a cosine value and a sine value of a first sample vector indicated by the first spectral feature of the sample noisy frequency signal, and a cosine value and a sine value of a second sample vector indicated by a second spectral feature of the sample target audio signal; and the sample reference audio signal and the sample target audio signal are respectively one of a sample noise signal or a sample clean speech signal in the sample noisy frequency signal.

[0187] The prediction unit 1204 is configured to input the first spectrum feature into the audio signal processing model to be trained to obtain reference phase estimation information that matches the sample noisy audio signal.

[0188] The training unit 1206 is configured to train the audio signal processing model according to the reference phase estimation information and the training loss determined by the training label.

[0189] The determining unit 1208 is configured to determine the trained audio signal processing model as a target neural network model when the training loss satisfies a target convergence condition.

[0190] Optionally, the training unit 1206 is configured to: input the first spectral feature into the audio signal processing model to be trained to obtain reference amplitude estimation information matching the sample noisy frequency signal, wherein the amplitude difference label is determined based on the amplitude difference between the sample noisy frequency signal and the sample target audio signal in multiple frequency components; and train the audio signal processing model based on the training loss determined by the reference amplitude estimation information, the reference phase estimation information, and the training label.

[0191] Optionally, the acquisition unit 1202 includes: a first acquisition module, configured to acquire the sample target audio signal and the sample reference audio signal, and obtain the sample noisy frequency signal by mixing the sample target audio signal and the sample reference audio signal; a second acquisition module, configured to respectively acquire the first spectral feature of the sample noisy frequency signal and the second spectral feature of the sample target audio signal; and a third acquisition module, configured to determine the amplitude difference label and the phase difference label based on the first spectral feature and the second spectral feature.

[0192] Optionally, the above-mentioned third acquisition module is used to: obtain the first sample amplitude value of the above-mentioned sample noisy frequency signal on the N above-mentioned frequency components according to the above-mentioned first spectral feature, and obtain the second sample amplitude value of the above-mentioned sample target audio signal on the N above-mentioned frequency components according to the above-mentioned second spectral feature, wherein the above-mentioned N is an integer greater than 1; determine the above-mentioned amplitude difference label according to the ratio between the N above-mentioned second sample amplitude values ​​and the respectively corresponding first sample amplitude values; obtain the first sample vector cosine value and the first sample vector sine value of the above-mentioned sample noisy frequency signal on the N above-mentioned frequency components according to the above-mentioned first spectral feature, and obtain the second sample vector cosine value and the second sample vector sine value of the above-mentioned sample target audio signal on the N above-mentioned frequency components according to the above-mentioned second spectral feature; determine the above-mentioned phase difference label according to the N above-mentioned first sample vector cosine values, the above-mentioned first sample vector sine values ​​and the respectively corresponding second sample vector cosine values, and the above-mentioned second sample vector sine values.

[0193] Optionally, the prediction unit 1204 is configured to: in the audio signal processing model, encode the first spectral feature through an encoding network to obtain a first encoding result; analyze the first encoding result through a recurrent neural network constructed based on a gated cyclic unit to obtain a first intermediate result carrying timing information; and input the first encoding result and the first intermediate result into a decoding network in the audio signal processing model to obtain the reference amplitude estimation information and the reference phase estimation information that match the sample noisy frequency signal, wherein the subnetwork in the decoding network is obtained based on adjustments to the subnetwork in the encoding network.

[0194] For a specific embodiment, please refer to the example shown in the training method of the audio signal processing model mentioned above, and this embodiment will not be described in detail here.

[0195] According to another aspect of the embodiments of the present application, an electronic device for implementing the above-mentioned training method for the audio signal processing model is also provided. This embodiment is illustrated using the electronic device as a terminal. As shown in Figure 13, the electronic device includes a memory 1302 and a processor 1304. The memory 1302 stores a computer program, and the processor 1304 is configured to execute the steps of any of the above-mentioned method embodiments through the computer program.

[0196] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0197] Alternatively, those skilled in the art will appreciate that the structure shown in FIG13 is for illustration only, and the electronic device may also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal device. FIG13 does not limit the structure of the electronic device. For example, the electronic device may include more or fewer components (such as a network interface, etc.) than those shown in FIG13, or may have a configuration different from that shown in FIG13.

[0198] Memory 1302 can be used to store software programs and modules, such as program instructions / modules corresponding to the audio signal processing model training method and apparatus in the embodiments of the present application. Processor 1304 executes the software programs and modules stored in memory 1302 to perform various functional applications and data processing, thereby implementing the aforementioned audio signal processing model training method. Memory 1302 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, memory 1302 may further include memory remotely located relative to processor 1304, which can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. Memory 1302 can specifically, but is not limited to, be used to store information such as the target audio signal. As an example, as shown in FIG13 , memory 1302 may, but is not limited to, include the acquisition unit 1202, prediction unit 1204, training unit 1206, and determination unit 1210 of the aforementioned audio signal processing model training apparatus. In addition, other module units in the training device of the above-mentioned audio signal processing model may also be included but not limited to, which will not be repeated in this example.

[0199] Optionally, the transmission device 1306 is configured to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one embodiment, the transmission device 1306 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 1306 is a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0200] In addition, the electronic device further includes: a display 1308 and a connection bus 1310 for connecting various module components in the electronic device.

[0201] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes via network communication. The nodes may form a point-to-point network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the point-to-point network.

[0202] According to one aspect of the present application, a computer program product is provided, comprising a computer program / instructions containing program code for executing the above-described method. In such an embodiment, the computer program can be downloaded and installed from a network via a communication component and / or installed from a removable medium. When the computer program is executed by a central processing unit, the various functions provided in the embodiments of the present application are performed.

[0203] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the above-mentioned training method of the audio signal processing model.

[0204] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0205] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling one or more electronic devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods of each embodiment of the present application.

[0206] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program having a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0207] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0208] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.

[0209] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0210] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0211] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0212] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A method for processing an audio signal, performed by an electronic device, comprising: Acquire a first frequency spectrum feature of a noisy audio signal, wherein the noisy audio signal includes a reference audio signal and a target audio signal to be extracted; the reference audio signal and the target audio signal are respectively one of a noise signal or a clean speech signal in the noisy audio signal; Inputting the first spectrum feature into a target neural network model to obtain reference phase estimation information matching the noisy audio signal, wherein the target neural network model is pre-trained based on a sample noisy audio signal and a sample target audio signal, and the reference phase estimation information is used to indicate a phase difference between the target audio signal and the noisy audio signal on multiple frequency components; the sample target audio signal is a sample noise signal or a sample clean speech signal in the sample noisy audio signal; Determine a second phase sine value and a second phase cosine value for determining a second spectrum feature according to a first phase sine value and a first phase cosine value indicated by the first spectrum feature, and a reference phase sine value and a reference phase cosine value indicated by the reference phase estimation information; and The second frequency spectrum feature is determined according to the second phase sine value and the second phase cosine value, and the target audio signal is determined according to the second frequency spectrum feature.

2. The method according to claim 1, wherein determining the second frequency spectrum feature according to the second phase sine value and the second phase cosine value comprises: Obtaining reference amplitude estimation information output by the target neural network model based on the first frequency spectrum feature, wherein the reference amplitude estimation information is used to indicate the amplitude difference between the target audio signal and the noisy audio signal in multiple frequency components; determining, according to the first audio amplitude indicated by the first frequency spectrum feature and the reference amplitude estimation information, a second audio amplitude for determining the second frequency spectrum feature; The second frequency spectrum feature is determined according to the second audio amplitude, the second phase sine value and the second phase cosine value.

3. The method according to claim 2, wherein determining the second audio amplitude for determining the second frequency spectrum feature according to the first audio amplitude indicated by the first frequency spectrum feature and the reference amplitude estimation information comprises: In a case where the reference amplitude estimation information is used to indicate a reference amplitude ratio of the target audio signal to the noisy audio signal on the N frequency components, obtaining first amplitude values ​​on the N frequency components indicated by the first frequency spectrum feature and N reference amplitude ratios indicated by the reference amplitude estimation information, wherein N is an integer greater than 1; The product values ​​of the N first amplitude values ​​and the corresponding reference amplitude ratios are determined as second amplitude values ​​on the N frequency components matching the second frequency spectrum characteristics.

4. The method according to claim 2 or 3, wherein determining the second phase sine value and the second phase cosine value for determining the second spectrum feature according to the first phase sine value and the first phase cosine value indicated by the first spectrum feature, and the reference phase sine value and the reference phase cosine value indicated by the reference phase estimation information comprises: In a case where the reference amplitude estimation information is used to indicate a phase difference between the target audio signal and the noisy audio signal on N frequency components, traversing the N frequency components indicated by the first frequency spectrum feature, and acquiring the first phase sine value and the first phase cosine value corresponding to the currently traversed frequency component, wherein N is an integer greater than 1; Acquire the reference phase sine value and the reference phase cosine value corresponding to the currently traversed frequency component from the reference phase estimation information; Obtaining a first product of the first phase sine value and the reference phase cosine value, and a second product of the first phase cosine value and the reference phase sine value, and determining a sum of the first product and the second product as the second phase sine value; Obtain a third product of the first phase cosine value and the reference phase cosine value, and the first phase sine The fourth product of the third product and the reference phase sine value is obtained, and the second phase cosine value is determined according to a difference between the third product and the fourth product.

5. The method according to any one of claims 2 to 4, wherein determining the second frequency spectrum feature according to the second audio amplitude, the second phase sine value and the second phase cosine value comprises: Determine a real part representation in the second frequency spectrum feature according to the second audio amplitude and the second phase cosine value; Determine the imaginary part representation in the second frequency spectrum feature according to the second audio amplitude and the second phase sine value; The second spectral characteristic is determined based on the real part representation and the imaginary part representation.

6. According to the method according to any one of claims 1 to 5, the step of inputting the first spectrum feature into a target neural network model to obtain reference amplitude estimation information and reference phase estimation information matching the noisy frequency signal comprises: In the target neural network model, encoding the first spectrum feature through an encoding network to obtain a first encoding result; Analyzing the first encoding result through a recurrent neural network constructed based on a gated recurrent unit to obtain a first intermediate result carrying timing information; The first encoding result and the first intermediate result are input into a decoding network in the target neural network model to obtain reference amplitude estimation information and reference phase estimation information that match the noisy frequency signal, wherein the sub-network in the decoding network is obtained based on an adjustment of the sub-network in the encoding network.

7. The method according to claim 6, wherein encoding the first spectrum feature through a coding network comprises: The first spectrum feature is encoded in sequence by M encoding sub-networks having a connection relationship in the encoding network to obtain M first encoding results, wherein each of the encoding sub-networks includes: a convolution layer, a normalization layer and an activation layer, and when the first spectrum feature corresponding to each frame is convoluted in the convolution layer, the first spectrum feature corresponding to the adjacent previous frame is referred to, and M is a natural number greater than or equal to 2; The step of analyzing the first coding result by a recurrent neural network constructed based on a gated recurrent unit to obtain a first intermediate result carrying time sequence information includes: inputting the first coding result output by the Mth coding subnetwork into the recurrent neural network to obtain the first intermediate result carrying time sequence information; The step of inputting the first encoding result and the first intermediate result into the decoding network in the target neural network model to obtain reference amplitude estimation information and reference phase estimation information matching the noisy frequency signal comprises: inputting the first intermediate result and the first encoding result output by the ith encoding subnetwork into the M-i+1th decoding subnetwork, wherein the decoding network comprises M decoding subnetworks having a connection relationship, each of the decoding subnetworks respectively comprising: a transposed convolution layer, a normalization layer and an activation layer associated with the convolution layer, the i being an integer greater than or equal to 1 and less than or equal to the M, a jump connection being arranged between the ith encoding subnetwork and the M-i+1th decoding subnetwork; obtaining the first decoding result output by the Mth encoding subnetwork; and determining the first decoding result as the reference amplitude estimation information and reference phase estimation information matching the noisy frequency signal.

8. The method according to claim 6 or 7, wherein obtaining the first frequency spectrum feature of the noisy frequency signal comprises: Acquire the noisy frequency signal; Resampling the noisy audio signal according to a target sampling rate to obtain a first reference audio signal; Performing time-domain frame division and windowing processing on the first reference audio signal according to a frame division parameter to obtain a plurality of first audio sub-signals; Discrete Fourier transform is performed on each of the first audio sub-signals to obtain the first frequency spectrum features respectively corresponding to the first audio sub-signals.

9. The method according to claim 8, wherein determining the target audio signal according to the second frequency spectrum feature comprises: Acquire a plurality of second frequency spectrum features determined according to the first frequency spectrum features respectively corresponding to a plurality of the first audio sub-signals; Performing an inverse short-time Fourier transform on the plurality of second frequency spectrum features to obtain a plurality of second audio sub-signals; The target audio signal is determined according to a splicing result of the plurality of second audio sub-signals.

10. The method according to any one of claims 1 to 9, after determining the second frequency spectrum feature according to the second phase sine value and the second phase cosine value, further comprising: Determining a reference frequency spectrum feature of the reference audio signal according to the first frequency spectrum feature and the second frequency spectrum feature; Performing an inverse short-time Fourier transform on the reference frequency spectrum feature to obtain the reference audio signal.

11. A method for training an audio signal processing model, performed by an electronic device, comprising: Acquire a first spectrum feature of a sample noisy audio signal and a training label matched with the sample noisy audio signal, wherein the sample noisy audio signal includes a sample reference audio signal and a sample target audio signal, and the training label includes a phase difference label, and the phase difference label is determined according to a first sample vector cosine value and a first sample vector sine value indicated by the first spectrum feature of the sample noisy audio signal, and a second sample vector cosine value and a second sample vector sine value indicated by a second spectrum feature of the sample target audio signal; the sample reference audio signal and the sample target audio signal are respectively one of a sample noise signal or a sample clean speech signal in the sample noisy audio signal; Inputting the first frequency spectrum feature into the audio signal processing model to be trained to obtain reference phase estimation information matching the sample noisy audio signal; Training the audio signal processing model according to the reference phase estimation information and the training loss determined by the training label; and When the training loss satisfies the target convergence condition, the trained audio signal processing model is determined as the target neural network model.

12. The method according to claim 11, wherein training the audio signal processing model according to the training loss determined by the reference phase estimation information and the training label comprises: Inputting the first spectrum feature into an audio signal processing model to be trained to obtain reference amplitude estimation information matching the sample noisy audio signal, wherein the amplitude difference label is determined according to the amplitude difference between the sample noisy audio signal and the sample target audio signal on multiple frequency components; The audio signal processing model is trained according to the reference amplitude estimation information, the reference phase estimation information, and a training loss determined by the training label.

13. The method according to claim 12, wherein obtaining a first spectrum feature of a sample noisy frequency signal and a training label matching the sample noisy frequency signal comprises: Acquire the sample target audio signal and the sample reference audio signal, and obtain the sample noisy audio signal by mixing the sample target audio signal and the sample reference audio signal; Respectively obtaining a first frequency spectrum feature of the sample noisy audio signal and a second frequency spectrum feature of the sample target audio signal; The amplitude difference label and the phase difference label are determined according to the first spectrum feature and the second spectrum feature.

14. The method according to claim 13, wherein determining the amplitude difference label and the phase difference label according to the first spectrum feature and the second spectrum feature comprises: Acquire a first sample amplitude value of the sample noisy audio signal on the N frequency components according to the first frequency spectrum feature, and acquire a second sample amplitude value of the sample target audio signal on the N frequency components according to the second frequency spectrum feature, wherein N is an integer greater than 1; Determining the amplitude difference label according to the ratio between the N second sample amplitude values ​​and the first sample amplitude values ​​corresponding to each of them; Acquire a first sample vector cosine value and a first sample vector sine value of the sample noisy audio signal on the N frequency components according to the first frequency spectrum feature, and acquire a second sample vector cosine value and a second sample vector sine value of the sample target audio signal on the N frequency components according to the second frequency spectrum feature; The phase difference label is determined according to N cosine values ​​of the first sample vectors, the sine values ​​of the first sample vectors and the corresponding cosine values ​​of the second sample vectors, and the sine values ​​of the second sample vectors.

15. The method according to claim 13 or 14, wherein the step of inputting the first spectrum feature into the audio signal processing model to be trained to obtain reference amplitude estimation information and reference phase estimation information matching the sample noisy audio signal comprises: In the audio signal processing model, encoding the first spectrum feature through an encoding network to obtain a first encoding result; Analyzing the first encoding result through a recurrent neural network constructed based on a gated recurrent unit to obtain a first intermediate result carrying timing information; The first encoding result and the first intermediate result are input into a decoding network in the audio signal processing model to obtain the reference amplitude estimation information and the reference phase estimation information that match the sample noisy audio signal, wherein the sub-network in the decoding network is obtained based on an adjustment of the sub-network in the encoding network.

16. An audio signal processing device, comprising: An acquisition unit is used to acquire a first frequency spectrum feature of a noisy audio signal, wherein the noisy audio signal includes a reference audio signal and a target audio signal to be extracted; the reference audio signal and the target audio signal are respectively one of a noise signal or a clean speech signal in the noisy audio signal; A processing unit, configured to input the first spectrum feature into a target neural network model to obtain matching information with the noisy audio signal and reference phase estimation information, wherein the target neural network model is pre-trained based on a sample noisy audio signal and a sample target audio signal, and the reference phase estimation information is used to indicate a phase difference between the target audio signal and the noisy audio signal on multiple frequency components; the sample target audio signal is a sample noise signal or a sample clean speech signal in the sample noisy audio signal; a first determining unit, configured to determine a second phase sine value and a second phase cosine value for determining a second spectrum feature according to a first phase sine value and a first phase cosine value indicated by the first spectrum feature, and a reference phase sine value and a reference phase cosine value indicated by the reference phase estimation information; and The second determining unit is configured to determine the second frequency spectrum feature according to the second phase sine value and the second phase cosine value, and determine the target audio signal according to the second frequency spectrum feature.

17. A training device for an audio signal processing model, comprising: An acquisition unit is used to acquire a first spectrum feature of a sample noisy audio signal and a training label matched with the sample noisy audio signal, wherein the sample noisy audio signal includes a sample reference audio signal and a sample target audio signal, and the training label includes a phase difference label, and the phase difference label is determined according to a first sample vector cosine value and a first sample vector sine value indicated by the first spectrum feature of the sample noisy audio signal, and a second sample vector cosine value and a second sample vector sine value indicated by a second spectrum feature of the sample target audio signal; the sample reference audio signal and the sample target audio signal are respectively one of a sample noise signal or a sample clean speech signal in the sample noisy audio signal; A prediction unit, configured to input the first frequency spectrum feature into an audio signal processing model to be trained, and obtain reference phase estimation information matching the sample noisy audio signal; a training unit, configured to train the audio signal processing model according to the reference phase estimation information and a training loss determined by the training label; and A determination unit is used to determine the trained audio signal processing model as a target neural network model when the training loss meets the target convergence condition.

18. A computer-readable storage medium, the computer-readable storage medium comprising a stored program, wherein: When the program is executed by a processor, the method described in any one of claims 1 to 10 or 11 to 15 is executed.

19. A computer program product comprising a computer program / instructions, which when executed by a processor implements the steps of the method according to any one of claims 1 to 10 or 11 to 15.

20. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the method according to any one of claims 1 to 10 or 11 to 15 through the computer program.

Citation Information

Patent Citations

  • Phase-dependent shared deep convolutional neural network speech enhancement method

    CN111081268A

  • Audio signal separation method, device, equipment, storage medium and program

    CN115731941A

  • Audio noise reduction model training method and device, and storage medium

    CN116597854A

  • Speech enhancement method

    KR102085739B1

  • Speech enhancement method and device using fast fourier convolution

    WO2023182765A1

Cited By

  • Signal processing model training method, signal processing method, device and equipment

    CN121662030A