Speech signal processing method and device, electronic equipment and storage medium

By using multi-channel filtering models and generative adversarial networks or deep learning networks to process far-field speech signals, the problem of low far-field speech quality is solved, achieving high-quality near-field speech conversion and improving the effect of voice communication.

CN119360868BActive Publication Date: 2025-12-09BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411273741.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-11
Publication Date
2025-12-09
Estimated Expiration
2044-09-11

AI Technical Summary

Technical Problem

In large spaces, far-field speech signals suffer from low quality due to the distance from the microphone array and environmental factors, which can easily cause auditory fatigue and reduce the efficiency of voice communication.

Method used

By acquiring multi-channel far-field speech signals, speech enhancement processing is performed using a preset multi-channel filtering model, and speech reconstruction is performed by combining generative adversarial networks or deep learning networks, converting the speech signals into high-quality near-field signals.

Benefits of technology

It improves the quality and intelligibility of far-field speech, enhances voice communication performance, reduces auditory fatigue, and increases the efficiency of voice communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360868B_ABST
    Figure CN119360868B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a speech signal processing method and device, electronic equipment and storage medium, and relates to the technical field of audio. The method comprises: acquiring a multi-channel far-field speech signal collected for a target sound source; inputting the multi-channel far-field speech signal into a preset multi-channel filtering model to perform speech enhancement processing to obtain target speech spectrum information; the preset multi-channel filtering model is a filtering model obtained by adaptively updating an original multi-channel filtering model according to the multi-channel far-field speech signal, or is a filtering model obtained by adaptively constructing according to a prediction result of a first neural network model for the multi-channel far-field speech signal; and inputting the target speech spectrum information into a preset near-field speech generation model to perform speech reconstruction processing to obtain a near-field speech signal corresponding to the multi-channel far-field speech signal. The technical scheme provided by the embodiments of the present disclosure can realize the transformation of speech hearing sensation and improve the quality and intelligibility of speech.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of audio technology, and in particular to a speech signal processing method and device, an electronic device, and a storage medium. BACKGROUND

[0002] With the development of communication technology, it is possible to conduct real-time voice communication in large spaces such as conference rooms and classrooms. In related applications, real-time voice communication in a far-field environment can be achieved using a microphone array-based far-field pickup technology.

[0003] However, due to the long distance between the speaker and the microphone array and the room transfer characteristics of the far-field environment, the collected far-field voice is significantly different from the voice collected in a near-field environment in terms of listening experience, and the quality of the collected far-field voice is low, which can easily cause auditory fatigue and reduce the efficiency of voice communication.

[0004] With the popularization of online real-time voice communication applications, there is an urgent need for an effective solution to convert the voice picked up in a far-field environment into high-quality voice with near-field effects in terms of listening experience. SUMMARY

[0005] The present disclosure provides a speech signal processing method, device, electronic device, and storage medium, which can convert the voice picked up in a far-field environment into voice with near-field effects in terms of listening experience, and improve the quality and intelligibility of the voice. The technical solutions of the present disclosure are as follows:

[0006] According to a first aspect of an embodiment of the present disclosure, a speech signal processing method is provided, comprising:

[0007] acquiring a multi-channel far-field voice signal collected for a target sound source;

[0008] inputting the multi-channel far-field voice signal into a preset multi-channel filter model to perform voice enhancement processing and obtain target spectrogram information; the preset multi-channel filter model is a filter model obtained by adaptively updating an original multi-channel filter model according to the multi-channel far-field voice signal, or a filter model obtained by adaptively constructing a prediction result of a first neural network model for the multi-channel far-field voice signal;

[0009] inputting the target spectrogram information into a preset near-field voice generation model to perform voice reconstruction processing and obtain a near-field voice signal corresponding to the multi-channel far-field voice signal.

[0010] In an optional embodiment, the preset near-field speech generation model is a vocoder model trained by a generative adversarial network, and the preset near-field speech generation model comprises a generation network; the target spectrogram information is input into the preset near-field speech generation model for speech reconstruction processing to obtain a near-field speech signal corresponding to the multi-channel far-field speech signal, and the method comprises:

[0011] determining target mel spectrum information corresponding to the target spectrogram information;

[0012] inputting the target mel spectrum information into an input convolutional layer of the generation network for up-sampling processing to obtain at least one original sampling point information; the at least one original sampling point information represents an amplitude and a phase corresponding to at least one time-frequency point;

[0013] inputting the at least one original sampling point information into a multi-receptive field fusion layer of the generation network for multi-receptive field-based fusion processing to obtain at least one target sampling point information;

[0014] inputting the at least one target sampling point information into an output convolutional layer of the generation network for time-frequency conversion processing to obtain the near-field speech signal.

[0015] In an optional embodiment, the preset near-field speech generation model is an audio codec model based on a deep learning network, and the preset near-field speech generation model comprises an encoder, a quantizer and a decoder; the target spectrogram information is input into the preset near-field speech generation model for speech reconstruction processing to obtain a near-field speech signal corresponding to the multi-channel far-field speech signal, and the method comprises:

[0016] inputting the target spectrogram information into the encoder for feature encoding processing to obtain hidden space feature information corresponding to the target spectrogram information;

[0017] inputting the hidden space feature information into the quantizer for feature quantization processing to obtain quantized feature information;

[0018] inputting the quantized feature information into the decoder for feature decoding processing to obtain the near-field speech signal.

[0019] In an optional embodiment, the multi-channel far-field speech signal comprises a plurality of far-field speech signals corresponding to a plurality of channels one by one; the multi-channel far-field speech signal is input into a preset multi-channel filtering model for speech enhancement processing to obtain target spectrogram information, and the method comprises:

[0020] determining spectrogram information corresponding to each far-field speech signal in the plurality of far-field speech signals;

[0021] determine a signal weight corresponding to each of the far-field voice signals based on the preset multi-channel filter model;

[0022] perform weighted summation processing based on the spectrum information corresponding to each of the far-field voice signals and the signal weight corresponding to each of the far-field voice signals to obtain the target spectrum information.

[0023] In an optional embodiment, the multi-channel far-field voice signal includes a plurality of far-field voice signals corresponding to a plurality of channels one by one, and the method further includes:

[0024] determine spectrum information corresponding to each of the plurality of far-field voice signals;

[0025] based on the spectrum information corresponding to each of the far-field voice signals, calculate a model update parameter, the model update parameter indicating an importance degree of each of the far-field voice signals;

[0026] based on the model update parameter, adaptively update the original multi-channel filter model to obtain the preset multi-channel filter model.

[0027] In an optional embodiment, the determination of the spectrum information corresponding to each of the plurality of far-field voice signals includes:

[0028] perform short-time Fourier transform processing on each of the plurality of far-field voice signals to obtain the spectrum information corresponding to each of the far-field voice signals.

[0029] In an optional embodiment, the calculation of the model update parameter based on the spectrum information corresponding to each of the far-field voice signals includes:

[0030] perform feature extraction processing on the spectrum information corresponding to each of the far-field voice signals to obtain time-frequency feature information corresponding to each of the far-field voice signals;

[0031] based on the time-frequency feature information corresponding to each of the far-field voice signals, calculate a first spatial covariance matrix and a second spatial covariance matrix; the first spatial covariance matrix indicates a correlation degree between the far-field voice signals; the second spatial covariance matrix indicates a correlation degree between noises of the far-field voice signals;

[0032] based on the first spatial covariance matrix and the second spatial covariance matrix, obtain a third spatial covariance matrix;

[0033] performing eigen-decomposition processing on the third spatial covariance matrix to obtain a steering vector, the steering vector indicating a direction of a target far-field speech signal in the multi-channel far-field speech signal, the target far-field speech signal being a far-field speech signal least affected by noise;

[0034] based on the steering vector and the second spatial covariance matrix, calculating the model update parameter.

[0035] In an optional embodiment, the calculating the model update parameter based on the spectrogram information corresponding to each far-field speech signal comprises:

[0036] performing feature extraction processing on the spectrogram information corresponding to each far-field speech signal to obtain time-frequency feature information corresponding to each far-field speech signal;

[0037] based on the time-frequency feature information corresponding to each far-field speech signal, calculating a first spatial covariance matrix and a second spatial covariance matrix; the first spatial covariance matrix indicating a degree of correlation between the far-field speech signals; the second spatial covariance matrix indicating a degree of correlation between noise components of the far-field speech signals;

[0038] inputting the first spatial covariance matrix and the second spatial covariance matrix into a second neural network model to perform parameter prediction processing to obtain the model update parameter.

[0039] In an optional embodiment, the first neural network model comprises a feature representation network, a feature extraction network and a feature mapping network, and the method further comprises:

[0040] inputting the multi-channel far-field speech signal into the feature representation network to perform feature representation processing to obtain first speech feature data;

[0041] inputting the first speech feature data into the feature extraction network to perform feature extraction processing to obtain second speech feature data;

[0042] inputting the second speech feature data into the feature mapping network to perform feature mapping processing to obtain a prediction result, the prediction result being a filter model construction parameter;

[0043] constructing the preset multi-channel filter model according to the prediction result.

[0044] In an optional embodiment, the method further comprises:

[0045] obtaining a clean speech sample signal and a near-field speech reference signal corresponding to the clean speech sample signal;

[0046] determine sample mel-spectrum information corresponding to the pure speech sample signal;

[0047] input the sample mel-spectrum information into a current generation network to perform speech reconstruction processing, and obtain a corresponding current near-field speech sample signal;

[0048] determine current first reconstruction loss data according to the current near-field speech sample signal and the near-field speech reference signal; the current first reconstruction loss data indicates a Manhattan distance between the current near-field speech sample signal and the near-field speech reference signal in a current training round;

[0049] input the current near-field speech sample signal into a current discrimination network to perform speech discrimination processing, and obtain a current first discrimination result, the current first discrimination result indicating whether the current near-field speech sample signal is discriminated as the corresponding near-field speech reference signal in the current training round;

[0050] determine current second reconstruction loss data based on the current first discrimination result;

[0051] determine current third reconstruction loss data according to the current first reconstruction loss data and the current second reconstruction loss data;

[0052] train the current generation network according to the current third reconstruction loss data, and obtain a next generation network;

[0053] input the near-field speech reference signal and the current near-field speech sample signal into the current discrimination network to perform speech discrimination processing, and obtain a current second discrimination result, the current second discrimination result indicating a discrimination accuracy degree of the near-field speech reference signal and the current near-field speech sample signal in the current training round;

[0054] determine current first discrimination loss data based on the current second discrimination result;

[0055] train the current discrimination network based on the current first discrimination loss data, and obtain a next discrimination network;

[0056] take the next generation network as a current generation network of a next training round, and take the next discrimination network as a current discrimination network of the next training round, and perform the training process cyclically;

[0057] end the training when the current generation network and the current discrimination network reach a preset convergence condition, and obtain the generation network and the discrimination network.

[0058] In an optional embodiment, the method further comprises:

[0059] obtain a multi-channel original speech sample signal and a pure speech sample signal corresponding to the multi-channel original speech sample signal;

[0060] input each original speech sample signal in the multi-channel original speech sample signal into an initial multi-channel filtering model to perform speech enhancement processing, and obtain first sample spectrum information;

[0061] determine second sample spectrum information corresponding to the pure speech sample signal;

[0062] based on the first sample spectrum information and the second sample spectrum information, determine first enhancement loss data; the first enhancement loss data indicates the spectrum difference degree of the first sample spectrum information and the second sample spectrum information in the norm space;

[0063] based on the first enhancement loss data, train the initial multi-channel filtering model to obtain the original multi-channel filtering model.

[0064] According to a second aspect of the embodiments of the present disclosure, a speech signal output device is provided, comprising:

[0065] an obtaining module configured to obtain a multi-channel far-field speech signal collected for a target sound source;

[0066] a speech enhancement module configured to input the multi-channel far-field speech signal into a preset multi-channel filtering model to perform speech enhancement processing, and obtain target spectrum information; the preset multi-channel filtering model is a filtering model obtained by adaptively updating an original multi-channel filtering model according to the multi-channel far-field speech signal, or is a filtering model obtained by adaptively constructing a prediction result of a first neural network model for the multi-channel far-field speech signal;

[0067] a speech reconstruction module configured to input the target spectrum information into a preset near-field speech generation model to perform speech reconstruction processing, and obtain a near-field speech signal corresponding to the multi-channel far-field speech signal.

[0068] In an optional embodiment, the preset near-field speech generation model is a vocoder model obtained by training a generative adversarial network, and the preset near-field speech generation model comprises a generative network, and the speech reconstruction module comprises:

[0069] a mel spectrum determination unit configured to determine target mel spectrum information corresponding to the target spectrum information;

[0070] an upsampling unit configured to perform input convolutional layer inputting the target mel spectrum information into the generation network, and perform upsampling processing to obtain at least one original sampling point information, wherein the at least one original sampling point information represents an amplitude and a phase corresponding to at least one time-frequency point;

[0071] a multi-receptive field fusion unit configured to perform multi-receptive field fusion layer inputting the at least one original sampling point information into the generation network, and perform multi-receptive field-based fusion processing to obtain at least one target sampling point information;

[0072] a time-frequency conversion unit configured to perform output convolutional layer inputting the at least one target sampling point information into the generation network, and perform time-frequency conversion processing to obtain the near-field speech signal.

[0073] In an optional embodiment, the preset near-field speech generation model is a deep learning network-based audio codec model, and the preset near-field speech generation model includes an encoder, a quantizer and a decoder, and the speech reconstruction module includes:

[0074] a feature encoding unit configured to perform inputting the target spectrum information into the encoder, and perform feature encoding processing to obtain hidden space feature information corresponding to the target spectrum information;

[0075] a feature quantization unit configured to perform inputting the hidden space feature information into the quantizer, and perform feature quantization processing to obtain quantized feature information;

[0076] a feature decoding unit configured to perform inputting the quantized feature information into the decoder, and perform feature decoding processing to obtain the near-field speech signal.

[0077] In an optional embodiment, the multi-channel far-field speech signal includes a plurality of far-field speech signals corresponding to a plurality of channels one by one, and the speech enhancement module includes:

[0078] a first signal spectrum determination unit configured to perform determining spectrum information corresponding to each far-field speech signal in the plurality of far-field speech signals;

[0079] a signal weight determination unit configured to perform determining a signal weight corresponding to each far-field speech signal based on the preset multi-channel filtering model;

[0080] a weighting unit configured to perform weighted sum processing according to the spectrum information corresponding to each far-field speech signal and the signal weight corresponding to each far-field speech signal to obtain the target spectrum information.

[0081] In an optional embodiment, the multi-channel far-field speech signal comprises a plurality of far-field speech signals corresponding to a plurality of channels, and the apparatus further comprises:

[0082] a second signal spectrum determination unit configured to determine spectrum information corresponding to each of the plurality of far-field speech signals;

[0083] a model update parameter calculation unit configured to calculate a model update parameter based on the spectrum information corresponding to each of the plurality of far-field speech signals, the model update parameter indicating an importance of each of the plurality of far-field speech signals;

[0084] an adaptive update unit configured to adaptively update the original multi-channel filtering model based on the model update parameter to obtain the preset multi-channel filtering model.

[0085] In an optional embodiment, either of the first signal spectrum determination unit or the second signal spectrum determination unit comprises:

[0086] a short-time Fourier transform sub-unit configured to perform short-time Fourier transform processing on each of the plurality of far-field speech signals to obtain the spectrum information corresponding to each of the plurality of far-field speech signals.

[0087] In an optional embodiment, the model update parameter calculation unit comprises:

[0088] a first feature extraction sub-unit configured to perform feature extraction processing on the spectrum information corresponding to each of the plurality of far-field speech signals to obtain time-frequency feature information corresponding to each of the plurality of far-field speech signals;

[0089] a covariance first calculation sub-unit configured to calculate a first spatial covariance matrix and a second spatial covariance matrix based on the time-frequency feature information corresponding to each of the plurality of far-field speech signals; the first spatial covariance matrix indicating a degree of correlation between the plurality of far-field speech signals; and the second spatial covariance matrix indicating a degree of correlation between noises of the plurality of far-field speech signals;

[0090] a covariance second calculation sub-unit configured to obtain a third spatial covariance matrix based on the first spatial covariance matrix and the second spatial covariance matrix;

[0091] a feature decomposition sub-unit configured to perform feature decomposition processing on the third spatial covariance matrix to obtain a steering vector, the steering vector indicating a direction of a target far-field speech signal in the multi-channel far-field speech signal, the target far-field speech signal being a far-field speech signal least affected by noise;

[0092] The model update parameter first calculation sub-unit is configured to calculate the model update parameter based on the steering vector and the second spatial covariance matrix.

[0093] In an optional embodiment, the model update parameter calculation unit comprises:

[0094] The first feature extraction sub-unit is configured to perform feature extraction processing on the spectrogram information corresponding to each far-field voice signal to obtain time-frequency feature information corresponding to each far-field voice signal.

[0095] The covariance third calculation sub-unit is configured to calculate a first spatial covariance matrix and a second spatial covariance matrix based on the time-frequency feature information corresponding to each far-field voice signal; the first spatial covariance matrix indicates the degree of correlation between the far-field voice signals; and the second spatial covariance matrix indicates the degree of correlation between noise components of the far-field voice signals.

[0096] The model update parameter second sub-unit is configured to input the first spatial covariance matrix and the second spatial covariance matrix into a second neural network model to perform parameter prediction processing to obtain the model update parameter.

[0097] In an optional embodiment, the device further comprises:

[0098] The feature representation unit is configured to input the multi-channel far-field voice signal into the feature representation network to perform feature representation processing to obtain first voice feature data.

[0099] The feature extraction unit is configured to input the first voice feature data into the feature extraction network to perform feature extraction processing to obtain second voice feature data.

[0100] The feature mapping unit is configured to input the second voice feature data into the feature mapping network to perform feature mapping processing to obtain a prediction result, the prediction result being a filter model construction parameter.

[0101] The model construction unit is configured to construct the preset multi-channel filter model according to the prediction result.

[0102] In an optional embodiment, the device further comprises:

[0103] The first training data acquisition unit is configured to acquire a clean voice sample signal and a near-field voice reference signal corresponding to the clean voice sample signal.

[0104] a sample mel spectrum determination unit configured to perform determination of sample mel spectrum information corresponding to the clean speech sample signal;

[0105] a sample speech reconstruction unit configured to perform input of the sample mel spectrum information into a current generation network, perform speech reconstruction processing, and obtain a corresponding current near-field speech sample signal;

[0106] a first reconstruction loss data determination unit configured to perform determination of current first reconstruction loss data according to the current near-field speech sample signal and the near-field speech reference signal; the current first reconstruction loss data indicating a Manhattan distance between the current near-field speech sample signal and the near-field speech reference signal in a current training round;

[0107] a first discrimination result determination unit configured to perform input of the current near-field speech sample signal into a current discrimination network, perform speech discrimination processing, and obtain a current first discrimination result; the current first discrimination result indicating whether the current near-field speech sample signal is discriminated as the corresponding near-field speech reference signal in the current training round;

[0108] a second reconstruction loss data determination unit configured to perform determination of current second reconstruction loss data based on the current first discrimination result;

[0109] a third reconstruction loss data determination unit configured to perform determination of current third reconstruction loss data according to the current first reconstruction loss data and the current second reconstruction loss data;

[0110] a first training unit configured to perform training of the current generation network according to the current third reconstruction loss data, and obtain a next generation network.

[0111] a second discrimination result determination unit configured to perform input of the near-field speech reference signal and the current near-field speech sample signal into the current discrimination network, perform speech discrimination processing, and obtain a current second discrimination result; the current second discrimination result indicating a discrimination accuracy degree of the near-field speech reference signal and the current near-field speech sample signal in the current training round;

[0112] a first discrimination loss data determination unit configured to perform determination of current first discrimination loss data based on the current second discrimination result;

[0113] a second training unit configured to perform training of the current discrimination network based on the current first discrimination loss data, and obtain a next discrimination network;

[0114] a cycle unit configured to perform the next generation network as a current generation network of a next training round and the next discriminative network as a current discriminative network of the next training round, and cyclically perform the training process described above;

[0115] a determination unit configured to perform ending the training in a case where the current generation network and the current discriminative network reach a preset convergence condition, to obtain the generation network and the discriminative network.

[0116] In an optional embodiment, the apparatus further comprises:

[0117] a second training data acquisition unit configured to perform acquiring a multi-channel original speech sample signal and a pure speech sample signal corresponding to the multi-channel original speech sample signal;

[0118] a sample speech enhancement unit configured to perform inputting each original speech sample signal in the multi-channel original speech sample signal into an initial multi-channel filter model, performing speech enhancement processing, and obtaining first sample spectrum information;

[0119] a sample signal spectrum determination unit configured to perform determining second sample spectrum information corresponding to the pure speech sample signal;

[0120] a first enhancement loss data determination unit configured to determine first enhancement loss data based on the first sample spectrum information and the second sample spectrum information; the first enhancement loss data indicating a spectrum difference degree of the first sample spectrum information and the second sample spectrum information in a norm space;

[0121] a third training unit configured to perform training the initial multi-channel filter model based on the first enhancement loss data, to obtain the original multi-channel filter model.

[0122] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the speech signal processing method according to any one of the first aspect of the embodiments of the present disclosure.

[0123] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can perform the speech signal processing method according to any one of the first aspect of the embodiments of the present disclosure.

[0124] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, comprising computer instructions which, when executed by a processor, implement the speech signal processing method according to any one of the first aspect of the embodiments of the present disclosure.

[0125] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0126] The embodiments of the present disclosure input the multi-channel far-field speech signal collected for the target sound source into a preset multi-channel filtering model to perform speech enhancement processing to obtain target speech spectrum information; the preset multi-channel filtering model is a filtering model obtained by adaptively updating an original multi-channel filtering model according to the multi-channel far-field speech signal, or is a filtering model obtained by adaptively constructing according to a prediction result of the multi-channel far-field speech signal by the first neural network model, which can more adaptively perform noise reduction and dereverberation and other speech enhancement processing on the multi-channel far-field speech signal, and improve the speech quality of the target speech spectrum information; the embodiments of the present disclosure input the target speech spectrum information into a preset near-field speech generation model to perform speech reconstruction processing, adjust the amplitude response and / or phase response of different frequencies in the target speech spectrum information, realize the change in speech listening experience, and can improve the speech damage that may be caused by the preset multi-channel filtering model after speech enhancement processing, thereby obtaining a near-field speech signal corresponding to the multi-channel far-field speech signal, improving the quality and intelligibility of the speech, and meeting the needs of more speech communication and applications.

[0127] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0128] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure, and do not constitute an undue limitation on the present disclosure.

[0129] Figure 1 is a schematic diagram of an application environment according to an exemplary embodiment;

[0130] Figure 2 is a flowchart of a speech signal processing method according to an exemplary embodiment;

[0131] Figure 3 is a flowchart of speech enhancement processing according to an exemplary embodiment;

[0132] Figure 4 is a flowchart of adaptively updating an original multi-channel filtering model according to an exemplary embodiment;

[0133] Figure 5is a flowchart of a speech reconstruction process according to an exemplary embodiment;

[0134] Figure 6 is a flowchart of another speech reconstruction process according to an exemplary embodiment;

[0135] Figure 7 is a flowchart of a model training method according to an exemplary embodiment;

[0136] Figure 8 is a flowchart of an adversarial training method according to an exemplary embodiment;

[0137] Figure 9 is a flowchart of a speech signal processing method in a specific application according to an exemplary embodiment;

[0138] Figure 10 is a block diagram of a speech signal processing apparatus according to an exemplary embodiment;

[0139] Figure 11 is a block diagram of a terminal for performing speech signal processing according to an exemplary embodiment;

[0140] Figure 12 is a block diagram of a server for performing speech signal processing according to an exemplary embodiment. DETAILED DESCRIPTION

[0141] In order for those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings.

[0142] It should be noted that the terms "first", "second", and the like in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. The implementation described in the following exemplary embodiments does not represent all implementations consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0143] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, analyzed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties.

[0144] In order to facilitate the understanding of the technical solutions and the technical effects produced by the embodiments of the present disclosure, the related professional terms involved in the embodiments of the present disclosure are explained:

[0145] Reverberation: refers to the phenomenon that the direct sound and the reflected sound of the voice signal superimposed after the voice signal encounters obstacles such as walls, ceilings, and floors.

[0146] STFT: short-time Fourier transform or short-term Fourier transform, full name short-time Fourier transform, is a mathematical transform closely related to Fourier transform, which is specially used to analyze the frequency and phase characteristics of time-varying signals in local regions. This transform adds a window function to the time domain of the signal, converting a long non-stationary signal into a short stationary signal matrix, and then performing a discrete Fourier transform on the windowed signal matrix to obtain a time-frequency spectrum matrix. STFT can provide frequency information of the signal at a specific time point, making up for the defect that Fourier transform cannot provide time resolution, making it possible to analyze non-stationary signals. In the fields of audio signal processing and speech recognition, STFT is widely used, and its results can be visually displayed as a sound spectrum graph, i.e. the modulus square of STFT energy.

[0147] Beamforming: the beamforming described in the embodiments of the present disclosure specifically refers to digital beamforming. Digital beamforming (DBF) is a kind of beamforming technology, which uses digital signal processing technology to process the signals received by the array antenna, compensates for the phase difference caused by the propagation path difference of sensors in different spatial positions, realizes in-phase superposition, and thus forms a beam with maximum energy reception in that direction. DBF technology is widely used in applications that require high directivity and high resolution, such as radar, sonar, and wireless communication.

[0148] Please refer to Figure 1 , which shows an application environment schematic diagram of a voice signal processing method according to an exemplary embodiment. The application environment can include a terminal 110 and a microphone array 120, and the terminal 110 and the microphone array 120 can be connected through a wired network or a wireless network.

[0149] The terminal 110 can be a smartphone, a tablet computer, a notebook computer, a desktop computer, etc. The terminal 110 can be an all-in-one machine equipped with the microphone array 120. The terminal 110 can be installed with an application program (Application, referred to as App for short). The application program can be an independent application program, or a subprogram in an independent application program. The user of the terminal 110 can log in to the application program through pre-registered user information, which can include an account and a password.

[0150] The microphone array 120 can also be referred to as a Microphone Array, which refers to an ordered arrangement of multiple microphones, which is composed of a certain number of microphones and is a system for sampling and processing a sound field. The configuration of the microphone array 120 is various, and can be divided into linear array, planar array and spatial array according to the geometric configuration. The embodiments of the present disclosure do not make any limitation on this. The multi-channel voice signal collected by the microphone array 120 retains the spatial phase characteristics of the signal, and provides more available information for dereverberation and echo cancellation.

[0151] In one specific embodiment, the microphone array 120 can be used to collect a multi-channel far-field voice signal of a target sound source in a space. The microphone array 120 transmits the collected multi-channel far-field voice signal to the terminal 110. The terminal 110 inputs the received multi-channel far-field voice signal into a preset multi-channel filter model to perform voice enhancement processing, and obtains target spectrogram information, wherein the preset multi-channel filter model is a filter model obtained by adaptively updating an original multi-channel filter model according to the multi-channel far-field voice signal, or is a filter model obtained by adaptively constructing a prediction result of the first neural network model for the multi-channel far-field voice signal; the terminal 110 inputs the target spectrogram information into a preset near-field voice generation model to perform voice reconstruction processing, and obtains a near-field voice signal corresponding to the multi-channel far-field voice signal, thereby realizing the transformation of the hearing experience. The terminal 110 can apply the near-field voice signal to a voice real-time communication type service to realize close-range interaction with a user outside the space.

[0152] In one specific embodiment, the microphone array 120 transmits the collected multi-channel far-field voice signal to a server to perform voice signal processing on the multi-channel far-field voice signal by the server, and obtains a corresponding near-field voice signal, and distributes the near-field voice signal to all terminals in a voice communication type service.

[0153] Figure 2 is a flowchart of a voice signal processing method according to an example embodiment, as shown in Figure 2 The voice signal processing method can include the following steps:

[0154] In step S210, a multi-channel far-field voice signal collected for a target sound source is obtained.

[0155] The voice signal processing method provided by the embodiments of the present disclosure is used to transform a far-field voice signal into a near-field voice signal, realize the transformation of the hearing experience of the voice signal, and optimize the hearing experience and voice communication effect of the user, which can provide services for the increasingly popular voice application.

[0156] In the embodiments of the present disclosure, the far-field voice signal refers to a voice signal generated by speaking from a position relatively far away from the position of the sound pickup device. A microphone array can be used as the sound pickup device to perform multi-channel data acquisition on the voice of the target sound source. The acquired multi-channel far-field voice signal includes a far-field voice signal corresponding to each channel of the multiple channels. The multi-channel far-field voice signal can effectively and accurately determine the relative distance and relative angle of the target sound source, thereby realizing the positioning and tracking of the target sound source and the directional pickup of the voice in the far-field environment.

[0157] In a specific embodiment, the target sound source includes at least one sound emitter.

[0158] In step S230, the multi-channel far-field voice signal is input into a preset multi-channel filtering model for voice enhancement processing to obtain target spectrogram information.

[0159] In the embodiments of the present disclosure, it can be understood that in a far-field environment, the target sound source is relatively far away from the sound pickup device and is in a movable state. The collected voice signal is subject to attenuation, scattering and reflection. In addition, the environment is noisy and there are many interference signals, which reduces the voice energy reaching the sound pickup device and causes it to be submerged in background noise and reverberation, thereby reducing the voice quality and intelligibility and causing auditory fatigue over a long period of time. At the same time, the voice recognition result is poor. Therefore, in order to improve the quality of the voice, the multi-channel far-field voice signal is input into a preset multi-channel filtering model for voice enhancement processing to perform time alignment, noise and reverberation removal and other voice quality processing on the multi-channel far-field voice signal.

[0160] In the embodiments of the present disclosure, the preset multi-channel filtering model is a filtering model obtained by adaptively updating an original multi-channel filtering model according to the multi-channel far-field voice signal, or a filtering model obtained by adaptively constructing a prediction result of the multi-channel far-field voice signal according to the first neural network model.

[0161] In a specific embodiment, the multi-channel far-field voice signal can provide richer spatial information than a single channel. In the voice enhancement processing process, the voice source can be more accurately positioned and separated by analyzing the differences and correlations between the multi-channel far-field voice signals, thereby improving the intelligibility of the voice.

[0162] In a specific embodiment, the multi-channel far-field voice signals from different microphones can be used to reduce the influence of environmental noise. In the voice enhancement processing process, the noise can be suppressed and the intelligibility and intelligibility of the voice can be improved by reasonably weighting and combining the multi-channel far-field voice signals.

[0163] In a specific embodiment, the multi-channel far-field speech signal can be subjected to interference suppression by an adaptive algorithm. In the speech enhancement processing, the time, frequency, phase, and other information of the multi-channel far-field speech signal can be used to model and remove noise and interference, thereby improving the purity and reliability of the speech signal.

[0164] In the embodiments of the present disclosure, the target spectrogram information represents the spectrogram features of the multi-channel far-field speech signal after speech enhancement processing, and can be represented in the format of a spectrogram, also known as a sonogram. The spectrogram is a graph representing the change of the frequency spectrum of a speech signal over time, with the horizontal axis representing time and the vertical axis representing frequency. The speech strength at a given time of any given frequency component is represented by the gray level or hue of the corresponding point, where the speech strength corresponds to the energy of the speech data, which can be represented by the amplitude of the speech signal or the real part and imaginary part of the complex spectrum of the speech signal.

[0165] In a specific embodiment, the multi-channel far-field speech signal includes a plurality of far-field speech signals corresponding to a plurality of channels, and the speech enhancement processing performed by the preset multi-channel filter model can include the following steps as shown in Figure 3

[0166] In step S310, the spectrogram information corresponding to each far-field speech signal in the plurality of far-field speech signals is determined.

[0167] In step S310, each far-field speech signal represented in the time domain waveform is subjected to time-frequency conversion to obtain the spectrogram information corresponding to each far-field speech signal. The spectrogram information corresponding to each far-field speech signal can indicate the change of the speech spectrum of the corresponding far-field speech signal over time.

[0168] In a specific embodiment, the plurality of far-field speech signals can be first time-aligned, and then each far-field speech signal in the time-aligned plurality of far-field speech signals is subjected to time-frequency conversion to obtain the spectrogram information corresponding to each far-field speech signal.

[0169] In step S320, based on the preset multi-channel filter model, the signal weight corresponding to each far-field speech signal is determined.

[0170] The preset multi-channel filter model is a filter model obtained by adaptively updating the original multi-channel filter model according to the multi-channel far-field speech signal, or a filter model obtained by adaptively constructing the prediction result of the first neural network model for the multi-channel far-field speech signal. The model parameters in the preset multi-channel filter model can indicate the signal weight corresponding to each channel, i.e., the signal weight corresponding to each far-field speech signal. The signal weight corresponding to each far-field speech signal indicates the importance of the corresponding far-field speech signal in the multi-channel far-field speech signal.

[0171] ​In step S330, according to the spectrum information corresponding to each far-field voice signal and the signal weight corresponding to each far-field voice signal, weighted sum processing is performed to obtain target spectrum information.

[0172] In one specific embodiment, the preset multi-channel filter model performs voice enhancement processing on the multi-channel far-field voice signal based on the principle of beamforming. Beamforming mainly aims at a microphone array, suppresses noise and interference directions by fusing voice data of multiple channels, and enhances voice signals in the target sound source direction. That is, by performing weighted sum processing on the spectrum information corresponding to each far-field voice signal, the voice component corresponding to the target sound source in the obtained target spectrum information is enhanced, and the noise component, echo component or reverberation component is suppressed.

[0173] In the above embodiment, by performing weighted sum processing on the spectrum information corresponding to each far-field voice signal and the signal weight corresponding to each far-field voice signal, enhancement of the voice component corresponding to the target sound source is realized, and the noise component, echo component or reverberation component can be effectively removed, and the voice quality corresponding to the target spectrum information is improved.

[0174] In one specific embodiment, the first neural network model includes a feature representation network, a feature extraction network and a feature mapping network, and the adaptive construction according to the prediction result of the first neural network model for the multi-channel far-field voice signal can include:

[0175] The multi-channel far-field voice signal is input into the feature representation network for feature representation processing to obtain first voice feature data.

[0176] The first voice feature data is input into the feature extraction network for feature extraction processing to obtain second voice feature data.

[0177] The second voice feature data is input into the feature mapping network for feature mapping processing to obtain a prediction result, which is a filter model construction parameter.

[0178] The preset multi-channel filter model is constructed according to the prediction result.

[0179] Specifically, the first voice feature data is feature data obtained after sampling, quantizing, encoding and embedding representation of the multi-channel far-field voice signal; the feature extraction network can include multiple hidden layers, and the feature extraction processing can map the input first voice feature data to a higher level of feature representation, thereby realizing nonlinear processing of the first voice feature data and obtaining second voice feature data; the feature mapping network is used to nonlinearly map the input second voice feature data to multiple model parameter prediction values.

[0180] In the above embodiment, the model parameters of the multi-channel far-field speech signal are predicted by the neural network model, so that the preset multi-channel filtering model constructed according to the prediction result can perform targeted speech enhancement processing on the multi-channel far-field speech signal, effectively improving the speech quality and also improving the construction efficiency of the preset multi-channel filtering model.

[0181] In one specific embodiment, as shown in Figure 4 The adaptive updating of the original multi-channel filtering model can include the following steps:

[0182] In step S410, the spectral information corresponding to each of the plurality of far-field speech signals is determined.

[0183] Step S410 can refer to step S310 in the foregoing embodiments, which will not be repeated here.

[0184] In step S420, the model update parameters are calculated based on the spectral information corresponding to each of the far-field speech signals, and the model update parameters indicate the importance of each of the far-field speech signals.

[0185] In one specific embodiment, the spectral information corresponding to each of the far-field speech signals can be analyzed based on a signal-to-noise ratio maximum criterion, a mean square error minimum criterion, a linearly constrained minimum variance criterion, a maximum likelihood criterion, or a neural network, to determine a signal existence probability corresponding to each of the far-field speech signals, which can represent the importance of the corresponding far-field speech signal, so that the model update parameters can be adaptively determined according to the signal existence probability corresponding to each of the far-field speech signals, so that an ideal speech quality gain can be obtained after the speech enhancement processing by the preset multi-channel filtering model.

[0186] In step S430, the original multi-channel filtering model is adaptively updated based on the model update parameters to obtain the preset multi-channel filtering model.

[0187] In one specific embodiment, the preset multi-channel filtering model obtained after adaptive updating can directly perform speech enhancement processing on the subsequently collected multi-channel far-field speech signal.

[0188] In the above embodiment, the parameters of the original multi-channel filtering model are adaptively updated based on the spectral statistical characteristics of each far-field speech signal itself, so that the obtained preset multi-channel filtering model can perform targeted speech enhancement processing on the multi-channel far-field speech signal according to the current far-field environment, effectively improving the speech quality.

[0189] In one specific embodiment, in steps S310 and S410, a short-time Fourier transform process can be performed on each of the plurality of far-field voice signals to obtain the corresponding spectrogram information of each far-field voice signal. The short-time Fourier transform process includes framing, windowing, and Fourier transforming each far-field voice signal to convert the time-domain far-field voice signal into an amplitude spectrum or a complex spectrum.

[0190] In the above embodiments, the time-domain far-field voice signal is converted into an amplitude spectrum or a complex spectrum by means of short-time Fourier transform, which preserves the time, frequency, phase, amplitude, and other characteristics of the far-field voice signal, providing a good data basis for voice enhancement and voice reconstruction.

[0191] In one specific embodiment, step S420 can be implemented as:

[0192] In step S421, feature extraction processing is performed on the spectrogram information corresponding to each far-field voice signal to obtain time-frequency feature information corresponding to each far-field voice signal.

[0193] In one specific embodiment, the spectrogram information corresponding to each far-field voice signal is discretized to extract spectrogram data corresponding to a plurality of preset time points and a plurality of preset frequency points as time-frequency feature information.

[0194] In step S423, based on the time-frequency feature information corresponding to each far-field voice signal, a first spatial covariance matrix and a second spatial covariance matrix are calculated; the first spatial covariance matrix indicates the degree of correlation between each far-field voice signal; and the second spatial covariance matrix indicates the degree of correlation between the noise of each far-field voice signal.

[0195] Alternatively, the covariance between each pair of time-frequency feature information corresponding to each far-field voice signal is calculated to obtain the first spatial covariance matrix.

[0196] Alternatively, for each component in the time-frequency feature information corresponding to each far-field voice signal, the probability of containing only noise is predicted, and the time-frequency feature information corresponding to each far-field voice signal is updated according to the predicted probability value; and the covariance between each pair of updated time-frequency feature information corresponding to each far-field voice signal is calculated to obtain the second spatial covariance matrix.

[0197] In step S425, based on the first spatial covariance matrix and the second spatial covariance matrix, a third spatial covariance matrix is obtained.

[0198] In the case that the dimensions of the rows and columns of the first spatial covariance matrix and the second spatial covariance matrix are the same, the components in the same row and column of the first spatial covariance matrix and the second spatial covariance matrix are subtracted to obtain a third spatial covariance matrix, which can indicate the correlation between the pure speech in each far-field speech signal to some extent.

[0199] In step S427, the third spatial covariance matrix is subjected to eigenvalue decomposition to obtain a steering vector, which indicates the direction of a target far-field speech signal in the multi-channel far-field speech signal, the target far-field speech signal being the far-field speech signal least affected by noise.

[0200] In the case that the dimensions of the rows and columns of the first spatial covariance matrix and the second spatial covariance matrix are the same, the components in the same row and column of the first spatial covariance matrix and the second spatial covariance matrix are subtracted to obtain a third spatial covariance matrix, which can indicate the correlation between the pure speech in each far-field speech signal to some extent.

[0201] In another specific embodiment, the steering vector can also be calculated according to the spatial phase difference between each far-field speech signal.

[0202] In step S429, the model update parameter is calculated based on the steering vector and the second spatial covariance matrix.

[0203] In one specific embodiment, the model update parameter is calculated based on the steering vector and the second spatial covariance matrix, and the model update parameter can indicate the importance of each far-field speech signal to adaptively update the original multi-channel filtering model.

[0204] In the above embodiment, the model update parameter indicating the importance of each far-field speech signal is determined according to the time-frequency characteristics of each far-field speech signal itself, so as to adaptively update the original multi-channel filtering model, thereby improving the adaptation degree of the obtained preset multi-channel filtering model after updating to the collected multi-channel far-field speech signal, and effectively improving the speech enhancement effect of the multi-channel far-field speech signal.

[0205] In one specific embodiment, step S420 can be implemented as:

[0206] In step S422, the spectral information corresponding to each far-field speech signal is subjected to feature extraction to obtain time-frequency characteristic information corresponding to each far-field speech signal.

[0207] In step S424, the first spatial covariance matrix and the second spatial covariance matrix are calculated based on the time-frequency feature information corresponding to each far-field voice signal; the first spatial covariance matrix indicates the correlation degree between each far-field voice signal; and the second spatial covariance matrix indicates the correlation degree between noise components of each far-field voice signal.

[0208] Steps S422 and S424 can refer to the foregoing embodiments, and details are not described herein.

[0209] In step S426, the first spatial covariance matrix and the second spatial covariance matrix are input into the second neural network model for parameter prediction processing to obtain model update parameters.

[0210] In one specific embodiment, the second neural network model can be a recurrent convolutional network or a deep learning network trained by using training sample data, and can directly output the model update parameters indicating the importance degree of each far-field voice signal based on the input first spatial covariance matrix and second spatial covariance matrix.

[0211] In the foregoing embodiments, the model update parameters indicating the importance degree of each far-field voice signal are predicted by using the second neural network model, which can improve the adaptation of the model update parameters to the multi-channel far-field voice signal, and further improve the adaptation of the updated preset multi-channel filtering model to the collected multi-channel far-field voice signal, thereby effectively improving the voice enhancement effect on the multi-channel far-field voice signal.

[0212] In step S250, the target spectrogram information is input into the preset near-field voice generation model for voice reconstruction processing to obtain a near-field voice signal corresponding to the multi-channel far-field voice signal.

[0213] Generally, a near-field voice signal refers to a voice signal received from a relatively close distance, such as a voice signal received by a microphone built in a device such as a mobile phone, a television, a computer, and the like. Although the near-field voice signal corresponding to the multi-channel far-field voice signal obtained by the voice reconstruction processing in the embodiments of the present disclosure is not a voice signal collected from a relatively close distance for a target sound source, it still has a near-field effect in terms of listening, which can avoid listener fatigue and improve the listening experience of the listener.

[0214] In the embodiments of the present disclosure, the target spectrum information is input into a preset near-field speech generation model for speech reconstruction processing. On the one hand, the target spectrum information can be converted into a time-domain waveform speech signal. On the other hand, the speech reconstruction function of the model is used to adjust the amplitude response and / or phase response of different frequencies in the target spectrum information, eliminate the influence of the spatial transmission performance of the far-field environment itself on the speech signal in the process of collecting the multi-channel far-field speech signal, and thus obtain a speech signal with a near-field listening experience. At the same time, in the process of adjusting the amplitude response and / or phase response of different frequencies in the target spectrum information, the excessive suppression of the amplitude and / or phase of the time-frequency point corresponding to the low signal-to-noise ratio in the speech enhancement stage can also be corrected, or residual noise can also be removed.

[0215] In a specific embodiment, the preset near-field speech generation model can be a vocoder model obtained by training a generative adversarial network such as HiFiGAN or NSF-HiFiGAN, or an audio codec model based on a deep learning network.

[0216] In a specific embodiment, as shown in Figure 5 the preset near-field speech generation model is a vocoder model obtained by training a generative adversarial network, and the preset near-field speech generation model includes a generation network. Step S250 can be implemented as:

[0217] In step S510, target mel spectrum information corresponding to the target spectrum information is determined.

[0218] Considering that the frequencies in the target spectrum information are linearly distributed, but the human ear is sensitive to the change of low frequencies and insensitive to the change of high frequencies, that is, the human ear is logarithmic to the frequency, so the target spectrum information with linearly distributed frequencies may appear "not useful enough" in the feature extraction process. Therefore, in order to adapt to the working principle of the vocoder model, it is necessary to convert the target spectrum information into corresponding target mel spectrum information. Specifically, the frequency values in the target spectrum information are logarithmically calculated to obtain corresponding mel frequencies.

[0219] In step S520, the target mel spectrum information is input into the input convolutional layer of the generation network for up-sampling processing to obtain at least one original sampling point information.

[0220] The at least one original sampling point information represents the amplitude and phase corresponding to at least one time-frequency point.

[0221] In step S530, the at least one original sampling point information is input into the multi-receptive field fusion layer of the generation network for multi-receptive field-based fusion processing to obtain at least one target sampling point information.

[0222] In step S540, the at least one target sampling point information is input into an output convolutional layer of the generation network to perform time-frequency conversion processing to obtain the near-field speech signal.

[0223] The generation network mainly includes an input convolutional layer, a multi-receptive field fusion layer, and an output convolutional layer. The input convolutional layer is specifically composed of one-dimensional transpose convolution; the multi-receptive field fusion layer is specifically composed of a residual network, and is mainly responsible for optimizing the at least one original sampling point information obtained by upsampling. The multi-receptive field fusion layer is a structure for improving the receptive field of the generation network by using a dilated convolution and a normal convolution. The dilation multiple of the dilated convolution gradually increases, and after each dilated convolution, a normal convolution with a kernel larger than 1 is followed, so as to realize the alternating use of the dilated convolution and the normal convolution. The input and output sizes of the dilated convolution and the normal convolution remain unchanged. After a round of dilated convolution and normal convolution is completed, the original input is connected to the result after convolution, thereby realizing a round of “multi-receptive field fusion”. By setting different convolution kernels and dilation multiples of the dilated convolution, the amplitude response and the phase response corresponding to the time-frequency point are adjusted implicitly. The output convolutional layer is mainly responsible for mapping the at least one target sampling point information after multi-receptive field fusion in the time domain until the length of the output sequence and the time resolution of the multi-channel far-field speech signal are matched, to obtain the above-mentioned near-field speech signal.

[0224] In the above embodiment, the preset near-field speech generation model constructed based on the generative adversarial network realizes the targeted adjustment of the amplitude response and / or the phase response of different frequencies in the target spectrogram information, obtains a speech signal with a near-field effect in hearing, meets the application requirement of near-field of the far-field speech signal, and further improves the quality and intelligibility of the speech. Meanwhile, the preset near-field speech generation model directly outputs the time-domain waveform speech signal, which can avoid the problem of insufficient reconstruction performance caused by the similar modalities of input and output.

[0225] In one specific embodiment, as shown in Figure 6 The preset near-field speech generation model is an audio codec model based on a deep learning network. The preset near-field speech generation model includes an encoder, a quantizer, and a decoder. Step S250 can be implemented as:

[0226] In step S610, the target spectrogram information is input into the encoder to perform feature encoding processing to obtain the latent space feature information corresponding to the target spectrogram information.

[0227] The latent space feature information represents the latent vector representation of the target spectrogram information in the latent space.

[0228] In the feature encoding processing stage, in addition to compressing the target spectrogram information into the latent space, the features of the target spectrogram information can also be extracted and encoded into the latent vector.

[0229] In step S620, the latent space feature information is input into a quantizer to perform feature quantization processing, and quantized feature information is obtained.

[0230] In step S630, the quantized feature information is input into a decoder to perform feature decoding processing, and a near-field speech signal is obtained.

[0231] The audio codec model based on the deep learning network can adopt an encoder-quantizer-decoder architecture, wherein the encoder-quantizer is used to output discrete quantized feature information, and the decoder is used to restore the discrete quantized feature information into a continuous audio signal, i.e., a near-field speech signal. Feasibly, the encoder mainly consists of convolutional layers, the quantizer adopts an RVQ layer to perform quantization processing on the latent space feature information, and the encoder and the decoder are mirror-imaged in structure.

[0232] In the above embodiment, the audio codec model based on the deep learning network is used as a preset near-field speech generation model. By encoding, quantizing and decoding the target spectrogram information in the latent space, the amplitude response and / or phase response of different frequencies in the target spectrogram information are adjusted implicitly, and a speech signal with a near-field effect in hearing is obtained, which meets the application requirement of near-field speech signal, and further improves the quality and intelligibility of the speech.

[0233] Figure 7 is a flowchart of a model training method according to an example embodiment, as shown in Figure 7 The method comprises the following steps:

[0234] In step S710, a multi-channel original speech sample signal and a pure speech sample signal corresponding to the multi-channel original speech sample signal are obtained.

[0235] In one specific embodiment, the multi-channel original speech sample signal can be obtained by randomly adding noise to the corresponding pure speech sample signal.

[0236] In one specific embodiment, the original speech sample signal and its corresponding pure speech sample signal are not limited to a speech signal with a far-field effect in hearing.

[0237] In step S720, each original speech sample signal in the multi-channel original speech sample signal is input into an initial multi-channel filtering model to perform speech enhancement processing, and a first sample spectrogram information is obtained.

[0238] In one specific embodiment, each of the plurality of original speech sample signals is input into an initial multi-channel filtering model to determine original spectrogram information corresponding to each of the original speech sample signals, determine a sample signal weight corresponding to each of the original speech sample signals based on the initial multi-channel filtering model, and perform weighted summation processing according to the original spectrogram information corresponding to each of the original speech sample signals and the sample signal weight corresponding to each of the original speech sample signals to obtain the first sample spectrogram information.

[0239] In step S730, second sample spectrogram information corresponding to the clean speech sample signal is determined.

[0240] In step S740, first enhancement loss data is determined based on the first sample spectrogram information and the second sample spectrogram information.

[0241] The first enhancement loss data indicates a spectrogram difference degree between the first sample spectrogram information and the second sample spectrogram information in a norm space.

[0242] In one specific embodiment, the first enhancement loss data is calculated according to formula (1), where X and Y represent the first sample spectrogram information and the second sample spectrogram information, respectively, a represents a weighting coefficient, which is exemplarily 0.7, and c represents a compression coefficient, which is exemplarily 0.3.

[0243]

[0244] In step S750, the initial multi-channel filtering model is trained based on the first enhancement loss data to obtain an original multi-channel filtering model.

[0245] Figure 8 is a flowchart of a model training method according to an exemplary embodiment, as shown in Figure 8 The method includes:

[0246] In step S810, a clean speech sample signal and a near-field speech reference signal corresponding to the clean speech sample signal are obtained.

[0247] In one specific embodiment, the clean speech sample signal can be an analog clean human voice signal, and the near-field speech reference signal can be a human voice signal obtained by speaking in a near-field environment according to the clean speech sample signal.

[0248] In step S820, sample mel-frequency spectrum information corresponding to the clean speech sample signal is determined.

[0249] In step S830, the sample mel-frequency spectrum information is input into a current generation network to perform speech reconstruction processing to obtain a corresponding current near-field speech sample signal.

[0250] Step S830 can be referred to the foregoing embodiments, and will not be repeated here.

[0251] In step S840, the current first reconstruction loss data is determined based on the current near-field speech sample signal and the near-field speech reference signal.

[0252] The current first reconstruction loss data indicates the Manhattan distance between the current near-field speech sample signal and the near-field speech reference signal in the current training round.

[0253] In step S850, the current near-field speech sample signal is input into the current discrimination network for speech discrimination processing to obtain the current first discrimination result.

[0254] The current first identification result indicates whether the current near-field speech sample signal is identified as the corresponding near-field speech reference signal in the current training round.

[0255] In one specific embodiment, the current discrimination network can employ a multi-scale discriminator to better enhance the discrimination capability of the generative network.

[0256] In step S860, the current second reconstruction loss data is determined based on the current first identification result.

[0257] Feasibly, the current second reconstruction loss data can be determined based on the difference between the current first discrimination result and the true result (the true result indicates that the signals input to the initial discrimination network are all near-field speech sample signals).

[0258] In step S870, the current third reconstruction loss data is determined based on the current first reconstruction loss data and the current second reconstruction loss data.

[0259] In a specific embodiment, the current third reconstruction loss data L is calculated according to formula (2). G Where L1 represents the current first reconstruction loss data, y represents the clean speech sample signal, G(y) represents the generated current near-field speech sample signal, and D i (G(y)) represents the discrimination result given by the i-th discriminator for G(y) in the current first discrimination result, and can be in the range of 0 to 1.

[0260]

[0261] In step S880, the current generator network is trained based on the current third reconstruction loss data to obtain the next generator network.

[0262] In one specific embodiment, the current generator network is trained based on the current third reconstruction loss data corresponding to multiple sets of samples to obtain the next generator network.

[0263] In step S910, the near-field speech reference signal and the current near-field speech sample signal are input into the current discrimination network for speech discrimination processing to obtain the current second discrimination result.

[0264] The current second discrimination result indicates whether the signal input to the current discrimination network is a near-field speech reference signal or a current near-field speech sample signal, which can indicate the accuracy of the distinction between the near-field speech reference signal and the current near-field speech sample signal in the current training round.

[0265] In step S920, the current first identification loss data is determined based on the current second identification result.

[0266] In a specific embodiment, the current first discrimination loss data L is calculated according to formula (3). D Where x represents the near-field speech reference signal, y represents the clean speech sample signal, G(y) represents the generated current near-field speech sample signal, and D... i (G(y)) represents the discrimination result given by the i-th discriminator for G(y) in the current second discrimination result, D i (x) represents the discrimination result given by the i-th discriminator for x in the current second discrimination result. The discrimination result can be in the range of 0 to 1.

[0267]

[0268] In step S930, the current discrimination network is trained based on the current first discrimination loss data to obtain the next discrimination network.

[0269] In step S940, the next generator network is used as the current generator network for the next training round, and the next discriminator network is used as the current discriminator network for the next training round, and the above training process is executed repeatedly.

[0270] In step S950, if the current generator network and the current discriminator network reach the preset convergence condition, the training ends, and the generator network and the discriminator network are obtained.

[0271] Figure 9 This is a schematic diagram illustrating a speech signal processing method according to an exemplary embodiment, such as... Figure 9As shown, after the multi-channel speech collected by the microphone array is filtered by the multi-channel filter, the corresponding target mel spectrum information or hidden space feature information is obtained, the target mel spectrum information or hidden space feature information is input into the generative model adapted thereto, speech reconstruction is performed, and the near-field speech signal of the time domain waveform can be obtained.

[0272] As can be seen from the technical solutions provided by the above embodiments of the present disclosure, the multi-channel far-field speech signal collected for the target sound source is input into the preset multi-channel filter model for speech enhancement processing to obtain target speech spectrum information; the preset multi-channel filter model is a filter model obtained by adaptively updating the original multi-channel filter model according to the multi-channel far-field speech signal, or a filter model obtained by adaptively constructing according to the prediction result of the first neural network model for the multi-channel far-field speech signal, which can more adaptively perform noise reduction and dereverberation and other speech enhancement processing on the multi-channel far-field speech signal, and improve the speech quality of the target speech spectrum information; the target speech spectrum information is input into the preset near-field speech generation model for speech reconstruction processing, the amplitude response and / or phase response of different frequencies in the target speech spectrum information are adjusted, the change in speech listening is realized, and the speech damage that may be caused by the preset multi-channel filter model after speech enhancement processing can be improved, so that the near-field speech signal corresponding to the multi-channel far-field speech signal is obtained, the quality and intelligibility of the speech are improved, and the demand for more speech communication and application can be met.

[0273] Figure 10 is a block diagram of a speech signal processing device according to an exemplary embodiment. Referring to Figure 10 , the device includes:

[0274] The acquisition module 1010 is configured to perform acquisition of the multi-channel far-field speech signal collected for the target sound source.

[0275] The speech enhancement module 1020 is configured to perform input of the multi-channel far-field speech signal into the preset multi-channel filter model for speech enhancement processing to obtain target speech spectrum information; the preset multi-channel filter model is a filter model obtained by adaptively updating the original multi-channel filter model according to the multi-channel far-field speech signal, or a filter model obtained by adaptively constructing according to the prediction result of the first neural network model for the multi-channel far-field speech signal.

[0276] The speech reconstruction module 1030 is configured to perform input of the target speech spectrum information into the preset near-field speech generation model for speech reconstruction processing to obtain the near-field speech signal corresponding to the multi-channel far-field speech signal.

[0277] In an optional embodiment, the preset near-field speech generation model is a vocoder model trained by a generative adversarial network, the preset near-field speech generation model comprises a generation network, and the speech reconstruction module 1030 comprises:

[0278] a mel spectrum determination unit configured to perform determination of target mel spectrum information corresponding to the target spectrogram information;

[0279] an up-sampling unit configured to perform input of the target mel spectrum information into an input convolutional layer of the generation network, perform up-sampling processing, and obtain at least one original sampling point information; the at least one original sampling point information represents amplitude and phase corresponding to at least one time-frequency point;

[0280] a multi-receptive field fusion unit configured to perform input of the at least one original sampling point information into a multi-receptive field fusion layer of the generation network, perform multi-receptive field-based fusion processing, and obtain at least one target sampling point information;

[0281] a time-frequency conversion unit configured to perform input of the at least one target sampling point information into an output convolutional layer of the generation network, perform time-frequency conversion processing, and obtain the near-field speech signal.

[0282] In an optional embodiment, the preset near-field speech generation model is an audio codec model based on a deep learning network, the preset near-field speech generation model comprises an encoder, a quantizer, and a decoder, and the speech reconstruction module 1030 comprises:

[0283] a feature encoding unit configured to perform input of the target spectrogram information into the encoder, perform feature encoding processing, and obtain hidden space feature information corresponding to the target spectrogram information;

[0284] a feature quantization unit configured to perform input of the hidden space feature information into the quantizer, perform feature quantization processing, and obtain quantized feature information;

[0285] a feature decoding unit configured to perform input of the quantized feature information into the decoder, perform feature decoding processing, and obtain the near-field speech signal.

[0286] In an optional embodiment, the multi-channel far-field speech signal comprises a plurality of far-field speech signals corresponding to a plurality of channels one by one, and the speech enhancement module 1020 comprises:

[0287] a first signal spectrogram determination unit configured to perform determination of spectrogram information corresponding to each far-field speech signal in the plurality of far-field speech signals;

[0288] a signal weight determination unit configured to determine a signal weight corresponding to each of the far-field voice signals based on the preset multi-channel filtering model;

[0289] a weighting unit configured to perform weighted summation processing according to the spectral information corresponding to each of the far-field voice signals and the signal weight corresponding to each of the far-field voice signals, to obtain the target spectral information.

[0290] In an optional embodiment, the multi-channel far-field voice signals include a plurality of far-field voice signals corresponding to a plurality of channels one by one, and the apparatus further includes:

[0291] a second signal spectral determination unit configured to determine spectral information corresponding to each of the plurality of far-field voice signals;

[0292] a model update parameter calculation unit configured to calculate a model update parameter based on the spectral information corresponding to each of the far-field voice signals, the model update parameter indicating an importance degree of each of the far-field voice signals;

[0293] an adaptive update unit configured to perform adaptive update on the original multi-channel filtering model based on the model update parameter, to obtain the preset multi-channel filtering model.

[0294] In an optional embodiment, the first signal spectral determination unit or the second signal spectral determination unit includes:

[0295] a short-time Fourier transform sub-unit configured to perform short-time Fourier transform processing on each of the plurality of far-field voice signals, to obtain the spectral information corresponding to each of the far-field voice signals.

[0296] In an optional embodiment, the model update parameter calculation unit includes:

[0297] a first feature extraction sub-unit configured to perform feature extraction processing on the spectral information corresponding to each of the far-field voice signals, to obtain time-frequency feature information corresponding to each of the far-field voice signals;

[0298] a covariance first calculation sub-unit configured to calculate a first spatial covariance matrix and a second spatial covariance matrix based on the time-frequency feature information corresponding to each of the far-field voice signals; the first spatial covariance matrix indicating a correlation degree between the far-field voice signals; the second spatial covariance matrix indicating a correlation degree between noises of the far-field voice signals;

[0299] a covariance second calculation subunit, configured to perform calculation of a third spatial covariance matrix based on the first spatial covariance matrix and the second spatial covariance matrix;

[0300] a feature decomposition subunit, configured to perform feature decomposition processing on the third spatial covariance matrix to obtain a steering vector, the steering vector indicating a direction of a target far-field voice signal in the multi-channel far-field voice signal, the target far-field voice signal being a far-field voice signal least affected by noise;

[0301] a model update parameter first calculation subunit, configured to perform calculation of the model update parameter based on the steering vector and the second spatial covariance matrix.

[0302] In an optional embodiment, the model update parameter calculation unit comprises:

[0303] a first feature extraction subunit, configured to perform feature extraction processing on spectrogram information corresponding to each far-field voice signal to obtain time-frequency feature information corresponding to each far-field voice signal;

[0304] a covariance third calculation subunit, configured to perform calculation of a first spatial covariance matrix and a second spatial covariance matrix based on the time-frequency feature information corresponding to each far-field voice signal; the first spatial covariance matrix indicating a degree of correlation between the far-field voice signals; the second spatial covariance matrix indicating a degree of correlation between noise components of the far-field voice signals;

[0305] a model update parameter second subunit, configured to perform parameter prediction processing on the first spatial covariance matrix and the second spatial covariance matrix by inputting the first spatial covariance matrix and the second spatial covariance matrix into a second neural network model to obtain the model update parameter.

[0306] In an optional embodiment, the device further comprises:

[0307] a feature representation unit, configured to perform feature representation processing on the multi-channel far-field voice signal by inputting the multi-channel far-field voice signal into the feature representation network to obtain first voice feature data;

[0308] a feature extraction unit, configured to perform feature extraction processing on the first voice feature data by inputting the first voice feature data into the feature extraction network to obtain second voice feature data;

[0309] a feature mapping unit, configured to perform feature mapping processing on the second voice feature data by inputting the second voice feature data into the feature mapping network to obtain a prediction result, the prediction result being a filter model construction parameter;

[0310] The model construction unit is configured to construct the preset multi-channel filtering model according to the prediction result.

[0311] In an optional embodiment, the apparatus further comprises:

[0312] The first training data acquisition unit is configured to acquire a clean speech sample signal and a near-field speech reference signal corresponding to the clean speech sample signal.

[0313] The sample mel spectrum determination unit is configured to determine sample mel spectrum information corresponding to the clean speech sample signal.

[0314] The sample speech reconstruction unit is configured to input the sample mel spectrum information into a current generation network to perform speech reconstruction processing, so as to obtain a corresponding current near-field speech sample signal.

[0315] The first reconstruction loss data determination unit is configured to determine current first reconstruction loss data according to the current near-field speech sample signal and the near-field speech reference signal, the current first reconstruction loss data indicating a Manhattan distance between the current near-field speech sample signal and the near-field speech reference signal in a current training round.

[0316] The first discrimination result determination unit is configured to input the current near-field speech sample signal into a current discrimination network to perform speech discrimination processing, so as to obtain a current first discrimination result, the current first discrimination result indicating whether the current near-field speech sample signal is discriminated as the corresponding near-field speech reference signal in the current training round.

[0317] The second reconstruction loss data determination unit is configured to determine current second reconstruction loss data based on the current first discrimination result.

[0318] The third reconstruction loss data determination unit is configured to determine current third reconstruction loss data according to the current first reconstruction loss data and the current second reconstruction loss data.

[0319] The first training unit is configured to train the current generation network according to the current third reconstruction loss data, so as to obtain a next generation network.

[0320] The second discrimination result determination unit is configured to input the near-field speech reference signal and the current near-field speech sample signal into the current discrimination network to perform speech discrimination processing, so as to obtain a current second discrimination result, the current second discrimination result indicating a discrimination accuracy degree of the near-field speech reference signal and the current near-field speech sample signal in the current training round.

[0321] a first discrimination loss data determination unit configured to determine current first discrimination loss data based on the current second discrimination result;

[0322] a second training unit configured to train the current discrimination network based on the current first discrimination loss data to obtain a next discrimination network;

[0323] a loop unit configured to loop the training process by taking the next generation network as a current generation network of a next training round and taking the next discrimination network as a current discrimination network of the next training round;

[0324] a judgment unit configured to end the training and obtain the generation network and the discrimination network when the current generation network and the current discrimination network reach a preset convergence condition.

[0325] In an optional embodiment, the apparatus further comprises:

[0326] a second training data acquisition unit configured to acquire a multi-channel original speech sample signal and a pure speech sample signal corresponding to the multi-channel original speech sample signal;

[0327] a sample speech enhancement unit configured to input each original speech sample signal in the multi-channel original speech sample signal into an initial multi-channel filter model to perform speech enhancement processing and obtain first sample spectrogram information;

[0328] a sample signal spectrogram determination unit configured to determine second sample spectrogram information corresponding to the pure speech sample signal;

[0329] a first enhancement loss data determination unit configured to determine first enhancement loss data based on the first sample spectrogram information and the second sample spectrogram information; the first enhancement loss data indicating a spectrogram difference degree of the first sample spectrogram information and the second sample spectrogram information in a norm space;

[0330] a third training unit configured to train the initial multi-channel filter model based on the first enhancement loss data to obtain the original multi-channel filter model.

[0331] As to the apparatus in the above embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and thus will not be described in detail here.

[0332] In an example embodiment, an electronic device is also provided, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the voice signal processing method as in the embodiments of the present disclosure.

[0333] Figure 11 is a block diagram of an electronic device for implementing a voice signal processing method according to an example embodiment. The electronic device can be a terminal, and its internal structure diagram can be as shown in Figure 11 . The electronic device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the electronic device is used to communicate with external terminals through network connection. The computer program is executed by the processor to implement a voice signal processing method. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad provided on the shell of the electronic device, or an external keyboard, touchpad or mouse, etc.

[0334] Figure 12 is a block diagram of an electronic device for implementing a voice signal processing method according to an example embodiment. The electronic device can be a terminal, and its internal structure diagram can be as shown in Figure 12 . The electronic device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the electronic device is used to communicate with external terminals through network connection. The computer program is executed by the processor to implement a voice signal processing method. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad provided on the shell of the electronic device, or an external keyboard, touchpad or mouse, etc.

[0335] Those skilled in the art can understand that the structures shown in Figure 11 and Figure 12 are only block diagrams of part of the structures related to the scheme of the present disclosure, and do not constitute a limitation on the electronic device to which the scheme of the present disclosure is applied. The specific electronic device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0336] In an example embodiment, a computer readable storage medium including instructions, which when executed by a processor of an electronic device, enable the electronic device to perform the voice signal processing method in the embodiments of the present disclosure is also provided.

[0337] In an example embodiment, a computer program product including computer instructions, which when executed by a processor, implement the voice signal processing method in the embodiments of the present disclosure is also provided.

[0338] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the computer program can include the processes of the above-mentioned embodiments of each method. Any reference to memory, storage, databases, or other media in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0339] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the present application cover any and all variations of the present disclosure that come within the scope of the present disclosure, including custom and practice of the present disclosure. The specification and examples are to be considered exemplary only, with the true scope and spirit of the present disclosure indicated by the following claims.

[0340] It should be understood that the present disclosure is not limited to the precise structures as herein described and illustrated in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A speech signal processing method, characterized by, The method comprises: acquiring a multi-channel far-field speech signal collected for a target sound source; the multi-channel far-field speech signal comprises a plurality of far-field speech signals corresponding to a plurality of channels; inputting the multi-channel far-field speech signal into a preset multi-channel filtering model to perform speech enhancement processing to obtain target spectrogram information; the preset multi-channel filtering model is a filtering model obtained by adaptively updating an original multi-channel filtering model according to the multi-channel far-field speech signal, or a filtering model obtained by adaptively constructing according to a prediction result of a first neural network model for the multi-channel far-field speech signal; inputting the target spectrogram information into a preset near-field speech generation model to perform speech reconstruction processing to obtain a near-field speech signal corresponding to the multi-channel far-field speech signal; wherein the inputting the multi-channel far-field speech signal into the preset multi-channel filtering model to perform speech enhancement processing to obtain target spectrogram information comprises: determining spectrogram information corresponding to each far-field speech signal in the plurality of far-field speech signals; determining a signal weight corresponding to each far-field speech signal based on the preset multi-channel filtering model; performing weighted summation processing according to the spectrogram information corresponding to each far-field speech signal and the signal weight corresponding to each far-field speech signal to obtain the target spectrogram information; the acquisition of the preset multi-channel filtering model comprises: calculating a model update parameter based on the spectrogram information corresponding to each far-field speech signal, the model update parameter indicating the importance of each far-field speech signal; adaptively updating the original multi-channel filtering model based on the model update parameter to obtain the preset multi-channel filtering model.

2. The method of claim 1, wherein, The preset near-field speech generation model is a vocoder model obtained by training a generative adversarial network, and the preset near-field speech generation model comprises a generation network, and the inputting the target spectrogram information into the preset near-field speech generation model to perform speech reconstruction processing to obtain a near-field speech signal corresponding to the multi-channel far-field speech signal comprises: determining target mel-frequency spectrum information corresponding to the target spectrogram information; inputting the target mel-frequency spectrum information into an input convolutional layer of the generation network to perform upsampling processing to obtain at least one original sampling point information; the at least one original sampling point information represents an amplitude and a phase corresponding to at least one time-frequency point; inputting the at least one original sampling point information into a multi-receptive field fusion layer of the generation network to perform multi-receptive field-based fusion processing to obtain at least one target sampling point information; inputting the at least one target sampling point information into an output convolutional layer of the generation network to perform time-frequency conversion processing to obtain the near-field speech signal.

3. The method of claim 1, wherein, The preset near-field speech generation model is an audio codec model based on a deep learning network, and the preset near-field speech generation model comprises an encoder, a quantizer, and a decoder, and the inputting the target spectrogram information into the preset near-field speech generation model to perform speech reconstruction processing to obtain a near-field speech signal corresponding to the multi-channel far-field speech signal comprises: input the target speech spectrum information into the encoder, perform feature encoding processing, and obtain hidden space feature information corresponding to the target speech spectrum information; input the hidden space feature information into the quantizer, perform feature quantization processing, and obtain quantized feature information; input the quantized feature information into the decoder, perform feature decoding processing, and obtain the near-field speech signal.

4. The method of claim 1, wherein, The method further includes: performing short-time Fourier transform processing on each of the plurality of far-field speech signals to obtain speech spectrum information corresponding to each of the plurality of far-field speech signals.

5. The method of claim 1, wherein, The method further includes: performing feature extraction processing on the speech spectrum information corresponding to each of the plurality of far-field speech signals to obtain time-frequency feature information corresponding to each of the plurality of far-field speech signals; based on the time-frequency feature information corresponding to each of the plurality of far-field speech signals, calculating a first spatial covariance matrix and a second spatial covariance matrix; the first spatial covariance matrix indicates a degree of correlation between the plurality of far-field speech signals; the second spatial covariance matrix indicates a degree of correlation between noise components of the plurality of far-field speech signals; based on the first spatial covariance matrix and the second spatial covariance matrix, obtaining a third spatial covariance matrix; performing feature decomposition processing on the third spatial covariance matrix to obtain a steering vector; the steering vector indicates a direction of a target far-field speech signal in the multi-channel far-field speech signal; the target far-field speech signal is a far-field speech signal that is least affected by noise; based on the steering vector and the second spatial covariance matrix, calculating the model update parameter.

6. The method of claim 1, wherein, The method further includes: performing feature extraction processing on the speech spectrum information corresponding to each of the plurality of far-field speech signals to obtain time-frequency feature information corresponding to each of the plurality of far-field speech signals; based on the time-frequency feature information corresponding to each of the plurality of far-field speech signals, calculating a first spatial covariance matrix and a second spatial covariance matrix; the first spatial covariance matrix indicates a degree of correlation between the plurality of far-field speech signals; the second spatial covariance matrix indicates a degree of correlation between noise components of the plurality of far-field speech signals; inputting the first spatial covariance matrix and the second spatial covariance matrix into a second neural network model to perform parameter prediction processing, and obtaining the model update parameter.

7. The method of claim 1, wherein, The first neural network model includes a feature representation network, a feature extraction network, and a feature mapping network, and the method further includes: inputting the multi-channel far-field speech signal into the feature representation network to perform feature representation processing and obtain first speech feature data; inputting the first speech feature data into the feature extraction network to perform feature extraction processing and obtain second speech feature data; inputting the second speech feature data into the feature mapping network to perform feature mapping processing and obtain a prediction result; the prediction result is a filter model construction parameter; and inputting the multi-channel far-field speech signal into the feature representation network to perform feature representation processing and obtain first speech feature data; inputting the first speech feature data into the feature extraction network to perform feature extraction processing and obtain second speech feature data; inputting the second speech feature data into the feature mapping network to perform feature mapping processing and obtain a prediction result; the prediction result is a filter model construction parameter. The preset multi-channel filtering model is constructed according to the prediction result.

8. The method of claim 2, wherein, The method further comprises: obtaining a pure speech sample signal and a near-field speech reference signal corresponding to the pure speech sample signal; determining sample mel spectrum information corresponding to the pure speech sample signal; inputting the sample mel spectrum information into a current generation network to perform speech reconstruction processing, and obtaining a corresponding current near-field speech sample signal; determining current first reconstruction loss data according to the current near-field speech sample signal and the near-field speech reference signal; the current first reconstruction loss data indicates a Manhattan distance between the current near-field speech sample signal and the near-field speech reference signal in a current training round; inputting the current near-field speech sample signal into a current discrimination network to perform speech discrimination processing, and obtaining a current first discrimination result; the current first discrimination result indicates whether the current near-field speech sample signal is discriminated as the corresponding near-field speech reference signal in the current training round; determining current second reconstruction loss data based on the current first discrimination result; determining current third reconstruction loss data according to the current first reconstruction loss data and the current second reconstruction loss data; training the current generation network according to the current third reconstruction loss data to obtain a next generation network; inputting the near-field speech reference signal and the current near-field speech sample signal into the current discrimination network to perform speech discrimination processing, and obtaining a current second discrimination result; the current second discrimination result indicates a discrimination accuracy degree of the near-field speech reference signal and the current near-field speech sample signal in the current training round; determining current first discrimination loss data based on the current second discrimination result; training the current discrimination network based on the current first discrimination loss data to obtain a next discrimination network; taking the next generation network as a current generation network of a next training round, and taking the next discrimination network as a current discrimination network of the next training round, and cyclically executing the training process; in a case where the current generation network and the current discrimination network reach a preset convergence condition, ending the training to obtain the generation network and the discrimination network.

9. The method of claim 1, wherein, The method further comprises: obtaining a multi-channel original speech sample signal and a pure speech sample signal corresponding to the multi-channel original speech sample signal; inputting each original speech sample signal in the multi-channel original speech sample signal into an initial multi-channel filtering model to perform speech enhancement processing, and obtaining first sample spectrum information; determining second sample spectrum information corresponding to the pure speech sample signal; determining first enhancement loss data based on the first sample spectrum information and the second sample spectrum information; the first enhancement loss data indicates a spectrum difference degree of the first sample spectrum information and the second sample spectrum information in a norm space; training the initial multi-channel filtering model based on the first enhancement loss data to obtain the original multi-channel filtering model.

10. A speech signal processing apparatus, characterized by comprising: The device comprises: The acquisition module is configured to acquire a multi-channel far-field speech signal collected for a target sound source; the multi-channel far-field speech signal includes a plurality of far-field speech signals corresponding to a plurality of channels; The speech enhancement module is configured to input the multi-channel far-field speech signal into a preset multi-channel filtering model to perform speech enhancement processing to obtain target spectrogram information; the preset multi-channel filtering model is a filtering model obtained by adaptively updating an original multi-channel filtering model according to the multi-channel far-field speech signal, or is a filtering model obtained by adaptively constructing a prediction result of a first neural network model for the multi-channel far-field speech signal; The speech reconstruction module is configured to input the target spectrogram information into a preset near-field speech generation model to perform speech reconstruction processing to obtain a near-field speech signal corresponding to the multi-channel far-field speech signal. The speech enhancement module includes: The first signal spectrogram determination unit is configured to determine spectrogram information corresponding to each far-field speech signal in the plurality of far-field speech signals; The signal weight determination unit is configured to determine a signal weight corresponding to each far-field speech signal based on the preset multi-channel filtering model; The weighting unit is configured to perform weighted summation processing according to the spectrogram information corresponding to each far-field speech signal and the signal weight corresponding to each far-field speech signal to obtain the target spectrogram information. The device further includes: The model update parameter calculation unit is configured to calculate a model update parameter based on the spectrogram information corresponding to each far-field speech signal, the model update parameter indicating the importance of each far-field speech signal; The adaptive updating unit is configured to adaptively update the original multi-channel filtering model based on the model update parameter to obtain the preset multi-channel filtering model.

11. An electronic device, comprising: It includes: A processor; A memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the speech signal processing method of any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, When the instructions in the computer readable storage medium are executed by the processor of the electronic device, the electronic device can perform the speech signal processing method of any one of claims 1 to 9.

13. A computer program product comprising computer instructions, characterized in that, The computer instructions are executed by the processor to implement the speech signal processing method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Audio encoder and decoder

    CN101925950A

  • Distributed optical fiber speech enhancement method based on GAN network and tunnel rescue system

    CN114898766A