Voice processing method and device, electronic equipment and storage medium
By adjusting the convolution kernel parameters of the neural network, real and imaginary parameters are generated based on the features of the speech input signal. This solves the problem of poor performance of existing audio enhancement techniques, achieves efficient feature enhancement and noise suppression of speech signals, and improves the quality of speech processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-03
AI Technical Summary
Existing audio enhancement technologies often use fixed vectors for calculation, resulting in poor enhancement effects and limited application scope. They cannot effectively eliminate echoes, noise, and reverberation, thus affecting the quality of conference calls.
By acquiring audio information from the speech input signal and the reference signal, the convolution parameters are determined, and the convolution kernel of the neural network, including the real part parameters and the imaginary part parameters, is adjusted. Based on the signal features, it is generated in different dimensions, and the adjusted neural network is used to enhance the audio information, thereby improving the generalization ability and noise suppression ability of the neural network.
It achieves precise processing of sudden noise and rapidly changing time-delay data, reduces noise residue, improves speech processing quality, expands the application scenarios of audio data enhancement, and enhances the feature enhancement effect of speech signals.
Smart Images

Figure CN121789653A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer application technology, and in particular to a voice processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] In recent years, the use of teleconferencing systems such as Microsoft Teams, Skype, and Zoom has increased significantly, becoming an indispensable part of remote work. Ensuring good call quality is crucial for these systems to provide end-users with an efficient and enjoyable experience. Noise and acoustic echo are major factors contributing to degraded conference call quality, significantly reducing speech intelligibility and hindering normal communication. This is especially pronounced in full-duplex communication modes when both parties are speaking simultaneously. Therefore, effective solutions to eliminate echo, noise, and reverberation are essential for achieving seamless communication.
[0003] The application of deep learning in audio tasks began with attempts to combine classical digital signal processing methods with neural networks. For example, combining adaptive filters with recurrent neural networks (RNNs) can achieve good performance in acoustic echo cancellation tasks. Meanwhile, another research path aims to build pure deep learning solutions, demonstrating convincing results on complex datasets. Among these, RNNs play a crucial role in solving the acoustic echo cancellation problem due to their powerful modeling capabilities for time-varying functions. Currently, a popular approach is a real-time deep speech quality enhancement scheme that combines acoustic echo cancellation, noise suppression, and dereverberation. The advantage of this model lies in its balance between key metrics such as computational complexity, echo return loss enhancement, and subjective opinion scoring for acoustic echo cancellation. However, current common audio data enhancement techniques often use fixed vectors for computation, resulting in poor enhancement effects. Summary of the Invention
[0004] The present invention provides several embodiments of a voice processing method, apparatus, electronic device, and storage medium, wherein at least one embodiment is used to solve the problem of poor audio enhancement effect and limited application scope.
[0005] According to one aspect of the present invention, a speech processing method is provided, wherein the method includes:
[0006] Acquire audio information from the voice input signal and the reference signal;
[0007] Convolutional parameters are determined based on the signal features of the speech input signal, and the convolutional kernel of the neural network is adjusted according to the convolutional parameters. The signal features include at least one of frequency domain features and time domain features, and the convolutional parameters include at least real part parameters and imaginary part parameters. The real part parameters and the imaginary part parameters are generated based on the signal features in different dimensions.
[0008] The audio information is enhanced using the adjusted neural network.
[0009] According to another aspect of the present invention, a voice processing apparatus is provided, wherein the apparatus comprises:
[0010] The feature extraction module is used to acquire audio information from the speech input signal and the reference signal;
[0011] A network adjustment module is used to obtain convolution parameters based on the signal features of the speech input signal, and adjust the convolution kernel of the neural network according to the convolution parameters. The signal features include at least one of frequency domain features and time domain features, and the convolution parameters include at least real part parameters and imaginary part parameters. The real part parameters and the imaginary part parameters are generated based on the signal features in different dimensions.
[0012] A feature enhancement module is used to enhance the audio information using the adjusted neural network.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the speech processing method according to any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the speech processing method described in any embodiment of the present invention.
[0018] The technical solution of this invention involves acquiring a speech input signal and a reference signal, determining the audio information of the speech input signal and the reference signal, determining convolution parameters according to the signal characteristics of the speech input signal, and adjusting the convolution kernel of the neural network according to the convolution parameters. These convolution parameters include at least real and imaginary parameters, which are generated based on the feature signal in different dimensions. The adjusted neural network is then used to enhance the audio information. This invention can enhance audio information based on a neural network that matches the signal characteristics of the speech input signal, improving the generalization ability of the neural network, expanding the application scenarios of audio data enhancement, reducing noise residue, improving the accuracy of audio data enhancement, enhancing the ability to process sudden noise and rapidly changing data, achieving precise interference suppression, and improving the quality of speech processing.
[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of a speech processing method according to an embodiment of the present invention;
[0022] Figure 2 This is a flowchart of another speech processing method provided according to an embodiment of the present invention;
[0023] Figure 3 This is a flowchart of another speech processing method provided according to an embodiment of the present invention;
[0024] Figure 4 This is a flowchart of another speech processing method provided according to an embodiment of the present invention;
[0025] Figure 5 This is a schematic diagram of the structure of a voice processing device according to an embodiment of the present invention;
[0026] Figure 6 This is a schematic diagram of the structure of an electronic device that implements a speech processing method according to an embodiment of the present invention. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] Figure 1 This is a flowchart of a voice processing method according to an embodiment of the present invention. This embodiment is applicable to reducing noise interference in scenarios such as online conferencing or voice playback. The method can be executed by a voice processing device, which can be implemented in hardware and / or software, and can be configured in a user terminal or server. Figure 1 As shown, the method includes:
[0030] Step 110: Obtain audio information of the voice input signal and the reference signal.
[0031] The speech input signal can be the original signal to be enhanced in terms of speech quality. It can be acquired by a sound acquisition device and may include environmental noise, system noise, and human-induced noise. The reference signal can be the original audio data stream being played or audio data output by a speech output device. The audio information can be the signal characteristics of the speech input signal and the reference signal, which may include, but is not limited to, complex spectrum characteristics, spectral characteristics, time-domain characteristics, deep learning characteristics, and modulation spectrum characteristics.
[0032] In this embodiment of the invention, feature extraction can be performed on the acquired speech input signal and reference signal to obtain the audio information corresponding to the speech input signal and reference signal, respectively. For example, the process of extracting audio complex spectrum features may include extracting signal data from the speech input signal and reference signal according to a time window, and performing a short-time Fourier transform on the extracted signal data to obtain the audio complex spectrum features.
[0033] Time windows are a technique used to collect and extract data features.
[0034] Step 120: Obtain convolution parameters based on the signal features of the speech input signal, and adjust the convolution kernel of the neural network according to the convolution parameters. The signal features include at least one of frequency domain features and time domain features, and the convolution parameters include at least real part parameters and imaginary part parameters. The real part parameters and imaginary part parameters are generated based on the signal features in different dimensions.
[0035] The convolution parameters can be the size information of the convolution kernel used for convolution processing. These parameters can include multi-dimensional information, such as two-dimensional, three-dimensional, and four-dimensional parameters. They can be determined by the signal features of the speech input signal, including but not limited to the number of time frames, frequency points, short-time energy, short-time zero-crossing rate, spectral flatness, spectral centroid, and spectral flux. The values of the convolution parameters can include at least real and imaginary parts, which can be determined by the feature signal in different dimensions. The neural network can be a network model for feature enhancement of audio information. Neural networks can include, but are not limited to, convolutional coding networks, recurrent convolutional networks, spatial attention networks, self-attention feature enhancement networks, generative adversarial networks, variational autoencoders, and depthwise separable convolutional networks.
[0036] In this embodiment of the invention, signal features of the speech input signal can be acquired, and these acquired signal features can be used as convolution parameters. For example, at least one of the time-domain and frequency-domain features of the speech input signal can be extracted using a feature extraction network as signal features. Alternatively, the time-domain and / or frequency-domain features of the speech input signal carried in the audio information can be extracted. The real and imaginary parts of the convolution parameters can be determined according to the values or distributions of the determined signal features in different dimensions. The convolution kernel of the pre-configured neural network can be adjusted according to the acquired convolution parameters, thereby readjusting the convolution kernel through the convolution parameters to enhance the matching degree of the neural network to the speech input signal.
[0037] Furthermore, in some embodiments, to enhance the accuracy of features and / or improve feature processing efficiency, the audio data may also undergo encoding, dimensionality reduction compression, decoding, and attention mechanism processing.
[0038] For example, in some embodiments, a pre-trained deep learning model can be used to process audio complex spectrum features. This deep learning model can be trained and generated based on speech data with interference and clean speech data. The deep speech quality enhancement model includes at least a convolutional mask model. This convolutional mask module can be used to extract the local / global feature capture capability through convolution, generating a dynamic, structured mask that can accurately filter clean speech signals, thereby suppressing interference. The deep learning model may include pooling layers, convolutional layers, and activation layers, etc. The convolutional kernel parameters of the convolutional mask module can be determined by the signal features of the speech input signal. For example, the first and second dimensions of the convolutional kernel can be set to the frequency points and frame numbers of the speech input signal, etc.
[0039] In one embodiment of a deep learning model, a convolutional layer serves as the starting point of the sequence, performing linear feature detection. Its input is a multidimensional tensor, which is locally connected and discretely convolved with the input through a set of learnable convolutional kernels to generate a linear response map. The activation layer follows immediately after the convolutional layer, applying an element-wise nonlinear transformation to the linear response. This layer receives the output of the convolutional layer as its input and performs a point-by-point mapping through a nonlinear activation function, producing an activated feature map. Pooling layers typically operate on the nonlinear feature map output by the activation layer, with the core purpose of spatial downsampling to obtain hierarchical and robust feature representations.
[0040] Step 130: Enhance audio information using the adjusted neural network.
[0041] In this embodiment of the invention, audio information can be enhanced by using a neural network adjusted by convolutional kernels, thereby improving the quality, discriminability, and adaptability of the audio information, which helps to improve the signal quality of the speech signal corresponding to the audio information.
[0042] This invention, in its embodiments, acquires a speech input signal and a reference signal, determines the audio information of the speech input signal and the reference signal, determines convolution parameters according to the signal characteristics of the speech input signal, and adjusts the convolution kernel of the neural network according to the convolution parameters. These convolution parameters include at least real and imaginary parameters, which are generated based on the feature signal in different dimensions. The adjusted neural network is then used to enhance the audio information. This invention can enhance audio information based on a neural network that matches the signal characteristics of the speech input signal, improving the generalization ability of the neural network, expanding the application scenarios of audio data enhancement, reducing noise residue, improving the accuracy of audio data enhancement, enhancing the ability to process sudden noise and rapidly changing data, achieving precise interference suppression, and improving the quality of speech processing.
[0043] The real and imaginary parameters are generated based on the feature signal in different dimensions. This means that, according to the set cutting and reshaping requirements, the feature signal is reshaped and segmented in the channel dimension, and parsed into the real and imaginary parts of the convolution, thereby constructing the real and imaginary parameters.
[0044] Figure 2 This is a flowchart of another speech processing method according to an embodiment of the present invention. This embodiment is a refinement based on the above embodiment, illustrating the extraction process of audio complex spectrum features. See [link to flowchart documentation]. Figure 2 The method provided in this embodiment of the invention specifically includes the following steps:
[0045] Step 210: Acquire the voice input signal collected by the microphone and the reference signal output by the speaker.
[0046] In embodiments of the present invention, audio signals collected by a microphone can be acquired as voice input signals, and audio signals output through a speaker can be acquired as reference signals. In some embodiments, the reference signal may be the currently playing audio data stream. In some embodiments, the audio data stream has not yet been output through an audio output device such as a speaker, and is used as the reference signal.
[0047] Step 220: Extract features from the speech input signal and reference signal based on a preset time window, and obtain time-frequency data by transforming the speech input signal and reference signal.
[0048] The preset time window can be a time-domain function for segmenting a continuous time-domain signal into frames. This time-domain function can obtain signal frames by multiplying it segment by segment with the original audio signal. The preset time window can include, but is not limited to, Hanning windows, square root Hanning windows, and Hamming windows. In one specific embodiment, a square root Hanning window is used as the preset time window, with a window length of 512 and a frame shift of 256.
[0049] In this embodiment of the invention, the voice input signal and the reference signal can be segmented into frames according to a preset time window. Features can be extracted from the segmented audio data frames, and the extracted features can be transformed. The transformation can include Fourier transform or short-time Fourier transform, thereby obtaining time-frequency data in complex form. The time-frequency data can include amplitude spectrum and phase spectrum. The amplitude spectrum can be composed of the amplitude of the time-frequency data, while the phase spectrum is composed of the phase of the time-frequency data. The dimensions of amplitude and phase can be consistent with the dimensions of the time-frequency data.
[0050] Step 230: Perform power-law compression complex spectrum processing on the amplitude spectrum of the time-frequency data to obtain audio information.
[0051] The amplitude spectrum can be a two-dimensional matrix composed of the magnitudes of each complex element in the time-frequency data. Each dimension of this two-dimensional matrix can correspond to a time point and a frequency point in the time-frequency data. Power-law compressed complex spectrum processing can perform a nonlinear power-law transformation on the amplitude dimension of the time-frequency data while preserving the integrity of the phase information in the time-frequency data.
[0052] Specifically, the time-frequency data can be split into amplitude spectrum and phase spectrum. Power-law compression is performed on the amplitude spectrum. The compressed amplitude spectrum and phase spectrum are then combined to reconstruct a compressed complex spectrum in complex form. The reconstructed compressed complex spectrum is then used as audio information.
[0053] Step 240: Obtain the number of time frames and frequency points of the audio information as signal features, and use the number of time frames and frequency points as the dimensions of the convolution parameters.
[0054] Specifically, feature information such as the number of time frames and the number of frequency points can be extracted from the audio information of the speech input signal as signal features. The number of time frames and the number of frequency points can be set as different dimensions of the convolution parameters used to construct the convolution kernel. That is, at least two dimensions of the convolution parameters used to construct the convolution kernel can be the number of time frames and the number of frequency points of the speech input signal, respectively. Furthermore, in some embodiments, the convolution parameters may also include a time-domain size and a frequency-domain size to limit the batch size of the audio data. The time-domain size and the frequency-domain size can be pre-configured.
[0055] Step 250: Construct a feature tensor based on the number of time frames and frequency points. Divide the feature tensor into real and imaginary parameters according to different dimensions, which will serve as the convolution kernel of the neural network.
[0056] In this embodiment of the invention, each dimension value within the convolution parameters can be extracted, and a feature tensor with a corresponding number of channels can be created based on all obtained dimension values. This feature tensor can be divided into real and imaginary parameters according to different dimensions, and the obtained real and imaginary parameters can be used as the convolution kernel of the neural network. It is understood that the dimension values within the convolution parameters are not limited to complex dimensions, but can also include other dimensions. The dimension values of other dimensions can include preset fixed values or dimension values determined according to reference signals, etc.
[0057] Step 260: Enhance audio information using the adjusted neural network.
[0058] In this embodiment of the invention, voice input signals are acquired through a microphone and reference signals are obtained through a speaker. Features are extracted from the voice input signals and reference signals according to a preset time window, and the voice input signals and reference signals are transformed to obtain time-frequency data. Power-law compression complex spectrum processing is performed on the amplitude spectrum of the time-frequency data to construct audio complex spectrum features. A dynamic convolution kernel is determined according to the signal characteristics of the voice input signal. The audio complex spectrum features are convolved based on the dynamic convolution kernel to obtain audio data. The number of time frames and the number of frequency points of the audio data are used as the dimension values of the convolution parameters to construct feature tensors with different dimension values. The feature tensors are divided into real part parameters and imaginary part parameters according to a preset dimension and then used as the convolution kernel of the neural network. The adjusted neural network is used to enhance the voice information. This invention converts continuous audio signals in the time domain into two-dimensional time-frequency data, enabling joint time-frequency analysis of non-stationary audio signals. It fully preserves the time and frequency characteristics of the audio signals, which helps to accurately locate the time-frequency position corresponding to noise and improves the speech enhancement effect. By compressing high amplitude values and amplifying low amplitude values in the compressed time-frequency data according to power-law compression complex spectrum processing, the frequency domain characteristics of the audio data can be optimized, and the model processing efficiency can be improved.
[0059] Compared to the unenhanced speech information, the enhanced speech information eliminates echoes, suppresses noise, and removes reverberation.
[0060] In some embodiments of the invention, the method further includes: preprocessing the audio information, wherein the preprocessing includes at least one of the following:
[0061] Encode the audio information of the voice input signal and the reference signal;
[0062] The audio information of the voice input signal and the reference signal is fused;
[0063] Compress the audio information of the voice input signal and the reference signal;
[0064] The audio information of the voice input signal and the reference signal is decoded.
[0065] In this embodiment of the invention, the audio information of the speech input signal and the reference signal can be encoded separately. The audio information of the speech input signal and the reference signal can be obtained based on different encoders. For example, different encoders can be used to encode the audio information of the speech input signal and the reference signal. Each encoding block can consist of a downsampling convolutional layer, a batch normalization layer, and an ELU activation function. The encoder that processes the audio information of the speech input signal can have a different structure than the encoder that processes the audio information of the reference signal, or the encoder's encoding rules or parameters can be different, etc.
[0066] Specifically, the audio information of the speech input signal and the reference signal is fused. This fusion method can include, but is not limited to, direct concatenation, weighted summation, cross-correlation fusion, and convolutional fusion. For example, the encoded audio information of the speech input signal and the reference signal can be input into the feature fusion module of a deep learning model. Within the feature fusion model, the audio information of the speech input signal and the reference signal can be directly concatenated, weighted summation, cross-correlation fusion, or convolutional fusion to achieve feature fusion. Information compression can also be performed on the audio information of the speech input signal and the reference signal to reduce the feature dimension of the audio information and make the compressed audio information more focused on core information. For example, recurrent neural networks, long short-term memory networks, gated recurrent networks, etc., can be used to process the audio information of the speech input signal and the reference signal, and / or linear transformations can be performed on the audio information of the speech input signal and the reference signal.
[0067] In this embodiment of the invention, the audio information of the voice input signal and the reference signal can be decoded. This decoding process can be the inverse operation of the encoding process. The decoding of the audio information of the voice input signal and the reference signal can be implemented using a decoder. The audio information of the voice input signal and the reference signal can be restored to feature information by the decoder.
[0068] For example, a 1×1 convolutional kernel can be used to reduce the dimensionality of the decoded features. The processed decoded features can be denoted as the dimensionality-reduced decoded features. These features can be divided according to a preset dimension, ensuring that the feature vectors of the dimensionality-reduced decoded features have the same dimension after splitting. For instance, when the dimensionality-reduced decoded features are split into two feature vectors, the preset dimension can be half the dimension of the dimensionality-reduced decoded features. Specifically, the feature vectors split according to the preset dimension can be denoted as the first sub-feature vector and the second sub-feature vector. Each feature vector can be reconstructed into a 4-dimensional tensor with dimensions (m+1), (2n+1), T, and F, respectively. It can be understood that the 4-dimensional tensor reconstructed from the first sub-feature vector can be used as the real part of the convolutional kernel parameters, while the 4-dimensional tensor reconstructed from the second sub-feature vector can be used as the imaginary part of the convolutional kernel parameters.
[0069] In other embodiments of the invention, the audio information of the voice input signal and the reference signal is fused, including:
[0070] In this embodiment of the invention, cross-attention can be used to fuse the audio information of the speech input signal and the reference signal. Attention weights can be determined between the audio information of the speech input signal and the reference signal. The speech input signal and the reference signal can be aligned according to the attention weights, and then the two aligned audio data can be fused into one audio data. This fusion can be performed by summation, product, or direct concatenation. The weights used for fusing the audio information of the speech input signal and the reference signal are not limited to cross-attention weights, but can also include self-attention weights, channel attention weights, correlation weights, entropy weights, etc.
[0071] Figure 3 This is a flowchart of another speech processing method according to an embodiment of the present invention. The embodiment of the present invention describes the data processing process of a deep speech quality enhancement model. See [link to flowchart documentation]. Figure 3 The method provided in this embodiment of the invention specifically includes the following steps:
[0072] Step 310: Obtain audio information of the voice input signal and the reference signal.
[0073] Step 320: Reduce the dimensionality of the audio information to obtain the reduced features; extract the time frame number and frequency point number of the reduced features as the dimension values of the convolution parameters.
[0074] In this embodiment of the invention, the audio information of the speech input signal and the reference signal can be dimensionality reduced. This dimensionality reduction can be achieved by convolution with a 1×1 kernel, or by principal component analysis. The dimensionality-reduced audio information can be recorded as dimensionality-reduced features. The number of time frames and the number of frequency points can be extracted from the dimensionality-reduced features as the time-domain and frequency-domain features of the speech input signal. The extracted number of time frames and the number of frequency points can be used as the dimensional values of different dimensions of the convolution parameters. It can be understood that the convolution parameters can include only two dimensions: the number of time frames and the number of frequency points. Alternatively, they can be extended to three, four, or even five dimensions according to fixed dimensional values. The fixed dimensional values can be pre-configured or determined based on the audio information.
[0075] Step 330: Divide the signal features of the audio information into a first sub-feature set and a second sub-feature set.
[0076] Specifically, the signal features of audio information can be divided into two different feature sets, denoted as the first sub-feature set and the second sub-feature set, respectively. The dimensions of the first sub-feature set and the second sub-feature set can be the same. For example, when the dimension reduction decoding feature is split into two feature vectors, the preset dimension can be half the dimension of the dimension reduction decoding feature. Of course, in some embodiments, the dimensions of the first sub-feature set and the second sub-feature set can be different.
[0077] Step 340: Reconstruct the first sub-feature set into a first tensor according to the first dimension as the real part parameter of the convolution kernel, and reconstruct the second sub-feature vector into a second tensor according to the second dimension as the imaginary part parameter of the convolution kernel.
[0078] In this embodiment of the invention, the first sub-feature set can be reconstructed into a first tensor according to a first preset dimension, and the reconstructed first tensor can be used as the real part parameter of the convolution kernel. The second sub-feature set can also be reconstructed into a second tensor according to a second preset dimension and used as the imaginary part parameter of the convolution kernel.
[0079] For example, in some embodiments, the first preset dimension and the second preset dimension can be associated with the number of time frames T and the number of frequency points F of the speech input signal. For instance, the first sub-feature set of the audio information can be split into a multi-dimensional tensor according to the number of time frames T and the number of frequency points F. This multi-dimensional tensor can be denoted as the first tensor and can be used as the real part parameter of the convolution kernel of the adjusted neural network. Similarly, the second sub-feature set of the audio information can be split into a multi-dimensional tensor according to the number of time frames T and the number of frequency points F as the second tensor, and this second tensor can be used as the imaginary part parameter of the convolution kernel of the adjusted neural network.
[0080] Step 350: Enhance audio information using the adjusted neural network.
[0081] This invention, through receiving audio information from a speech input signal and a reference signal, reduces the dimensionality of the audio information as a dimensionality-reduced feature. The number of time frames and frequency points of the dimensionality-reduced feature are extracted as the dimensionality values of the convolution parameters. The signal features of the audio information are divided into a first sub-feature set and a second sub-feature set. The first sub-feature set is reconstructed into a first tensor according to a first preset dimension as the real part parameter of the convolution kernel, and the second sub-feature set is reconstructed into a second tensor according to a second preset dimension as the imaginary part parameter of the convolution kernel to adjust the convolution kernel of the neural network. The adjusted neural network enhances the audio information. This invention, through encoding processing to extract the complex spectral features of the audio signal, adjusts the convolution kernel of the neural network based on the time-domain and frequency-domain features of the speech input signal. This improves the adaptability of the neural network to the audio information, enhances the processing effect of audio data, reduces noise residue, improves the neural network's ability to handle sudden noise and rapidly changing time-delay data, achieves precise interference suppression, and improves the quality of speech processing.
[0082] In some embodiments of the invention, the audio information is enhanced using the adjusted neural network, including:
[0083] For each time-frequency point corresponding to each feature data in the audio information, local neighborhood blocks are formed by extracting neighboring feature data within the audio information that are located before the time-frequency point in the time domain and on both sides of the time-frequency point in the frequency domain, according to the time-domain size and frequency-domain size of the convolution kernel. The time-domain size and the frequency-domain size are pre-configured. For each neighboring feature data in each local neighborhood block, a first product of the real part element of the neighboring feature data and the real part parameter of the convolution kernel, and a second product of the imaginary part element of the neighboring feature data and the imaginary part parameter of the convolution kernel are determined. The sum of the first product and the second product is used as the complex number operation result of the neighboring feature data. The cumulative sum of the complex number operation results of each neighboring feature data in each local neighborhood block is used as the augmented data of the feature data corresponding to the local neighborhood block in the audio information.
[0084] Among them, the time and frequency points can be the time domain points and frequency domain points corresponding to each feature data in the audio information. The time domain size and frequency domain size can be pre-configured in the convolution parameters to control the data processing scale in the time domain and frequency domain of the audio data. The time domain size and frequency domain size can be pre-configured or determined according to the time domain characteristics and frequency domain characteristics of the speech input signal.
[0085] In this embodiment of the invention, the temporal and frequency domain dimensions of the convolution kernel are extracted. For each feature data in the audio data, based on the time-frequency point of the feature data, the temporal dimension range before the time-frequency point of the feature data is extracted in the audio information according to the temporal and frequency domain dimensions of the convolution kernel. The feature data within the frequency domain dimension range on both sides of the time-frequency point of the feature data are recorded as neighboring feature data. The data set consisting of multiple neighboring feature data corresponding to each feature data can be used as the local neighborhood block of the feature data. For each local neighborhood block, the first product of the real part element of each neighboring feature data and the real part parameter of the convolution kernel, and the second product of the imaginary part element of the neighborhood feature data and the imaginary part parameter of the convolution kernel can be determined. The sum of the first and second products is determined as the complex number operation result of the neighboring feature data. The cumulative sum of the complex number operation results of each neighboring feature data in each local neighborhood block is used as the enhanced data of the feature data corresponding to the local neighborhood block in the audio information.
[0086] Figure 4 This is a flowchart of another speech processing method according to an embodiment of the present invention. The embodiment of the present invention describes the speech enhancement signal processing procedure in the above embodiment. See [link to flowchart]. Figure 4 The method provided in this embodiment of the invention specifically includes the following steps:
[0087] Step 410: Obtain audio information of the voice input signal and the reference signal.
[0088] Step 420: Obtain convolution parameters based on the signal features of the speech input signal, and adjust the convolution kernel of the neural network according to the convolution parameters. The signal features include at least one of frequency domain features and time domain features, and the convolution parameters include at least real part parameters and imaginary part parameters. The real part parameters and imaginary part parameters are generated based on the signal features in different dimensions.
[0089] Step 430: Enhance audio information using the adjusted neural network.
[0090] Step 440: Perform an inverse transform on the enhanced audio information to obtain a time-domain signal, and then overlap and add the time-domain signals of each frame to obtain the speech enhancement signal.
[0091] In this embodiment of the invention, an inverse transform can be performed on the complex spectrum data to convert it into a time-domain signal. This inverse transform can be achieved through inverse Fourier transform or inverse short-time Fourier transform, etc. The time-domain signals of each frame can be fused by overlapping and adding them together to obtain the speech enhancement signal. The size of the overlap between the time-domain signals of each frame can correspond to the preset time window used in the audio complex spectrum feature extraction process. For example, if the audio complex spectrum features are extracted using a square root Hanning window, the overlap between the time-domain signals of each frame is 50%, thereby ensuring that the amplitude of the overlapping region of the speech enhancement signal after overlapping and adding is smooth.
[0092] This invention, in its embodiments, acquires a speech input signal and a reference signal, determines the audio data of the speech input signal and the reference signal, adjusts the convolutional parameters of the neural network based on the signal features of the speech input signal, enhances the audio data according to the adjusted neural network, obtains a time-domain signal through inverse transformation of the audio data, and then overlaps and adds the various time-domain signals to generate a speech enhancement signal corresponding to the speech input signal. This invention, based on a convolutional kernel adapted to the speech input signal, enhances the audio data, improves the accuracy of separating interference information within the speech data, reduces the degree of speech signal distortion, and improves the quality of speech processing.
[0093] One embodiment of this invention addresses the shortcomings of existing methods that construct complex-valued masks using fixed weights by providing a speech processing method based on a Residual Convolutional Neural Network (ResNet) and a Recurrent Neural Network (RNN). Specifically, the model first extracts local spatial features of the input speech signal in the time-frequency domain using residual convolutional layers; then, it captures the long-term contextual dependencies of the signal over time using convolutional layers. Above this, a cross-attention module is introduced to achieve precise soft alignment of near-end and far-end signals. Finally, a dedicated convolutional reconstruction module uses the high-level features fused by the cross-attention model as input to directly generate a set of refined and adaptive complex-valued masks through a single complex-valued convolution operation. This mask is multiplied by the complex spectrum of the noisy speech, thereby simultaneously eliminating echo, suppressing noise, and removing reverberation in one step, achieving efficient end-to-end joint optimization.
[0094] In one embodiment of the present invention, two input signals are acquired: one is a near-end microphone signal, which includes near-end speech, ambient noise, reverberation, and far-end echo generated by the speaker; the other is a far-end reference signal, which serves as the reference input for echo cancellation. The near-end microphone signal is used as the target signal to be processed. Specifically, the basic architecture model adopts clean speech acoustic modeling. Based on existing clean speech corpus data, a real-time deep speech quality enhancement model is constructed. The input feature is a power-law compressed complex spectrum calculated using a square root Hamming window, and the output is the enhanced target speech signal. The speech processing method provided by this embodiment of the present invention may include the following steps:
[0095] Step 1: Signal Acquisition and Preprocessing
[0096] 1. Real-time capture of two audio streams
[0097] (1) Near-end microphone signal x(n): sampling rate 16kHz, length 32ms per frame (i.e. 512 sampling points). This signal is the target signal, expressed as x(n)=s(n)+n(n)+r(n)+d(n), where s(n) is the near-end speech, n(n) is the ambient noise, r(n) is the room reverberation, and d(n) is the echo generated by the speaker playing the far-end signal.
[0098] (2) Far-end reference signal y(n): Also with a sampling rate of 16kHz and a frame of 32ms. The far-end reference signal can be the audio data stream that is about to be played by the speaker or is currently being played.
[0099] 2. Feature Extraction
[0100] For each frame of near-end and far-end signals, a square root Hanning window (SFT) is applied to perform a short-time Fourier transform (STFT) with a window length of 512 and a frame shift of 256. The STFT transforms X(t,f) into the frequency domain of the near-end microphone signal and Y(t,f) into the frequency domain of the far-end reference signal. Power-law compressed complex spectra are then used as model input features: taking the near-end microphone signal X(t,f) as an example, the amplitude spectrum is power-law compressed (|X|^0.5), and then the power-law compressed amplitude spectrum is combined with the phase spectrum to form the complex input features X_frames of the model, where X_frames∈C^(T×F), where T is the number of frames and F is the number of frequency points.
[0101] Step 2: Voice Quality Enhancement
[0102] The complex spectral features of the two signals are input into the pre-trained real-time deep speech quality enhancement model. The model's forward propagation process is as follows:
[0103] 1. Feature Encoding: The encoder consists of a microphone and a far-end branch. The microphone branch has five coding blocks, while the far-end branch has only two. Following this is an alignment module. Each coding block consists of a downsampling convolutional layer, batch normalization (BatchNorm), and an ELU activation function. The complex spectra of the near-end signal X and the far-end signal Y are respectively input into the corresponding branches. This step focuses on extracting local spectral patterns (such as phonemes and formants) in the time-frequency domain.
[0104] 2. Cross-Attention Alignment: After the complex spectra of the near-end signal X and the far-end signal Y are processed by two coding blocks, the two outputs are fed together into the cross-attention module to achieve precise soft alignment of the far-end reference signal and the near-end mixed signal in the time-frequency domain, effectively solving the time delay problem. The aligned far-end and microphone features are concatenated and fed into the third coding block in the microphone branch for further encoding.
[0105] 3. Bottleneck Layer: Located between the encoder and decoder, the bottleneck consists of a recurrent layer and a linear projection. The feature map from the encoder is input into the recurrent layer and flattened along the channel and frequency dimensions. A gated recurrent unit (GRU) is used to reduce model complexity. Using a linear projection after the recurrent layer reduces the number of hidden units in the recurrent layer, thereby improving performance and training stability.
[0106] 4. Feature Decoding: The decoder consists of five decoding blocks. The residual blocks are omitted in the first and last decoding blocks to save more computation. The remaining modules are constructed by stacking skip blocks, residual blocks, subpixel convolutional blocks, batch normalization (BatchNorm), and ELU activation functions.
[0107] 5. Complex Convolutional Mask Block: The complex convolutional mask block consists of two sequentially executed stages: a parameter generation stage and a complex filtering stage. Its core idea is that the neural network dynamically generates a complex convolutional kernel that adapts to time and frequency based on the characteristics of the input signal. This kernel is then used to perform local convolution operations on the complex spectrum of the input microphone signal to accurately recover the complex spectrum of the target speech.
[0108] (1) First stage: Parameter generation of complex convolution kernel
[0109] The high-level feature map output from the decoder is used as input, and a parameter generator network is used to infer all parameters of the complex-valued convolutional kernels used for filtering in real time. The specific operation steps are as follows:
[0110] Step 1: Input the feature tensor H output by the decoder (its dimension is C × T × F, where C is the number of channels, T is the number of time frames, and F is the number of frequency points) into a 1×1 convolutional layer, and output a new feature tensor H' (its dimension is K × T × F), where K is the number of feature channels after dimensionality reduction.
[0111] Step 2: Divide the feature tensor H' obtained in Step 1 into two parts along the channel dimension (K dimensions). Specifically, extract the data from the first K / 2 channels and reshape it into a four-dimensional tensor, denoted as the real part parameter M_real of the convolution kernel (dimension (m+1) × (2n+1) × T × F). Similarly, extract the data from the last K / 2 channels and reshape it, denoted as the imaginary part parameter M_imag of the convolution kernel (dimension (m+1) × (2n+1) × T × F), where m is the temporal dimension of the convolution kernel, n is the frequency dimension of the convolution kernel, and m and n are integers.
[0112] In this embodiment of the invention, step 1 achieves feature dimensionality reduction and extraction of specific information, which prepares for generating convolution kernel parameters. The value of K is determined by the total number of parameters required in subsequent steps. Step 2 clarifies the complex properties of the convolution kernel, preparing for subsequent filtering operations in the complex domain.
[0113] (2) Second stage: Complex filtering based on dynamic convolution kernel
[0114] Using the complex-valued convolution kernel generated in the first stage, a local convolution operation is performed on the input complex spectrum, directly outputting the enhanced complex spectrum. The specific steps are as follows:
[0115] Step 1: For the target time-frequency point (t, f), extract a local neighborhood block of size (m+1)×(2n+1) from the input microphone complex spectrum X (dimension T × F). To ensure causality, this local neighborhood block only contains the current frame and the past m frames on the time axis, that is, the value range of the time domain index i of the local neighborhood block is [-m, 0]; and it contains the current frequency point and the n points to its left and right on the frequency axis, that is, the value range of the frequency domain index j of the local neighborhood block is [-n, n].
[0116] Step 2: Perform a complex multiplication operation on each complex element X(t+i, f+j) in the neighborhood block extracted in Step 1, and the convolutional kernel weight M(i, j, t, f) = M_real(i, j, t,f)+j×M_imag(i,j,t,f) generated in Step 2 of the first stage corresponding to the same time-frequency point (t, f). Then, sum the results for all points in the neighborhood to obtain the enhanced time-frequency point Ŝ(t,f). This operation is strictly defined by the following formula:
[0117]
[0118] Where i∈[-m, 0], j∈[-n, n].
[0119] In this embodiment of the invention, step 1 obtains local context information, i.e., a local neighborhood block centered at (t, f), which serves as input data for convolutional filtering. Step 2 is equivalent to using a content-adaptive, localized two-dimensional filter to filter the spectrum, which can fully utilize the context information of adjacent time-frequency units, thus being more effective than traditional point-to-point masks.
[0120] 6. Inverse Transform: Perform an Inverse Short-Time Fourier Transform (iSTFT) on Ŝ and use the Overlap-Add method to reconstruct the final enhanced time-domain clean speech signal ŝ(n).
[0121] In this embodiment of the invention, the adaptive characteristic stems from its input dependence. Each parameter of the convolutional kernel M is inferred by the network in real time based on the current input audio features H, rather than being a pre-set fixed value. Therefore, for different input signals, such as mixed speech with different speakers, different noise types, and different reverberation levels, the features H output by the decoder will also be different. Consequently, the convolutional kernel M generated after passing through a 1×1 convolutional layer will also change accordingly. This means that when processing frames dominated by noise, the network will generate a filter kernel that excels at noise reduction; when processing frames dominated by speech, the network will generate a filter kernel that preserves speech fidelity; and for frequencies f1 and f2, the network will also generate different filter weights.
[0122] Figure 5 This is a schematic diagram of the structure of a voice processing device according to an embodiment of the present invention, as shown below. Figure 5 As shown, the device includes:
[0123] The feature extraction module 510 is used to acquire audio information of the speech input signal and the reference signal.
[0124] The network adjustment module 520 is used to obtain convolution parameters based on the signal features of the speech input signal, and adjust the convolution kernel of the neural network according to the convolution parameters. The signal features include at least one of frequency domain features and time domain features, and the convolution parameters include at least real part parameters and imaginary part parameters. The real part parameters and the imaginary part parameters are generated based on the signal features in different dimensions.
[0125] Feature enhancement module 530 is used to enhance the audio information using the adjusted neural network.
[0126] In this embodiment of the invention, a feature extraction module acquires a speech input signal and a reference signal, and determines the audio information of the speech input signal and the reference signal. A network adjustment module determines convolution parameters according to the signal characteristics of the speech input signal, and adjusts the convolution kernel of the neural network according to the convolution parameters. A feature enhancement module uses the adjusted neural network to enhance the audio information. This embodiment of the invention can perform feature enhancement on audio information based on a neural network that matches the signal characteristics of the speech input signal. This can improve the generalization ability of the neural network, expand the application scenarios of audio data enhancement, reduce noise residue, improve the accuracy of audio data enhancement, enhance the ability to process sudden noise and rapidly changing data, achieve precise interference suppression, and improve the quality of speech processing.
[0127] Optionally, the voice processing device may also include at least one of the following:
[0128] An encoder is used to encode the audio information of the voice input signal and the reference signal;
[0129] A feature fusion module is used to fuse the audio information of the voice input signal and the reference signal;
[0130] An information compression module is used to compress the audio information of the voice input signal and the reference signal;
[0131] A decoder is used to decode the audio information of the voice input signal and the reference signal.
[0132] Based on the above embodiments of the invention, the network adjustment module 520 includes:
[0133] The signal feature unit is used to obtain the number of time frames and the number of frequency points of the audio information of the speech input signal as the signal features, and to use the number of time frames and the number of frequency points as the dimension values of the convolution parameters, respectively.
[0134] The convolution kernel adjustment unit is used to construct a feature tensor based on the number of time frames and the number of frequency points, and to divide the feature tensor into real part parameters and imaginary part parameters according to different dimensions, which serve as the convolution kernel of the neural network.
[0135] In some embodiments of the invention, the decoder includes at least five sequentially connected decoding blocks, wherein the first and fifth decoding blocks omit residual blocks, and the other decoding blocks include stacked skip blocks, residual blocks, subpixel convolution blocks, batch normalization, and ELU activation functions.
[0136] Based on the above embodiments of the invention, the decoder used by the complex spectrum convolution unit consists of five decoding blocks. The first and fifth decoding blocks can be the first and last decoding blocks in a series of five sequentially connected decoding blocks. The first and fifth decoding blocks may not include residual blocks, while the other decoding blocks may include skip blocks, residual blocks, subpixel convolution blocks, batch normalization (BatchNorm), and ELU activation functions. It is understood that the first and fifth decoding blocks may stack skip blocks, subpixel convolution blocks, batch normalization (BatchNorm), and ELU activation functions.
[0137] The speech processing device provided in one embodiment of the present invention can execute the speech processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0138] Figure 6 This is a schematic diagram of an electronic device implementing a voice processing method according to an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0139] like Figure 6As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0140] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0141] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as speech processing methods.
[0142] In some embodiments, the voice processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the voice processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the voice processing method by any other suitable means (e.g., by means of firmware).
[0143] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0144] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0145] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0146] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0147] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0148] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0149] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0150] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A speech processing method, characterized in that, The method includes: Acquire audio information from the voice input signal and the reference signal; Convolutional parameters are obtained based on the signal features of the speech input signal, and the convolutional kernel of the neural network is adjusted according to the convolutional parameters. The signal features include at least one of frequency domain features and time domain features, and the convolutional parameters include at least real part parameters and imaginary part parameters. The real part parameters and the imaginary part parameters are generated based on the signal features in different dimensions. The audio information is enhanced using the adjusted neural network.
2. The speech processing method according to claim 1, characterized in that, The acquisition of audio information from the voice input signal and the reference signal includes: Acquire the voice input signal collected by the microphone and the reference signal output by the speaker; Feature extraction is performed on the voice input signal and the reference signal based on a preset time window, and time-frequency data is obtained by transforming the voice input signal and the reference signal. The amplitude spectrum of the time-frequency data is subjected to power-law compression complex spectrum processing to obtain the audio information.
3. The speech processing method according to claim 1 or 2, characterized in that, The step of obtaining convolution parameters based on the signal features of the speech input signal and adjusting the convolution kernel of the neural network according to the convolution parameters includes: The number of time frames and the number of frequency points of the audio information of the voice input signal are obtained as the signal features, and the number of time frames and the number of frequency points are respectively used as the dimension values of the convolution parameters; A feature tensor is constructed based on the number of time frames and the number of frequency points. The feature tensor is then divided into real and imaginary parameters according to different dimensions, which serve as the convolution kernel of the neural network.
4. The method according to claim 1 or 2, characterized in that, The process also includes preprocessing the audio information, wherein the preprocessing includes at least one of the following: The audio information of the voice input signal and the reference signal is encoded; The audio information of the voice input signal and the reference signal are fused; The audio information of the voice input signal and the reference signal is compressed; The audio information of the voice input signal and the reference signal is decoded.
5. The speech processing method according to claim 3, characterized in that, The step of acquiring the number of time frames and frequency points of the audio information as the signal features, and using the number of time frames and frequency points as the dimension values of the convolution parameters, includes: The audio information is then dimensionality reduced to obtain dimensionality-reduced features; The number of time frames and the number of frequency points of the dimensionality-reduced features are extracted as the dimension values of the convolution parameters.
6. The method according to claim 3, characterized in that, The step of constructing a feature tensor based on the number of time frames and the number of frequency points, and then segmenting the feature tensor according to different dimensions to form real and imaginary parameters, which serve as the convolution kernel of the neural network, includes: The signal features of the audio information are divided into a first sub-feature set and a second sub-feature set; The first sub-feature set is reconstructed into a first tensor according to a first preset dimension and used as the real part parameter of the convolution kernel. The second sub-feature set is reconstructed into a second tensor according to a second preset dimension and used as the imaginary part parameter of the convolution kernel.
7. The speech processing method according to claim 1, characterized in that, Also includes: The enhanced audio information is inversely transformed to obtain a time-domain signal, and the time-domain signals of each frame are overlapped and added together to obtain the speech enhancement signal.
8. The speech processing method according to claim 1, characterized in that, The enhancement of the audio information using the adjusted neural network includes: For each time-frequency point corresponding to each feature data in the audio information, local neighborhood blocks are formed by extracting the neighboring feature data in the audio information that are located before the time-frequency point in the time domain and on both sides of the time-frequency point in the frequency domain, according to the time domain size and frequency domain size of the convolution kernel. For each of the neighboring feature data of each of the local neighborhood blocks, a first product of the real part element of the neighboring feature data and the real part parameter of the convolution kernel, and a second product of the imaginary part element of the neighboring feature data and the imaginary part parameter of the convolution kernel are determined, and the sum of the first product and the second product is taken as the complex number operation result of the neighboring feature data; The sum of the complex number operations of each neighboring feature data within each local neighborhood block is used as the augmented data of the feature data corresponding to the local neighborhood block in the audio information.
9. A voice processing device, characterized in that, The device includes: The feature extraction module is used to acquire audio information from the speech input signal and the reference signal; A network adjustment module is used to obtain convolution parameters based on the signal features of the speech input signal, and adjust the convolution kernel of the neural network according to the convolution parameters. The signal features include at least one of frequency domain features and time domain features, and the convolution parameters include at least real part parameters and imaginary part parameters. The real part parameters and the imaginary part parameters are generated based on the signal features in different dimensions. A feature enhancement module is used to enhance the audio information using the adjusted neural network.
10. The voice processing apparatus according to claim 9, characterized in that, The device further includes at least one of the following: An encoder is used to encode the audio information of the voice input signal and the reference signal; A feature fusion module is used to fuse the audio information of the voice input signal and the reference signal; An information compression module is used to compress the audio information of the voice input signal and the reference signal; A decoder is used to decode the audio information of the voice input signal and the reference signal.
11. The speech processing apparatus according to claim 9, characterized in that, The network adjustment module includes: The signal feature unit is used to obtain the number of time frames and the number of frequency points of the audio information as the signal features, and to use the number of time frames and the number of frequency points as the dimension values of the convolution parameters, respectively. The convolution kernel adjustment unit is used to construct a feature tensor based on the number of time frames and the number of frequency points, and to divide the feature tensor into real part parameters and imaginary part parameters according to different dimensions, which serve as the convolution kernel of the neural network.
12. The voice processing apparatus according to claim 10, characterized in that, The decoder includes at least five sequentially connected decoding blocks. The first and fifth decoding blocks omit the residual blocks. The other decoding blocks include stacked skip blocks, residual blocks, subpixel convolution blocks, batch normalization, and ELU activation functions.
13. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the speech processing method according to any one of claims 1-8.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the speech processing method according to any one of claims 1-8.