Speech processing method, conference speech enhancement method, and speech model training method

By synchronizing the feature of speech data in the time-frequency domain, the problem of poor noise removal effect in the prior art is solved, and high-accurate speech processing and enhancement effect are achieved.

WO2025092406A1PCT designated stage expired Publication Date: 2025-05-08ALIBABA (CHINA) CO LTD

Patent Information

Application Number
PCT/CN2024/124577
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-03
Filing Date
2024-10-12
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively remove noise when processing voice data, resulting in a decline in voice quality, and the inability to fully explore the feature correlations on the time domain and frequency domain channels, affecting the voice enhancement effect.

Method used

By obtaining voice data, encoded as feature sequences of time domain channels and frequency domain channels, perform time-frequency domain synchronization feature processing, obtain target voice feature vectors, and decode based on this to improve speech processing results.

Benefits of technology

It realizes high accuracy processing of voice data, fully explores the feature correlations on time and frequency domain channels, and improves voice enhancement effect and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024124577_08052025_PF_FP_ABST
    Figure CN2024124577_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present specification provide a speech processing method, a conference speech enhancement method, and a speech model training method. The speech processing method comprises: acquiring speech data to be processed; encoding the speech data to obtain a speech feature vector, the speech feature vector comprising feature sequences of a time domain channel and a frequency domain channel; performing time-frequency domain synchronization feature processing on the feature sequences of the time domain channel and the frequency domain channel to obtain a target speech feature vector; and, on the basis of the target speech feature vector, decoding to obtain a speech processing result. In the time domain channel and the frequency domain channel, feature processing is synchronously carried out, and the close correlation between the features on the time domain channel and the frequency domain channel is fully exploited, such that effective context understanding is achieved, a highly-accurate target speech feature vector is obtained for decoding, a highly-accurate speech processing result is obtained, and the effectiveness of speech processing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Speech processing method, conference speech enhancement method and speech model training method

[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on November 3, 2023, with application number 202311462977.0 and application name “Speech Processing Method, Conference Speech Enhancement Method and Speech Model Training Method”, the entire contents of which are incorporated by reference into this disclosure. Technical Field

[0002] The embodiments of this specification relate to the fields of speech processing methods, conference speech enhancement methods, or speech model training technology, and in particular to a speech processing method. Background Art

[0003] With the development of Internet technology, voice data has been widely used in production and life scenarios such as social communication, online meetings and video production.

[0004] Currently, voice data inevitably carries noise during its generation, transmission, and reception, significantly impacting its usefulness. Speech enhancement technology improves voice quality, addresses noise interference, and enhances the effectiveness of voice data. Therefore, extracting high-quality target speech feature vectors to perform speech enhancement is a pressing issue.

[0005] Summary of the Invention

[0006] In view of this, embodiments of this specification provide a speech processing method. One or more embodiments of this specification also relate to a conference speech enhancement method, a speech model training method, a speech processing device, a conference speech enhancement device, a speech model training device, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.

[0007] According to a first aspect of the embodiments of this specification, a speech processing method is provided, including:

[0008] Obtaining voice data to be processed;

[0009] Encoding the speech data to obtain a speech feature vector, wherein the speech feature vector includes a feature sequence of a time domain channel and a frequency domain channel;

[0010] Perform time-domain and frequency-domain synchronous feature processing on the feature sequences of the time domain channel and the frequency domain channel to obtain the target speech feature vector;

[0011] Based on the target speech feature vector, the speech processing result is obtained by decoding.

[0012] According to a second aspect of the embodiments of this specification, a conference speech enhancement method is provided, which is applied to a cloud-side device, including:

[0013] Get conference voice data;

[0014] Encoding the conference voice data to obtain a voice feature vector, wherein the voice feature vector includes a feature sequence of a time domain channel and a frequency domain channel;

[0015] Perform time-domain and frequency-domain synchronous feature processing on the feature sequences of the time domain channel and the frequency domain channel to obtain the target speech feature vector;

[0016] Based on the target speech feature vector, the enhanced conference speech data is decoded;

[0017] Sends enhanced conference voice data to the front end.

[0018] According to a third aspect of the embodiments of this specification, a speech model training method is provided, which is applied to a cloud-side device, including:

[0019] Acquire a sample set, wherein the sample set includes sample speech data and labeled speech data;

[0020] Inputting the sample speech data into the encoding module of the speech model and encoding the sample speech data to obtain the sample speech feature vector, wherein the speech model includes an encoding module, a processing module and a decoding module;

[0021] The sample speech feature vector is input into the processing module, and the feature sequences of the time domain channel and the frequency domain channel are subjected to time-frequency domain synchronization feature processing to obtain the predicted speech feature vector;

[0022] The predicted speech feature vector is input into the decoding module and decoded to obtain the predicted speech data;

[0023] Calculate the loss value based on the predicted speech data and the labeled speech data;

[0024] Based on the loss value, the model parameters of the speech model are adjusted, and when the preset training end conditions are met, a trained speech model is obtained;

[0025] Send the model parameters of the voice model to the end-side device.

[0026] According to a fourth aspect of the embodiments of this specification, there is provided a speech processing device, including:

[0027] A first acquisition module is configured to acquire voice data to be processed;

[0028] A first encoding module is configured to encode the speech data to obtain a speech feature vector, wherein the speech feature vector includes a feature sequence of a time domain channel and a frequency domain channel;

[0029] The first processing module is configured to perform time-frequency domain synchronization feature processing on the feature sequences of the time domain channel and the frequency domain channel to obtain a target speech feature vector;

[0030] The first decoding module is configured to decode and obtain a speech processing result based on the target speech feature vector.

[0031] According to a fifth aspect of the embodiments of this specification, a conference voice enhancement device is provided, which is applied to a cloud-side device, including:

[0032] A second acquisition module is configured to acquire conference voice data;

[0033] A second encoding module is configured to encode the conference voice data to obtain a voice feature vector, wherein the voice feature vector includes a feature sequence of a time domain channel and a frequency domain channel;

[0034] The second processing module is configured to perform time-frequency domain synchronization feature processing on the feature sequences of the time domain channel and the frequency domain channel to obtain a target speech feature vector;

[0035] A second decoding module is configured to decode the enhanced conference voice data based on the target voice feature vector;

[0036] The data sending module is configured to send the enhanced conference voice data to the front end.

[0037] According to a sixth aspect of the embodiments of this specification, a speech model training device is provided, which is applied to a cloud-side device, including:

[0038] A third acquisition module is configured to acquire a sample set, wherein the sample set includes sample speech data and labeled speech data;

[0039] a third encoding module configured to input the sample speech data into an encoding module of a speech model and encode the sample speech data to obtain a sample speech feature vector, wherein the speech model includes an encoding module, a processing module, and a decoding module;

[0040] A third processing module is configured to input the sample speech feature vector into the processing module, perform time-frequency domain synchronization feature processing on the feature sequences of the time domain channel and the frequency domain channel, and obtain a predicted speech feature vector;

[0041] a third decoding module, configured to input the predicted speech feature vector into the decoding module and decode to obtain predicted speech data;

[0042] A loss calculation module is configured to calculate a loss value based on the predicted speech data and the labeled speech data;

[0043] The training adjustment module is configured to adjust the model parameters of the speech model based on the loss value, and obtain a trained speech model when a preset training end condition is met;

[0044] The model sending module is configured to send the model parameters of the speech model to the end-side device.

[0045] According to a seventh aspect of the embodiments of this specification, a computing device is provided, including:

[0046] memory and processor;

[0047] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above method are implemented.

[0048] According to an eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and the steps of the above method are implemented when the instructions are executed by a processor.

[0049] According to a ninth aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above method.

[0050] In one embodiment of the present specification, speech data to be processed is obtained; the speech data is encoded to obtain a speech feature vector, wherein the speech feature vector includes a feature sequence of a time domain channel and a frequency domain channel; based on the speech feature vector, attention calculation is performed on the speech feature vector in the time domain channel and the frequency domain channel to obtain a target speech feature vector; based on the target speech feature vector, the speech processing result is obtained by decoding. When the speech data is feature-encoded to obtain a speech feature vector including a feature sequence of a time domain channel and a frequency domain channel, feature processing is performed simultaneously on the time domain channel and the frequency domain channel, fully exploiting the close correlation between the features in the time domain channel and the frequency domain channel, achieving effective contextual understanding of the speech data, capturing the complex interaction in the time domain and the frequency domain, obtaining a highly accurate target speech feature vector for decoding, and obtaining a highly accurate speech processing result, thereby improving the effectiveness of speech processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] FIG1 is a flow chart of a speech processing method provided by one embodiment of this specification;

[0052] FIG2 is a schematic diagram of the structure of a speech model in a speech processing method provided by one embodiment of this specification;

[0053] FIG3 is a schematic diagram of the structure of a processing module in a speech processing method provided by one embodiment of this specification;

[0054] FIG4 is a schematic diagram of the structure of an attention module in a speech processing method provided by one embodiment of this specification;

[0055] FIG5 is a flow chart of a conference speech enhancement method provided by one embodiment of this specification;

[0056] FIG6 is a flow chart of a speech model training method provided by one embodiment of this specification;

[0057] FIG7 is a flowchart of a processing process of a voice processing method applied to a conference voice scenario provided by one embodiment of this specification;

[0058] FIG8 is a schematic structural diagram of a speech processing device provided by one embodiment of this specification;

[0059] FIG9 is a schematic structural diagram of a conference voice enhancement device provided by one embodiment of this specification;

[0060] FIG10 is a schematic diagram of the structure of a speech model training device provided by one embodiment of this specification;

[0061] FIG11 is a structural block diagram of a computing device provided in one embodiment of this specification. DETAILED DESCRIPTION

[0062] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0063] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0064] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0065] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0066] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model. It is pre-trained on a large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as a large language model (LLM) and a multi-modal pre-training model.

[0067] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0068] First, the terms involved in one or more embodiments of this specification are explained.

[0069] Speech enhancement: By reducing noise interference, it enhances the audibility and intelligibility of real speech data, as well as the back-end speech recognition effect.

[0070] Attention: A widely used technique in deep learning, it originates from neuroscience research on how humans allocate attention. It simulates how humans focus on different pieces of information and how they allocate attention. Specifically, it manifests as a weighted feature that assigns weights to input information and is used in a series of deep learning networks based on the Transformer model.

[0071] Fourier transform: A signal processing method used to convert a signal from the time domain (i.e., the time domain) to the frequency domain (i.e., the frequency domain). The Fourier transform can represent a signal as a linear combination of sine and cosine functions, making it easy to analyze the frequency characteristics of the signal.

[0072] Short-Time Fourier Transform (STFT): It is an effective Fourier transform variant for speech signals and is used to process stationary signals in a short time.

[0073] Loop filtering: A widely used technique in signal processing that exploits the cyclic nature of a signal to filter it. The basic idea behind loop filtering is to treat each signal sample as a loop, concatenating the loops to form a long signal. This long signal is then filtered using a filter, effectively removing noise and other interference. The main advantage of loop filtering is that it effectively removes noise and other interference from the signal, improving model performance and accuracy.

[0074] RNNs (Recurrent Neural Networks) are a special type of neural network whose structure allows for variable-length input sequences and memorizes previously input information. RNNs can model dependencies between sequential data by transitioning state within the network while continuously updating hidden states.

[0075] LSTM (Long Short-Term Memory) is a specially designed RNN that can better capture long-term dependencies. It features a special memory module that determines which information to retain and which to discard. It also has a complex gating system that controls the flow of information into and out of the memory module.

[0076] The Gated Recurrent Unit (GRU) is another specialized RNN designed to simplify LSTM networks and improve their efficiency. It uses only two gates, rather than three: a reset gate and an update gate. These two gates control the updating of short-term memory in the network, enabling the GRU to more effectively capture long-term dependencies.

[0077] At present, speech enhancement based on deep learning networks mainly relies on processing feature sequences in the time domain channel and the frequency domain channel respectively, analyzing the speech features in the time domain channel and the speech features in the frequency domain channel, and performing corresponding speech enhancement processing.

[0078] However, this dual-channel processing method lacks cross-channel interaction between features on the two channels, cannot fully exploit the close correlation between features on the time domain channel and the frequency domain channel, and cannot effectively understand the context of speech data to capture the complex interaction in the time domain and frequency domain. There is a bottleneck in speech enhancement performance and the speech processing effect is insufficient.

[0079] In response to the above problems, this specification provides a speech processing method. This specification also involves a conference speech enhancement method, a speech model training method, a speech processing device, a conference speech enhancement device, a speech model training device, a computing device, a computer-readable storage medium and a computer program, which are described in detail one by one in the following embodiments.

[0080] Referring to FIG1 , FIG1 shows a flow chart of a speech processing method provided by an embodiment of this specification, including the following specific steps:

[0081] Step 102: Acquire voice data to be processed.

[0082] The embodiments of this specification are applied to voice processing applications, websites or mini-programs with voice enhancement functions. They can be the client of the application, website or mini-program, or the server of the application, website or mini-program, without limitation here.

[0083] The voice data to be processed is the voice data of the voice task to be processed. For example, the voice data of a voice enhancement task in an online meeting includes the voice data of the speaker in the meeting. Another example is the voice data of a voice noise reduction task in social software, including the voice data of each speaker in the social software. Another example is the voice data of a voice-to-text conversion task, including the input voice data.

[0084] The speech data is the original speech data to be enhanced, including real speech data and noise data, which are composed of real speech data and noise data by sound superposition. Specifically, x=s+n∈R 1×L, where x is speech data, s is real speech data, and n is noise data. Speech data is in the form of a unidirectional numerical sequence. The goal of speech enhancement is to predict the real speech data from the speech data. For example, the voice data is the original voice data to be enhanced collected by the speaker's peripheral device, including the speaker's real voice data and noise data caused by the peripheral device and the environment.

[0085] Acquiring the voice data to be processed may be receiving the voice data to be processed uploaded by a voice acquisition device, or may be acquiring the voice data to be processed from a database, which is not limited here.

[0086] For example, user A logs in to an online conference application, creates a conference room in the application, conducts an online audio conference, and receives user A's speech data x∈R uploaded in real time by a voice acquisition device. 1×L .

[0087] Obtain the speech data to be processed. This lays the foundation for subsequent feature encoding.

[0088] Step 104: Encode the speech data to obtain a speech feature vector, wherein the speech feature vector includes feature sequences of the time domain channel and the frequency domain channel.

[0089] The speech feature vector is the feature encoding vector of the speech data, which is a high-dimensional vector. Specifically, the speech data is x, and the corresponding speech feature vector is X∈R T×F×C , where T is the time domain channel dimension, F is the frequency domain channel dimension, and C is the feature dimension.

[0090] The time domain channel processes speech data in the time domain, specifically analyzing the temporal variations of speech data. The frequency domain channel processes speech data in the frequency domain, specifically analyzing the distribution of speech data in the frequency domain. The feature sequence of the time domain channel is the feature encoding sequence of the temporal variations of speech data. The feature sequence of the frequency domain channel is the feature encoding sequence of the distribution of speech data in the frequency domain.

[0091] The speech data is encoded to obtain a speech feature vector, specifically by extracting speech features across the time domain and frequency domain of the speech data to obtain a speech feature vector.

[0092] For example, the speech data x is subjected to feature extraction across the time domain T and frequency domain F to obtain the speech feature vector X∈R T×F×C .

[0093] Speech data is encoded to obtain a speech feature vector, which includes feature sequences from both the time and frequency domains. This provides a high-dimensional feature encoding and lays the foundation for subsequent exploration of the close correlation between features in the time and frequency domains, enabling effective contextual understanding of speech data and capturing the complex interactions between the time and frequency domains.

[0094] Step 106: Perform time-frequency domain synchronization feature processing on the feature sequences of the time domain channel and the frequency domain channel to obtain a target speech feature vector.

[0095] Since speech feature vectors include feature sequences of time and frequency domain channels, for example, speech feature vectors are multidimensional feature vectors, including time and frequency domain channels. In a two-dimensional feature vector, the X dimension represents the time domain channel, and the Y dimension represents the frequency domain channel. Matrix multiplication of the two-dimensional feature vectors with a 1×M matrix performs feature processing in the time domain, while matrix multiplication of the two-dimensional feature vectors with an N×1 matrix performs feature processing in the frequency domain. Matrix multiplication of the two-dimensional feature vectors with a 1×M matrix and an N×1 matrix, respectively, and further feature processing of the two results, achieves synchronized feature processing in the time and frequency domains.

[0096] Time-frequency domain synchronous feature processing is a feature processing technology for mining deep key features in time domain channels and frequency domain channels in deep learning, including but not limited to attention calculation. For example, the self-attention feature of the input feature vector O is an N×M feature matrix. The input feature vector O is multiplied by the N×1 weight matrix Wq and the N×1 weight matrix Wk to obtain two N×1 feature vectors of query feature Q and key feature K. This can realize self-attention calculation on the time domain channel N and obtain the self-attention feature on the time domain channel U. Similarly, the input feature vector O is multiplied by the 1×M weight matrix Wq' and the 1×M weight matrix Wk' to obtain two 1×M feature vectors of query feature Q' and key feature K'. This can realize self-attention calculation on the frequency domain channel N and obtain the self-attention feature on the frequency domain channel V. Further, the input feature vector O is weighted by U and V to obtain the output feature vector O', realizing time-frequency domain synchronous feature processing.

[0097] Attention calculation is based on the attention mechanism to perform weighted processing on the input feature vector, focusing on key features in the input data and achieving contextual understanding. It includes but is not limited to: self-attention calculation and cross-attention calculation. It is calculated by the feature vector before the query feature (Q, Query), key feature (K, Key), and value feature (V, Value), see Formula 1:

[0098] Among them, Q is the query feature, K is the key feature, V is the value feature, d k It is a diffusion factor. When Q, K, and V are the same, it is a self-attention calculation. When Q, K, and V are different, it is a cross-attention calculation.

[0099] The target speech feature vector is a feature coding vector of complex speech interactions in the time domain and frequency domains, which is mined based on the close correlation between the features in the time domain channel and the frequency domain channel. It is a high-dimensional vector. For example, the speech feature vector is X∈R T×F×C , the target speech feature vector is O∈R T×F’×C , where the frequency domain channel dimension F' is half the frequency domain channel dimension F. The target speech feature vector represents the feature encoding vector of the mined speech data in different feature dimensions, such as phonological structure, intonation, and semantic associations in the time and frequency domains. This highly accurate feature encoding vector of the target speech features can be subsequently decoded to produce highly accurate speech processing results.

[0100] Performing time-domain synchronous feature processing on the feature sequences of the time and frequency domain channels to obtain a target speech feature vector is performed by performing an attention calculation on the speech feature vector in the time and frequency domain channels based on the speech feature vector to obtain the target speech feature vector. More specifically, based on the feature sequences of the time and frequency domain channels, query features, key features, and value features are determined on the time and frequency domain channels, and attention calculation is performed on the query features, key features, and value features to obtain the target speech feature vector.

[0101] For example, based on the speech feature vector X, the speech feature vector is self-attentionally calculated in the time domain channel and the frequency domain channel to obtain the target speech feature vector O∈R T×F’×C .

[0102] Synchronous feature processing in the time and frequency domains is performed on the feature sequences of the time and frequency domains to obtain the target speech feature vector. This synchronous feature processing of the speech feature vector in both the time and frequency domains fully exploits the close correlation between features in the time and frequency domains, enabling effective contextual understanding of speech data and capturing the complex interactions between the time and frequency domains. This yields a highly accurate target speech feature vector, ensuring highly accurate speech processing results in subsequent decoding.

[0103] Step 108: Based on the target speech feature vector, decode to obtain a speech processing result.

[0104] The speech processing result is determined based on speech data after speech enhancement. For example, this could be the speech data of a speaker in an online meeting after speech enhancement. Another example could be the speech data of each speaker in a social networking app after speech noise reduction. Another example could be text data converted from a speech-to-text conversion task.

[0105] Based on the target speech feature vector, the speech processing result is obtained by decoding. Specifically, based on the target speech feature vector, the target spectrogram is obtained by decoding. Based on the target spectrogram, the enhanced target speech data is determined. Based on the target speech data, the speech processing result is determined.

[0106] Exemplarily, based on the target speech feature vector O, the target spectrogram is decoded to obtain Based on the target spectrogram, determine the enhanced speech data of user A

[0107] In an embodiment of the present specification, speech data to be processed is obtained; the speech data is encoded to obtain a speech feature vector, wherein the speech feature vector includes a feature sequence of a time domain channel and a frequency domain channel; based on the speech feature vector, attention calculation is performed on the speech feature vector in the time domain channel and the frequency domain channel to obtain a target speech feature vector; based on the target speech feature vector, the speech processing result is obtained by decoding. When the speech data is feature-encoded to obtain a speech feature vector including a feature sequence of a time domain channel and a frequency domain channel, feature processing is performed on the speech feature vector simultaneously in the time domain channel and the frequency domain channel, fully exploiting the close correlation between the features in the time domain channel and the frequency domain channel, achieving effective contextual understanding of the speech data, capturing the complex interactions in the time domain and the frequency domain, obtaining a highly accurate target speech feature vector for decoding, obtaining a highly accurate speech processing result, and improving the effectiveness of speech processing.

[0108] In an optional embodiment of the present specification, before step 104, the following specific steps are further included:

[0109] Convert the speech data to obtain a spectrogram corresponding to the speech data;

[0110] Correspondingly, step 104 includes the following specific steps:

[0111] Feature extraction is performed on the spectrogram on the time domain channel and frequency domain channel corresponding to the spectrogram to obtain the speech feature vector.

[0112] The spectrogram corresponding to the speech data is an image used to represent the frequency characteristics of the speech data, including the real part, imaginary part and amplitude. It is used to determine the time characteristics, frequency characteristics and energy distribution of the speech data, realize the conversion of the speech data, and obtain the spectrogram corresponding to the speech data.

[0113] The speech data is converted to obtain a spectrogram corresponding to the speech data. Specifically, the speech data is subjected to Fourier transform to obtain a spectrogram corresponding to the speech data. More specifically, the speech data is subjected to Short-Time Fourier Transform (STFT) to obtain a spectrogram corresponding to the speech data. For example, the speech signal x is converted to a spectrogram X = STFT(x)∈C T×F , where STFT() is short-time Fourier transform. The real part, imaginary part and amplitude of the spectrum stack are represented by X in =[X R ,X I ,|X| P ]∈R T×F×3 , where X R is the real part, X I is the imaginary part, |X| P is the amplitude and P is a compression factor.

[0114] Feature extraction is performed on the spectrogram in both the time and frequency domains corresponding to the spectrogram to obtain a speech feature vector. Specifically, the encoding module extracts features from the spectrogram in both the time and frequency domains, and determines the speech feature vector based on the extracted speech features across the time and frequency domains. The encoding module is a network module that can encode features in both the time and frequency domains. It consists of two convolutional blocks and a dilated network. Each block consists of a convolutional layer (Conv2d), instance normalization (IN), and a parametric rectified linear unit (PReLU) activation. The first block increases the feature dimension to C, while the second block halves the frequency dimension (F' = 1 / 2F). The dilation module contains four similar dense blocks, each containing a dilated convolutional layer and a feedforward sequential memory network (FSMN) block. The dilation factors of the dilated convolutional layers are set to {1, 2, 4, 8}, in sequence. The specific speech feature extraction formula is shown in Formula 2: X = G-Conv2d(Linear(PReLU(Linear(X in )))) Formula 2

[0115] Among them, X is the speech coding vector, X inIt represents the real, imaginary, and amplitude parts of the stacked spectrograms. G-Conv2d() is a convolutional layer, which acts as a memory layer to capture long-range speech features across time and frequency domains. Linear() is a linear layer, and PReLU() is an activation layer.

[0116] For example, the speech data x is subjected to a short-time Fourier transform to obtain a spectrum corresponding to the speech data X = STFT(x)∈C T×F , using the time-frequency encoding module TF Encoder, feature extraction is performed on the spectrogram on the time domain channel T and frequency domain channel F corresponding to the spectrogram X, and the speech feature vector X is determined based on the extracted speech features across the time domain and frequency domain.

[0117] The speech data is converted to obtain a corresponding spectrogram. Feature extraction is performed on the spectrogram's corresponding time and frequency domain channels to obtain a speech feature vector. By more accurately extracting high-dimensional speech features in both the time and frequency domains, the close correlation between features in these two channels is fully explored, enabling effective contextual understanding of speech data and capturing the complex interactions between the time and frequency domains.

[0118] In an optional embodiment of the present specification, before step 106, the following specific steps are further included:

[0119] The feature sequences of the time domain channel and the frequency domain channel are filtered respectively to obtain updated speech feature vectors.

[0120] The feature sequences of the time domain channel and the frequency domain channel contain a variety of noise information, which need to be filtered separately to obtain updated speech feature vectors to ensure the accuracy of the synchronous time-frequency domain processing in step 10.

[0121] The updated speech feature vector is a feature coding vector of speech data obtained by filtering based on the speech feature vector. It is a high-dimensional vector that more fully integrates the cyclic pattern of speech data globally and locally compared to the speech feature vector.

[0122] The feature sequences of the time domain channel and the frequency domain channel are filtered respectively to obtain an updated speech feature vector. Specifically, based on the cyclic pattern of the speech feature vector, the feature sequences of the time domain channel and the frequency domain channel are filtered respectively to obtain an updated speech feature vector.

[0123] For example, based on the cyclic pattern of the speech feature vector X, the feature sequences of the time domain channel and the frequency domain channel are filtered respectively to obtain an updated speech feature vector O'∈R T×F’×C .

[0124] The feature sequences of the time and frequency domain channels are filtered separately to obtain an updated speech feature vector. This process extracts the speech features of the speech feature vector itself, filters out noise, and completes the update of the speech feature vector. This ensures that the close relationship between the features in the time and frequency domain channels is further explored, achieving effective contextual understanding of the speech data and capturing the complex interactions in the time and frequency domains.

[0125] In an optional embodiment of the present specification, the speech feature vector includes multiple feature dimensions;

[0126] Correspondingly, filtering the feature sequences of the time domain channel and the frequency domain channel respectively to obtain an updated speech feature vector includes the following specific steps:

[0127] For any feature dimension of the speech feature vector, the feature sequences of the time domain channel and the frequency domain channel within any feature dimension are subjected to cyclic filtering respectively;

[0128] Based on the loop filtering processing results corresponding to each feature dimension, an updated speech feature vector is obtained.

[0129] Circular filtering involves filtering the feature sequences of time and frequency domain channels based on the cyclic patterns of speech feature vectors. Specifically, the process takes speech feature vectors as input and uses a specific structure to capture the cyclic patterns between them. For example, in speech recognition tasks, the sound waveform can be split into a series of frames, each containing a set of acoustic features such as frequency and intensity. Through a cyclic mechanism, each hidden state is progressively updated from the previous frame to the next, thereby better capturing long-term dependencies in continuous speech data. This can be achieved using RNNs, LSTMs, GRUs, and other feature processing modules with cyclic mechanisms.

[0130] The feature dimensions are different speech feature dimensions. For example, speech data of different speakers represent different feature dimensions. The phonological structure, intonation, and semantic association of different speech data represent different feature dimensions, which are not limited here.

[0131] For any feature dimension of a speech feature vector, cyclic filtering is performed on the feature sequences of the time and frequency domain channels within each feature dimension. An updated speech feature vector is obtained based on the cyclic filtering results corresponding to each feature dimension. Specifically, a cyclic module is used to perform cyclic filtering on the global and local cyclic patterns of the feature sequences of the time and frequency domain channels within each feature dimension. An updated speech feature vector is obtained based on the cyclic filtering results corresponding to each feature dimension. The above steps are implemented using a cyclic module, a network module that can extract cyclic patterns across feature dimensions and perform the corresponding cyclic filtering. The cyclic module comprises two convolutional blocks, a feedforward sequential memory network block, and a transposed convolutional layer (T-Conv1d), organized in a gated architecture. All convolutional blocks share the same structure. The convolutional blocks include a normalization layer, a linear processing layer, a SiLU (Sigmoid-Weighted Linear Unit) activation layer, and a depthwise convolutional layer (D-Conv1d). Skip connections connect the input and output of the depthwise convolutional layer. For the speech feature vector X input to the convolution block i ∈R T×F’×C In the case of i ∈R T’×F’×C As shown in Formula 3:

[0132] Among them, LayerNorm() is the normalization process, SiLU is the SiLU() activation process, D-Conv() is the depth convolution process, and Yi is the convolution block output.

[0133] In the recurrent module, the feature dimension is doubled and the frequency dimension is halved (C' = 2C). The convolution block complements the feedforward sequential memory network block by extracting local recurrent patterns within the recurrent structure. The feedforward sequential memory network block is used to simulate recurrent patterns on the feature sequence of the time domain channel or the feature sequence of the frequency domain channel within each feature dimension. This block consists of two linear layers and a deep convolutional layer (G-Conv2d) with a PReLU activation after the first linear layer. The deep convolutional layer (G-Conv2d) acts as a memory layer, enabling the block to capture long-range recurrent patterns.

[0134] In the embodiment of this specification, the processing process from the speech feature vector to the updated speech feature vector is shown in Formula 4:

[0135]

[0136] Where X is the speech feature vector, O' is the updated speech feature vector, the transposed convolution layer (T-Conv1d) is used to restore the C dimension from the C' dimension, and the circle product represents the element-by-element matrix multiplication.

[0137] The above-mentioned cyclic module is used to implement cyclic filtering processing to obtain an updated speech feature vector, determine the local cyclic patterns in different feature dimensions, ensure that the cyclic patterns in phonological structure, intonation and semantic association are effectively integrated, and obtain a highly accurate speech feature vector.

[0138] Exemplarily, a recurrent module is used to perform recurrent filtering based on the global and local recurrent patterns of the feature sequences of the time domain channel and the frequency domain channel in any feature dimension, and an updated speech feature vector O' is obtained based on the recurrent filtering results corresponding to each feature dimension.

[0139] For each dimension of the speech feature vector, cyclic filtering is performed on the feature sequences of the time and frequency domain channels within each dimension. Based on the cyclic filtering results corresponding to each dimension, an updated speech feature vector is obtained. By extracting the cyclic patterns of the speech feature vector itself in the time and frequency domains at a fine-grained level within the feature dimensions, the speech feature vector is updated. This ensures that the close correlation between features in the time and frequency domains is further fully explored, achieving effective contextual understanding of the speech data and capturing the complex interactions in the time and frequency domains.

[0140] In an optional embodiment of this specification, step 106 includes the following specific steps:

[0141] Based on the speech feature vector, self-attention calculation is performed on the speech feature vector in the time domain channel and the frequency domain channel to obtain the self-attention feature;

[0142] Based on the self-attention feature, the speech feature vector is weighted to obtain the target speech feature vector.

[0143] Self-attention is a technique used to focus on different parts of the input, focusing on the most relevant parts of the input feature vector and ignoring less important parts. Self-attention typically consists of three parts: query features, key features, and value features, all of which are obtained by transforming the same input feature vector using a weight matrix. Self-attention features are weighted features obtained through self-attention, calculated from the query, key, and value features.

[0144] Weighted processing is to assign different self-attention features (weights) to different variables to enhance the representativeness of the data. In this method, the model assigns different weights to different parts of the input to highlight the most important parts of the input while ignoring other unimportant parts. This technology can help the model better understand the input data and extract more valuable information. This can be achieved through gating technology.

[0145] Based on the speech feature vector, self-attention calculation is performed on the speech feature vector in the time domain channel and the frequency domain channel to obtain self-attention features. The specific method is as follows: based on the speech feature vector, self-attention calculation is performed on the speech feature vector in the time domain channel and the frequency domain channel synchronously to obtain self-attention features on the time domain channel and self-attention features on the frequency domain channel. Based on the self-attention features on the time domain channel and the self-attention features on the frequency domain channel, the self-attention features are determined. For the specific calculation method on different channels, please refer to the description in step 106 above. The self-attention calculation formula is shown in Formula 5: U′, V′=TF Att (Z,U,U T ,V,V T ) Formula 5

[0146] Among them, Z, U and V are the input speech features, TF Att (), synchronous self-attention calculation is performed in the time domain channel and the frequency domain channel, U T and V T is the transpose of the input speech features, and U′, V′ are self-attention features.

[0147] Exemplarily, based on the speech feature vector Z, synchronous self-attention calculation is performed on the speech feature vector in the time domain channel and the frequency domain channel to obtain self-attention features U' and V'. Based on the self-attention features, the speech feature vector is weighted to obtain the target speech feature vector O.

[0148] Based on the speech feature vector, self-attention calculations are performed on the speech feature vector in both the time and frequency domains to obtain self-attention features. Based on the self-attention features, the speech feature vector is updated to obtain the target speech feature vector. Simultaneously performing self-attention calculations on the speech feature vector in both the time and frequency domains fully exploits the close correlation between features in the time and frequency domains, enabling effective contextual understanding of speech data and capturing the complex interactions between the time and frequency domains. This yields a more accurate target speech feature vector, ensuring highly accurate speech processing results in subsequent decoding.

[0149] In an optional embodiment of the present specification, before performing self-attention calculation on the speech feature vector in the time domain channel and the frequency domain channel based on the speech feature vector to obtain the self-attention feature, the following specific steps are also included:

[0150] Convolution processing is performed on the local features in the speech feature vector to obtain a speech feature vector with updated local features.

[0151] In an embodiment of the present specification, based on the speech feature vector, the speech feature vector is self-attentionally calculated in the time domain channel and the frequency domain channel to obtain the self-attention feature, specifically in the following manner: using the attention module, the local features in the speech feature vector are processed to obtain the speech feature vector with updated local features, and based on the speech feature vector, the speech feature vector is self-attentionally calculated in the time domain channel and the frequency domain channel to obtain the self-attention feature. Among them, the attention module is a network module with attention calculation function, which integrates three gate mechanisms to minimize the computational load of TF (Time-Frequency, time-frequency) attention. These gate mechanisms enhance the ability of the attention module to capture long-distance correlation relationships, thereby eliminating the requirement of multi-head self-attention (MHSA, Multiple Head Self Attention). First, the local features in the speech feature vector are processed by three convolution blocks (Conv-U). The central block provides a shared representation for the synchronized time-frequency attention module, while the other two blocks provide two value features for the gate mechanism. These convolution blocks supplement the attention on the time domain channel and the frequency domain channel, and can extract complex local feature patterns. The overall processing of the attention module is shown in Formula 6: U = Conv-U(O′) V = Conv-U(O′) Z = Conv-U(O′)

[0152] Where ring multiplication represents element-by-element matrix multiplication, Z∈R T×F’×d is a shared representation where d is a parameter factor.

[0153] The local features in the speech feature vector are processed to obtain the speech feature vector with updated local features, specifically in the following manner: the local features in the speech feature vector are convoluted to obtain the speech feature vector with updated local features.

[0154] Exemplarily, convolution processing is performed on the local features in the speech feature vector O' to obtain a speech feature vector Z with updated local features.

[0155] The local features in the speech feature vector are processed to obtain a speech feature vector with updated local features. By processing the local features in the speech feature vector, attention in the time and frequency domain channels is supplemented, complex local feature patterns can be extracted, and the accuracy of subsequent self-attention calculations is improved.

[0156] In an optional embodiment of the present specification, based on the speech feature vector, self-attention calculation is performed on the speech feature vector in the time domain channel and the frequency domain channel to obtain the self-attention feature, including the following specific steps:

[0157] performing transposition processing on the speech feature vector to obtain a transposed speech feature vector;

[0158] Perform self-attention calculation on the speech feature vector and the transposed speech feature vector respectively to obtain the secondary attention features of the time domain channel and the frequency domain channel;

[0159] Based on the secondary attention features of the time domain channel and the frequency domain channel, the self-attention features are determined.

[0160] The quadratic attention feature is a weighted feature widely used in deep learning. It extracts key features from the input data by calculating the quadratic relationship between each feature and other features in the input data. The quadratic attention feature focuses more on local features. The quadratic attention feature is to first calculate the attention feature by multiplying matrices of different lengths. For example, the query feature is an M×1 feature matrix, the key feature is a 1×M feature matrix, and the value feature is a 1×M feature matrix. The quadratic attention feature is to first calculate the transpose of the query feature and the key feature to obtain M 2 ×1 feature matrix, and then multiply it with the value feature to get M 2 ×N feature matrix, the attention feature obtained is the secondary attention feature. For such secondary attention features, it is necessary to perform higher-dimensional feature activation through secondary activation, such as ReLU 2 ()deal with.

[0161] Transposition refers to reversing the order of elements of a feature coding vector to form a new vector. In the embodiment of this specification, the speech feature vector includes a feature sequence of the time domain channel and a feature sequence of the frequency domain channel. The attention calculation is the multiplication between the sequences. Therefore, the time-frequency domain conversion can be achieved by transposition. For example, the row sequence of the speech feature vector X is the feature sequence of the time domain channel, and the column sequence is the feature sequence of the frequency domain channel. The multiplication between the sequences is the row sequence of one vector multiplied by the column sequence of another vector. By transposing the speech feature vector, the synchronous self-attention calculation of the time domain channel and the frequency domain channel is achieved.

[0162] The speech feature vector and the transposed speech feature vector are respectively self-attention calculated to obtain the secondary attention features of the time domain channel and the frequency domain channel. The specific method is as follows: using the attention module, the speech feature vector and the transposed speech feature vector are respectively self-attention calculated to obtain the secondary attention features of the time domain channel and the frequency domain channel. The mask setting can make the attention focus on the local features on the time domain channel or the frequency domain channel. Among them, the attention module is a network module that performs self-attention calculation on the time domain channel and the frequency domain channel simultaneously, including a query feature and a key feature conversion network pair. The input speech feature vector Z is scalared and offset in each dimension to obtain shared query features Ql and Qq, as well as shared key features Kl and Kq. Through transposition processing, transposed Qq,T and Kq,T are obtained. Subsequently, in U T and V T The second self-attention is calculated as shown in Formula 7: U′ q,T =A T U T V′ q,T =A T V T Formula 7

[0163] Among them, ReLU2() is the secondary activation process, γ T is a scaling factor, and the mask matrix M is an all-1 matrix with zeros on its diagonal. Its purpose is to prevent redundant attention calculations on elements on the time-frequency cross path, which participate in the secondary attention calculations of U and V, as shown in Formula 8:

[0164] Where γ is a scaling factor, the secondary attention on the time domain channel of Uq,T' and Vq,T' and the secondary attention on the frequency domain channel of Uq' and Vq' determine the self-attention feature, as shown in Formula 9: U′ q =U′ q +Reshape(U′ q,T ) V′ q =V′ q +Reshape(V′ q,T ) Formula 9

[0165] Among them, Reshape() is the transposition process.

[0166] It's important to note that we chose a quadratic ReLU2 activation instead of a softmax to compute the attention matrix. This choice is crucial for achieving synchronized time-frequency (TF) attention. Unlike the softmax function used in multi-head self-attention, which generates a probability distribution that requires weights to sum to 1, the quadratic ReLU2 activation relaxes this restriction. Furthermore, using a quadratic ReLU2 activation within a gated architecture has been shown to optimize attention performance.

[0167] For example, the speech feature vectors U and V are transposed to obtain the transposed speech feature vector U T and V T , through mask processing, the speech feature vector U, V and the transposed speech feature U T and V T The vectors are self-attention calculated respectively to obtain the secondary attention features Uq'Vq' of the time domain channel and the frequency domain channel. Based on the secondary attention features of the time domain channel and the frequency domain channel, the self-attention features U' and V' are determined.

[0168] The speech feature vector is transposed to obtain a transposed speech feature vector. Self-attention is calculated on both the speech feature vector and the transposed speech feature vector to obtain secondary attention features for the time and frequency domains. Based on these secondary attention features, the self-attention features are determined. Performing secondary attention calculations on both the time and frequency domains strengthens local attention in these channels, further extracting complex local feature patterns and improving the accuracy of subsequent self-attention calculations.

[0169] In an optional embodiment of the present specification, before determining the self-attention feature based on the secondary attention features of the time domain channel and the frequency domain channel, the following specific steps are also included:

[0170] Perform self-attention calculation on the speech feature vector to obtain linear attention features;

[0171] Based on the secondary attention features of the time domain channel and the frequency domain channel, the self-attention features are determined, including:

[0172] The linear attention features, the secondary attention features of the time domain channel and the frequency domain channel are fused to obtain the self-attention features.

[0173] Linear attention is a weighted feature widely used in deep learning. It extracts key features from the input data by calculating the linear relationship between each feature in the input data and other features. The linear attention feature focuses more on global features. The linear attention feature first calculates the attention feature by multiplying matrices of the same length. For example, the query feature is an M×1 feature matrix, the key feature is a 1×M feature matrix, and the value feature is a 1×M feature matrix. The linear attention feature first calculates the transpose of the key feature and the value feature to obtain a 1×1 feature matrix, and then multiplies it with the query feature to obtain an M×1 feature matrix. The resulting attention feature is the linear attention feature. For such a linear attention feature, it is necessary to perform lower-dimensional feature activation through linear activation, such as ReLU() processing.

[0174] In the embodiment of this specification, the calculation formula of linear attention is shown in Formula 10:

[0175] Where Ql is the linear query feature and Kl,T is the transpose of the linear key feature. β is a scaling factor. We add the quadratic and linear attention to form the final output of the TF attention block, as shown in Equation 11: U′ = U′ q +U′ l V′=V′ q +V′ l Formula 11

[0176] Exemplarily, self-attention calculation is performed on the speech feature vectors U and V to obtain linear attention features Ul' and Vl'; feature fusion is performed on the linear attention features Ul' and Vl', and the secondary attention features Uq' and Vq' of the time domain channel and the frequency domain channel to obtain self-attention features.

[0177] Self-attention is calculated on the speech feature vector to obtain linear attention features. This linear attention feature is then fused with secondary attention features from the time and frequency domains to obtain self-attention features. This linear attention calculation strengthens global attention and further improves the accuracy of subsequent self-attention calculations.

[0178] In an optional embodiment of this specification, step 108 includes the following specific steps:

[0179] Based on the target speech feature vector, an amplitude mask is predicted, and based on the target speech feature vector, a spectrum parameter is predicted;

[0180] Generate target spectrum graph according to amplitude mask and spectrum parameters;

[0181] Based on the target spectrogram, the speech processing result is determined.

[0182] In order to achieve better noise reduction effect, decoding can be performed in two different ways: one is amplitude mask product, and the other is spectrum estimation. This can take both methods into account and make up for the shortcomings of a single decoder.

[0183] The amplitude mask is a mask for masking a specific noise amplitude. The spectrum parameter is a spectrum parameter of the speech data on the frequency graph, including real part parameters and imaginary part parameters.

[0184] Based on the target speech feature vector, an amplitude mask is predicted, and based on the target speech feature vector, spectrum parameters are predicted. Based on the amplitude mask and spectrum parameters, a target spectrogram is generated. Based on the target spectrogram, the speech processing result is determined. The specific method is as follows: using a decoding module, based on the target speech feature vector, an amplitude mask is predicted, and based on the target speech feature vector, spectrum parameters are predicted. Based on the amplitude mask and spectrum parameters, a target spectrogram is generated. Based on the target spectrogram, the speech processing result is determined. Among them, the decoding module is a dual decoder module. The first decoder predicts the amplitude mask, and the second decoder predicts the spectrum parameters of the real part and the imaginary part. Based on the amplitude mask and spectrum parameters, a target spectrogram is generated. The specific calculation formula is shown in Formula 12:

[0185] in, is the target spectrum, α and β are the preset weights, is the amplitude mask, is the real part parameter, is the imaginary part parameter.

[0186] Exemplarily, based on the target speech feature vector X, the amplitude mask is predicted And based on the target speech feature vector X, predict the spectrum parameters Generate target spectrum according to amplitude mask and spectrum parameters Based on the target spectrogram, determine the enhanced speech data of user A

[0187] Based on the target speech feature vector, the amplitude mask is predicted, and the spectral parameters are also predicted based on the target speech feature vector. A target spectrogram is generated based on the amplitude mask and spectral parameters. Based on the target spectrogram, the speech processing result is determined. This improves the accuracy of speech processing.

[0188] In an optional embodiment of this specification, step 104 includes the following specific steps:

[0189] Inputting speech data into an encoding module of a speech model and encoding the speech data to obtain a speech feature vector, wherein the speech model includes an encoding module, a processing module, and a decoding module;

[0190] Correspondingly, step 106 includes the following specific steps:

[0191] The speech feature vector is input into the processing module, and based on the speech feature vector, self-attention calculation is performed on the speech feature vector in the time domain channel and the frequency domain channel to obtain self-attention features. Based on the self-attention features, the speech feature vector is weighted to obtain the target speech feature vector;

[0192] Correspondingly, step 108 includes the following specific steps:

[0193] The target speech feature vector is input into the decoding module and decoded to obtain the speech processing result.

[0194] The speech model is a deep learning network model with speech enhancement processing capabilities, including an encoding module, a processing module, and a decoding module. The processing module includes a loop module and an attention module. The specific speech model structure is shown in Figure 2. Figure 2 shows a schematic diagram of the structure of a speech model in a speech processing method provided in one embodiment of this specification, as shown in Figure 2:

[0195] The speech data undergoes short-time Fourier transform, enters the time-frequency domain encoding module, undergoes N cycles of processing in the synchronous attention module, and enters the decoding module composed of dual decoders. The decoding module includes a time-frequency domain amplitude mask decoding module and a time-frequency domain spectrum parameter module. After feature fusion, it undergoes inverse short-time Fourier transform to obtain the speech processing result.

[0196] The structure of the processing module is shown in Figure 3. Figure 3 shows a schematic diagram of the structure of the processing module in a speech processing method provided by an embodiment of this specification, as shown in Figure 3:

[0197] The speech feature vectors are input into the first branch (convolution block and feedforward sequential memory network) and the second branch (convolution block) of the cyclic module respectively. After feature fusion, they are processed through the feature dimension recovery network and cyclic filtering to obtain the updated speech feature vectors. The speech feature vectors are input into the three convolution blocks respectively to obtain feature Z, feature U and feature V respectively. Feature U and feature V are input into the transposition network to obtain U. T and V T , based on features Z, U, V, U T and V T , using formula 5, the self-attention features U' and V' are calculated. Based on the self-attention features, the speech feature vector is updated and, after feature fusion, input into the convolution block. Combined with the updated speech feature vector, the target speech feature vector is determined.

[0198] The structure of the attention module is shown in FIG4. FIG4 shows a schematic diagram of the structure of the attention module in a speech processing method provided by an embodiment of this specification, as shown in FIG4:

[0199] The speech feature vector is input into the query feature and key feature conversion networks to obtain secondary query features, secondary key features, linear query features, and linear key features, respectively. In the secondary self-attention branch, the secondary query features and secondary key features are transposed and passed through the network, then passed through the mathematical product network. They are then input into the secondary attention calculation network, along with the local feature vector of the updated speech feature vector transpose and the updated speech feature vector. After passing through the mathematical product network, the secondary self-attention features are determined. In the linear self-attention branch, the linear query features and the local feature vector of the updated speech feature vector are passed through the mathematical product network. This is then combined with the linear key features and the mathematical product network to determine the linear self-attention features.

[0200] For the specific methods of the embodiments of this specification, please refer to the specific methods of the embodiments of the above specification, which will not be repeated here.

[0201] In an optional embodiment of the present specification, the speech model further includes: a loop module and a local convolution module;

[0202] Before the speech feature vector is input into the processing module and attention calculation is performed on the speech feature vector in the time domain channel and the frequency domain channel to obtain the target speech feature vector, the following specific steps are also included:

[0203] The speech feature vector is input into the cyclic module, and the feature sequences of the time domain channel and the frequency domain channel are filtered respectively to obtain an updated speech feature vector;

[0204] Inputting the speech feature vector into the local convolution module, performing convolution processing on the local features in the speech feature vector, and obtaining a speech feature vector with updated local features;

[0205] Correspondingly, the target speech feature vector is input into the decoding module, and the decoding module obtains the speech processing result, which includes the following specific steps:

[0206] Inputting the target speech feature vector into the amplitude mask decoder of the decoding module, and predicting the amplitude mask based on the target speech feature vector, and inputting the target speech feature vector into the spectrum parameter decoder of the decoding module, and predicting the spectrum parameters based on the target speech feature vector;

[0207] Perform weighted processing on the amplitude mask and spectrum parameters to generate the target spectrum graph;

[0208] The target spectrogram is converted to obtain the speech processing result.

[0209] For the specific methods of the embodiments of this specification, please refer to the specific methods of the embodiments of the above specification, which will not be repeated here.

[0210] Referring to FIG5 , FIG5 shows a flowchart of a conference speech enhancement method provided by an embodiment of this specification. The method is applied to a cloud-side device and includes the following specific steps:

[0211] Step 502: Acquire conference voice data.

[0212] Step 504: Encode the conference voice data to obtain a voice feature vector, wherein the voice feature vector includes feature sequences of the time domain channel and the frequency domain channel.

[0213] Step 506: Perform time-frequency domain synchronization feature processing on the feature sequences of the time domain channel and the frequency domain channel to obtain a target speech feature vector.

[0214] Step 508: Based on the target speech feature vector, decode to obtain enhanced conference speech data.

[0215] Step 510: Send the enhanced conference voice data to the front end.

[0216] The embodiments of this specification apply to a network cloud device, which is a virtual device, and is located on the server side of a voice processing application, website, or mini-program with conference voice enhancement capabilities. The front-end is a physical device, and is located on the client side of a webpage, application, or mini-program with conference voice enhancement capabilities that users log into. The cloud-side device and the front-end are connected via a network transmission channel for data transmission. The computing power and storage performance of the cloud-side device are superior to those of the front-end.

[0217] The embodiment of this specification and the embodiment of the specification of FIG. 1 are based on the same inventive concept. For the specific methods of steps 502 to 508 , refer to the specific methods of steps 102 to 108 above, which will not be repeated here.

[0218] In an embodiment of the present specification, conference voice data is obtained; the conference voice data is encoded to obtain a voice feature vector, wherein the voice feature vector includes a feature sequence of a time domain channel and a frequency domain channel; the feature sequence of the time domain channel and the frequency domain channel is subjected to time-frequency synchronous feature processing to obtain a target voice feature vector; based on the target voice feature vector, enhanced conference voice data is decoded; and the enhanced conference voice data is sent to the front end. When the conference voice data is feature-encoded to obtain a voice feature vector including a feature sequence of a time domain channel and a frequency domain channel, the voice feature vector is feature-processed synchronously on the time domain channel and the frequency domain channel, fully exploiting the close correlation between the features on the time domain channel and the frequency domain channel, achieving effective contextual understanding of the conference voice data, capturing the complex interaction in the time domain and the frequency domain, obtaining a highly accurate target voice feature vector for decoding, obtaining highly accurate voice-enhanced conference voice data, and feeding it back to the front end, thereby improving the effectiveness of conference voice enhancement and enhancing the user experience. At the same time, conference voice enhancement is completed using cloud-side devices with high computing performance and high storage performance, further improving the efficiency and accuracy of conference voice enhancement.

[0219] Referring to FIG. 6 , FIG. 6 shows a flowchart of a speech model training method provided by one embodiment of this specification. The method is applied to a cloud-side device and includes the following specific steps:

[0220] Step 602: Acquire a sample set, where the sample set includes sample speech data and labeled speech data.

[0221] Step 604: Input the sample speech data into the encoding module of the speech model, encode the sample speech data to obtain a sample speech feature vector, wherein the speech model includes an encoding module, a processing module and a decoding module.

[0222] Step 606: Input the sample speech feature vector into the processing module, perform time-frequency domain synchronization feature processing on the feature sequences of the time domain channel and the frequency domain channel, and obtain a predicted speech feature vector.

[0223] Step 608: Input the predicted speech feature vector into a decoding module and decode it to obtain predicted speech data.

[0224] Step 610: Calculate a loss value based on the predicted speech data and the labeled speech data.

[0225] Step 612: Based on the loss value, adjust the model parameters of the speech model, and obtain a trained speech model when the preset training end condition is met.

[0226] Step 614: Send the model parameters of the speech model to the end-side device.

[0227] The embodiments of this specification apply to network cloud devices with speech model training capabilities, which are virtual devices. The end-side device is a physical device, representing the client terminal for a webpage, application, or mini-program with speech model training capabilities that users log into. The cloud-side device and the end-side device are connected via a network transmission channel for data transmission. The cloud-side device has higher computing power and storage performance than the end-side device.

[0228] The embodiment of this specification and the embodiment of the specification of FIG. 1 are based on the same inventive concept. For the specific methods of steps 602 to 608 , refer to the specific methods of steps 102 to 108 above, which will not be repeated here.

[0229] The loss value measures the difference between the predicted speech data and the labeled speech data and is used to evaluate the model performance of the speech model, including but not limited to: cross entropy loss, mean square error loss, average error loss, and logarithmic loss.

[0230] The training stop condition is a pre-set judgment condition for stopping training, including but not limited to: a preset number of iterations, a preset loss value threshold, a preset training time and a preset model convergence condition.

[0231] In an embodiment of the present specification, a sample set is obtained, wherein the sample set includes sample speech data and labeled speech data; the sample speech data is input into the encoding module of the speech model, and the sample speech data is encoded to obtain a sample speech feature vector, wherein the speech model includes an encoding module, a processing module and a decoding module; the sample speech feature vector is input into the processing module, and time-frequency domain synchronization feature processing is performed on the feature sequences of the time domain channel and the frequency domain channel to obtain a predicted speech feature vector; the predicted speech feature vector is input into the decoding module, and the predicted speech data is obtained by decoding; a loss value is calculated based on the predicted speech data and the labeled speech data; based on the loss value, the model parameters of the speech model are adjusted, and when the preset training end conditions are met, a trained speech model is obtained; the model parameters of the speech model are sent to the end-side device. When feature encoding is performed on the sample speech data to obtain a sample speech feature vector including a feature sequence of the time domain channel and the frequency domain channel, feature processing is performed on the sample speech feature vector simultaneously on the time domain channel and the frequency domain channel, fully exploiting the close correlation between the features on the time domain channel and the frequency domain channel, achieving effective contextual understanding of the sample speech data, capturing the complex interaction in the time domain and frequency domain, and obtaining a highly accurate predicted speech feature vector for decoding, obtaining highly accurate predicted speech data, completing highly accurate supervised training, improving the model performance of the trained speech model, and providing a model foundation for subsequent speech processing. At the same time, the conference speech model training is completed using cloud-side devices with high computing performance and high storage performance, further improving the efficiency and accuracy of speech model training.

[0232] The following further illustrates the speech processing method provided in this specification using the application of the speech processing method in a conference speech scenario as an example, in conjunction with FIG7 . FIG7 shows a flowchart of a speech processing method applied to a conference speech scenario provided by one embodiment of this specification, including the following specific steps:

[0233] Step 702: Acquire conference voice data.

[0234] Step 704: Convert the conference voice data to obtain a spectrogram corresponding to the conference voice data, extract features from the spectrogram on the time domain channel and frequency domain channel corresponding to the spectrogram, and determine a voice feature vector based on the extracted features.

[0235] Step 706: For any feature dimension of the speech feature vector, perform cyclic filtering on the feature sequences of the time domain channel and the frequency domain channel within any feature dimension, and obtain an updated speech feature vector based on the cyclic filtering results corresponding to each feature dimension.

[0236] Step 708: Process the local features in the speech feature vector to obtain a speech feature vector with updated local features.

[0237] Step 710: Transpose the speech feature vector to obtain a transposed speech feature vector. Perform self-attention calculations on the speech feature vector and the transposed speech feature vector to obtain secondary attention features for the time domain channel and the frequency domain channel. Determine a self-attention feature based on the secondary attention features for the time domain channel and the frequency domain channel. Perform self-attention calculations on the speech feature vector to obtain a linear attention feature.

[0238] Step 712: Perform feature fusion on the linear attention features, the secondary attention features of the time domain channel and the frequency domain channel to obtain the self-attention features.

[0239] Step 714: Based on the target speech feature vector, predict the amplitude mask and spectrum parameters, generate a target spectrogram according to the amplitude mask and spectrum parameters, and determine the enhanced conference speech data based on the target spectrogram.

[0240] In the embodiments of this specification, by simultaneously operating the quadratic self-attention and linear self-attention on the time domain channel and the frequency domain channel, the self-attention features are determined, and the close correlation between the features on the time domain channel and the frequency domain channel is fully exploited, and the close correlation between the features on the time domain channel and the frequency domain channel is fully exploited, thereby achieving effective contextual understanding of conference voice data, capturing the complex interactions in the time domain and frequency domain, and obtaining a highly accurate target voice feature vector for decoding, thereby obtaining highly accurate voice-enhanced conference voice data, improving the effectiveness of conference voice enhancement, and improving user experience.

[0241] Corresponding to the above method embodiment, this specification also provides an embodiment of a speech processing device. FIG8 shows a schematic diagram of the structure of a speech processing device provided by one embodiment of this specification. As shown in FIG8 , the device includes:

[0242] A first acquisition module 802 is configured to acquire voice data to be processed;

[0243] A first encoding module 804 is configured to encode the speech data to obtain a speech feature vector, wherein the speech feature vector includes a feature sequence of a time domain channel and a frequency domain channel;

[0244] The first processing module 806 is configured to perform time-frequency domain synchronization feature processing on the feature sequences of the time domain channel and the frequency domain channel to obtain a target speech feature vector;

[0245] The first decoding module 808 is configured to decode and obtain a speech processing result based on the target speech feature vector.

[0246] Optionally, the device further comprises:

[0247] a spectrum conversion module, configured to convert the voice data to obtain a spectrum graph corresponding to the voice data;

[0248] Correspondingly, the first encoding module 804 is further configured to:

[0249] Feature extraction is performed on the spectrogram on the time domain channel and frequency domain channel corresponding to the spectrogram to obtain the speech feature vector.

[0250] Optionally, the device further comprises:

[0251] The loop filtering module is configured to filter the feature sequences of the time domain channel and the frequency domain channel respectively to obtain an updated speech feature vector.

[0252] Optionally, the speech feature vector includes multiple feature dimensions;

[0253] Correspondingly, the loop filtering module is further configured as follows:

[0254] For any feature dimension of the speech feature vector, the feature sequences of the time domain channel and the frequency domain channel within any feature dimension are subjected to cyclic filtering respectively; based on the cyclic filtering processing results corresponding to each feature dimension, an updated speech feature vector is obtained.

[0255] Optionally, the first processing module 806 is further configured to:

[0256] Based on the speech feature vector, self-attention calculation is performed on the speech feature vector in the time domain channel and the frequency domain channel to obtain the self-attention feature; based on the self-attention feature, the speech feature vector is weighted to obtain the target speech feature vector.

[0257] Optionally, the device further comprises:

[0258] The local processing module is configured to perform convolution processing on the local features in the speech feature vector to obtain a speech feature vector with updated local features.

[0259] Optionally, the first processing module 806 is further configured to:

[0260] The speech feature vector is transposed to obtain a transposed speech feature vector; self-attention calculation is performed on the speech feature vector and the transposed speech feature vector respectively to obtain secondary attention features of the time domain channel and the frequency domain channel; based on the secondary attention features of the time domain channel and the frequency domain channel, the self-attention feature is determined.

[0261] Optionally, the device further comprises:

[0262] A linear processing module is configured to perform self-attention calculation on the speech feature vector to obtain a linear attention feature;

[0263] Correspondingly, the first processing module 806 is further configured to:

[0264] The linear attention features, the secondary attention features of the time domain channel and the frequency domain channel are fused to obtain the self-attention features.

[0265] Optionally, the first decoding module 808 is further configured to:

[0266] Based on the target speech feature vector, an amplitude mask is predicted, and based on the target speech feature vector, spectrum parameters are predicted; a target spectrogram is generated according to the amplitude mask and the spectrum parameters; and based on the target spectrogram, a speech processing result is determined.

[0267] Optionally, the first encoding module 804 is further configured to:

[0268] Inputting speech data into an encoding module of a speech model and encoding the speech data to obtain a speech feature vector, wherein the speech model includes an encoding module, a processing module, and a decoding module;

[0269] Correspondingly, the first processing module 806 is further configured to:

[0270] The speech feature vector is input into the processing module, and based on the speech feature vector, self-attention calculation is performed on the speech feature vector in the time domain channel and the frequency domain channel to obtain self-attention features. Based on the self-attention features, the speech feature vector is weighted to obtain the target speech feature vector;

[0271] Correspondingly, the first decoding module 808 is further configured to:

[0272] The target speech feature vector is input into the decoding module and decoded to obtain the speech processing result.

[0273] Optionally, the device further comprises:

[0274] The loop module is configured to input the speech feature vector into the loop module, filter the feature sequences of the time domain channel and the frequency domain channel respectively, and obtain an updated speech feature vector; input the speech feature vector into the local convolution module, convolve the local features in the speech feature vector, and obtain a speech feature vector with updated local features;

[0275] Correspondingly, the first decoding module 808 is further configured to:

[0276] The target speech feature vector is input into the amplitude mask decoder of the decoding module, and the amplitude mask is predicted based on the target speech feature vector. The target speech feature vector is also input into the spectrum parameter decoder of the decoding module, and the spectrum parameters are predicted based on the target speech feature vector. The amplitude mask and the spectrum parameters are weighted to generate a target spectrogram. The target spectrogram is converted to obtain a speech processing result.

[0277] In the embodiments of this specification, when feature encoding is performed on speech data to obtain a speech feature vector including a feature sequence of a time domain channel and a frequency domain channel, feature processing is performed on the speech feature vector simultaneously on the time domain channel and the frequency domain channel, fully exploiting the close relationship between the features on the time domain channel and the frequency domain channel, achieving effective contextual understanding of the speech data, capturing the complex interactions on the time domain and the frequency domain, obtaining a highly accurate target speech feature vector for decoding, obtaining a highly accurate speech processing result, and improving the effectiveness of speech processing.

[0278] The above is a schematic diagram of a speech processing device according to this embodiment. It should be noted that the technical solution of the speech processing device and the technical solution of the speech processing method described above are based on the same concept. For details not described in detail in the technical solution of the speech processing device, please refer to the description of the technical solution of the speech processing method described above.

[0279] Corresponding to the above method embodiment, this specification also provides an embodiment of a conference voice enhancement device. FIG9 shows a schematic diagram of the structure of a conference voice enhancement device provided by one embodiment of this specification. As shown in FIG9, the device includes:

[0280] The second acquisition module 902 is configured to acquire conference voice data;

[0281] The second encoding module 904 is configured to encode the conference voice data to obtain a voice feature vector, wherein the voice feature vector includes a feature sequence of a time domain channel and a frequency domain channel;

[0282] The second processing module 906 is configured to perform attention calculation on the speech feature vector in the time domain channel and the frequency domain channel based on the speech feature vector to obtain a target speech feature vector;

[0283] A second decoding module 908 is configured to decode the enhanced conference voice data based on the target voice feature vector;

[0284] The data sending module 910 is configured to send the enhanced conference voice data to the front end.

[0285] In the embodiments of this specification, when feature encoding is performed on conference voice data to obtain a voice feature vector including a feature sequence of a time domain channel and a frequency domain channel, feature processing is performed on the voice feature vector simultaneously on the time domain channel and the frequency domain channel, and the close relationship between the features on the time domain channel and the frequency domain channel is fully exploited, thereby achieving effective contextual understanding of the conference voice data, capturing the complex interactions on the time domain and the frequency domain, and obtaining a highly accurate target voice feature vector for decoding, thereby obtaining highly accurate voice-enhanced conference voice data, and feeding it back to the front end, thereby improving the effectiveness of conference voice enhancement and enhancing the user experience. At the same time, conference voice enhancement is completed using cloud-side devices with high computing performance and high storage performance, further improving the efficiency and accuracy of conference voice enhancement.

[0286] The above is a schematic diagram of a conference voice enhancement device according to this embodiment. It should be noted that the technical solution of the conference voice enhancement device and the technical solution of the conference voice enhancement method described above are based on the same concept. For details not described in detail in the technical solution of the conference voice enhancement device, please refer to the description of the technical solution of the conference voice enhancement method described above.

[0287] Corresponding to the above method embodiment, this specification also provides an embodiment of a speech model training device. FIG10 shows a schematic diagram of the structure of a speech model training device provided by one embodiment of this specification. As shown in FIG10 , the device includes:

[0288] The third acquisition module 1002 is configured to acquire a sample set, wherein the sample set includes sample speech data and labeled speech data;

[0289] A third encoding module 1004 is configured to input the sample speech data into an encoding module of a speech model, encode the sample speech data to obtain a sample speech feature vector, wherein the speech model includes an encoding module, a processing module, and a decoding module;

[0290] The third processing module 1006 is configured to input the sample speech feature vector into the processing module, perform attention calculation on the sample speech feature vector in the time domain channel and the frequency domain channel, and obtain a predicted speech feature vector;

[0291] The third decoding module 1008 is configured to input the predicted speech feature vector into the decoding module and decode it to obtain predicted speech data;

[0292] The loss calculation module 1010 is configured to calculate a loss value based on the predicted speech data and the labeled speech data;

[0293] The training adjustment module 1012 is configured to adjust the model parameters of the speech model based on the loss value, and obtain a trained speech model when a preset training end condition is met;

[0294] The model sending module 1014 is configured to send the model parameters of the speech model to the terminal-side device.

[0295] In the embodiments of this specification, when feature encoding is performed on sample speech data to obtain a sample speech feature vector including a feature sequence of a time domain channel and a frequency domain channel, feature processing is performed on the sample speech feature vector simultaneously on the time domain channel and the frequency domain channel, fully exploiting the close correlation between the features on the time domain channel and the frequency domain channel, achieving effective contextual understanding of the sample speech data, capturing the complex interactions on the time domain and frequency domains, obtaining a highly accurate predicted speech feature vector for decoding, obtaining highly accurate predicted speech data, completing highly accurate supervised training, improving the model performance of the trained speech model, and providing a model basis for subsequent speech processing. At the same time, cloud-side devices with high computing performance and high storage performance are used to complete conference speech model training, further improving the efficiency and accuracy of speech model training.

[0296] The above is a schematic diagram of a speech model training device according to this embodiment. It should be noted that the technical solution of the speech model training device and the technical solution of the speech model training method described above are based on the same concept. For details not described in detail in the technical solution of the speech model training device, please refer to the description of the technical solution of the speech model training method described above.

[0297] Figure 11 shows a block diagram of a computing device according to one embodiment of this specification. Components of the computing device 1100 include, but are not limited to, a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.

[0298] The computing device 1100 also includes an access device 1140 that enables the computing device 1100 to communicate via one or more networks 1160. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1140 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.

[0299] In one embodiment of the present specification, the aforementioned components of the computing device 1100 and other components not shown in FIG11 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG11 is for illustrative purposes only and does not limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0300] Computing device 1100 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1100 may also be a mobile or stationary server.

[0301] Among them, the processor 1120 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned speech processing method, conference speech enhancement method or speech model training method.

[0302] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of the computing device is based on the same concept as the technical solutions of the aforementioned speech processing method, conference speech enhancement method, and speech model training method. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the aforementioned speech processing method, conference speech enhancement method, or speech model training method.

[0303] An embodiment of the present specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned speech processing method, conference speech enhancement method, or speech model training method.

[0304] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium is based on the same concept as the technical schemes of the aforementioned speech processing method, conference speech enhancement method, and speech model training method. For details not described in detail in the technical scheme of the storage medium, please refer to the description of the technical schemes of the aforementioned speech processing method, conference speech enhancement method, or speech model training method.

[0305] An embodiment of the present specification also provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned speech processing method, conference speech enhancement method or speech model training method.

[0306] The above is a schematic scheme of a computer program of this embodiment. It should be noted that the technical scheme of this computer program is based on the same concept as the technical schemes of the aforementioned speech processing method, conference speech enhancement method, and speech model training method. For details not described in detail in the technical scheme of the computer program, please refer to the description of the technical schemes of the aforementioned speech processing method, conference speech enhancement method, or speech model training method.

[0307] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or sequential order to achieve the desired results. In certain embodiments, multi-language processing and parallel processing are also possible or may be advantageous.

[0308] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0309] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0310] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

Claims

1. A speech processing method, comprising: Obtaining voice data to be processed; Encoding the speech data to obtain a speech feature vector, wherein the speech feature vector includes a feature sequence of a time domain channel and a frequency domain channel; Performing time-domain and frequency-domain synchronous feature processing on the feature sequences of the time-domain channel and the frequency-domain channel to obtain a target speech feature vector; Based on the target speech feature vector, a speech processing result is obtained by decoding.

2. The method according to claim 1, before encoding the speech data to obtain the speech feature vector, further comprising: Converting the speech data to obtain a spectrogram corresponding to the speech data; The step of encoding the speech data to obtain a speech feature vector comprises: Feature extraction is performed on the spectrogram on the time domain channel and the frequency domain channel corresponding to the spectrogram to obtain the speech feature vector.

3. The method according to claim 1, before performing time-frequency domain synchronization feature processing on the feature sequences of the time domain channel and the frequency domain channel to obtain the target speech feature vector, further comprises: The feature sequences of the time domain channel and the frequency domain channel are filtered respectively to obtain updated speech feature vectors.

4. The method according to claim 3, wherein the speech feature vector comprises a plurality of feature dimensions; The filtering process is performed on the feature sequences of the time domain channel and the frequency domain channel respectively to obtain an updated speech feature vector, including: For any feature dimension of the speech feature vector, performing loop filtering processing on the feature sequences of the time domain channel and the frequency domain channel in the any feature dimension respectively; Based on the loop filtering processing results corresponding to each feature dimension, an updated speech feature vector is obtained.

5. According to the method according to any one of claims 1 to 4, the step of performing time-frequency domain synchronous feature processing on the feature sequences of the time domain channel and the frequency domain channel to obtain the target speech feature vector comprises: Based on the speech feature vector, performing self-attention calculation on the speech feature vector in the time domain channel and the frequency domain channel to obtain a self-attention feature; Based on the self-attention feature, the speech feature vector is weighted to obtain a target speech feature vector.

6. The method according to claim 5, before performing self-attention calculation on the speech feature vector in the time domain channel and the frequency domain channel based on the speech feature vector to obtain the self-attention feature, further comprises: The local features in the speech feature vector are subjected to convolution processing to obtain the speech feature vector with updated local features.

7. The method according to claim 5, wherein based on the speech feature vector, performing self-attention calculation on the speech feature vector in the time domain channel and the frequency domain channel to obtain the self-attention feature comprises: Transposing the speech feature vector to obtain a transposed speech feature vector; The speech feature vector and the transposed speech feature vector are respectively subjected to self-attention calculation to obtain the Secondary attention features of the time domain channel and the frequency domain channel; Based on the secondary attention features of the time domain channel and the frequency domain channel, a self-attention feature is determined.

8. The method according to claim 7, wherein determining the self-attention feature based on the secondary attention feature of the time domain channel and the frequency domain channel comprises: Performing self-attention calculation on the speech feature vector to obtain a linear attention feature; The linear attention feature, the secondary attention features of the time domain channel and the frequency domain channel are fused to obtain a self-attention feature.

9. The method according to any one of claims 1 to 8, wherein decoding based on the target speech feature vector to obtain a speech processing result comprises: Based on the target speech feature vector, an amplitude mask is predicted, and based on the target speech feature vector, a spectrum parameter is predicted; Generate a target spectrum diagram according to the amplitude mask and the spectrum parameters; Based on the target spectrogram, a speech processing result is determined.

10. The method according to any one of claims 1 to 8, wherein encoding the speech data to obtain a speech feature vector comprises: Inputting the speech data into a coding module of a speech model, encoding the speech data to obtain a speech feature vector, wherein the speech model includes the coding module, a processing module and a decoding module; The step of performing time-frequency domain synchronization feature processing on the speech feature vector based on the feature sequence of the time domain channel and the frequency domain channel to obtain a target speech feature vector includes: Inputting the speech feature vector into the processing module, performing self-attention calculation on the speech feature vector in the time domain channel and the frequency domain channel based on the speech feature vector to obtain a self-attention feature, and performing weighted processing on the speech feature vector based on the self-attention feature to obtain a target speech feature vector; The decoding to obtain a speech processing result based on the target speech feature vector includes: The target speech feature vector is input into the decoding module and decoded to obtain a speech processing result.

11. The method according to claim 10, wherein the speech model further comprises: Recurrent modules and local convolution modules; Before inputting the speech feature vector into the processing module and performing attention calculation on the speech feature vector in the time domain channel and the frequency domain channel to obtain the target speech feature vector, the method further includes: Inputting the speech feature vector into the circulation module, filtering the feature sequences of the time domain channel and the frequency domain channel respectively, to obtain an updated speech feature vector; Inputting the speech feature vector into the local convolution module, performing convolution processing on the local features in the speech feature vector, and obtaining the speech feature vector with updated local features; The step of inputting the target speech feature vector into the decoding module and decoding to obtain a speech processing result includes: Inputting the target speech feature vector into the amplitude mask decoder of the decoding module, predicting the amplitude mask based on the target speech feature vector, and inputting the target speech feature vector into the spectrum parameter decoder of the decoding module, predicting the spectrum parameter based on the target speech feature vector; Performing weighted processing on the amplitude mask and the spectrum parameters to generate a target spectrum graph; The target spectrogram is converted to obtain a speech processing result.

12. A conference voice enhancement method, applied to a cloud-side device, comprising: Get conference voice data; Encoding the conference voice data to obtain a voice feature vector, wherein the voice feature vector includes a feature sequence of a time domain channel and a frequency domain channel; Performing time-domain and frequency-domain synchronous feature processing on the feature sequences of the time-domain channel and the frequency-domain channel to obtain a target speech feature vector; Based on the target speech feature vector, decoding is performed to obtain enhanced conference speech data; The enhanced conference voice data is sent to the front end.

13. A speech model training method, applied to a cloud-side device, comprising: Acquire a sample set, wherein the sample set includes sample speech data and labeled speech data; Inputting the sample speech data into a coding module of a speech model, encoding the sample speech data to obtain a sample speech feature vector, wherein the speech model includes the coding module, the processing module and the decoding module; Inputting the sample speech feature vector into the processing module, performing time-frequency domain synchronization feature processing on the feature sequences of the time domain channel and the frequency domain channel in the sample speech feature vector, and obtaining a predicted speech feature vector; Inputting the predicted speech feature vector into the decoding module, and decoding to obtain predicted speech data; Calculating a loss value based on the predicted speech data and the labeled speech data; Based on the loss value, adjusting the model parameters of the speech model, and obtaining a trained speech model when a preset training end condition is met; The model parameters of the speech model are sent to the terminal side device.

14. A computing device comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method described in any one of claims 1 to 13 are implemented.

15. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 13.

16. A computer program, wherein: When the computer program is executed in a computer, the computer is caused to execute the steps of implementing the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Audio noise reduction and audio noise reduction model processing method, device, equipment and medium

    CN113763979A

  • Voice noise reduction method and device, equipment and storage medium

    CN114067826A

  • Speech enhancement model, electronic device, storage medium and related method

    CN114333895A

  • Single-channel speech enhancement method based on interactive time-frequency attention mechanism

    CN115295002A

  • Speech processing method, conference speech enhancement method and speech model training method

    CN117524241A

Cited By

  • Speech recognition method and system based on artificial intelligence

    CN120564724A

  • Lightweight causal audio and video voice separation method and device based on multi-branch SRU

    CN121122309A