Voice processing method and device, computer equipment and storage medium

By building a comprehensive model, combining text encoder and speech encoder, the integration of speech recognition and speech synthesis tasks is achieved, solving the problems of complex model training and high computing resources in the existing technology, and achieving efficient speech processing and knowledge transfer.

CN120089142APending Publication Date: 2025-06-03PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510272344.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The use of independent ASR and TTS models in financial business scenarios has resulted in complex model training, high computing resources and time consumption, and it is difficult to achieve efficient knowledge transfer between speech recognition and speech synthesis tasks.

Method used

A speech processing method is proposed, and by building a comprehensive model including text encoder, speech encoder and corresponding decoder, it realizes a high degree of integration of speech recognition and text-to-speech synthesis tasks. This method adopts a pre-trained speech recognition and synthesis model to extract feature of text and speech data to be processed, and fuses features of different modes through task vectors to reduce the redundancy of model parameters and reduce the consumption of computing resources and time.

Benefits of technology

The fusion of cross-modal features is realized, the model's understanding of the semantics and speech characteristics of the input data is improved, the accuracy and generalization capabilities of speech recognition and speech synthesis tasks are improved, and the computing resources and time consumption of model training is significantly reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089142A_ABST
    Figure CN120089142A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to a voice processing method and device, computer equipment and a storage medium, and the method comprises the steps: employing a text encoder to pre-network extract a target text feature vector of text data through a voice recognition and synthesis model; extracting a target voice feature vector of the voice data through a voice encoder pre-network; combining the target text feature vector, the target voice feature vector and the task vector to obtain a target task fusion vector; processing the target task fusion vector by adopting a text decoder post-network, and outputting target synthetic text data; and processing the target task fusion vector by adopting a voice decoder post-network, and outputting target synthetic voice data. In addition, the invention further relates to a block chain technology, and data such as text data and voice data can be stored in a block chain. According to the invention, the computing resource and time consumption required by the voice processing model can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a voice processing method, apparatus, computer device, and storage medium. Background Art

[0002] With the rapid development of financial business and the in-depth digital transformation, intelligent voice customer service, as an advanced form of customer service, relies on the core technologies behind it - automatic speech recognition (ASR) and text-to-speech synthesis (TTS), which play a crucial role in financial business scenarios. Traditionally, the automatic speech recognition (ASR) and text-to-speech synthesis (TTS) tasks are achieved by training independent ASR models and TTS models separately, mainly due to the significant differences in input-output modalities and learning objectives between these tasks. In previous practices, the ASR model and the TTS model were usually trained separately, each with its own learning objective, training data, and model parameters, and finally constructed into two independent large networks.

[0003] In financial business scenarios, intelligent voice customer service needs to handle a large number of customer inquiries and transaction requests. With the increase in business volume and the diversification of customer needs, for example, when dealing with complex financial problems, the ASR model may not accurately recognize the customer's voice input, resulting in misunderstandings or incorrect processing; while the TTS model may also affect the customer experience due to lack of sufficient naturalness and fluency when synthesizing speech. To address this, recent research has focused on developing unified audio-text models to achieve smooth fusion between text and audio. These models can be classified into two categories: unimodal and cross-modal according to the output modality. Unimodal models such as Whisper and Google USM can only produce text outputs; while cross-modal models, such as Viola, SpeechT5, and SpeechGPT, can generate outputs in both speech and text modalities. However, although cross-modal models such as SpeechT5 are pre-trained with cross-modal objectives and can handle both text and speech as input and output modalities, they still need to be fine-tuned separately when applied to downstream tasks (such as ASR and TTS). This not only increases the complexity of the training process but also raises the consumption of computing resources and time.

[0004] Therefore, there is an urgent need to develop new voice processing technologies to reduce the computing resources and time consumption required for model training. Summary of the Invention

[0005] The purpose of the embodiments of this application is to propose a voice processing method, apparatus, computer device, and storage medium, which can reduce the computing resources and time consumption required for voice processing models.

[0006] To solve the above technical problems, an embodiment of the present application provides a voice processing method, which adopts the following technical solutions:

[0007] Obtain the text data to be processed and the voice data to be processed;

[0008] Input the text data to be processed and the voice data to be processed into a pre-trained voice recognition and synthesis model, where the voice recognition and synthesis model includes a text encoder pre-network, a voice encoder pre-network, a text decoder post-network, and a voice decoder post-network;

[0009] Use the text encoder pre-network to extract features from the text data to be processed to obtain a target text feature vector;

[0010] Use the voice encoder pre-network to extract features from the voice data to be processed to obtain a target voice feature vector;

[0011] Obtain a predefined task vector, and combine the target text feature vector, the target voice feature vector, and the task vector to obtain a target task fusion vector;

[0012] Use the text decoder post-network to process the target task fusion vector and output target synthesized text data;

[0013] Use the voice decoder post-network to process the target task fusion vector and output target synthesized voice data.

[0014] To solve the above technical problems, an embodiment of the present application also provides a voice processing device, which adopts the following technical solutions:

[0015] A data acquisition module for obtaining the text data to be processed and the voice data to be processed;

[0016] A model input module for inputting the text data to be processed and the voice data to be processed into a pre-trained voice recognition and synthesis model, where the voice recognition and synthesis model includes a text encoder pre-network, a voice encoder pre-network, a text decoder post-network, and a voice decoder post-network;

[0017] A text extraction module for using the text encoder pre-network to extract features from the text data to be processed to obtain a target text feature vector;

[0018] A voice extraction module for using the voice encoder pre-network to extract features from the voice data to be processed to obtain a target voice feature vector;

[0019] A vector combination module, configured to obtain a predefined task vector, and combine the target text feature vector, the target speech feature vector, and the task vector to obtain a target task fusion vector;

[0020] A first synthesis module, configured to process the target task fusion vector by using the network after the text decoder, and output target synthesized text data;

[0021] A second synthesis module, configured to process the target task fusion vector by using the network after the speech decoder, and output target synthesized speech data.

[0022] To solve the above technical problem, an embodiment of the present application further provides a computer device, which adopts the following technical solution: The computer device includes a memory and a processor. A computer-readable instruction is stored in the memory, and when the processor executes the computer-readable instruction, the steps of the speech processing method described in any one of the above are implemented.

[0023] To solve the above technical problem, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable instruction is stored on the computer-readable storage medium, and when the computer-readable instruction is executed by a processor, the steps of the speech processing method described in any one of the above are implemented.

[0024] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:

[0025] The solution of the embodiment of the present application constructs a speech recognition and synthesis model. By using the text encoder pre-network and the speech encoder pre-network, feature extraction is respectively performed on the text to be processed and the speech data to be processed, generating high-quality feature vectors. By introducing a task vector, a target task fusion vector is obtained by combining the target text feature vector, the target speech feature vector, and the task vector, realizing the fusion of cross-modal features. This fusion mechanism helps the model better understand the semantic and speech characteristics of the input data, thereby more accurately completing the speech recognition and speech synthesis tasks. By using the text decoder post-network and the speech decoder post-network to process the target task fusion vector respectively, the target synthesized text data and the target synthesized speech data are output, enabling the simultaneous generation of two-modal outputs of synthesized text and synthesized speech, realizing the knowledge transfer between the speech recognition and speech synthesis tasks, and improving the generalization ability of the model. Based on the solution of the present application, by constructing an integrated model including a text encoder, a speech encoder, and corresponding decoders, a high degree of integration of the two tasks of speech recognition and text-to-speech synthesis is achieved. This not only simplifies the model architecture but also enables the model to share knowledge when processing different tasks by sharing the feature extraction and encoding processes, improving the data processing effect. By adopting the multi-task learning and parameter sharing strategies, while ensuring the model performance, the redundancy of the model parameters is reduced, significantly reducing the consumption of computing resources and time. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] To more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0027] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;

[0028] Figure 2 is a flowchart of an embodiment of the speech processing method of the present application;

[0029] Figure 3 is a structural diagram of an embodiment of the speech processing apparatus of the present application;

[0030] Figure 4 is a structural diagram of an embodiment of the computer device of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.

[0032] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the application. The phrase does not necessarily refer to the same embodiment at every occurrence in the specification, nor is it an independent or alternative embodiment mutually exclusive of other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0033] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0034] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0035] Users can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.

[0036] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop 1011, tablet computer 1012, or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, a desktop computer, and the like.

[0037] The server 103 can be a server that provides various services, such as a background server that supports the pages displayed on the terminal device 101.

[0038] It should be noted that the speech processing method provided by the embodiments of the present application is generally executed by the server / terminal device. Correspondingly, the speech processing device is generally arranged in the server / terminal device.

[0039] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in

[0040] Continue to refer to Figure 2 , which shows a flowchart of an embodiment of the question-and-answer method according to the present application. The speech processing method includes the following steps:

[0041] Step S201, obtain the text data to be processed and the speech data to be processed;

[0042] In this embodiment, the electronic device (such as Figure 1 the server / terminal device shown) on which the speech processing method runs can obtain the text data to be processed and the speech data to be processed through a wired connection or a wireless connection. It should be noted that the above wireless connection methods can include, but are not limited to, 2G / 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra wideband) connections, and other currently known or future-developed wireless connection methods.

[0043] Specifically, obtain the text data to be processed and the speech data to be processed. Herein, the text data to be processed refers to the text information that needs to be semantically processed or analyzed, which may include, but is not limited to, the current user input, the current web content, the current document file, etc.; the speech data to be processed refers to the audio information that needs to be speech processed or analyzed, which may include, but is not limited to, the current recorded audio, the current phone call, the current video file, etc.

[0044] In some alternative implementation manners, before the above-mentioned obtaining of the text data to be processed and the speech data to be processed, the following steps may further be included:

[0045] Obtain the pre-collected training samples, where the training samples include a number of training text data and training speech data;

[0046] Specifically, obtain the pre-collected training samples, which include a number of training text data and training speech data. Herein, the training text data refers to the historical text information used for semantic processing or analysis; the training speech data refers to the historical audio information used for speech processing or analysis.

[0047] Input the training samples into the pre-trained speech recognition and synthesis model, where the pre-trained speech recognition and synthesis model includes a text encoder pre-network, a speech encoder pre-network, a text decoder post-network, and a speech decoder post-network;

[0048] Specifically, load the pre-trained speech recognition and synthesis model, which is composed of a text encoder pre-network, a speech encoder pre-network, a text decoder post-network, and a speech decoder post-network. Among them, the text encoder pre-network is responsible for converting the input text data (such as characters, words, or tokens) into a fixed-dimensional word embedding form (i.e., a vector), and capturing the semantic information in the text data through the word embedding. The speech encoder pre-network is responsible for converting the input speech signal into a high-dimensional vector representation, and capturing the acoustic features and semantic information in the speech data through the vector. The text decoder post-network is responsible for converting the encoded text data back into a readable text form. The speech decoder post-network is responsible for converting the encoded speech signal back into a recognizable speech waveform or text form. Input the obtained training samples (including training text data and training speech data) into the pre-trained speech recognition and synthesis model (hereinafter referred to as the pre-model) for the pre-model to train.

[0049] Use the text encoder pre-network to extract features from the training text data to obtain training text feature vectors;

[0050] Specifically, a text encoder pre-network is used to extract features from the training text data. For example, a series of non-linear transformations are sequentially performed on the training text data through an embedding layer, a fully connected layer (or encoder layer), and an activation function layer to extract the training text feature vector. The training text feature vector can include vocabulary, grammar, semantics, etc. captured from the training text data.

[0051] More specifically, the training text data is input into the text encoder pre-network of the pre-model. Through the embedding layer in the text encoder pre-network, the training text data is converted into an embedding vector. For example, each word or character in the training text data is mapped to a vector space with a fixed dimension through the embedding layer. The fixed dimension can be 768 dimensions. The semantic relationship between words or characters is captured through this vector space to realize the conversion of the training text data into an embedding vector. The embedding vector is input into the encoder layer in the text encoder pre-network. The deep feature information in the embedding vector is captured through the encoder layer, and a corresponding hidden state is generated according to the deep feature information (that is, the hidden state contains the deep feature information of the training data text). The hidden state is used as the training text feature vector.

[0052] The training speech data is subjected to feature extraction using the speech encoder pre-network to obtain a training speech feature vector;

[0053] Specifically, the training speech data is subjected to feature extraction using the speech encoder pre-network. For example, feature extraction and transformation are sequentially performed on the training speech data through a convolutional layer, a normalization layer, and a linear layer to obtain a training speech feature vector. The training speech feature vector can include pitch, timbre, volume, Mel Frequency Cepstral Coefficients (MFCC), fundamental frequency, formants, etc. captured from the training speech data.

[0054] In some alternative implementation manners, the above-mentioned step of subjecting the training speech data to feature extraction using the speech encoder pre-network to obtain a training speech feature vector may include the following steps:

[0055] The convolutional layer in the speech encoder pre-network is used to perform feature extraction on the training speech data to obtain a first speech feature vector;

[0056] Specifically, through multiple convolutional layers in the speech encoder pre-network, convolutional operations are used to extract local features of the training speech data, and the multiple extracted features are superimposed to obtain a preliminarily extracted first speech feature vector. The first speech feature vector is a lower-dimensional vector containing the key information of the training speech data.

[0057] A predefined activation function is obtained, and the activation function is used to perform non-linear transformation on the first speech feature vector to obtain a second speech feature vector;

[0058] Specifically, a predefined activation function is obtained, and the activation function can be a Gaussian Error Linear Unit (GELU) function. The GELU function is a non-linear activation function that can endow the network with stronger non-linear expression ability. The activation function is used to perform non-linear transformation on each element in the first speech feature vector, making the elements in the first speech feature vector sparser and increasing the non-linear expression ability of the network, thereby obtaining a second speech feature vector.

[0059] The normalization layer in the pre-network of the speech encoder is used to adjust the scale of the second speech feature vector to obtain a third speech feature vector;

[0060] Specifically, through the normalization layer in the pre-network of the speech encoder, the mean and standard deviation of the second speech feature vector are calculated, and operations such as scaling and translation are performed on the second speech feature vector according to the calculated mean and standard deviation to adjust the size of the second speech feature vector, thereby obtaining a third speech feature vector.

[0061] The linear layer in the pre-network of the speech encoder is used to perform dimensional upsampling on the third speech feature vector to obtain a fourth speech feature vector as the training speech feature vector.

[0062] Specifically, through the linear layer in the pre-network of the speech encoder, the dimension of the third speech feature vector is upsampled from a lower dimension to a higher dimension. For example, the third speech feature vector is multiplied by a weight matrix and added with a bias vector to realize upsampling the dimension of the third speech feature vector from a lower dimension to 768 dimensions, thereby obtaining a fourth speech feature vector. This fourth speech feature vector is used as the training speech feature vector.

[0063] In the embodiment of the present application, through the convolutional layer in the pre-network of the speech encoder, convolutional operations can efficiently extract local features in the training speech data; by introducing an activation function, the non-linear expression ability of the network is increased, enabling the network to process more complex speech data and improving the generalization ability of the model; through the normalization layer, the size of the feature vector is adjusted, making the decibels of the feature vector more stable, which helps to accelerate the training process of the network; through the linear layer, the dimension of the feature vector is upsampled, making the finally obtained training speech feature vector have a higher dimension. Among them, a higher dimension means that the feature vector can contain more information, enhancing the richness of the feature vector. Based on the solution of the present application, through the sequential processing of the convolutional layer, activation function, normalization layer, and linear layer, technical effects such as efficient feature extraction, enhanced non-linear expression ability, stable feature vector scale, dimensional upsampling, and feature richness are achieved, thereby improving the performance and adaptability of the model.

[0064] Obtain a predefined task vector, and combine the training text feature vector, the training speech feature vector, and the task vector to obtain a training task fusion vector;

[0065] Specifically, obtain a predefined task vector, which is a vector representing specific task features. The task vector contains information related to a specific task and is used to guide the behavior of the model when processing the specific task. The task vector can include vectors for speech recognition tasks, vectors for speech synthesis tasks, etc. Combine the extracted training text feature vector and training speech feature vector with the task vector. For example, perform weighted summation or vector concatenation on the training text feature vector, the training speech feature vector, and the task vector to obtain a training task fusion vector. This training task fusion vector can reflect the extracted features of the input data (i.e., training text data and training speech data) and the task requirements included in the task vector.

[0066] In some optional implementation manners, the above-mentioned combining the training text feature vector, the training speech feature vector, and the task vector to obtain a training task fusion vector may include the following steps:

[0067] Concatenate the training text feature vector, the training speech feature vector, and the task vector in a specific format to obtain a concatenated vector;

[0068] Specifically, determine the specific format of vector concatenation. For example, the specific format is to perform concatenation in a specific order. Concatenate the training text feature vector, the training speech feature vector, and the task vector in the determined specific format. For example, first concatenate the training text feature vector and the training speech feature vector in a specific order, and then concatenate the task vector to obtain a concatenated vector.

[0069] Use the fully connected layer in the pre-trained speech recognition and synthesis model to perform dimensionality conversion on the concatenated vector according to specific dimensionality requirements to obtain a training task fusion vector.

[0070] Specifically, through the fully connected layer in the pre-trained speech recognition and synthesis model, obtain specific dimensionality requirements. For example, the specific dimensionality requirements can be that the input and output dimensions are the same; configure corresponding weight and bias parameters according to the specific dimensionality requirements; obtain the dimension of the concatenated vector according to the specific dimensionality requirements, use the dimension of the concatenated vector as the input of the fully connected layer, multiply by the weight and add the bias parameter to obtain the output vector after dimensionality conversion, and use the output vector after dimensionality conversion as the training task fusion vector.

[0071] In the embodiments of the present application, by combining the training text feature vector, the training speech feature vector, and the task vector, a training task fusion vector is obtained, such that the training task fusion vector not only contains the joint information of the text and the speech, but also incorporates specific information related to the task. Through the dimensionality transformation of the fully connected layer, it is ensured that the feature vectors from different sources can be fused and processed in a unified dimensional space, thereby improving the performance and adaptability of the model.

[0072] The post-network of the text decoder is used to process the training task fusion vector and output the synthetic text data for training;

[0073] Specifically, the post-network of the text decoder of the pre-model is used to perform text conversion according to the training task fusion vector. For example, through the Transformer decoder, the training task fusion vector is gradually used to generate the text sequence of a specific speech, and the finally generated text sequence is output as the synthetic text data for training.

[0074] In some optional implementation manners, the above-mentioned use of the post-network of the text decoder to process the training task fusion vector and output the synthetic text data for training may include the following steps:

[0075] Obtain the historical text output sequence, and use the Transformer decoder to generate the hidden state of the next word according to the training task fusion vector and the historical text output sequence;

[0076] Specifically, obtain the historical text output sequence, which refers to the word sequence that has been generated before the current word is generated during the text generation process. This word sequence is gradually constructed from the initial state, that is, starting from the start symbol, and each time a new word is generated, the generated word is added to the end of the word sequence to form the historical text output sequence. Through the Transformer decoder in the pre-model using the self-attention mechanism and positional encoding, the historical text output sequence and the training task fusion vector are encoded to generate the hidden state of the next word, and this hidden state contains the information required to generate the next word.

[0077] Use the post-network of the text decoder to project the hidden state of the next word into the vocabulary space, and use the softmax function to convert the hidden state of the next word into a word probability distribution;

[0078] Specifically, through the post-network of the text decoder, project the hidden state of the next word into the vocabulary space, where each vocabulary entry in this space corresponds to a possible word. Obtain a predefined normalized exponential function (such as the softmax function), pass the mapped hidden state to the normalized exponential function, and use the normalized exponential function to normalize the mapped hidden state so that the sum of the probabilities of all possible words is 1, generating a word probability distribution.

[0079] Select the word with the highest probability according to the word probability distribution and add it to the historical text output sequence; return to execute the step of obtaining the historical text output sequence, and use the Transformer decoder to generate the hidden state of the next word according to the training task fusion vector and the historical text output sequence.

[0080] Until the stop condition is met, output the historical text output sequence as the synthetic text data for training.

[0081] Specifically, traverse the probabilities of all words in the word probability distribution, select the word with the highest probability as the next word to be added to the historical text input sequence, and add the selected word to the end of the historical text input sequence. Return to the above step of obtaining the historical text output sequence, continue to execute using the Transformer decoder to generate the hidden state of the next word according to the training task fusion vector and the historical text output sequence, continue to generate the next word, and update the historical text output sequence. Until the stop condition is met, the stop condition can be reaching the preset maximum text length, the generated word is a special end symbol, or other custom stop conditions, and output the final historical text output sequence as the synthetic text data for training.

[0082] In the embodiment of the present application, by using the Transformer decoder and combining the training task fusion vector and the historical text output sequence, the hidden state of the next word can be efficiently generated; by projecting the hidden state into the vocabulary space and using the normalized exponential function to convert it into a word probability distribution, the next word can be intelligently selected based on the context information of the historical text output sequence and the training task fusion vector, which helps to generate coherent and grammatically correct text.

[0083] Process the training task fusion vector using the post-network of the speech decoder and output the synthetic speech data for training.

[0084] Specifically, after adopting the speech decoder, the network performs speech conversion based on the training task fusion vector. For example, the Transformer decoder is used to gradually generate a sequence of speech signals from the training task fusion vector, and a vocoder is used to convert this sequence of speech signals into speech audio data, which is output as the synthetic speech data for training.

[0085] In some alternative implementation manners, the above-mentioned process of the network adopting the speech decoder to process the training task fusion vector and output the synthetic speech data for training may include the following steps:

[0086] Obtain the embedding vector of the target speaker, where the target speaker refers to the speaker included in the training speech data;

[0087] Specifically, to obtain the embedding vector of the target speaker, which refers to the speaker included in the training speech data, and the embedding vector contains the unique speech features of the target speaker. More specifically, a number of historical speech data of historical speakers are collected in advance, where the historical speakers include not only the speakers in the training speech data but also the speakers in the speech data to be processed. The x-vectors technology is used to extract the embedding vectors of the historical speakers from the historical speech data, and the extracted embedding vectors are mapped into the vector space of the pre-model. During the training process of the pre-model, the speaker (i.e., the target speaker) is determined according to the training speech data, and the embedding vector of the target speaker is obtained from the vector space of the pre-model.

[0088] Combine the embedding vector and the training task fusion vector to obtain a first combined vector;

[0089] Specifically, combine the embedding vector of the target speaker and the training task fusion vector. For example, splice the embedding vector to the end position of the training task fusion vector to obtain a first combined vector, which contains both the speech features of the target speaker and the task requirements for the speech synthesis task and the context information of the training speech data.

[0090] Adopt the first linear layer in the network after the speech decoder to perform a linear transformation on the first combined vector in combination with a predefined activation function to obtain a second combined vector;

[0091] Specifically, through the first linear layer in the network after the speech decoder, a linear transformation is performed on the first combined vector, that is, a weighted sum processing is performed on the first combined vector using preset weight values to obtain a linearly transformed vector; a predefined activation function is obtained, and this activation function can be a Rectified Linear Unit (ReLU) function. The ReLU activation function is used to perform a non-linear process on the linearly transformed vector to increase the ability of the model to learn more complex feature representations, and a second combined vector is obtained.

[0092] Using the second linear layer in the network after the speech decoder, an upsampling in dimension is performed on the second combined vector to obtain a third combined vector;

[0093] Specifically, through the second linear layer in the network after the speech decoder, the dimension of the second combined vector is upsampled to a higher dimension (such as 768 dimensions) to obtain a third combined vector.

[0094] A historical speech output sequence is obtained, and a Transformer decoder is used to generate the hidden state of the next speech word according to the third combined vector and the historical speech output sequence;

[0095] Specifically, a historical speech output sequence is obtained. This historical speech output sequence refers to the sequence of speech words that have been generated before the current speech word during the speech synthesis process. This speech word sequence is gradually constructed from the initial state, that is, starting from the start symbol, after each new speech word is generated, the generated speech word is added to the end of the speech word sequence to form a historical speech output sequence. Through the Transformer decoder in the pre-model using the self-attention mechanism and positional encoding, encoding is performed according to the historical speech output sequence and the training task fusion vector to generate the hidden state of the next speech word, and this hidden state contains the information required to generate the next speech word.

[0096] Using the third linear layer in the network after the speech decoder, an initial Mel spectrogram is predicted according to the hidden state of the next speech word;

[0097] Specifically, through the third linear layer in the network after the speech decoder, an initial Mel spectrogram is predicted according to the hidden state of the next speech word. Among them, the Mel spectrogram refers to the result obtained by converting the sound signal into a frequency domain representation through a series of mathematical transformations and performing weighted averaging based on the Mel scale.

[0098] Using the convolutional layer in the network after the speech decoder, a residual vector is generated according to the initial Mel spectrogram, and the residual vector is added to the initial Mel spectrogram to obtain a target Mel spectrogram;

[0099] Specifically, through the convolutional layer in the network after the speech decoder, feature extraction is performed on the initial Mel spectrogram to form one or more feature maps; the feature maps of the ideal Mel spectrogram are obtained, and according to the feature maps of the ideal Mel spectrogram and the feature maps of the initial Mel spectrogram, a residual vector is extracted, and this residual vector represents the difference or error between the initial Mel spectrogram and the ideal Mel spectrogram. The residual vector is added back to the initial Mel spectrogram to correct or enhance the initial Mel spectrogram, such as increasing or decreasing the energy of certain frequency components, adjusting the smoothness or sharpness of the spectrogram, etc., to obtain the target Mel spectrogram.

[0100] Using the fourth linear layer in the network after the speech decoder, the hidden state of the next speech word is converted into a scalar value to predict the stop speech word identifier;

[0101] Specifically, through the fourth linear layer in the network after the speech decoder, a linear transformation is performed on the hidden state of the next speech word to compress the high-dimensional hidden state into a lower-dimensional scalar value, and this scalar value is used to indicate whether more speech words (i.e., the next speech word) should be stopped from being generated after the current speech word. If the scalar value is greater than a preset threshold, it is predicted that the generation of speech words should be stopped. At this time, the next speech word is determined as the stop speech word identifier; otherwise, the next speech word is continued to be generated.

[0102] Speech synthesis is performed according to the target Mel spectrogram, and the speech synthesis process is ended according to the stop speech word identifier to obtain the synthetic speech data for training.

[0103] Specifically, when the target Mel spectrogram and the stop speech word identifier are generated, the target Mel spectrogram and the stop speech word identifier are passed to the vocoder configured by the pre-model. Speech synthesis is performed according to the target Mel spectrogram, and it is determined when to end the speech synthesis process according to the stop speech word identifier. When the speech synthesis process ends, the synthetic speech data for training is output.

[0104] In the embodiments of the present application, by combining the embedding vector of the target speaker with the training task fusion vector, a first combined vector is obtained, realizing the effective fusion of different features; through a first linear layer and a predefined activation function, a linear transformation is performed on the first combined vector to obtain a second combined vector, further enhancing the feature expression ability; through a second linear layer, dimensional upsampling is performed on the second combined vector to provide a suitable input dimension for the Transformer decoder; by using the Transformer decoder to generate the hidden state of the next speech word based on the historical speech output sequence, the temporal characteristics of the speech sequence are fully considered, which helps to maintain the coherence and naturalness of the generated speech in terms of time sequence; through a third linear layer, the initial Mel spectrogram is predicted according to the hidden state of the next speech word, providing key audio features for speech synthesis; a convolutional layer is used to further generate a residual vector based on the initial Mel spectrogram and add it to the initial Mel spectrogram to capture more subtle audio features, which helps to improve the quality of speech synthesis and make the generated speech clearer and more realistic; through a fourth linear layer, the hidden state of the next speech word is converted into a scalar value to predict the stop speech word identifier, which can accurately determine when to end the speech synthesis process and avoid long or incomplete speech output.

[0105] Obtain a predefined target loss function, and combine the synthetic text data for training, the synthetic speech data for training, and the target loss function to perform model training on the pre-trained speech recognition and synthesis model to obtain a trained speech recognition and synthesis model.

[0106] In some optional implementation manners, the above target loss function includes a first loss function and a second loss function. The above step of performing model training on the pre-trained speech recognition and synthesis model to obtain a trained speech recognition and synthesis model may include the following steps:

[0107] Use the first loss function to calculate the loss of the synthetic text data for training and the synthetic speech data for training to obtain a first loss function value for the speech recognition task;

[0108] Specifically, the synthetic text data for training and the synthetic speech data for training are passed to the first loss function, which is a loss function for the speech recognition task and may include a cross-entropy loss function and a connectionist temporal classification (CTC) loss function, etc.; use the first loss function to calculate the loss of the synthetic text data for training and the synthetic speech data for training to obtain a first loss function value for the speech recognition task.

[0109] Exemplarily, the loss function for the speech recognition task (i.e., the first loss function) is shown in Formula 1 below:

[0110] l asr = α·L ce + β·L ctc (1);

[0111] wherein, L asr is the first loss function value of the automatic speech recognition task (ASR); L ce is the cross-entropy loss function, and L ctc is the CTC loss function; both α and β are weight factors used to balance the proportions of the cross-entropy loss function and the CTC loss function.

[0112] Using the second loss function, calculate the loss of the synthetic text data for training and the synthetic speech data for training to obtain the second loss function value for the text-to-speech task;

[0113] Specifically, transfer the synthetic text data for training and the synthetic speech data for training to the second loss function, which is a loss function for the text-to-speech task and may include the L1 loss function (used to minimize the distance between the target mel spectrogram and the ideal mel spectrogram), the binary cross-entropy loss function (used to predict the generated stop speech word identifier), and the guided attention loss function (used to accelerate training convergence), etc.; use the second loss function to calculate the loss of the synthetic text data for training and the synthetic speech data for training to obtain the second loss function value for the text-to-speech task.

[0114] Combine the first loss function value and the second loss function value to obtain the total loss value, and perform normalization processing on the total loss value to obtain the normalized total loss value;

[0115] Specifically, to implement the joint training of the automatic speech recognition task (ASR) and the text-to-speech task (TTS), it is necessary to ensure that the convergence speeds of the automatic speech recognition task and the text-to-speech task are similar. Combine the first loss function value and the second loss function value to obtain the total loss value, wherein the combination method is shown in Formula 2 below:

[0116] L = L asr + L tts (2);

[0117] wherein, L is the total loss value, L asr is the first loss function value of the automatic speech recognition task (ASR), and L tts is the second loss function value of the text-to-speech task (TTS).

[0118] Specifically, perform normalization processing on the total loss value to obtain the normalized total loss value to ensure that the contributions of the automatic speech recognition task (ASR) and the text-to-speech task (TTS) to the total loss value are balanced.

[0119] Calculate the gradient according to the normalized total loss value, and update the parameters of the pre-trained speech recognition and synthesis model according to the gradient. Iteratively train the pre-trained speech recognition and synthesis model until the model converges to obtain a trained speech recognition and synthesis model.

[0120] Specifically, according to the normalized total loss value, use the gradient descent method to calculate the gradient, and update the parameters of the pre-model according to this gradient to minimize the total loss value. Repeat the above steps until the preset number of training rounds is reached or other stopping conditions are met, and then stop training the pre-model to obtain a trained speech recognition and synthesis model.

[0121] In the embodiment of the present application, the speech recognition task is optimized through the first loss function (usually the weighted sum of cross-entropy loss and CTC loss), which can improve the accuracy of the model in sequence-level classification and frame-by-frame alignment, thereby improving the accuracy and robustness of the speech recognition system; the speech synthesis task is optimized through the second loss function (including L1 loss, binary cross-entropy loss, and guided attention loss), which not only improves the quality of the generated speech but also significantly speeds up the training speed; by combining the task losses of speech recognition and speech synthesis into the total loss and normalizing it in each training step, it is ensured that the model can be stably and consistently optimized during the joint training process, avoiding a certain task dominating the training process and causing the performance of the model on another task to decline, thus achieving the balance of multi-task learning.

[0122] Step S202: Input the to-be-processed text data and the to-be-processed speech data into a pre-trained speech recognition and synthesis model, where the speech recognition and synthesis model includes a text encoder pre-network, a speech encoder pre-network, a text decoder post-network, and a speech decoder post-network;

[0123] Specifically, load the pre-trained speech recognition and synthesis model from a storage medium (such as a hard disk, cloud storage, etc.). This speech recognition and synthesis model is composed of a text encoder pre-network, a speech encoder pre-network, a text decoder post-network, and a speech decoder post-network. Among them, the explanations of the text encoder pre-network, the speech encoder pre-network, the text decoder post-network, and the speech decoder post-network are as shown above and will not be elaborated here. Initialize the pre-trained speech recognition and synthesis model, and input the obtained to-be-processed text data and to-be-processed speech data into this pre-trained speech recognition and synthesis model (hereinafter referred to as the model) for the model to process.

[0124] Step S203: Use the text encoder pre-network to extract features from the to-be-processed text data to obtain a target text feature vector;

[0125] Specifically, a text encoder pre-network is used to extract features from the text data to be processed. For example, a series of non-linear transformations are sequentially performed on the text data to be processed through an embedding layer, a fully connected layer (or an encoder layer), and an activation function layer to extract the target text feature vector. The target text feature vector may include vocabulary, grammar, semantics, etc. captured from the text data to be processed.

[0126] Step S204: Use the speech encoder pre-network to extract features from the speech data to be processed to obtain a target speech feature vector.

[0127] Specifically, a speech encoder pre-network is used to extract features from the speech data to be processed. For example, feature extraction and transformation are sequentially performed on the speech data to be processed through a convolutional layer, a normalization layer, and a linear layer to obtain a target speech feature vector. The target speech feature vector may include pitch, timbre, volume, Mel Frequency Cepstral Coefficients (MFCC), fundamental frequency, formants, etc. captured from the speech data to be processed.

[0128] Step S205: Obtain a predefined task vector, and combine the target text feature vector, the target speech feature vector, and the task vector to obtain a target task fusion vector.

[0129] Specifically, a predefined task vector is obtained. The explanation of the task vector is as shown above and will not be elaborated here. The extracted target text feature vector and target speech feature vector are combined with the task vector. For example, the target text feature vector, the target speech feature vector, and the task vector are weighted and summed or vector concatenated to obtain a target task fusion vector. The target task fusion vector can reflect the extracted features of the input data (i.e., the text data to be processed and the speech data to be processed) and the task requirements included in the task vector.

[0130] Step S206: Use the text decoder post-network to process the target task fusion vector and output target synthesized text data.

[0131] Specifically, the text decoder post-network performs text conversion based on the target task fusion vector. For example, the Transformer decoder gradually generates a text sequence of the target speech from the target task fusion vector, and the finally generated text sequence is output as the target synthesized text data.

[0132] Step S207: Use the speech decoder post-network to process the target task fusion vector and output target synthesized speech data.

[0133] Specifically, after adopting a speech decoder, the network performs speech conversion according to the target task fusion vector. For example, the Transformer decoder is used to gradually generate a sequence of speech signal from the target task fusion vector, and a vocoder is used to convert the speech signal sequence into speech audio data, which is output as the target synthesized speech data.

[0134] Through the above solutions, the embodiments of the present application specifically construct a speech recognition and synthesis model, and use the text encoder pre-network and the speech encoder pre-network to extract features from the text to be processed and the speech data to be processed respectively, generating high-quality feature vectors. By introducing a task vector, the target task fusion vector is obtained by combining the target text feature vector, the target speech feature vector and the task vector, realizing the fusion of cross-modal features. This fusion mechanism helps the model better understand the semantics and speech characteristics of the input data, so as to more accurately complete the speech recognition and speech synthesis tasks. By using the text decoder post-network and the speech decoder post-network to process the target task fusion vector respectively, the target synthesized text data and the target synthesized speech data are output, enabling the simultaneous generation of two-modal outputs of synthesized text and synthesized speech, realizing the knowledge transfer between the speech recognition and speech synthesis tasks, and improving the generalization ability of the model. Based on the solution of the present application, by constructing an integrated model including a text encoder, a speech encoder and corresponding decoders, the high integration of the two tasks of speech recognition and text-to-speech synthesis is realized, which not only simplifies the model architecture, but also enables the model to share knowledge when processing different tasks by sharing the feature extraction and encoding processes, improving the data processing effect. By adopting the multi-task learning and parameter sharing strategy, while ensuring the model performance, the redundancy of the model parameters is reduced, significantly reducing the consumption of computing resources and time.

[0135] Exemplarily, with the rapid development of fintech, intelligent customer service has become an important tool for the financial industry to improve service quality and customer experience. In order to process customer inquiries and transaction requests more efficiently and accurately, this example proposes an intelligent customer service system based on unified speech processing technology, which adopts a new speech processing method. First, through the speech interface and text input box of the financial platform, the voice input and text input of customers are captured in real time as the text data to be processed and the voice data to be processed. These inputs may include customer account inquiries, transaction requests, complaints and suggestions, etc. The captured text data and voice data to be processed are input into a pre-trained speech recognition and synthesis model. This model is a unified network structure that integrates a text encoder pre-network and a voice encoder pre-network, as well as a text decoder post-network and a voice decoder post-network. The text encoder pre-network is used to extract features from the text data to be processed to generate a target text feature vector. At the same time, the voice encoder pre-network is used to extract features from the voice data to be processed to generate a target voice feature vector. According to the input content of the customer and the requirements of the financial platform, the system generates a predefined task vector. This vector indicates the specific type of the current processing task, such as speech recognition, speech synthesis, or a combination of both. The target text feature vector, the target voice feature vector, and the task vector are fused to generate a target task fusion vector. This step realizes the effective integration of text and voice features and the clear indication of task information. The text decoder post-network is used to process the target task fusion vector to generate target synthesized text data. These data may include the text reply of the system to the customer, transaction confirmation information, etc. At the same time, the voice decoder post-network is used to process the target task fusion vector to generate target synthesized voice data. These data may include the voice reply of the system to the customer, operation guidance, etc. Then, based on the generated target synthesized text data and target synthesized voice data, the system provides real-time feedback on the processing results to the customer through the text display box and voice playback interface of the financial platform. In addition, the system also collects the satisfaction and opinions of customers on the intelligent customer service through the feedback channels of the financial platform to continuously optimize the service quality and improve the customer experience.

[0136] Through the above content, the intelligent customer service system based on unified speech processing technology proposed in this example realizes seamless interaction between customers and the financial system by integrating text and voice processing functions. The system can accurately understand the voice input and text input of customers and provide accurate and timely responses in natural language or voice form, significantly improving the service quality and customer experience of the financial industry.

[0137] It should be emphasized that, to further ensure the privacy and security of the above-mentioned information such as the text data to be processed, the voice data to be processed, the target synthesized text data, and the target synthesized voice data, as well as other relevant data, the above-mentioned information such as the text data to be processed, the voice data to be processed, the target synthesized text data, and the target synthesized voice data, as well as other relevant data can also be stored in a node of a blockchain.

[0138] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.

[0139] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), etc., or a random access memory (RAM), etc.

[0140] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0141] Further reference Figure 3 to Figure 2 As an implementation of the method shown above, an embodiment of a voice processing device is provided in this application. This device embodiment corresponds to the method embodiment shown in Figure 2 and this device can be specifically applied to various electronic devices.

[0142] Such as Figure 3As shown in the figure, the speech processing device 400 described in this embodiment includes: a data acquisition module 401, a model input module 402, a text extraction module 403, a speech extraction module 404, a vector combination module 405, a first synthesis module 406, and a second synthesis module 407. Among them:

[0143] The data acquisition module 401 is used to acquire text data to be processed and speech data to be processed;

[0144] The model input module 402 is used to input the text data to be processed and the speech data to be processed into a pre-trained speech recognition and synthesis model, and the speech recognition and synthesis model includes a text encoder pre-network, a speech encoder pre-network, a text decoder post-network, and a speech decoder post-network;

[0145] The text extraction module 403 is used to extract features from the text data to be processed by using the text encoder pre-network to obtain a target text feature vector;

[0146] The speech extraction module 404 is used to extract features from the speech data to be processed by using the speech encoder pre-network to obtain a target speech feature vector;

[0147] The vector combination module 405 is used to obtain a predefined task vector, combine the target text feature vector, the target speech feature vector, and the task vector to obtain a target task fusion vector;

[0148] The first synthesis module 406 is used to process the target task fusion vector by using the text decoder post-network and output target synthesized text data;

[0149] The second synthesis module 407 is used to process the target task fusion vector by using the speech decoder post-network and output target synthesized speech data.

[0150] In this embodiment, the data acquisition module 401 is used to acquire text data to be processed and speech data to be processed. The text data to be processed refers to text information that needs to be semantically processed or analyzed, and may include, but is not limited to, current user input, current web page content, current document files, etc.; the speech data to be processed refers to audio information that needs to be speech processed or analyzed, and may include, but is not limited to, current recorded audio, current phone calls, current video files, etc.

[0151] Specifically, the model input module 402 is used to load a pre-trained speech recognition and synthesis model from a storage medium (such as a hard disk, cloud storage, etc.). The speech recognition and synthesis model is composed of a text encoder pre-network, a speech encoder pre-network, a text decoder post-network, and a speech decoder post-network. Among them, the text encoder pre-network is responsible for converting the input text data (such as characters, words, or tokens) into a fixed-dimensional word embedding form (i.e., a vector), and capturing the semantic information in the text data through the word embedding. The speech encoder pre-network is responsible for converting the input speech signal into a high-dimensional vector representation, and capturing the acoustic features and semantic information in the speech data through the vector. The text decoder post-network is responsible for converting the encoded text data back into a readable text form. The speech decoder post-network is responsible for converting the encoded speech signal back into a recognizable speech waveform or text form. Initialize the pre-trained speech recognition and synthesis model, and input the obtained text data to be processed and speech data to be processed into the pre-trained speech recognition and synthesis model (hereinafter referred to as the model) for the model to process.

[0152] Specifically, the text extraction module 403 is used to extract features from the text data to be processed by using the text encoder pre-network. For example, a series of non-linear transformations are performed on the text data to be processed through an embedding layer, a fully connected layer (or an encoder layer), and an activation function layer in sequence to extract the target text feature vector. The target text feature vector can include vocabulary, grammar, semantics, etc. captured from the text data to be processed.

[0153] Specifically, the speech extraction module 404 is used to extract features from the speech data to be processed by using the speech encoder pre-network. For example, feature extraction and transformation are performed on the speech data to be processed through a convolutional layer, a normalization layer, and a linear layer in sequence to obtain the target speech feature vector. The target speech feature vector can include pitch, timbre, volume, Mel-frequency cepstral coefficients (MFCC), fundamental frequency, formants, etc. captured from the speech data to be processed.

[0154] Specifically, the vector combination module 405 is used to obtain a predefined task vector, which is a vector representing specific task features. The task vector contains information related to a specific task and is used to guide the behavior of the model when processing the specific task. The task vector can include vectors for speech recognition tasks, vectors for speech synthesis tasks, etc. Combine the extracted target text feature vector and target speech feature vector with the task vector. For example, perform weighted summation or vector concatenation on the target text feature vector, target speech feature vector, and task vector to obtain the target task fusion vector. The target task fusion vector can reflect the extracted features of the input data (i.e., the text data to be processed and the speech data to be processed) and the task requirements contained in the task vector.

[0155] Specifically, the first synthesis module 406 is configured to perform text conversion on the basis of the fusion vector of the target task by using the post-network of the text decoder. For example, the Transformer decoder is used to gradually generate the text sequence of the target speech from the fusion vector of the target task, and the finally generated text sequence is output as the target synthesized text data.

[0156] Specifically, the second synthesis module 407 is configured to perform voice conversion on the basis of the fusion vector of the target task by using the post-network of the voice decoder. For example, the Transformer decoder is used to gradually generate the voice signal sequence from the fusion vector of the target task, and the vocoder is used to convert the voice signal sequence into voice audio data, which is output as the target synthesized voice data.

[0157] The voice processing device 400 according to the embodiment of the present application constructs a speech recognition and synthesis model, and uses the pre-network of the text encoder and the pre-network of the voice encoder to extract features from the text to be processed and the voice data to be processed respectively, generating high-quality feature vectors. By introducing the task vector and combining the target text feature vector, the target voice feature vector and the task vector to obtain the target task fusion vector, the cross-modal feature fusion is realized. This fusion mechanism helps the model to better understand the semantic and voice characteristics of the input data, so as to more accurately complete the speech recognition and speech synthesis tasks. By using the post-network of the text decoder and the post-network of the voice decoder to process the target task fusion vector respectively and output the target synthesized text data and the target synthesized voice data, it is possible to generate two-modal outputs of synthesized text and synthesized voice at the same time, realizing the knowledge transfer between the speech recognition and speech synthesis tasks and improving the generalization ability of the model. Based on the solution of the present application, by constructing an integrated model including a text encoder, a voice encoder and corresponding decoders, the high-degree integration of the two tasks of speech recognition and text-to-speech synthesis is realized, which not only simplifies the model architecture, but also enables the model to share knowledge when processing different tasks by sharing the feature extraction and encoding processes, improving the data processing effect. By adopting the multi-task learning and parameter sharing strategy, while ensuring the model performance, the redundancy of the model parameters is reduced, and the consumption of computing resources and time is significantly reduced.

[0158] In some optional implementation manners of this embodiment, the voice processing device 400 may further include a sample acquisition module, a sample input module, a first extraction module, a second extraction module, a vector fusion module, a first processing module, a second processing module and a model training module. Among them:

[0159] The sample acquisition module is configured to acquire pre-collected training samples, and the training samples include a plurality of training text data and training voice data;

[0160] A sample input module for inputting the training samples into a pre-trained speech recognition and synthesis model, which includes a text encoder pre-network, a speech encoder pre-network, a text decoder post-network, and a speech decoder post-network;

[0161] A first extraction module for extracting features from the training text data using the text encoder pre-network to obtain a training text feature vector;

[0162] A second extraction module for extracting features from the training speech data using the speech encoder pre-network to obtain a training speech feature vector;

[0163] A vector fusion module for obtaining a predefined task vector, combining the training text feature vector, the training speech feature vector, and the task vector to obtain a training task fusion vector;

[0164] A first processing module for processing the training task fusion vector using the text decoder post-network to output synthetic text data for training;

[0165] A second processing module for processing the training task fusion vector using the speech decoder post-network to output synthetic speech data for training;

[0166] A model training module for obtaining a predefined target loss function, combining the synthetic text data for training, the synthetic speech data for training, and the target loss function to perform model training on the pre-trained speech recognition and synthesis model to obtain a trained speech recognition and synthesis model.

[0167] Specifically, a sample acquisition module for acquiring pre-collected training samples, which include a number of training text data and training speech data. Among them, the training text data refers to historical text information for semantic processing or analysis; the training speech data refers to historical audio information for speech processing or analysis.

[0168] Specifically, a sample input module for loading a pre-trained speech recognition and synthesis model, which is composed of a text encoder pre-network, a speech encoder pre-network, a text decoder post-network, and a speech decoder post-network. The explanations of the text encoder pre-network, the speech encoder pre-network, the text decoder post-network, and the speech decoder post-network are as shown above and will not be elaborated here. Input the obtained training samples (including training text data and training speech data) into this pre-trained speech recognition and synthesis model (hereinafter referred to as the pre-model) for the pre-model to train.

[0169] Specifically, the first extraction module is used to extract features from the training text data by using a text encoder pre-network. For example, a series of non-linear transformations are performed on the training text data through an embedding layer, a fully connected layer (or an encoder layer), and an activation function layer in sequence to extract the training text feature vector. The training text feature vector can include vocabulary, grammar, semantics, etc. captured from the training text data.

[0170] Specifically, the second extraction module is used to extract features from the training speech data by using a speech encoder pre-network. For example, feature extraction and transformation are performed on the training speech data through a convolutional layer, a normalization layer, and a linear layer in sequence to obtain the training speech feature vector. The training speech feature vector can include pitch, timbre, volume, Mel Frequency Cepstral Coefficients (MFCC), fundamental frequency, formants, etc. captured from the training speech data.

[0171] Specifically, the vector fusion module is used to obtain a predefined task vector, the explanation of which is as shown above and will not be elaborated here. The extracted training text feature vector and training speech feature vector are combined with the task vector. For example, the training text feature vector, the training speech feature vector, and the task vector are weighted and summed or vector concatenated to obtain the training task fusion vector. The training task fusion vector can reflect the extraction features of the input data (i.e., the training text data and the training speech data) and the task requirements included in the task vector.

[0172] Specifically, the first processing module is used to perform text conversion according to the training task fusion vector by using the text decoder post-network of the pre-model. For example, the training task fusion vector is gradually generated into a text sequence of a specific speech through a Transformer decoder, and the finally generated text sequence is output as the synthetic text data for training.

[0173] Specifically, the second processing module is used to perform speech conversion according to the training task fusion vector by using the speech decoder post-network. For example, the training task fusion vector is gradually generated into a sequence of speech signals through a Transformer decoder, and a vocoder is used to convert the sequence of speech signals into speech audio data, which is output as the synthetic speech data for training.

[0174] Specifically, the model training module is used to obtain a predefined target loss function, which can include a loss function for the speech recognition task and a loss function for the speech synthesis task. Combining the synthetic text data for training, the synthetic speech data for training, and the target loss function, the pre-trained speech recognition and synthesis model is trained to obtain the trained speech recognition and synthesis model.

[0175] In some alternative implementation manners of this embodiment, the above-mentioned second extraction module may include a feature extraction sub-module, a first conversion sub-module, a size adjustment sub-module, and a first sampling sub-module. Among them:

[0176] The feature extraction sub-module is configured to use the convolutional layer in the pre-network of the speech encoder to extract features from the training speech data to obtain a first speech feature vector;

[0177] The first conversion sub-module is configured to obtain a predefined activation function and use the activation function to perform a non-linear conversion on the first speech feature vector to obtain a second speech feature vector;

[0178] The size adjustment sub-module is configured to use the normalization layer in the pre-network of the speech encoder to adjust the scale of the second speech feature vector to obtain a third speech feature vector;

[0179] The first sampling sub-module is configured to use the linear layer in the pre-network of the speech encoder to perform dimensional up-sampling on the third speech feature vector to obtain a fourth speech feature vector as the training speech feature vector.

[0180] In some alternative implementation manners of this embodiment, the above-mentioned vector fusion module may include a task splicing sub-module and a second conversion sub-module. Among them:

[0181] The task splicing sub-module is configured to splice the training text feature vector, the training speech feature vector, and the task vector in a specific format to obtain a spliced vector;

[0182] The second conversion sub-module is configured to use the fully connected layer in the pre-trained speech recognition and synthesis model to perform dimensional conversion on the spliced vector according to specific dimensional requirements to obtain a training task fusion vector.

[0183] In some alternative implementation manners of this embodiment, the above-mentioned first processing module may include a first decoding sub-module, a state projection sub-module, a word addition sub-module, and a text output sub-module. Among them:

[0184] The first decoding sub-module is configured to obtain a historical text output sequence and use a Transformer decoder to generate a hidden state of the next word according to the training task fusion vector and the historical text output sequence;

[0185] The state projection sub-module is configured to use the post-network of the text decoder to project the hidden state of the next word into the vocabulary space and use the softmax function to convert the hidden state of the next word into a word probability distribution;

[0186] A word addition sub-module, configured to select the word with the highest probability according to the word probability distribution and add it to the historical text output sequence; return the step of performing the obtaining of the historical text output sequence, and using a Transformer decoder to generate the hidden state of the next word according to the training task fusion vector and the historical text output sequence.

[0187] A text output sub-module, configured to output the historical text output sequence as synthetic text data for training until a stop condition is satisfied.

[0188] In some optional implementation manners of this embodiment, the second processing module may include an object acquisition sub-module, a vector combination sub-module, a linear transformation sub-module, a second sampling sub-module, a second decoding sub-module, an image prediction sub-module, a residual addition sub-module, a stop prediction sub-module, and a speech synthesis sub-module. Among them:

[0189] The object acquisition sub-module is configured to acquire the embedding vector of the target speaking object, where the target speaking object refers to the speaking object included in the training speech data.

[0190] The vector combination sub-module is configured to combine the embedding vector and the training task fusion vector to obtain a first combined vector.

[0191] The linear transformation sub-module is configured to use the first linear layer in the post-network of the speech decoder to perform a linear transformation on the first combined vector in combination with a predefined activation function to obtain a second combined vector.

[0192] The second sampling sub-module is configured to use the second linear layer in the post-network of the speech decoder to perform dimensional upsampling on the second combined vector to obtain a third combined vector.

[0193] The second decoding sub-module is configured to obtain a historical speech output sequence, and use a Transformer decoder to generate the hidden state of the next speech word according to the third combined vector and the historical speech output sequence.

[0194] The image prediction sub-module is configured to use the third linear layer in the post-network of the speech decoder to predict an initial Mel spectrogram according to the hidden state of the next speech word.

[0195] The residual addition sub-module is configured to use the convolutional layer in the post-network of the speech decoder to generate a residual vector according to the initial Mel spectrogram, and add the residual vector to the initial Mel spectrogram to obtain a target Mel spectrogram.

[0196] A stop prediction sub-module, which uses the fourth linear layer in the post-network of the speech decoder to convert the hidden state of the next speech word into a scalar value to predict a stop speech word identifier;

[0197] A speech synthesis sub-module, which is used to perform speech synthesis based on the target Mel spectrogram and end the speech synthesis process according to the stop speech word identifier to obtain synthetic speech data for training.

[0198] In some alternative implementation manners of this embodiment, the model training module may include a first calculation sub-module, a second calculation sub-module, a third calculation sub-module, and a model training sub-module.

[0199] Among them:

[0200] The first calculation sub-module is used to calculate the loss of the training synthetic text data and the training synthetic speech data by using the first loss function to obtain a first loss function value for the speech recognition task;

[0201] The second calculation sub-module is used to calculate the loss of the training synthetic text data and the training synthetic speech data by using the second loss function to obtain a second loss function value for the speech synthesis task;

[0202] The third calculation sub-module is used to combine the first loss function value and the second loss function value to obtain a total loss value, and perform normalization processing on the total loss value to obtain a normalized total loss value;

[0203] The model training sub-module is used to calculate the gradient according to the normalized total loss value, update the parameters of the pre-trained speech recognition and synthesis model according to the gradient, and iteratively train the pre-trained speech recognition and synthesis model until the model converges to obtain a trained speech recognition and synthesis model.

[0204] To solve the above technical problems, an embodiment of the present application also provides a computer device. Specifically, please refer to Figure 4 , Figure 4 which is the basic structural block diagram of the computer device in this embodiment.

[0205] The computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are communicatively connected to each other via a system bus. It should be noted that only the computer device 6 having the memory 61, the processor 62, and the network interface 63 is shown in the figure. However, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0206] The computer device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote control, a touchpad, a voice control device, etc.

[0207] The memory 61 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 61 may be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 may also be an external storage device of the computer device 6, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 6. Of course, the memory 61 may also include both the internal storage unit of the computer device 6 and its external storage device. In this embodiment, the memory 61 is generally used to store the operating system and various application software installed on the computer device 6, such as computer-readable instructions of the voice processing method. In addition, the memory 61 may also be used to temporarily store various types of data that have been output or will be output.

[0208] In some embodiments, the processor 62 may be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 62 is generally used to control the overall operation of the computer device 6. In this embodiment, the processor 62 is used to run the computer-readable instructions stored in the memory 61 or process data, such as running the computer-readable instructions of the voice processing method.

[0209] The network interface 63 may include a wireless network interface or a wired network interface. The network interface 63 is generally used to establish a communication connection between the computer device 6 and other electronic devices.

[0210] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor, so that the at least one processor executes the steps of the voice processing method as described above.

[0211] The computer device, computer-readable storage medium, and computer-readable instructions provided in the embodiments of the present application build a voice recognition and synthesis model through the execution of the processor. By using the text encoder pre-network and the voice encoder pre-network, the text to be processed and the voice data to be processed are respectively subjected to feature extraction, generating high-quality feature vectors. By introducing a task vector, a target task fusion vector is obtained by combining the target text feature vector, the target voice feature vector, and the task vector, realizing the fusion of cross-modal features. This fusion mechanism helps the model better understand the semantic and voice characteristics of the input data, thereby more accurately completing the voice recognition and voice synthesis tasks. By using the text decoder post-network and the voice decoder post-network to process the target task fusion vector respectively, the target synthesized text data and the target synthesized voice data are output, enabling the simultaneous generation of two-modal outputs of synthesized text and synthesized voice, realizing the knowledge transfer between the voice recognition and voice synthesis tasks, and improving the generalization ability of the model. Based on the solution of the present application, by constructing an integrated model including a text encoder, a voice encoder, and corresponding decoders, a high degree of integration of the two tasks of voice recognition and text-to-speech synthesis is achieved. This not only simplifies the model architecture but also enables the model to share knowledge when processing different tasks by sharing the feature extraction and encoding processes, improving the data processing effect. By adopting the multi-task learning and parameter sharing strategies, while ensuring the model performance, the redundancy of the model parameters is reduced, significantly reducing the consumption of computing resources and time.

[0212] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on this understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0213] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The accompanying drawings show preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments or equivalently replace some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is equally within the scope of the patent protection of the present application.

[0214] The non-company software tools or components appearing in the embodiments of the present application are only introduced by way of example and do not represent actual use.

Claims

1. A speech processing method, characterized in that: The steps include: Obtaining text data and voice data to be processed; Inputting the to-be-processed text data and the to-be-processed speech data into a pre-trained speech recognition and synthesis model, wherein the speech recognition and synthesis model comprises a text encoder pre-network, a speech encoder pre-network, a text decoder post-network, and a speech decoder post-network; Using the text encoder pre-network to extract features from the text data to be processed to obtain a target text feature vector; Using the speech encoder pre-network to extract features from the speech data to be processed to obtain a target speech feature vector; Obtaining a predefined task vector, combining the target text feature vector, the target speech feature vector and the task vector to obtain a target task fusion vector; The target task fusion vector is processed by the text decoder post-network to output target synthetic text data; The speech decoder back network is used to process the target task fusion vector and output target synthesized speech data.

2. The speech processing method according to claim 1, characterized in that: Before the step of obtaining the text data to be processed and the voice data to be processed, the method further includes: Acquire pre-collected training samples, wherein the training samples include a plurality of training text data and training voice data; Inputting the training sample into a pre-trained speech recognition and synthesis model, wherein the pre-trained speech recognition and synthesis model includes a text encoder pre-network, a speech encoder pre-network, a text decoder post-network, and a speech decoder post-network; Using the text encoder pre-network to perform feature extraction on the training text data to obtain a training text feature vector; Using the speech encoder pre-network to extract features from the training speech data to obtain a training speech feature vector; Obtaining a predefined task vector, combining the training text feature vector, the training speech feature vector and the task vector to obtain a training task fusion vector; The text decoder back network is used to process the training task fusion vector and output synthetic text data for training; The speech decoder back network is used to process the training task fusion vector and output the synthetic speech data for training; A predefined target loss function is obtained, and model training is performed on the pre-trained speech recognition and synthesis model in combination with the synthetic text data for training, the synthetic speech data for training and the target loss function to obtain a trained speech recognition and synthesis model.

3. The speech processing method according to claim 2, characterized in that: The step of extracting features from the training speech data using the speech encoder pre-network to obtain a training speech feature vector comprises: Using the convolutional layer in the speech encoder pre-network, extracting features from the training speech data to obtain a first speech feature vector; Obtaining a predefined activation function, and using the activation function to perform nonlinear transformation on the first speech feature vector to obtain a second speech feature vector; Using a normalization layer in the speech encoder pre-network, rescaling the second speech feature vector to obtain a third speech feature vector; The third speech feature vector is dimensionally upsampled using a linear layer in the speech encoder pre-network to obtain a fourth speech feature vector as a training speech feature vector.

4. The speech processing method according to claim 2, characterized in that: The step of combining the training text feature vector, the training speech feature vector and the task vector to obtain a training task fusion vector comprises: Performing vector splicing on the training text feature vector, the training speech feature vector and the task vector according to a specific format to obtain a spliced ​​vector; The fully connected layer in the pre-trained speech recognition and synthesis model is used to perform dimension conversion on the concatenated vector according to specific dimensional requirements to obtain a training task fusion vector.

5. The speech processing method according to claim 2, characterized in that: The step of using the text decoder back network to process the training task fusion vector and outputting synthetic text data for training includes: Obtain a historical text output sequence, and use a Transformer decoder to generate a hidden state of the next word according to the training task fusion vector and the historical text output sequence; Using the text decoder post-network, projecting the hidden state of the next word into the vocabulary space, and converting the hidden state of the next word into a word probability distribution using a normalized exponential function; According to the word probability distribution, the word with the highest probability is selected and added to the historical text output sequence; returning to execute the step of obtaining the historical text output sequence, using a Transformer decoder to generate the hidden state of the next word according to the training task fusion vector and the historical text output sequence; Until the stopping condition is met, the historical text output sequence is output as synthetic text data for training.

6. The speech processing method according to claim 2, characterized in that: The step of using the speech decoder back network to process the training task fusion vector and outputting synthetic speech data for training includes: Obtaining an embedding vector of a target speaking object, wherein the target speaking object refers to a speaking object included in the training speech data; Combining the embedding vector with the training task fusion vector to obtain a first combined vector; Using the first linear layer in the network after the speech decoder, combined with a predefined activation function, linearly transform the first combined vector to obtain a second combined vector; Using the second linear layer in the network after the speech decoder, upsampling the second combined vector to obtain a third combined vector; Obtaining a historical speech output sequence, and using a Transformer decoder to generate a hidden state of a next speech word according to the third combined vector and the historical speech output sequence; Using the third linear layer in the network after the speech decoder, predicting an initial Mel-spectrogram according to the hidden state of the next speech word; Using the convolutional layer in the network after the speech decoder, generating a residual vector according to the initial Mel-spectrogram, and adding the residual vector to the initial Mel-spectrogram to obtain a target Mel-spectrogram; Using a fourth linear layer in the network after the speech decoder, converting the hidden state of the next speech word into a scalar value to predict a stop speech word flag; Speech synthesis is performed according to the target mel-spectrogram, and the speech synthesis process is terminated according to the stop speech word identifier to obtain synthesized speech data for training.

7. The speech processing method according to claim 2, characterized in that: The target loss function includes a first loss function and a second loss function. The step of performing model training on the pre-trained speech recognition and synthesis model to obtain a trained speech recognition and synthesis model includes: Using the first loss function, performing loss calculation on the synthetic text data for training and the synthetic speech data for training to obtain a first loss function value for a speech recognition task; Using the second loss function, performing loss calculation on the synthetic text data for training and the synthetic speech data for training to obtain a second loss function value for the speech synthesis task; Combining the first loss function value and the second loss function value to obtain a total loss value, and normalizing the total loss value to obtain a normalized total loss value; The gradient is calculated according to the normalized total loss value, and the parameters of the pre-trained speech recognition and synthesis model are updated according to the gradient. The pre-trained speech recognition and synthesis model is iteratively trained until the model converges to obtain a trained speech recognition and synthesis model.

8. A speech processing device, characterized in that: The device comprises: A data acquisition module, used to acquire text data and voice data to be processed; A model input module, used for inputting the to-be-processed text data and the to-be-processed speech data into a pre-trained speech recognition and synthesis model, wherein the speech recognition and synthesis model comprises a text encoder pre-network, a speech encoder pre-network, a text decoder post-network and a speech decoder post-network; A text extraction module, used for performing feature extraction on the text data to be processed by using the text encoder pre-network to obtain a target text feature vector; A speech extraction module, used for extracting features of the speech data to be processed by using the speech encoder pre-network to obtain a target speech feature vector; A vector combining module, used to obtain a predefined task vector, combine the target text feature vector, the target speech feature vector and the task vector, and obtain a target task fusion vector; A first synthesis module, used to process the target task fusion vector using the text decoder post-network and output target synthesized text data; The second synthesis module is used to process the target task fusion vector using the post-speech decoder network and output target synthesized speech data.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the speech processing method according to any one of claims 1 to 7 when executing the computer-readable instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the speech processing method according to any one of claims 1 to 7 are implemented.