Speech processing model training method, speech processing method, and speech translation method
By training the speech processing model with training data for speech processing tasks and sub-tasks, the problem of inaccuracy of neural network models in speech processing is solved, and more accurate speech processing results are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ALIBABA (CHINA) CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-29
Smart Images

Figure CN122116879A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and in particular to a method for training a speech processing model. One or more embodiments of this specification also relate to a method for training a speech translation model, a speech processing method, a speech translation method, a computing device, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] With the continuous development of artificial intelligence technology, neural network models are being applied to various scenarios to perform tasks. For example, neural network models are being applied to speech processing scenarios to perform speech processing tasks.
[0003] Current neural network models often encounter inaccuracies in speech processing results due to the complexity of speech processing tasks and the inability of neural network models to accurately handle sub-tasks. Therefore, improving the accuracy of speech processing results from neural network models has become an urgent problem to be solved. Summary of the Invention
[0004] In view of the above, embodiments of this specification provide a method for training a speech processing model. One or more embodiments of this specification also relate to a speech translation model training method, a speech processing method, a speech translation method, a speech processing model training device, a speech translation model training device, a speech processing device, a speech translation device, a computing device, an electronic device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a method for training a speech processing model is provided, comprising:
[0006] Determine the first speech training data corresponding to the speech processing task and the second speech training data corresponding to the speech processing subtask, wherein the speech processing subtask is a subtask of the speech processing task;
[0007] Based on the second speech training data, the speech processing network layer in the initial speech processing model is trained to obtain the trained initial speech processing model, wherein the speech processing network layer is related to the speech processing subtask.
[0008] Based on the first speech training data, the trained initial speech processing model is trained to obtain the target speech processing model.
[0009] According to a second aspect of the embodiments of this specification, a speech processing model training apparatus is provided, comprising:
[0010] The data determination module is configured to determine the first speech training data corresponding to the speech processing task and the second speech training data corresponding to the speech processing subtask, wherein the speech processing subtask is a subtask of the speech processing task.
[0011] The first training module is configured to train the speech processing network layer in the initial speech processing model based on the second speech training data to obtain the trained initial speech processing model, wherein the speech processing network layer is related to the speech processing subtask.
[0012] The second training module is configured to train the initial speech processing model after training based on the first speech training data to obtain the target speech processing model.
[0013] According to a third aspect of the embodiments of this specification, a method for training a speech translation model is provided, comprising:
[0014] Determine the first speech training data corresponding to the speech translation task and the second speech training data corresponding to the speech recognition task, wherein the speech recognition task is a subtask of the speech translation task;
[0015] Based on the second speech training data, the speech processing network layer in the initial speech translation model is trained to obtain the trained initial speech translation model, wherein the speech processing network layer is related to the speech recognition task;
[0016] Based on the first speech training data, the trained initial speech translation model is trained to obtain the target speech translation model.
[0017] According to a fourth aspect of the embodiments of this specification, a speech translation model training apparatus is provided, comprising:
[0018] The data determination module is configured to determine first speech training data corresponding to the speech translation task and second speech training data corresponding to the speech recognition task, wherein the speech recognition task is a subtask of the speech translation task;
[0019] The first training module is configured to train the speech processing network layer in the initial speech translation model based on the second speech training data to obtain the trained initial speech translation model, wherein the speech processing network layer is related to the speech recognition task.
[0020] The second training module is configured to train the initial speech translation model based on the first speech training data to obtain the target speech translation model.
[0021] According to a fifth aspect of the embodiments of this specification, a speech processing method is provided, comprising:
[0022] Identify the speech to be processed;
[0023] The speech to be processed is input into a target speech processing model to obtain the speech processing result corresponding to the speech to be processed. The target speech processing model is obtained by training an initial speech processing model based on a first speech training data. The initial speech processing model is obtained by training the speech processing network layer in the initial speech processing model based on a second speech training data. The first speech training data corresponds to the speech processing task, and the second speech training data corresponds to the speech processing sub-task. The speech processing sub-task is a sub-task of the speech processing task, and the speech processing network layer is related to the speech processing sub-task.
[0024] According to a sixth aspect of the embodiments of this specification, a voice processing apparatus is provided, comprising:
[0025] The voice determination module is configured to determine the voice to be processed;
[0026] A speech processing module is configured to input the speech to be processed into a target speech processing model to obtain a speech processing result corresponding to the speech to be processed. The target speech processing model is obtained by training an initial speech processing model based on first speech training data. The initial speech processing model is obtained by training the speech processing network layer in the initial speech processing model based on second speech training data. The first speech training data corresponds to a speech processing task, and the second speech training data corresponds to a speech processing subtask. The speech processing subtask is a subtask of the speech processing task, and the speech processing network layer is related to the speech processing subtask.
[0027] According to a seventh aspect of the embodiments of this specification, a speech translation method is provided, applied to a cloud-side device, comprising:
[0028] The voice to be translated sent by the receiving device;
[0029] The speech to be translated is input into the target speech translation model to obtain the speech translation result corresponding to the speech to be translated. The target speech translation model is obtained by training an initial speech translation model based on a first speech training data. The initial speech translation model is obtained by training the speech processing network layer in the initial speech translation model based on a second speech training data. The first speech training data corresponds to the speech translation task, and the second speech training data corresponds to the speech recognition task. The speech recognition task is a subtask of the speech translation task, and the speech processing network layer is related to the speech recognition task.
[0030] The speech translation result is sent to the terminal device.
[0031] According to an eighth aspect of the embodiments of this specification, a voice translation device is provided, applied to a cloud-based device, comprising:
[0032] The voice receiving module is configured to receive the voice to be translated sent by the receiving device.
[0033] A speech processing module is configured to input the speech to be translated into a target speech translation model to obtain a speech translation result corresponding to the speech to be translated. The target speech translation model is obtained by training an initial speech translation model based on first speech training data. The initial speech translation model is obtained by training the speech processing network layer in the initial speech translation model based on second speech training data. The first speech training data corresponds to the speech translation task, the second speech training data corresponds to the speech recognition task, the speech recognition task is a subtask of the speech translation task, and the speech processing network layer is related to the speech recognition task.
[0034] The voice translation result sending module is configured to send the voice translation result to the end-side device.
[0035] According to a ninth aspect of the embodiments of this specification, a computing device is provided, comprising:
[0036] Memory and processor;
[0037] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of any of the above methods.
[0038] According to a tenth aspect of the embodiments of this specification, an electronic device is provided, comprising:
[0039] A memory and a processor, the memory and the processor being connected via a bus;
[0040] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of any of the above methods.
[0041] According to an eleventh aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.
[0042] According to a twelfth aspect of an embodiment of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.
[0043] This specification provides one or more embodiments of a speech processing model training method. Considering the inaccuracy of speech processing results output by neural network models, this method, during the training of the speech processing model, determines a first speech training data corresponding to a large number of complex speech processing tasks and a second speech training data corresponding to speech processing sub-tasks. The speech processing sub-tasks are sub-tasks of the speech processing task. The speech processing model is trained using the second speech training data and the first speech training data, thereby obtaining a target speech processing model that can accurately perform fine processing on speech processing tasks containing sub-tasks and output accurate speech processing results. This improves the accuracy of the speech processing results of the neural network model and avoids the problem of inaccurate speech processing results due to the complexity of the speech processing task. Attached Figure Description
[0044] Figure 1 This is a schematic diagram of a speech processing technology solution provided in one embodiment of this specification;
[0045] Figure 2 This is a schematic diagram illustrating the application of a speech processing model training method provided in one embodiment of this specification;
[0046] Figure 3 This is a flowchart illustrating a speech processing model training method provided in one embodiment of this specification;
[0047] Figure 4 This is a schematic diagram of the processing procedure of a speech processing model training method provided in one embodiment of this specification;
[0048] Figure 5 This is a schematic diagram of an encoder output provided in one embodiment of this specification;
[0049] Figure 6This is a schematic diagram of the Audio Encoder output in a speech processing model training method provided in one embodiment of this specification;
[0050] Figure 7 This is a schematic diagram of the task indicator results in a speech processing model training method provided in one embodiment of this specification;
[0051] Figure 8 This is a flowchart illustrating a speech translation model training method provided in one embodiment of this specification;
[0052] Figure 9 This is a flowchart illustrating a speech processing method provided in one embodiment of this specification;
[0053] Figure 10 This is a flowchart illustrating a speech translation method provided in one embodiment of this specification;
[0054] Figure 11 This is a structural block diagram of a computing device provided in one embodiment of this specification;
[0055] Figure 12 This is a structural block diagram of an electronic device provided in one embodiment of this specification. Detailed Implementation
[0056] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0057] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0058] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0059] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0060] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.
[0061] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0062] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0063] S2TT (Speech to Text Translation): refers to the translation of speech (src) into text.
[0064] MT (Machine Translation): refers to machine translation.
[0065] Cross-modal: refers to cross-modal communication.
[0066] End2End Audio Translation LLM (End2End Audio Translation Large Language Model): refers to an end-to-end large-scale audio translation model, which is further trained from LLM to obtain an end-to-end audio translation model focused on translation functions.
[0067] LoRA (Low-Rank Adaptation of Large Language Models) is a technique for effectively fine-tuning large language models.
[0068] ASR (Automatic Speech Recognition) refers to speech recognition, which can convert human speech into computer-readable input.
[0069] Input Projector: An input projector is used to transform input data (such as text, images, or other types of signals) into a format and dimension suitable for subsequent processing. In multimodal systems, data from different modalities (such as text, images, and audio) typically have different feature representations. The input projector can help map these different modalities of data to a common feature space (usually text feature tokens), enabling the model to fuse and process them more effectively.
[0070] Audio Encoder: A speech encoder used to process and extract features from audio signals. It transforms the raw audio waveform into a high-level feature representation, which typically captures more important information in the audio, such as pitch, rhythm, and timbre. Audio encoders play a crucial role in large multimodal models, enabling the model to understand and utilize the meaning of audio data within its context.
[0071] ACC (Accuracy): refers to the accuracy of the model, that is, the proportion of samples correctly predicted by the model out of the total number of samples.
[0072] BLEU (Bilingual Evaluation Understudy) is a metric used to evaluate the quality of machine translation.
[0073] src: refers to the address or source of the audio file, used to specify the loading path of the audio file.
[0074] Stable Diffusion: A deep learning model for generating detailed images conditioned on textual descriptions.
[0075] With the continuous development of artificial intelligence technology, neural network models are being applied to various scenarios to perform tasks, such as speech processing. However, due to the complexity of speech data, current neural network models often produce inaccurate results during speech processing. Therefore, improving the accuracy of speech processing results from neural network models has become a pressing issue.
[0076] For example, current speech translation technology is mainly based on cascaded speech translation technology, where speech is first recognized into text by an ASR module, and then the text is translated into a target text by a translation model. This cascaded technology suffers from high latency, error propagation, and poor controllability during speech translation.
[0077] To address the aforementioned problems, this specification provides a speech processing technical solution. Figure 1 This is a schematic diagram of a speech processing technology solution provided in one embodiment of this specification, based on... Figure 1 It can be seen that this scheme freezes the audio-conditional large language model (Audio-Conditioned LLM) but does not freeze the audio encoder; and uses speech as training data. Figure 1 The audio encoder is trained using the phrase "thank you" from the example; multi-task instruction tuning is employed during training. However, this approach suffers from training instability and poor modal alignment.
[0078] Based on this, this specification provides a speech processing model training method. One or more embodiments of this specification also relate to a speech translation model training method, a speech processing method, a speech translation method, a speech processing model training device, a speech translation model training device, a speech processing device, a speech translation device, a computing device, an electronic device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0079] See Figure 2 , Figure 2This diagram illustrates an application of a speech processing model training method according to an embodiment of this specification. The speech processing model training method includes:
[0080] Determine the first speech training data corresponding to the speech processing task and the second speech training data corresponding to the speech processing subtask, wherein the speech processing subtask is a subtask of the speech processing task;
[0081] Based on the second speech training data, the speech processing network layer in the initial speech processing model is trained to obtain the trained initial speech processing model, wherein the speech processing network layer is related to the speech processing subtask.
[0082] Based on the first speech training data, the trained initial speech processing model is trained to obtain the target speech processing model.
[0083] Specifically, this method takes into account the inaccuracy of speech processing results output by neural network models. Therefore, during the training of the speech processing model, a large number of complex speech processing tasks are identified as first speech training data and speech processing sub-tasks as second speech training data. By training the model with the training data of speech processing tasks and speech processing sub-tasks, the model can be equipped with the ability to perform refined processing of speech processing sub-tasks and the overall processing capability of speech processing tasks, thereby improving the performance of the model.
[0084] Based on this, during model training, the speech processing network layer in the initial speech processing model is trained using the second speech training data to obtain the trained initial speech processing model. The speech processing network layer is related to the speech processing sub-task. This trained initial speech processing model has the ability to accurately process the speech processing sub-task. Then, the trained initial speech processing model is trained using the first speech training data to obtain the target speech processing model. This results in a target speech processing model capable of accurately and precisely processing speech processing tasks containing sub-tasks, improving the accuracy of the speech processing results from the neural network model and avoiding the problem of inaccurate speech processing results due to the complexity of the speech data.
[0085] For example, to meet users' speech translation needs, training data (i.e., first speech training data and second speech training data) can be determined for the large speech translation model (i.e., the speech processing model). Specifically, the first speech training data corresponding to the speech translation task and the second speech training data corresponding to the speech recognition task are determined; the Audio Encoder and / or Input Projector (i.e., the speech processing network layer) in the large speech translation model are trained using the second speech training data to obtain a pre-trained large speech translation model (i.e., the initial speech processing model after training); then, the pre-trained large speech translation model is trained using the first speech training data corresponding to the speech translation task to obtain the trained large speech translation model (i.e., the target speech processing model).
[0086] After training the large-scale speech translation model, the server 104 can input the speech data sent by the client 102 into the large-scale speech translation model for translation, obtain the translated text, and send the translated text back to the client 102, thereby meeting the user's speech translation needs.
[0087] See Figure 3 , Figure 3 A flowchart of a speech processing model training method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0088] Step 302: Determine the first speech training data corresponding to the speech processing task and the second speech training data corresponding to the speech processing subtask, wherein the speech processing subtask is a subtask of the speech processing task.
[0089] Among them, speech processing tasks can be understood as tasks that process speech data. These speech processing tasks can be executed by the target speech processing model. These speech processing tasks include, but are not limited to, speech translation tasks and speech conversion operations. Speech translation tasks can be understood as translating speech data in one language into text data in another language; for example, translating Chinese speech data into English text.
[0090] The speech processing task can be understood as a subtask that needs to be performed during the execution of the speech processing task; in the case of a speech translation task, the speech processing subtask can be a speech recognition task; the speech recognition task refers to the task of recognizing the speech content contained in the speech data; for example, recognizing the language content contained in Chinese speech data, thereby obtaining the language text corresponding to the Chinese speech data.
[0091] The first and second speech training data can be understood as the training data for model training.
[0092] Step 304: Based on the second speech training data, train the speech processing network layer in the initial speech processing model to obtain the trained initial speech processing model, wherein the speech processing network layer is related to the speech processing subtask.
[0093] A speech processing model can be understood as a model used to process speech data. This speech processing model can be a speech translation model, a speech recognition model, or a large language model (LLM); for example, it can be an end-to-end speech translation model. The initial speech processing model is the one that needs to be trained using second speech training data.
[0094] A speech processing network layer can be understood as a network layer in the initial speech processing model used to perform speech processing sub-tasks. In the initial speech processing model architecture, different network layers can perform different operations; this speech processing network layer is the network layer used to perform speech processing sub-tasks. For example, this speech processing network layer can be a speech encoder (AudioEncoder) and / or an input projector.
[0095] In one or more embodiments provided in this specification, the second speech training data includes a second speech training sample, a second speech sample label corresponding to the second speech training sample, and second speech processing prompt information corresponding to the speech processing subtask.
[0096] The step of training the speech processing network layer in the initial speech processing model based on the second speech training data to obtain the trained initial speech processing model includes:
[0097] The speech processing network layer is determined from the initial speech processing model;
[0098] The second speech training sample and the second speech processing prompt information are input into the initial speech processing model for speech processing to obtain the initial speech processing result, wherein the initial speech processing result is related to the speech processing subtask;
[0099] Based on the initial speech processing results and the second speech sample labels, the parameters of the speech processing network layer are adjusted to obtain the trained initial speech processing model.
[0100] The second speech training sample can be speech data used as training samples. This second speech training sample can be human voice, conference voice, song, etc., without any specific restrictions.
[0101] The second speech sample label can be understood as a sample label associated with the speech processing subtask; when the speech processing subtask is a speech recognition task, the second speech sample label can be the real speech text corresponding to the speech training sample, for example, the language text corresponding to the Chinese speech data in the above embodiment.
[0102] The second speech processing prompt can be understood as a prompt associated with the speech processing subtask; when the speech processing subtask is a speech recognition task, the second speech processing prompt can be "speech transcription", "speech recognition", "speech transcription", etc.
[0103] Specifically, this method can determine the speech processing network layer from the initial speech processing model, and input the second speech training sample and the second speech processing prompt information into the initial speech processing model for speech processing to obtain the initial speech processing result. The initial speech processing result is related to the speech processing sub-task; for example, the language processing sub-task can be a speech recognition task, and the initial speech processing result is a speech recognition result.
[0104] A loss function is calculated based on the initial speech processing results and the labels of the second speech samples; for example, this loss function could be the cross-entropy loss function. Then, this loss function is used to adjust the parameters of the speech processing network layer in the initial speech processing model, thereby obtaining the trained initial speech processing model.
[0105] In the above embodiments, before training the initial speech processing model, the initial speech processing model is pre-trained using the second speech training data corresponding to the speech processing sub-task, thereby improving the performance of the language processing model and avoiding the problem of unstable training.
[0106] In one or more embodiments provided in this specification, the speech processing network layer includes a speech feature processing network layer;
[0107] Determining the speech processing network layer from the initial speech processing model includes:
[0108] Based on the second speech training data, the speech feature processing network layer is determined from the initial speech processing model. This speech feature processing network layer performs feature transformation on the initial speech recognition features output by the speech recognition network layer in the initial speech processing model to obtain target speech recognition features. These target speech recognition features are then input into the speech translation network layer in the initial speech processing model for speech translation, resulting in a speech processing outcome.
[0109] The speech feature processing network layer can be understood as a network layer used to process speech characteristics. For example, the speech feature processing network layer can be an Input Projector.
[0110] The speech recognition network layer can be understood as a network layer used to process speech recognition; the speech recognition network layer is used to perform speech recognition on input data (such as first speech training data, second speech training data, and speech to be processed) to obtain initial speech recognition features.
[0111] A speech translation network layer can be understood as a network layer used for speech translation processing. This layer translates input data (e.g., target speech recognition features, or target speech recognition features and speech processing prompts) to obtain a speech processing result or an initial speech processing result. The speech processing prompts include first or second prompts. For example, this speech translation network layer is an LLM (Local Level Management) layer.
[0112] In one or more embodiments provided in this specification, the step of inputting the second speech training sample and the second speech processing prompt information into the initial speech processing model for speech processing to obtain the initial speech processing result includes:
[0113] The second speech training sample and the second speech processing prompt information are input into the initial speech processing model, wherein the initial speech processing model includes a speech recognition network layer, a speech feature processing network layer and a speech translation network layer;
[0114] The speech recognition network layer is used to perform speech recognition on the second speech training sample to obtain initial speech recognition features;
[0115] The initial speech recognition features are transformed using the speech feature processing network layer to obtain target speech recognition features, wherein the modality of the target speech recognition features is consistent with the modality of the input data of the speech translation network layer.
[0116] Using a speech translation network layer, the target speech recognition features are translated based on the second speech processing prompt information to obtain the initial speech processing result.
[0117] Specifically, this method can identify the speech feature processing network layer in the initial speech processing model as the speech processing network layer whose parameters are to be adjusted, and perform the first stage of model training on the speech processing network layer.
[0118] The second speech training sample and the second speech processing prompt information are input into the initial speech processing model, wherein the initial speech processing model includes a speech recognition network layer, a speech feature processing network layer and a speech translation network layer;
[0119] The speech recognition network layer is used to perform speech recognition on the second speech training sample to obtain initial speech recognition features; the speech feature processing network layer is used to perform feature transformation on the initial speech recognition features to obtain target speech recognition features; and the speech translation network layer is used to perform speech translation on the target speech recognition features based on the second speech processing prompt information to obtain the initial speech processing result.
[0120] Then, based on the initial speech processing results and the second speech sample labels, a loss function is calculated, and all or some of the model parameters in the speech feature processing network layer are adjusted based on the loss function to obtain the trained initial speech processing model.
[0121] The speech recognition network layer can be an Audio Encoder; the speech translation network layer can be an LLM.
[0122] Taking the speech processing model training method in one or more embodiments of this specification as an example in a speech translation scenario, the speech processing model training method is explained below. The speech processing model is an end-to-end speech translation model, the speech processing network layer can be an Input Projector, and the second speech training data is the speech training data corresponding to the speech recognition task. Based on this, this method can use the speech training data to train the end-to-end speech translation model. The model training is divided into three stages. By performing alignment training on the end-to-end speech translation model through these three stages, a trained end-to-end speech translation model can be obtained. The first stage uses the speech training data to train the input projector in the end-to-end speech translation model. Specifically:
[0123] In the first stage, the parameters of the Input Projector can be enabled for training, while all other model parameters are frozen. The audio signal (the first speech training sample) is input into the Audio Encoder for speech encoding to obtain the speech feature encoding (i.e., the initial speech recognition features) corresponding to the audio signal.
[0124] The speech feature encoding is input into the Input Projector for projection processing, thereby obtaining the audio representation (i.e., the target speech feature encoding) through the Input Projector;
[0125] The audio representation and the text prompt's embedding representation (i.e., the second speech processing prompt information) are concatenated together to obtain the model input data for the first stage, and this model input data is then fed into the Large Speech Model (LLM).
[0126] Using a large speech model, speech transcription (speech recognition) is performed on the audio representation based on the Prompt, thereby obtaining the text data corresponding to the audio signal. This text data records the language information contained in the audio signal, such as "thank you" and other language information.
[0127] The training is performed using the cross-entropy loss function between the prediction results (text data) of the large speech model and the correct answer (label of the second speech sample). During training, only the parameters of the InputProjector are updated with gradients, while the parameters of other network layers in the language processing model are frozen.
[0128] This method pre-warms the Input Projector parameters, avoiding the instability caused by training a randomly initialized Input Projector and a well-pretrained Encoder / LLM together in the initial stage. Here, the well-pretrained Encoder refers to a pre-trained Audio Encoder; this training instability includes, but is not limited to, the model failing to converge and exhibiting hallucinations.
[0129] In one or more embodiments provided in this specification, the speech processing network layer includes a speech feature processing network layer and a target speech recognition network layer in the initial speech processing model;
[0130] Determining the speech processing network layer from the initial speech processing model includes:
[0131] Based on the second speech training data, the speech feature processing network layer and multiple initial speech recognition network layers are determined from the initial speech processing model, wherein the multiple initial speech recognition network layers are used to perform speech recognition on the speech training samples to obtain initial speech recognition features;
[0132] Determine the positional relationship between each initial speech recognition network layer and the speech feature processing network layer;
[0133] Based on the positional relationship, a target speech recognition network layer corresponding to the speech feature processing network layer is determined from the plurality of initial speech recognition network layers, wherein the target speech recognition network layer is some or all of the network layers in the plurality of initial speech recognition network layers.
[0134] The target speech recognition network layer can be a network layer used to perform speech processing tasks; for example, the target speech recognition network layer can be an Audio Encoder.
[0135] Specifically, during the model training process using the second speech training data, this method can also determine the speech feature processing network layer and multiple initial speech recognition network layers from the initial speech processing model, and determine the positional relationship between each initial speech recognition network layer and the speech feature processing network layer; wherein, the positional relationship can be determined based on the number of network layers corresponding to the initial speech recognition network layer and the speech feature processing network layer in the initial speech processing model.
[0136] Based on their positional relationships, a predetermined number of target speech recognition network layers are determined from multiple initial speech recognition network layers. These target layers are located close to the speech feature processing network layer. The number of network layers between the target layer and the speech feature processing network layer can be less than the number of network layers between other initial speech recognition network layers and the speech feature processing network layer. These other initial speech recognition network layers refer to all initial speech recognition network layers other than the target layer. In other words, the target layer is closer to the speech feature processing network layer than other initial speech recognition network layers.
[0137] In the above embodiments, by selecting the Audio Encoder and Input Projector modules for model training, the trained Audio Encoder and Input Projector modules are more focused on speech and text modality alignment, thereby improving the modality alignment effect. This makes it easier for the subsequent LLM large model to obtain accurate input data, thereby improving the speech translation accuracy of the LLM large model.
[0138] In one or more embodiments provided in this specification, the step of inputting the second speech training sample and the second speech processing prompt information into the initial speech processing model for speech processing to obtain the initial speech processing result includes:
[0139] The second speech training sample and the second speech processing prompt information are input into the initial speech processing model, wherein the initial speech processing model includes a speech recognition network layer, a speech feature processing network layer and a speech translation network layer;
[0140] The speech recognition network layer is used to perform speech recognition on the second speech training sample to obtain initial speech recognition features;
[0141] The initial speech recognition features are transformed using the speech feature processing network layer to obtain target speech recognition features, wherein the modality of the target speech recognition features is consistent with the modality of the input data of the speech translation network layer.
[0142] Using a speech translation network layer, the target speech recognition features are translated based on the second speech processing prompt information to obtain the initial speech processing result.
[0143] Specifically, after determining the speech feature processing network layer and the target speech recognition network layer in the initial speech processing model as the speech processing network layers whose parameters are to be adjusted, the first stage of model training is performed on the speech processing network layers.
[0144] The second speech training sample and the second speech processing prompt information are input into the initial speech processing model;
[0145] The speech recognition network layer is used to perform speech recognition on the second speech training sample to obtain initial speech recognition features;
[0146] The initial speech recognition features are transformed using the speech feature processing network layer to obtain target speech recognition features, wherein the modality of the target speech recognition features is consistent with the modality of the input data of the speech translation network layer; for example, the input data of the speech translation network layer is a token, and the target speech recognition features are also tokens.
[0147] Using a speech translation network layer, the target speech recognition features are translated based on the second speech processing prompt information to obtain the initial speech processing result;
[0148] Based on the initial speech processing results and the second speech sample labels, a loss function is calculated, and all or some of the model parameters in the speech feature processing network layer and the speech recognition network layer are adjusted based on the loss function to obtain the trained initial speech processing model.
[0149] Following the previous example, in the second stage of model training, while training with the Input Projector parameters enabled, the Audio Encoder parameters are gradually enabled. For example, the last few network layers of the Audio encoder can be enabled first, and then all network layers of the Audio encoder can be enabled.
[0150] It should be noted that in the second stage of model training, the input and output of the end-to-end speech translation model are the same in the first node. The purpose of using the same input and output is to avoid the problem of training instability caused by turning on all the parameters of the entire Audio Encoder at once, which would lead to the model getting stuck in local optima.
[0151] Furthermore, the two stages mentioned above only use ASR speech recognition task data for training, not translation task data. The prompt is also only the speech recognition task prompt, such as "Speech Transcription:". Because the speech processing model architecture of this method uses the "Input Projector" module to convert the output data of the "Audio Encoder" into the input data (token) of the "LLM", cross-modal alignment between the "Audio Encoder" and the "LLM" is achieved. Therefore, through the training of the above two stages, the Audio Encoder and Input Projector modules focus more on speech and text modal alignment, thereby improving the modal alignment effect. This facilitates the subsequent large LLM model obtaining accurate input data, thus improving the speech translation accuracy of the large LLM model.
[0152] Step 306: Based on the first speech training data, train the trained initial speech processing model to obtain the target speech processing model.
[0153] In one or more embodiments provided in this specification, the speech processing task is a speech translation task, the speech processing subtask is a speech recognition task, the speech recognition task is a subtask of the speech translation task, and the speech processing model is a speech translation model;
[0154] The step of training the initial speech processing model based on the first speech training data to obtain the target speech processing model includes:
[0155] Based on the first speech training data, the trained initial speech translation model is trained to obtain the target speech translation model.
[0156] The explanation of this embodiment can be found in the corresponding or relevant explanations in one or more of the above embodiments, and will not be repeated here.
[0157] In one or more embodiments provided in this specification, the step of training the initial speech processing model based on the first speech training data to obtain the target speech processing model includes:
[0158] Using the first speech training data and the second speech training data, the target model parameters in the trained initial speech processing model are adjusted to obtain the target speech processing model, wherein the target model parameters are all or part of the model parameters of the trained initial speech processing model.
[0159] Following the previous example, during model training, in addition to enabling the Input Projector and Encoder (i.e., AudioEncoder), the model parameters of LLM are enabled for training, and both speech recognition and speech translation tasks are trained simultaneously. The model input and output formats are similar to the previous two stages, the difference being that the speech translation task requires a corresponding prompt, such as "Translate into Chinese:". Enabling LLM in this stage allows it to better leverage its advantages in cross-language instruction understanding and generation.
[0160] In one or more embodiments provided in this specification, the first speech training data includes a first speech training sample, a first speech sample label corresponding to the first speech training sample, and a first speech processing prompt information corresponding to the speech processing task; the second speech training data includes a second speech training sample, a second speech sample label corresponding to the second speech training sample, and a second speech processing prompt information corresponding to the speech processing subtask.
[0161] The step of adjusting the target model parameters in the trained initial speech processing model using the first speech training data and the second speech training data to obtain the target speech processing model includes:
[0162] The first speech training sample and the first speech processing prompt information are input into the trained initial speech processing model for speech processing to obtain a first speech processing result, wherein the first speech processing result is related to the speech processing task.
[0163] The second speech training sample and the second speech processing prompt information are input into the trained initial speech processing model for speech processing to obtain the second speech processing result, wherein the second speech processing result is related to the speech processing subtask.
[0164] Based on the first speech processing result, the second speech processing result, the first speech sample label, and the second speech sample label, the target model parameters of the trained initial speech processing model are adjusted to obtain the target speech processing model.
[0165] The first speech training sample can be speech data used as training samples, such as human voice, conference recordings, songs, etc., without specific limitations. It should be noted that the first speech training sample and the second speech training sample can be the same or different.
[0166] The first speech sample label can be understood as a sample label associated with a speech processing task; in the case that the speech processing task is a speech translation task, the first speech sample label can be the real speech translation corresponding to the first speech training sample, for example, the English translation corresponding to the Chinese speech data in the above embodiment.
[0167] The first voice processing prompt can be understood as a prompt associated with the voice processing task; in the case that the voice processing task is a voice translation task, the first voice processing prompt can be "voice translation", "Translate into Chinese:", "Translate into English:", etc.
[0168] Specifically, firstly, inputting the first speech training sample and the first speech processing prompt information into the trained initial speech processing model for speech processing to obtain the first speech processing result includes:
[0169] The first speech training sample and the first speech processing prompt information are input into the trained initial speech processing model.
[0170] The speech recognition network layer is used to perform speech recognition on the first speech training sample to obtain initial speech recognition features;
[0171] The initial speech recognition features are transformed using the speech feature processing network layer to obtain target speech recognition features, wherein the modality of the target speech recognition features is consistent with the modality of the input data of the speech translation network layer.
[0172] Using a speech translation network layer, the target speech recognition features are translated based on the first speech processing prompt information to obtain the first speech processing result.
[0173] Secondly, the step of inputting the second speech training sample and the second speech processing prompt information into the trained initial speech processing model for speech processing to obtain the second speech processing result includes:
[0174] The second speech training sample and the second speech processing prompt information are input into the trained initial speech processing model;
[0175] The speech recognition network layer is used to perform speech recognition on the second speech training sample to obtain initial speech recognition features;
[0176] The initial speech recognition features are transformed using the speech feature processing network layer to obtain target speech recognition features, wherein the modality of the target speech recognition features is consistent with the modality of the input data of the speech translation network layer.
[0177] Using a speech translation network layer, the target speech recognition features are translated based on the second speech processing prompt information to obtain the second speech processing result.
[0178] Finally, the step of adjusting the target model parameters of the trained initial speech processing model based on the first speech processing result, the second speech processing result, the first speech sample label, and the second speech sample label to obtain the target speech processing model includes:
[0179] Based on the first speech processing result, the second speech processing result, the first speech sample label, and the second speech sample label, a loss function (e.g., cross-entropy loss function) is calculated, and the loss function is used to adjust all or part of the model parameters of the trained initial speech processing model to obtain the target speech processing model.
[0180] Following the previous example, in the third stage of model training, based on enabling the parameters of the Audio Encoder and Input Projector, the LLM parameters are enabled using LoRa fine-tuning for training. Furthermore, both speech recognition and speech translation tasks are trained simultaneously. The input and output formats of the end-to-end speech translation model are similar to those in the first two stages, the difference being that the speech translation task requires a corresponding prompt, such as "Translate into Chinese." Specifically:
[0181] 1. In the third stage, the parameters of Input Projector, Audio Encoder, and LLM can be turned on for training; the audio signal (first speech training sample or second speech training sample) is input into Audio Encoder for speech encoding to obtain the speech feature code corresponding to the audio signal.
[0182] 2. Input the speech feature encoding into the Input Projector for projection processing, thereby obtaining the audio representation (i.e., the initial speech feature encoding) through the Input Projector.
[0183] 3. Concatenate the audio representation and the text prompt's embedding representation (i.e., the first speech processing prompt or the second speech processing prompt) together to obtain the model input data for the first stage, and then input this model input data into the Large Speech Model (LLM).
[0184] 4. Using a large speech model, the audio representation is transcribed (translated) based on the Prompt to obtain the text data corresponding to the audio signal. This text data records the language information contained in the audio signal, such as "thank you".
[0185] 5. The model is trained using the cross-entropy loss function between the prediction results (text data) and the correct answers (labels of the first and second speech samples) of the large speech model as the target. During the training process, the speech processing model is updated with gradients until the model training is complete.
[0186] In the third stage, by enabling LLM, we can better leverage its advantages in cross-language instruction understanding and generation, and obtain a model with better performance.
[0187] This specification provides one or more embodiments of a speech processing model training method. Considering the inaccuracy of speech processing results output by neural network models, this method, during the training of the speech processing model, determines a first speech training data corresponding to a large number of complex speech processing tasks and a second speech training data corresponding to speech processing sub-tasks. The speech processing sub-tasks are sub-tasks of the speech processing tasks. The speech processing model is trained using the second speech training data and the first speech training data, thereby obtaining a target speech processing model that can accurately perform fine processing on speech processing tasks containing sub-tasks, improving the accuracy of the speech processing results of the neural network model, and avoiding the problem of inaccurate speech processing results due to the complexity of the speech data.
[0188] The following is in conjunction with the appendix Figure 4 Taking the application of the speech processing model training method provided in this specification in speech translation as an example, the speech processing model training method will be further explained. Among other things, Figure 4 This document illustrates a schematic diagram of the processing procedure for a speech processing model training method provided in one embodiment of this specification.
[0189] based on Figure 4As can be seen, the structure of the end-to-end speech large model provided in this method includes: a pre-trained Audio Encoder and a pre-trained LLM; the Audio Encoder and LLM are connected through an Input Projector layer, and the output of the Input Projector and the prompt embedding are concatenated together as input to the LLM for processing.
[0190] This method divides the training process of the end-to-end speech translation model into three stages. By aligning and training the end-to-end speech translation model through these three stages, a fully trained end-to-end speech translation model can be obtained.
[0191] The first stage uses training data to train the input projector in the end-to-end speech translation model; the specific method is as follows:
[0192] 1. In the first stage, you can enable the parameters of the Input Projector for training, while freezing all other model parameters. Input the audio signal (speech training sample) into the Audio Encoder for speech encoding to obtain the speech feature encoding corresponding to the audio signal.
[0193] 2. Input the encoded speech features into the Input Projector for projection processing, thereby obtaining the audio representation (i.e., speech feature encoding) through the Input Projector.
[0194] 3. Concatenate the audio representation and the text prompt's embedding representation (i.e., the second speech processing prompt information) together to obtain the model input data for the first stage, and then input this model input data into the Large Speech Model (LLM).
[0195] 4. Using a large speech model, the audio representation is transcribed (speech recognition) based on the Prompt to obtain the text data corresponding to the audio signal. This text data records the language information contained in the audio signal, such as "thank you" and other language information.
[0196] 5. The training is performed using the cross-entropy loss function between the prediction results (text data) and the correct answer (second speech sample label) of the large speech model as the target. During the training process, only the parameters of the InputProjector are updated with gradients, while the parameters of other network layers in the language processing model are frozen.
[0197] This method pre-warms the Input Projector parameters, avoiding the instability caused by training a randomly initialized Input Projector and a well-pretrained Encoder / LLM together in the initial stage. Here, the well-pretrained Encoder refers to a pre-trained Audio Encoder; this training instability includes, but is not limited to, the model failing to converge and exhibiting hallucinations.
[0198] The second stage uses training data to train the input projector and audio encoder in the end-to-end speech translation model; the specific method is as follows:
[0199] In the second stage of model training, while training with the Input Projector parameters enabled, the Audio Encoder parameters are gradually enabled. For example, the last few network layers of the Audio Encoder can be enabled first, and then all network layers of the Audio Encoder can be enabled.
[0200] It should be noted that in the second stage of model training, the input and output of the end-to-end speech translation model are the same as in the first stage. The purpose of using the same input and output is to avoid the problem of training instability caused by turning on all the parameters of the entire Audio Encoder at once, which could lead to the model getting stuck in local optima.
[0201] Furthermore, the two stages mentioned above are trained using only ASR speech recognition task data, not translation task data, and the prompts are only from the speech recognition task, such as "Speech Transcription". This allows the Audio Encoder and Input Projector modules to focus more on speech and text modality alignment, thereby improving the modality alignment effect.
[0202] The third stage uses training data to train the input projector, audio encoder, and LLM in the end-to-end speech translation model; the specific method is as follows:
[0203] In the third stage of model training, in addition to enabling the parameters of Audio Encoder and Input Projector, the parameters of LLM can be enabled for training, and both speech recognition and speech translation tasks can be trained simultaneously.
[0204] It should be noted that in the third stage of model training, either LoRa fine-tuning can be used to enable some parameters of the LLM for training, or full fine-tuning can be used to enable all parameters of the LLM for training; no specific restrictions are made here. That is to say, in the third stage of model training, all or some parameters of the LLM can be enabled for training.
[0205] The training samples corresponding to the two tasks of speech recognition and speech translation are input into the end-to-end speech translation model for processing, and the cross-entropy loss function is calculated based on the model output and the real results (sample labels).
[0206] Based on this cross-entropy loss function, the parameters of the input projector, audio encoder, and LLM are adjusted until the training stopping condition is met, thus obtaining the trained end-to-end speech translation model.
[0207] It should be noted that the input and output formats of the end-to-end speech translation large model are similar to those of the first two stages. The difference is that the speech translation task requires a corresponding prompt, such as "Translate into Chinese:". In the third stage, by enabling LLM, its advantages in cross-language instruction understanding and generation can be better utilized.
[0208] Based on the above, the speech processing model training method in one or more embodiments of this specification provides a multi-stage cross-modal speech translation alignment method based on a large model. This method explores how to upgrade some classic cross-modal tasks from traditional cascaded systems to multi-modal large models in the context of a multi-modal large model.
[0209] This method achieves end-to-end speech translation in applications by performing cross-modal alignment of a large model in stages, enabling the model to directly translate speech into target text. Specifically, to leverage the powerful speech understanding capabilities of a pre-trained speech encoder and the text understanding and generation capabilities of a pre-trained LLM, this method first aligns the speech encoder and LLM modally, and then trains the speech translation task.
[0210] Figure 5 This is a schematic diagram of an encoder output provided in one embodiment of this specification; Figure 6 This is a schematic diagram of the Audio Encoder output in a speech processing model training method provided in one embodiment of this specification; based on Figure 5It can be seen that by visualizing the speech representations (encoder outputs) of different languages, it is found that before optimization, cross-language tasks forcibly aggregate the encoder output representations together, resulting in unclear encoder outputs and an inability to align them; based on Figure 6 It can be seen that after optimization using this method, the Audio Encoder has better differentiation for modal representations of different languages, and the output results are easier to align.
[0211] It should be noted that, regarding the issue of training instability caused by the random initialization of the Input Projector and its co-optimization with the AudioEncoder during model training, this method adopts a method of training the Input Projector first and then co-training, thus avoiding the problem of training instability. Furthermore, to address the issue that freezing LLM training for cross-language tasks injects translation capabilities into the Encoder, thereby weakening the modality alignment effect, this method overcomes the problem of weakened modality alignment by training S2TT using the LoRa method.
[0212] This method can be applied to modal alignment of speech, text, and images, helping to achieve better results in end-to-end multimodal tasks, such as speech translation, conference captioning, financial report presentations, AI translation, e-commerce live streaming translation, and translated film captioning. For example, the speech processing model in this solution can include an Input Projector, an Audio Encoder, and an LLM. In this model architecture, the Audio Encoder encodes the speech to be processed to obtain a speech representation. Using the "Input Projector" module, the output data (speech representation) of the "Audio Encoder" is converted into the input data (token) of the "LLM," and the "LLM" is used to generate text content such as conference captions, financial report presentations, AI translation, e-commerce live streaming translation, and translated film captioning, thereby achieving cross-modal alignment between the "Audio Encoder" and the "LLM" (i.e., cross-modal alignment between speech and text).
[0213] For example, the speech processing model in this solution may include an Input Projector, an Audio Encoder, and an image generation network layer (Stable Diffusion). In this model architecture, the Audio Encoder encodes the speech to be processed to obtain a speech representation. The output data (speech representation) of the Audio Encoder is converted into the input data (description text) of the Stable Diffusion through the "Input Projector" module. The Stable Diffusion then generates corresponding image data based on the description text, thereby achieving cross-modal alignment between the Audio Encoder and Stable Diffusion (i.e., cross-modal alignment between speech and image).
[0214] For task performance metrics results across various tasks, please refer to [link / reference]. Figure 7 ; Figure 7 This is a schematic diagram of the task indicator results in a speech processing model training method provided in one embodiment of this specification; Figure 7 In this context, A represents Baseline (ASR+S2TT), which refers to a basic model that has both ASR and S2TT capabilities. Figure 7 The "B" in the text indicates phased training (ASR first, then ASR+S2TT), which means training the model in stages. In this training process, the ASR module is trained first, and then the ASR+S2TT module is trained. Figure 7 In this context, "C" represents "staged training + LLM enabled," referring to a combination of staged training and LLM training strategies for model training (this method). Figure 7 It can be seen that there are differences in BLEU among the three training methods A, B, and C in the translation task; and differences in ACC among the three training methods A, B, and C in the speech recognition task; S2TT translation improves BLEU by 4 points and ASR accuracy by 8 points.
[0215] See Figure 8 , Figure 8 A flowchart of a speech translation model training method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0216] Step 802: Determine the first speech training data corresponding to the speech translation task and the second speech training data corresponding to the speech recognition task, wherein the speech recognition task is a subtask of the speech translation task;
[0217] Step 804: Based on the second speech training data, train the speech processing network layer in the initial speech translation model to obtain the trained initial speech translation model, wherein the speech processing network layer is related to the speech recognition task;
[0218] Step 806: Based on the first speech training data, train the trained initial speech translation model to obtain the target speech translation model.
[0219] This specification provides one or more embodiments of a speech translation model training method. Considering the inaccuracy of speech translation results output by neural network models, this method, during the training of the speech translation model, determines a large amount of first speech training data corresponding to a relatively complex speech translation task and a second speech training data corresponding to a speech recognition task, where the speech recognition task is a subtask of the speech translation task. The speech translation model is trained using the second speech training data and the first speech training data, thereby obtaining a target speech translation model that can accurately process speech translation tasks containing subtasks and output accurate speech translation results. This improves the accuracy of the speech translation results of the neural network model and avoids the problem of inaccurate speech translation results due to the complexity of the speech translation task processing.
[0220] The above is an illustrative scheme of a speech translation model training method according to this embodiment. It should be noted that the technical solution of this speech translation model training method belongs to the same concept as the technical solution of the speech processing model training method described above. Details not described in detail in the technical solution of the speech translation model training method can be found in the description of the technical solution of the speech processing model training method described above.
[0221] See Figure 9 , Figure 9 A flowchart of a speech processing method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0222] Step 902: Determine the speech to be processed;
[0223] Step 904: Input the speech to be processed into the target speech processing model to obtain the speech processing result corresponding to the speech to be processed. The target speech processing model is obtained by training the initial speech processing model based on the first speech training data. The initial speech processing model is obtained by training the speech processing network layer in the initial speech processing model based on the second speech training data. The first speech training data corresponds to the speech processing task, and the second speech training data corresponds to the speech processing sub-task. The speech processing sub-task is a sub-task of the speech processing task, and the speech processing network layer is related to the speech processing sub-task.
[0224] In one or more embodiments provided in this specification, the step of inputting the speech to be processed into a target speech processing model to obtain the speech processing result corresponding to the speech to be processed includes:
[0225] The speech to be processed is input into the target speech processing model, wherein the target speech processing model includes a speech recognition network layer, a speech feature processing network layer, and a speech translation network layer;
[0226] The speech recognition network layer is used to perform speech recognition on the speech to be processed to obtain initial speech recognition features;
[0227] The initial speech recognition features are transformed using the speech feature processing network layer to obtain target speech recognition features, wherein the modality of the target speech recognition features is consistent with the modality of the input data of the speech translation network layer.
[0228] The speech recognition features of the target speech are translated using a speech translation network layer to obtain the speech processing result.
[0229] For explanations of the above embodiments, please refer to the corresponding explanations in one or more embodiments of the above speech processing model training method, which will not be repeated here.
[0230] This specification provides one or more embodiments of a speech processing method. Considering the inaccuracy of speech processing results output by neural network models, this method, during the training of the speech processing model, determines a first speech training data corresponding to a large number of complex speech processing tasks and a second speech training data corresponding to speech processing sub-tasks. The speech processing sub-tasks are sub-tasks of the speech processing task. The speech processing model is trained using the second speech training data and the first speech training data to obtain a target speech processing model. This target speech processing model can accurately perform fine processing on speech processing tasks containing sub-tasks and output accurate speech processing results, thereby improving the accuracy of the speech processing results of the neural network model and avoiding the problem of inaccurate speech processing results due to the complexity of the speech processing task.
[0231] The above is an illustrative scheme of a speech processing method according to this embodiment. It should be noted that the technical solution of this speech processing method belongs to the same concept as the technical solution of the speech processing model training method described above. For details not described in detail in the technical solution of the speech processing method, please refer to the description of the technical solution of the speech processing model training method described above.
[0232] See Figure 10 , Figure 10 A flowchart of a speech translation method according to an embodiment of this specification is shown. The speech translation method is applied to a cloud-side device and specifically includes the following steps.
[0233] Step 1002: The receiving end device sends the voice to be translated;
[0234] Step 1004: Input the speech to be translated into the target speech translation model to obtain the speech translation result corresponding to the speech to be translated. The target speech translation model is obtained by training the initial speech translation model based on the first speech training data. The initial speech translation model is obtained by training the speech processing network layer in the initial speech translation model based on the second speech training data. The first speech training data corresponds to the speech translation task, the second speech training data corresponds to the speech recognition task, the speech recognition task is a subtask of the speech translation task, and the speech processing network layer is related to the speech recognition task.
[0235] Step 1006: Send the speech translation result to the terminal device.
[0236] In one or more embodiments provided in this specification, the cloud-side device can be a central cloud device in a distributed architecture or an edge cloud device in a distributed architecture. The cloud-side device can be a cloud-side device with a cloud desktop system or cloud desktop software installed and deployed, such as a cloud server or cloud host. The endpoint device can be understood as any terminal that interacts with the cloud-side device. This terminal can be a laptop, desktop computer, tablet, smart device, server, etc.
[0237] This specification provides one or more embodiments of a speech translation method applied to cloud-side devices. Considering the inaccuracy of speech translation results output by neural network models, this method, during the training of the speech translation model, determines a large amount of complex first speech training data corresponding to a speech translation task and a second speech training data corresponding to a speech recognition task, where the speech recognition task is a subtask of the speech translation task. The speech translation model is trained using the second speech training data and the first speech training data to obtain a target speech translation model. This target speech translation model can accurately perform fine processing on speech processing tasks containing subtasks and output accurate speech translation results, thereby improving the accuracy of the speech translation results of the neural network model and avoiding the problem of inaccurate speech translation results due to the complexity of the speech translation task processing.
[0238] The above is an illustrative scheme of a speech translation method according to this embodiment. It should be noted that the technical solution of this speech translation method belongs to the same concept as the technical solution of the speech processing model training method described above. For details not described in detail in the technical solution of the speech translation method, please refer to the description of the technical solution of the speech processing model training method described above.
[0239] Corresponding to the above method embodiments, this specification also provides an embodiment of a speech processing model training device, which includes:
[0240] The data determination module is configured to determine the first speech training data corresponding to the speech processing task and the second speech training data corresponding to the speech processing subtask, wherein the speech processing subtask is a subtask of the speech processing task.
[0241] The first training module is configured to train the speech processing network layer in the initial speech processing model based on the second speech training data to obtain the trained initial speech processing model, wherein the speech processing network layer is related to the speech processing subtask.
[0242] The second training module is configured to train the initial speech processing model after training based on the first speech training data to obtain the target speech processing model.
[0243] Optionally, the second speech training data includes a second speech training sample, a second speech sample label corresponding to the second speech training sample, and a second speech processing prompt information corresponding to the speech processing subtask;
[0244] The first training module is also configured as follows:
[0245] The speech processing network layer is determined from the initial speech processing model;
[0246] The second speech training sample and the second speech processing prompt information are input into the initial speech processing model for speech processing to obtain the initial speech processing result, wherein the initial speech processing result is related to the speech processing subtask;
[0247] Based on the initial speech processing results and the second speech sample labels, the parameters of the speech processing network layer are adjusted to obtain the trained initial speech processing model.
[0248] Optionally, the speech processing network layer includes a speech feature processing network layer;
[0249] The first training module is also configured as follows:
[0250] Based on the second speech training data, the speech feature processing network layer is determined from the initial speech processing model. This speech feature processing network layer performs feature transformation on the initial speech recognition features output by the speech recognition network layer in the initial speech processing model to obtain target speech recognition features. These target speech recognition features are then input into the speech translation network layer in the initial speech processing model for speech translation, resulting in a speech processing outcome.
[0251] Optionally, the speech processing network layer includes a speech feature processing network layer and a target speech recognition network layer in the initial speech processing model;
[0252] The first training module is also configured as follows:
[0253] Based on the second speech training data, the speech feature processing network layer and multiple initial speech recognition network layers are determined from the initial speech processing model, wherein the multiple initial speech recognition network layers are used to perform speech recognition on the second speech training samples to obtain initial speech recognition features.
[0254] Determine the positional relationship between each initial speech recognition network layer and the speech feature processing network layer;
[0255] Based on the positional relationship, a target speech recognition network layer corresponding to the speech feature processing network layer is determined from the plurality of initial speech recognition network layers, wherein the target speech recognition network layer is some or all of the network layers in the plurality of initial speech recognition network layers.
[0256] Optionally, the first training module is further configured to:
[0257] The second speech training sample and the second speech processing prompt information are input into the initial speech processing model, wherein the initial speech processing model includes a speech recognition network layer, a speech feature processing network layer and a speech translation network layer;
[0258] The speech recognition network layer is used to perform speech recognition on the second speech training sample to obtain initial speech recognition features;
[0259] The initial speech recognition features are transformed using the speech feature processing network layer to obtain target speech recognition features, wherein the modality of the target speech recognition features is consistent with the modality of the input data of the speech translation network layer.
[0260] Using a speech translation network layer, the target speech recognition features are translated based on the second speech processing prompt information to obtain the initial speech processing result.
[0261] Optionally, the second training module is further configured as follows:
[0262] Using the first speech training data and the second speech training data, the target model parameters in the trained initial speech processing model are adjusted to obtain the target speech processing model, wherein the target model parameters are all or part of the model parameters of the trained initial speech processing model.
[0263] Optionally, the first speech training data includes a first speech training sample, a first speech sample label corresponding to the first speech training sample, and a first speech processing prompt information corresponding to the speech processing task; the second speech training data includes a second speech training sample, a second speech sample label corresponding to the second speech training sample, and a second speech processing prompt information corresponding to the speech processing subtask.
[0264] The second training module is also configured as follows:
[0265] The first speech training sample and the first speech processing prompt information are input into the trained initial speech processing model for speech processing to obtain a first speech processing result, wherein the first speech processing result is related to the speech processing task.
[0266] The second speech training sample and the second speech processing prompt information are input into the trained initial speech processing model for speech processing to obtain the second speech processing result, wherein the second speech processing result is related to the speech processing subtask.
[0267] Based on the first speech processing result, the second speech processing result, the first speech sample label, and the second speech sample label, the target model parameters of the trained initial speech processing model are adjusted to obtain the target speech processing model.
[0268] Optionally, the speech processing task is a speech translation task, the speech processing subtask is a speech recognition task, the speech recognition task is a subtask of the speech translation task, and the speech processing model is a speech translation model;
[0269] The second training module is also configured as follows:
[0270] Based on the first speech training data, the trained initial speech translation model is trained to obtain the target speech translation model.
[0271] This specification provides one or more embodiments of a speech processing model training device. Considering the inaccuracy of speech processing results output by neural network models, this method, during the training of the speech processing model, determines a first speech training data corresponding to a large number of complex speech processing tasks and a second speech training data corresponding to speech processing sub-tasks, where the speech processing sub-tasks are sub-tasks of the speech processing task. The speech processing model is trained using the second speech training data and the first speech training data, thereby obtaining a target speech processing model that can accurately perform fine processing on speech processing tasks containing sub-tasks and output accurate speech processing results. This improves the accuracy of the speech processing results of the neural network model and avoids the problem of inaccurate speech processing results due to the complexity of the speech processing task.
[0272] The above is an illustrative scheme of a speech processing model training device according to this embodiment. It should be noted that the technical solution of this speech processing model training device and the technical solution of the speech processing model training method described above belong to the same concept. For details not described in detail in the technical solution of the speech processing model training device, please refer to the description of the technical solution of the speech processing model training method described above.
[0273] Corresponding to the above method embodiments, this specification also provides an embodiment of a speech translation model training device, which includes:
[0274] The data determination module is configured to determine first speech training data corresponding to the speech translation task and second speech training data corresponding to the speech recognition task, wherein the speech recognition task is a subtask of the speech translation task;
[0275] The first training module is configured to train the speech processing network layer in the initial speech translation model based on the second speech training data to obtain the trained initial speech translation model, wherein the speech processing network layer is related to the speech recognition task.
[0276] The second training module is configured to train the initial speech translation model based on the first speech training data to obtain the target speech translation model.
[0277] This specification provides one or more embodiments of a speech translation model training device. Considering the inaccuracy of speech translation results output by neural network models, this method, during the training of the speech translation model, determines a large amount of first speech training data corresponding to a relatively complex speech translation task and a second speech training data corresponding to a speech recognition task, where the speech recognition task is a subtask of the speech translation task. The speech translation model is trained using the second speech training data and the first speech training data, thereby obtaining a target speech translation model that can accurately process speech translation tasks containing subtasks and output accurate speech translation results. This improves the accuracy of the speech translation results of the neural network model and avoids the problem of inaccurate speech translation results due to the complexity of the speech translation task processing.
[0278] The above is a schematic scheme of a speech translation model training device according to this embodiment. It should be noted that the technical solution of this speech translation model training device and the technical solution of the speech translation model training method described above belong to the same concept. For details not described in detail in the technical solution of the speech translation model training device, please refer to the description of the technical solution of the speech translation model training method described above.
[0279] Corresponding to the above method embodiments, this specification also provides embodiments of a voice processing device, which includes:
[0280] The voice determination module is configured to determine the voice to be processed;
[0281] A speech processing module is configured to input the speech to be processed into a target speech processing model to obtain a speech processing result corresponding to the speech to be processed. The target speech processing model is obtained by training an initial speech processing model based on first speech training data. The initial speech processing model is obtained by training the speech processing network layer in the initial speech processing model based on second speech training data. The first speech training data corresponds to a speech processing task, and the second speech training data corresponds to a speech processing subtask. The speech processing subtask is a subtask of the speech processing task, and the speech processing network layer is related to the speech processing subtask.
[0282] Optionally, the voice processing module is further configured to:
[0283] The speech to be processed is input into the target speech processing model, wherein the target speech processing model includes a speech recognition network layer, a speech feature processing network layer, and a speech translation network layer;
[0284] The speech recognition network layer is used to perform speech recognition on the speech to be processed to obtain initial speech recognition features;
[0285] The initial speech recognition features are transformed using the speech feature processing network layer to obtain target speech recognition features, wherein the modality of the target speech recognition features is consistent with the modality of the input data of the speech translation network layer.
[0286] The speech recognition features of the target speech are translated using a speech translation network layer to obtain the speech processing result.
[0287] This specification provides one or more embodiments of a speech processing device. Considering the inaccuracy of speech processing results output by neural network models, this method, during the training of the speech processing model, determines a first speech training data corresponding to a large number of complex speech processing tasks and a second speech training data corresponding to speech processing sub-tasks. The speech processing sub-tasks are sub-tasks of the speech processing tasks. The speech processing model is trained using the second speech training data and the first speech training data to obtain a target speech processing model. This target speech processing model can accurately perform fine processing on speech processing tasks containing sub-tasks and output accurate speech processing results, thereby improving the accuracy of the speech processing results of the neural network model and avoiding the problem of inaccurate speech processing results due to the complexity of the speech processing tasks.
[0288] The above is an illustrative scheme of a voice processing device according to this embodiment. It should be noted that the technical solution of this voice processing device and the technical solution of the above-described voice processing method belong to the same concept. For details not described in detail in the technical solution of the voice processing device, please refer to the description of the technical solution of the above-described voice processing method.
[0289] Corresponding to the above method embodiments, this specification also provides embodiments of a voice translation device applied to cloud-side equipment, including:
[0290] The voice receiving module is configured to receive the voice to be translated sent by the receiving device.
[0291] A speech processing module is configured to input the speech to be translated into a target speech translation model to obtain a speech translation result corresponding to the speech to be translated. The target speech translation model is obtained by training an initial speech translation model based on first speech training data. The initial speech translation model is obtained by training the speech processing network layer in the initial speech translation model based on second speech training data. The first speech training data corresponds to the speech translation task, the second speech training data corresponds to the speech recognition task, the speech recognition task is a subtask of the speech translation task, and the speech processing network layer is related to the speech recognition task.
[0292] The voice translation result sending module is configured to send the voice translation result to the end-side device.
[0293] This specification provides one or more embodiments of a speech translation device applied to cloud-side devices. Considering the inaccuracy of speech translation results output by neural network models, this method, during the training of the speech translation model, determines a large amount of complex speech translation task-related first speech training data and a speech recognition task-related second speech training data, where the speech recognition task is a subtask of the speech translation task. The speech translation model is trained using the second speech training data and the first speech training data to obtain a target speech translation model. This target speech translation model can accurately perform fine processing on speech processing tasks containing subtasks and output accurate speech translation results, thereby improving the accuracy of the speech translation results of the neural network model and avoiding the problem of inaccurate speech translation results due to the complexity of the speech translation task processing.
[0294] The above is an illustrative scheme of a voice translation device according to this embodiment. It should be noted that the technical solution of this voice translation device and the technical solution of the above-described voice translation method belong to the same concept. For details not described in detail in the technical solution of the voice translation device, please refer to the description of the technical solution of the above-described voice translation method.
[0295] Figure 11 A structural block diagram of a computing device 1100 provided according to one embodiment of this specification is shown.
[0296] The computing device 1100 includes:
[0297] Memory 1110 and processor 1120;
[0298] The memory 1110 is used to store computer programs / instructions, and the processor 1120 is used to execute the computer programs / instructions, which, when executed by the processor 1120, implement the steps of any of the above methods.
[0299] In one or more embodiments of this specification, the computing device can be understood as an integrated smart terminal, including but not limited to a server, desktop computer, PC (Personal Computer), all-in-one model machine, mobile phone, tablet computer or other portable smart terminal, etc., and the computing device may have the model described in the above embodiments of this application pre-installed.
[0300] Specifically, this computing device can pre-install various types of models, including but not limited to models in natural language processing, visual processing, speech processing, code processing, and multimodal task processing, thus providing diverse model selection. In different product forms, this computing device can support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference, and application. In some product forms, this computing device also supports model management, including but not limited to multi-type model management (supporting the management of discriminative, generative, and other model types), model version control (supporting the control of different model versions), and model evaluation (evaluating model performance and effectiveness based on model evaluation tools). In other product forms, this computing device can also create applications based on models, providing API (Application Programming Interface) calling capabilities. Users can call models into created applications through the API interface, and application management tools are also provided for application management and monitoring.
[0301] Furthermore, the computing device may also include data management (supporting the creation and management of model tuning datasets), a training center (providing abundant training resources to help users learn and master AI (Artificial Intelligence) technology), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, it provides a comprehensive and integrated device for AI development, training, deployment, and application.
[0302] Figure 12 A structural block diagram of an electronic device 1200 provided according to one embodiment of this specification is shown.
[0303] The memory 1210 and the processor 1220 are connected via a bus 1230;
[0304] The memory 1210 is used to store computer programs / instructions, and the processor 1220 is used to execute the computer programs / instructions, which, when executed by the processor 1220, implement the steps of the method.
[0305] Specifically, the components of the electronic device 1200 include, but are not limited to, a memory 1210 and a processor 1220. The processor 1220 is connected to the memory 1210 via a bus 1230, and the database 1250 is used to store data.
[0306] Electronic device 1200 also includes access device 1240, which enables electronic device 1200 to communicate via one or more networks 1260. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. Access device 1240 may include one or more of any type of wired or wireless network interface (e.g., network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Wi-MAX (Worldwide Interoperability for Microwave Access) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC) interface, and so on.
[0307] In one embodiment of this specification, the above-described components of the electronic device 1200 and Figure 12 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 12 The block diagram of the electronic device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0308] Electronic device 1200 can be any type of stationary or mobile electronic device, including mobile computers or mobile electronic devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable electronic devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary electronic devices such as desktop computers or personal computers (PCs). Electronic device 1200 can also be a mobile or stationary server.
[0309] The above is an illustrative scheme of an electronic device according to this embodiment. It should be noted that the technical solution of this electronic device belongs to the same concept as the technical solution of any of the above methods, and any details not described in detail in the technical solution of the electronic device can be referred to the description of the technical solution of any of the above methods.
[0310] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.
[0311] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are basically similar to any of the above-described method embodiments, so the description is relatively simple; relevant parts can be referred to in the description of any of the above-described method embodiments.
[0312] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.
[0313] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the technical solution of any of the above methods, and any details not described in detail in the technical solution of the computer program product can be referred to the description of the technical solution of any of the above methods.
[0314] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0315] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0316] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0317] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0318] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A method for training a speech processing model, comprising: Determine the first speech training data corresponding to the speech processing task and the second speech training data corresponding to the speech processing subtask, wherein the speech processing subtask is a subtask of the speech processing task; Based on the second speech training data, the speech processing network layer in the initial speech processing model is trained to obtain the trained initial speech processing model, wherein the speech processing network layer is related to the speech processing subtask. Based on the first speech training data, the trained initial speech processing model is trained to obtain the target speech processing model.
2. The speech processing model training method according to claim 1, wherein the second speech training data includes a second speech training sample, a second speech sample label corresponding to the second speech training sample, and second speech processing prompt information corresponding to the speech processing subtask; The step of training the speech processing network layer in the initial speech processing model based on the second speech training data to obtain the trained initial speech processing model includes: The speech processing network layer is determined from the initial speech processing model; The second speech training sample and the second speech processing prompt information are input into the initial speech processing model for speech processing to obtain the initial speech processing result, wherein the initial speech processing result is related to the speech processing subtask; Based on the initial speech processing results and the second speech sample labels, the parameters of the speech processing network layer are adjusted to obtain the trained initial speech processing model.
3. The speech processing model training method according to claim 2, wherein the speech processing network layer includes a speech feature processing network layer; Determining the speech processing network layer from the initial speech processing model includes: Based on the second speech training data, the speech feature processing network layer is determined from the initial speech processing model.
4. The speech processing model training method according to claim 2, wherein the speech processing network layer includes a speech feature processing network layer and a target speech recognition network layer in the initial speech processing model; Determining the speech processing network layer from the initial speech processing model includes: Based on the second speech training data, the speech feature processing network layer and multiple initial speech recognition network layers are determined from the initial speech processing model, wherein the multiple initial speech recognition network layers are used to perform speech recognition on the second speech training samples to obtain initial speech recognition features. Determine the positional relationship between each initial speech recognition network layer and the speech feature processing network layer; Based on the positional relationship, a target speech recognition network layer corresponding to the speech feature processing network layer is determined from the plurality of initial speech recognition network layers, wherein the target speech recognition network layer is some or all of the network layers in the plurality of initial speech recognition network layers.
5. The speech processing model training method according to any one of claims 2 to 4, wherein inputting the second speech training sample and the second speech processing prompt information into the initial speech processing model for speech processing to obtain the initial speech processing result includes: The second speech training sample and the second speech processing prompt information are input into the initial speech processing model, wherein the initial speech processing model includes a speech recognition network layer, a speech feature processing network layer and a speech translation network layer; The speech recognition network layer is used to perform speech recognition on the second speech training sample to obtain initial speech recognition features; The initial speech recognition features are transformed using the speech feature processing network layer to obtain target speech recognition features, wherein the modality of the target speech recognition features is consistent with the modality of the input data of the speech translation network layer. Using a speech translation network layer, the target speech recognition features are translated based on the second speech processing prompt information to obtain the initial speech processing result.
6. The speech processing model training method according to any one of claims 1 to 4, wherein the step of training the initial speech processing model based on the first speech training data to obtain a target speech processing model includes: Using the first speech training data and the second speech training data, the target model parameters in the trained initial speech processing model are adjusted to obtain the target speech processing model, wherein the target model parameters are all or part of the model parameters of the trained initial speech processing model.
7. The speech processing model training method according to claim 6, wherein the first speech training data includes a first speech training sample, a first speech sample label corresponding to the first speech training sample, and a first speech processing prompt information corresponding to the speech processing task; and the second speech training data includes a second speech training sample, a second speech sample label corresponding to the second speech training sample, and a second speech processing prompt information corresponding to the speech processing sub-task. The step of adjusting the target model parameters in the trained initial speech processing model using the first speech training data and the second speech training data to obtain the target speech processing model includes: The first speech training sample and the first speech processing prompt information are input into the trained initial speech processing model for speech processing to obtain a first speech processing result, wherein the first speech processing result is related to the speech processing task. The second speech training sample and the second speech processing prompt information are input into the trained initial speech processing model for speech processing to obtain the second speech processing result, wherein the second speech processing result is related to the speech processing subtask. Based on the first speech processing result, the second speech processing result, the first speech sample label, and the second speech sample label, the target model parameters of the trained initial speech processing model are adjusted to obtain the target speech processing model.
8. The speech processing model training method according to any one of claims 1 to 4, wherein the speech processing task is a speech translation task, the speech processing subtask is a speech recognition task, the speech recognition task is a subtask of the speech translation task, and the speech processing model is a speech translation model; The step of training the initial speech processing model based on the first speech training data to obtain the target speech processing model includes: Based on the first speech training data, the trained initial speech translation model is trained to obtain the target speech translation model.
9. A method for training a speech translation model, comprising: Determine the first speech training data corresponding to the speech translation task and the second speech training data corresponding to the speech recognition task, wherein the speech recognition task is a subtask of the speech translation task; Based on the second speech training data, the speech processing network layer in the initial speech translation model is trained to obtain the trained initial speech translation model, wherein the speech processing network layer is related to the speech recognition task; Based on the first speech training data, the trained initial speech translation model is trained to obtain the target speech translation model.
10. A speech processing method, comprising: Identify the speech to be processed; The speech to be processed is input into a target speech processing model to obtain the speech processing result corresponding to the speech to be processed. The target speech processing model is obtained by training an initial speech processing model based on a first speech training data. The initial speech processing model is obtained by training the speech processing network layer in the initial speech processing model based on a second speech training data. The first speech training data corresponds to the speech processing task, and the second speech training data corresponds to the speech processing subtask. The speech processing subtask is a subtask of the speech processing task, and the speech processing network layer is related to the speech processing subtask.
11. The speech processing method according to claim 10, wherein inputting the speech to be processed into a target speech processing model to obtain the speech processing result corresponding to the speech to be processed includes: The speech to be processed is input into the target speech processing model, wherein the target speech processing model includes a speech recognition network layer, a speech feature processing network layer, and a speech translation network layer; The speech recognition network layer is used to perform speech recognition on the speech to be processed to obtain initial speech recognition features; The initial speech recognition features are transformed using the speech feature processing network layer to obtain target speech recognition features, wherein the modality of the target speech recognition features is consistent with the modality of the input data of the speech translation network layer. The speech recognition features of the target speech are translated using a speech translation network layer to obtain the speech processing result.
12. A speech translation method, applied to cloud-side devices, comprising: The voice to be translated sent by the receiving device; The speech to be translated is input into the target speech translation model to obtain the speech translation result corresponding to the speech to be translated. The target speech translation model is obtained by training an initial speech translation model based on a first speech training data. The initial speech translation model is obtained by training the speech processing network layer in the initial speech translation model based on a second speech training data. The first speech training data corresponds to the speech translation task, and the second speech training data corresponds to the speech recognition task. The speech recognition task is a subtask of the speech translation task, and the speech processing network layer is related to the speech recognition task. The speech translation result is sent to the terminal device.
13. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 12.
14. An electronic device, comprising: A memory and a processor, the memory and the processor being connected via a bus; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 12.
15. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 12.
16. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 12.