Language processing model training method and device, electronic equipment and readable storage medium
By introducing style feature loss function in the speech synthesis large model and optimizing the training of language processing models, the problem of poor performance of existing models when processing speech style information is solved, and more accurate speech style reproduction and user experience improvement are achieved.
Patent Information
- Application Number
- CN202510196407.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-27
AI Technical Summary
Existing large-scale speech synthesis models perform poorly when processing style information, making it difficult to accurately reproduce the style characteristics of speech such as rhythm and emotion.
By obtaining training text data and real phonetic marking data, input it into the language processing model, extracting the style features in the real and predicting the phonetic marking data, and building a total loss function, introducing style feature loss, and optimizing the training language processing model to improve its performance in processing speech style information.
By introducing a style feature loss function, we ensure that the language processing model can reproduce the voice style more accurately, thereby improving the performance of speech synthesis large models in processing voice style information and improving user experience.
Smart Images

Figure CN120048242A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech synthesis, and in particular, to a method, apparatus, electronic device, and computer-readable storage medium for training a language processing model. Background Art
[0002] With the rapid development of information technology, as a technology for converting text into natural and fluent speech, text-to-speech (TTS) has become an important part of human-computer interaction and directly affects the user experience. TTS is widely used in fields such as outbound calls, voice navigation, and audiobooks.
[0003] In recent years, the training of large speech synthesis models based on neural networks has become a research hotspot in the TTS field. However, existing large speech synthesis models perform poorly in processing style information (such as prosody, emotion, etc.). Summary of the Invention
[0004] The purpose of the present invention is to provide a method, apparatus, electronic device, and readable storage medium for training a language processing model to solve the problem that existing large speech synthesis models perform poorly in processing style information.
[0005] In a first aspect, the present invention provides a method for training a language processing model, the method comprising:
[0006] Obtaining training text data and real speech label data;
[0007] Inputting the training text data and the real speech label data into a language processing model to obtain predicted speech label data output by the language processing model;
[0008] Obtaining real style features according to the real speech label data, and obtaining predicted style features according to the predicted speech label data;
[0009] Obtaining a total loss function according to the real speech label data, the predicted speech label data, the real style features, and the predicted style features;
[0010] Optimally training the language processing model according to the total loss function.
[0011] In an optional embodiment, the language processing model includes a style encoding module, and obtaining real style features according to the real speech label data, and obtaining predicted style features according to the predicted speech label data includes:
[0012] Performing style extraction on the real speech label data through the style encoding module to obtain the real style features;
[0013] The style extraction module extracts the style from the predicted speech tag data to obtain the predicted style features.
[0014] In an optional embodiment, obtaining the total loss function according to the true speech tag data, the predicted speech tag data, the true style features, and the predicted style features includes:
[0015] Obtaining a prediction loss function according to the true speech tag data and the predicted speech tag data;
[0016] Obtaining a style loss function according to the true style features and the predicted style features;
[0017] Obtaining the total loss function according to the prediction loss function and the style loss function.
[0018] In an optional embodiment, obtaining the prediction loss function according to the true speech tag data and the predicted speech tag data includes:
[0019] Calculating the distribution distance between the true speech tag data and the predicted speech tag data to obtain the prediction loss function.
[0020] In an optional embodiment, obtaining the style loss function according to the true style features and the predicted style features includes:
[0021] Calculating the similarity between the true style features and the predicted style features to obtain the style loss function.
[0022] In an optional embodiment, obtaining the total loss function according to the prediction loss function and the style loss function includes:
[0023] Performing weighted calculation on the prediction loss function and the style loss function to obtain the total loss function.
[0024] In a second aspect, the present invention provides a language processing model training device, and the device includes:
[0025] A predicted speech tag data acquisition module, configured to acquire training text data and true speech tag data; input the training text data and the true speech tag data into a language processing model to obtain the predicted speech tag data output by the language processing model;
[0026] A style feature acquisition module, configured to obtain true style features according to the true speech tag data, and obtain predicted style features according to the predicted speech tag data;
[0027] A loss function obtaining module, configured to obtain a total loss function according to the true speech label data, the predicted speech label data, the true style feature, and the predicted style feature;
[0028] An optimization training module, configured to optimize and train the language processing model according to the total loss function.
[0029] In an optional embodiment, the language processing model includes a style encoding module, and the style feature obtaining module is further configured to perform style extraction on the true speech label data through the style encoding module to obtain the true style feature; perform style extraction on the predicted speech label data through the style encoding module to obtain the predicted style feature.
[0030] In a third aspect, the present invention provides an electronic device, including a processor and a memory, where the memory stores a computer program that can be executed by the processor, and the computer program executable by the processor is used to implement the language processing model training method according to any one of the foregoing embodiments.
[0031] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the language processing model training method according to any one of the foregoing embodiments.
[0032] The language processing model training method, device, electronic device, and readable storage medium provided by the embodiments of the present invention, the method includes: acquiring training text data and true speech label data and inputting them into a language processing model to obtain predicted speech label data, respectively extracting style features from the true and predicted speech label data, then constructing a total loss function according to the true speech label data, the predicted speech label data, the true style feature, and the predicted style feature, and then optimizing and training the language processing model according to the total loss function, introducing a style feature loss into the total loss function to ensure that the language processing model can learn style information, so that the language processing model can more accurately reproduce the speech style, thereby improving the performance of the language processing model in processing speech style information, and the speech processing model is a key part of the speech synthesis model, and thus can improve the performance of the speech synthesis large model in processing speech style information. Description of the Drawings
[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained according to these drawings without creative efforts.
[0034] Figure 1 It shows a schematic flowchart of a method for training a language processing model provided by an embodiment of the present invention;
[0035] Figure 2 It shows another schematic flowchart of a method for training a language processing model provided by an embodiment of the present invention;
[0036] Figure 3 It shows a schematic framework diagram of a large speech synthesis model provided by an embodiment of the present invention;
[0037] Figure 4 It shows yet another schematic flowchart of a method for training a language processing model provided by an embodiment of the present invention;
[0038] Figure 5 It shows a block diagram of a device for training a language processing model provided by an embodiment of the present invention;
[0039] Figure 6 It shows a schematic block diagram of an electronic device provided by an embodiment of the present invention.
[0040] Icons: 100 - Electronic device; 110 - Memory; 120 - Processor; 130 - Communication module; 200 - Device for training a language processing model; 210 - Module for obtaining predicted speech marker data; 220 - Module for obtaining style features; 230 - Module for obtaining a loss function; 240 - Optimization training module. Detailed implementation manners
[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0042] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0043] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0044] With the rapid development of information technology, text-to-speech (TTS) technology, as a technology that converts text into natural and fluent speech, has become an important part of human-computer interaction and directly affects the user experience. TTS is widely used in fields such as outbound calling announcements, voice navigation, and audiobooks.
[0045] The existing main models of text-to-speech large models consist of three parts: a large language model (LLM), an acoustic feature prediction model, and a vocoder model. In the prediction speech token data output by the large language model in the first stage, it mainly includes semantic information and some style information. The loss function of the existing large language model mainly focuses on semantic information and pays less attention to style information. Therefore, when the text-to-speech large model performs speech synthesis, it cannot reproduce the speech style well, resulting in poor performance of the text-to-speech large model in processing style information (such as prosody, emotion, etc.).
[0046] To address the above problems, the inventors proposed a method, apparatus, server, and storage medium for training a large language model provided in the embodiments of the present invention. By introducing a style feature loss and optimizing and training the large language model according to the total loss function, the performance of the large language model in processing speech style information is improved, and further the performance of the text-to-speech large model in processing speech style information is improved. Among them, the specific method for training the large language model will be described in detail in the subsequent embodiments.
[0047] Please refer to Figure 1 , Figure 1 which shows a schematic flowchart of a method for training a large language model provided in an embodiment of the present invention. This method for training a large language model can be applied to an electronic device. Taking the electronic device as an example, the specific process of this embodiment will be described below. The following will elaborate in detail on the Figure 1 shown process. The method for training the large language model may specifically include the following steps:
[0048] Step 110: Obtain training text data and true speech marker data.
[0049] Among them, the training text data refers to text sentences that have been sorted and annotated. Inputting the training text data into a language processing model, the language processing model can learn how to generate corresponding predicted speech marker data based on the training text data and the true speech marker data. The true speech marker data is discrete markers extracted from the reference speech. The true speech marker data includes semantic information and partial style information of the reference speech. The content of the reference speech can be the same as the content of the training text data, or the content of the reference speech can be different from the content of the training text data.
[0050] In some embodiments, the electronic device can obtain the reference speech, preprocess the reference speech, for example, cleaning processing, denoising processing, and unification processing, etc., then extract features from the preprocessed reference speech to obtain speech features, and mark each frame of speech features to obtain the true speech marker data.
[0051] Step 120: Input the training text data and the true speech marker data into the language processing model to obtain the predicted speech marker data output by the language processing model.
[0052] Among them, the language processing model can be an LLM model.
[0053] In some embodiments, the electronic device can obtain the predicted speech marker data output by the language processing model through the following formula.
[0054] speech_token predict = LLM(speech_token truth , text)
[0055] Among them, speech_token predict represents the predicted speech marker data, LLM represents the LMM model, speech_token truth represents the true speech marker data, and text represents the training text data.
[0056] Step 130: Obtain the true style feature according to the true speech marker data, and obtain the predicted style feature according to the predicted speech marker data.
[0057] In some embodiments, the electronic device can extract the style from the true speech marker data to obtain the true style feature, and the electronic device can also extract the style from the predicted speech marker data to obtain the predicted style feature.
[0058] Step 140: Obtain the total loss function based on the real speech token data, predicted speech token data, real style features, and predicted style features.
[0059] In some embodiments, the electronic device can construct the total loss function according to the real speech token data, predicted speech token data, real style features, and predicted style features.
[0060] Step 150: Optimize and train the language processing model according to the total loss function.
[0061] In some embodiments, the language processing model can gradually adjust its internal parameters according to the feedback of the total loss function to improve its ability to generate predicted speech token data and capture style information. The language processing model can continuously iterate according to the total loss function until the total loss function reaches a threshold or the total loss function tends to be stable.
[0062] The language processing model training method provided by the embodiments of the present invention obtains training text data and real speech token data and inputs them into the language processing model to obtain predicted speech token data. The style features in the real and predicted speech token data are respectively extracted, and then the total loss function is constructed according to the real speech token data, predicted speech token data, real style features, and predicted style features. Then, the language processing model is optimized and trained according to the total loss function, and the style feature loss is introduced into the total loss function to ensure that the language processing model can learn style information, so that the language processing model can more accurately reproduce the speech style, thereby improving the performance of the language processing model in processing speech style information. And the speech processing model is a key part of the speech synthesis model, and thus can improve the performance of the large speech synthesis model in processing speech style information.
[0063] To ensure that the language processing model can learn style information, a style encoding module is introduced into the language processing model. The language processing model includes a style encoding module. As Figure 2 shown, step 130 specifically includes the following steps:
[0064] Step 131: Extract the style of the real speech token data through the style encoding module to obtain the real style features.
[0065] In the architecture of the large speech synthesis model with a style encoding module as Figure 3 shown, the training text data (text) and the real speech token data (speech_token truth ) are input into the language processing model (LLM) in the first stage of the large speech synthesis model to obtain the predicted speech token data (speech_token predict), the style encoder extracts the style from the real speech token data (speech_token truth ), obtaining the real style feature (style truth). The style encoder extracts the style from the predicted speech token data (speech_token predict ), obtaining the predicted style feature (style predict). The predicted speech token data (speech_token predict ) output by the language processing model is input into the acoustic feature prediction model (Token-to-mel) in the second stage of the speech synthesis large model, and then the data output by the acoustic feature prediction model (Token-to-mel) is input into the vocoder model (Mel-to-wav).
[0066] In some embodiments, the style encoder can use Linear Predictive Coding (LPC) technology to extract the style from speech tokens and obtain style features.
[0067] It should be noted that the LPC model represents the spectral characteristics of speech token data by extracting linear prediction coefficients (LPC coefficients). The LPC coefficients can reflect the formant positions and shapes of speech token data, thereby indirectly representing the style features of speech, such as intonation, speech rate, and timbre.
[0068] In some embodiments, the electronic device can obtain the real style feature through the following formula.
[0069] style_hidden truth =Style_encoder(speech_token truth )
[0070] where style_hidden truth represents the real style feature, Style_encoder represents the style encoder, and speech_token truth represents the real speech token data.
[0071] Please continue to refer to Figure 2 , step 132: The style encoder extracts the style from the predicted speech token data to obtain the predicted style feature.
[0072] In some embodiments, the electronic device can obtain the predicted style feature through the following formula.
[0073] style_hidden predict= Style_encoder(speech_token predict )
[0074] where style_hidden predict represents the predicted style feature, Style_encoder represents the style encoding module, and speech_token predict represents the predicted speech token data.
[0075] In some embodiments, the style encoding module included in the language processing model is a pre-trained model. During the training process of the language processing model, the style encoding module serves as a part for supervising the training of the language processing model.
[0076] In other embodiments, the style encoding module included in the language processing model is untrained, and the parameters in the style encoding module can be adjusted simultaneously during the training process of the language processing model.
[0077] It can be understood that by using the unified style encoding module to extract styles from the real speech token data and the predicted speech token data respectively, it can ensure that the real style feature and the predicted style feature are compared under the same standard, thereby improving the accuracy of calculating the style loss function during the training process of the language processing model, and further improving the performance of the language processing model in dealing with language styles.
[0078] To improve the performance of the language processing model in dealing with language styles, a style loss function is introduced. As Figure 4 shown, step 140 specifically includes the following steps:
[0079] Step 141: Obtain a prediction loss function based on the real speech token data and the predicted speech token data.
[0080] where the prediction loss function is used to measure the difference between the real speech token data and the predicted speech token data.
[0081] In some embodiments, the electronic device can calculate the distribution distance between the real speech token data and the predicted speech token data to obtain the prediction loss function.
[0082] In some embodiments, the electronic device can calculate the cross-entropy between the real speech token data and the predicted speech token data through the following formula to calculate the distribution distance between the real speech token data and the predicted speech token data, and obtain the prediction loss function.
[0083] Loss ce = Cross_Entropy(speech_token truth , speech_tokenpredict )
[0084] Among them, Loss ce represents the predicted loss function, Cross_Entropy(speech_token truth , speech_token predict ) represents the cross-entropy loss between the true speech token data and the predicted speech token data, and speech_token truth represents the true speech token data, and speech_token predict represents the predicted speech token data.
[0085] Step 142: Obtain a style loss function according to the true style feature and the predicted style feature.
[0086] In some embodiments, calculate the similarity between the true style feature and the predicted style feature to obtain a style loss function.
[0087] Among them, the electronic device can calculate the similarity between the true style feature and the predicted style feature through techniques such as cosine similarity, KL divergence, and Dynamic Time Warping (DTW).
[0088] In some embodiments, the electronic device can calculate the cosine similarity between the true style feature and the predicted style feature through the following formula to obtain a predicted loss function.
[0089] Loss style = Cosine_Embedding(style_hidden truth , style_hidden predict )
[0090] Among them, Lossstyle represents the style loss function, and Cosine_Embedding(style_hidden truth , style_hidden predict ) represents the cosine similarity between the true style feature and the predicted style feature, style_hidden truth represents the true style feature, and style_hidden predict represents the predicted style feature.
[0091] Step 143: Obtain a total loss function according to the predicted loss function and the style loss function.
[0092] In some embodiments, perform weighted calculation on the predicted loss function and the style loss function to obtain a total loss function.
[0093] In some embodiments, preset weighted values are pre-set and stored in the electronic device, and the electronic device can perform weighted calculation on the prediction loss function and the style loss function according to the following formula to obtain the total loss function.
[0094] Loss llm= α×Loss ce +(1-α)×Loss style
[0095] Where Loss llm represents the total loss function, α represents the preset weighted value, Loss ce represents the prediction loss function, and Loss style represents the style loss function.
[0096] It should be noted that the prediction loss function optimizes the hard matching between the real speech label data and the predicted speech label data. When the difference between the real speech label data and the predicted speech label data is small, the style information contained in the speech label data may not be accurately predicted. Therefore, weighted calculation is used to make a trade-off in terms of the difference between semantic information and style information.
[0097] The language processing model training method provided by the embodiments of the present invention ensures that the speech generated by the speech synthesis large model is consistent with the style of the training text data and the real speech label data (the speech label data corresponding to the reference speech) in terms of both semantics and style by introducing an additional style encoding module and a style loss function, thereby realizing more accurate speech style control and further improving the user experience.
[0098] To execute the corresponding steps in the above embodiments and each possible way, an implementation manner of a language processing model training device is given below. Further, please refer to Figure 5 , the figure is a functional module diagram of a language processing model training device provided by an embodiment of the present invention. It should be noted that the basic principle and the technical effects generated by the language processing model training device provided in this embodiment are the same as those in the above embodiments. For a brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above embodiments. The language processing model training device 200 includes: a predicted speech label data obtaining module 210, a style feature obtaining module 220, a loss function obtaining module 230, and an optimization training module 240, where:
[0099] The predicted speech label data obtaining module 210 is configured to obtain training text data and real speech label data; input the training text data and the real speech label data into the language processing model to obtain the predicted speech label data output by the language processing model.
[0100] A style feature acquisition module 220 is configured to acquire true style features based on true speech tagging data and acquire predicted style features based on predicted speech tagging data.
[0101] A loss function acquisition module 230 is configured to acquire a total loss function based on the true speech tagging data, the predicted speech tagging data, the true style features, and the predicted style features.
[0102] An optimization training module 240 is configured to perform optimization training on the language processing model according to the total loss function.
[0103] Optionally, the language processing model includes a style encoding module. The style feature acquisition module is further configured to perform style extraction on the true speech tagging data through the style encoding module to acquire true style features; and perform style extraction on the predicted speech tagging data through the style encoding module to acquire predicted style features.
[0104] Optionally, the loss function acquisition module 230 is specifically further configured to acquire a predicted loss function according to the true speech tagging data and the predicted speech tagging data; acquire a style loss function according to the true style features and the predicted style features; and acquire the total loss function according to the predicted loss function and the style loss function.
[0105] Optionally, the loss function acquisition module 230 is specifically further configured to calculate a distribution distance between the true speech tagging data and the predicted speech tagging data to acquire the predicted loss function.
[0106] Optionally, the loss function acquisition module 230 is specifically further configured to calculate a similarity between the true style features and the predicted style features to acquire the style loss function.
[0107] Optionally, the loss function acquisition module 230 is specifically further configured to perform weighted calculation on the predicted loss function and the style loss function to acquire the total loss function.
[0108] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0109] In several embodiments provided by the present invention, the coupling between modules may be electrical, mechanical, or other forms of coupling.
[0110] In addition, in each embodiment of the present invention, the functional modules may be integrated into one processing module, or each module may exist physically alone, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0111] Please refer to Figure 6 , which is a block diagram of the electronic device 100 provided by an embodiment of the present invention. The electronic device 100 includes a memory 110, a processor 120, and a communication module 130. The components of the memory 110, the processor 120, and the communication module 130 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0112] Among them, the memory 110 is used to store programs or data. The memory 110 can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0113] The processor 120 is used to read / write the data or programs stored in the memory and execute corresponding functions. For example, when the computer program stored in the memory 110 is executed by the processor 120, the language processing model training methods disclosed in the above embodiments can be implemented.
[0114] The communication module 130 is used to establish a communication connection between the electronic device 100 and the cloud server through a network and is used to send and receive data through the network.
[0115] It should be understood that Figure 6 the structure shown is only a schematic diagram of the structure of the electronic device, and the electronic device may further include more or fewer components than those shown in Figure 6 . Figure 6 The components shown in can be implemented by hardware, software, or a combination thereof.
[0116] An embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executable by a processor, the language processing model training method described in the above method embodiment can be implemented.
[0117] A computer-readable storage medium may be an electronic memory such as a flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium has a storage space for program code that executes any of the method steps in the above-described methods. These program codes may be read out from or written into one or more computer program products. The program codes may be compressed in a suitable form, for example.
[0118] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods may also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.
[0119] If a function is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a portable hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program code.
[0120] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A language processing model training method, characterized in that: The method comprises: Obtain training text data and real speech label data; Inputting the training text data and the real speech mark data into a language processing model to obtain predicted speech mark data output by the language processing model; Obtaining real style features according to the real speech tag data, and obtaining predicted style features according to the predicted speech tag data; Obtaining a total loss function according to the real speech tag data, the predicted speech tag data, the real style features and the predicted style features; The language processing model is optimized and trained according to the total loss function.
2. The method according to claim 1, characterized in that The language processing model includes a style encoding module, and the method of obtaining the real style features according to the real speech tag data and obtaining the predicted style features according to the predicted speech tag data includes: Performing style extraction on the real speech mark data by the style encoding module to obtain the real style feature; The style encoding module is used to perform style extraction on the predicted speech tag data to obtain the predicted style features.
3. The method according to claim 1, characterized in that The obtaining of a total loss function according to the real speech tag data, the predicted speech tag data, the real style feature and the predicted style feature comprises: Obtaining a prediction loss function according to the real speech tag data and the predicted speech tag data; Obtaining a style loss function according to the true style feature and the predicted style feature; The total loss function is obtained according to the prediction loss function and the style loss function.
4. The method according to claim 3, characterized in that The step of obtaining a prediction loss function according to the real speech tag data and the predicted speech tag data comprises: The distribution distance between the real speech tag data and the predicted speech tag data is calculated to obtain the prediction loss function.
5. The method according to claim 3, characterized in that: The obtaining of a style loss function according to the real style feature and the predicted style feature includes: The similarity between the true style feature and the predicted style feature is calculated to obtain the style loss function.
6. The method according to claim 3, characterized in that: The obtaining the total loss function according to the prediction loss function and the style loss function includes: The prediction loss function and the style loss function are weightedly calculated to obtain the total loss function.
7. A language processing model training device, characterized in that: The device comprises: A predicted speech mark data acquisition module is used to obtain training text data and real speech mark data; input the training text data and the real speech mark data into a language processing model to obtain the predicted speech mark data output by the language processing model; A style feature acquisition module, used to acquire real style features according to the real speech mark data, and to acquire predicted style features according to the predicted speech mark data; A loss function obtaining module, used for obtaining a total loss function according to the real speech mark data, the predicted speech mark data, the real style feature and the predicted style feature; An optimization training module is used to optimize the language processing model according to the total loss function.
8. The device according to claim 7, characterized in that The language processing model includes a style encoding module, and the style feature acquisition module is further used to perform style extraction on the real speech tag data through the style encoding module to obtain the real style feature; and perform style extraction on the predicted speech tag data through the style encoding module to obtain the predicted style feature.
9. An electronic device, characterized in that: It comprises a processor and a memory, wherein the memory stores a computer program that can be executed by the processor, and the computer program executable by the processor is used to implement the language processing model training method described in any one of claims 1-6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the language processing model training method according to any one of claims 1 to 6 is implemented.