Spoken-text generation method, speech synthesis method, and related devices
By obtaining the target written text and prompts, using the oral text generation model to generate more colloquialized targeted oral text, the problem that the pronunciation synthesis method in the prior art is difficult to generate colloquialized pronunciation in personalized scenarios, and a more anthropomorphic pronunciation synthesis effect is achieved.
Patent Information
- Application Number
- PCT/CN2024/084194
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-29
- Filing Date
- 2024-03-27
- Publication Date
- 2025-08-07
AI Technical Summary
Existing speech synthesis methods are difficult to generate more colloquial speech texts in personalized human-computer interaction and virtual dialogue scenarios.
By obtaining the target written text and prompts, the oral text generation model is used to generate a more colloquialized target spoken text according to the indicated content of the prompts, and synthesize oral pronunciation based on the target spoken text.
It realizes the generation of more colloquial voice texts in personalized human-computer interaction and virtual dialogue scenarios, and improves the anthropomorphic effect of voice synthesis.
Smart Images

Figure CN2024084194_07082025_PF_FP_ABST
Abstract
Description
Spoken text generation method, speech synthesis method and related device
[0001] This application claims priority to Chinese patent application No. 2024101259854, filed on January 29, 2024, entitled “Spoken text generation method, speech synthesis method and related device”, which is incorporated herein by reference in its entirety.
Technical field
[0002] The present application relates to the field of natural language processing technology, and in particular to a spoken text generation method, a speech synthesis method, and related devices. [Background Technology]
[0003] As interactive scenarios become increasingly widespread, researchers are exploring ways to make these interactions more human-like, such as in human-machine speech interaction, enabling machines to produce more human-like speech. The machine speech synthesis process consists of a speech synthesis front-end and a speech synthesis back-end. The front-end primarily converts text sequences in various languages into phoneme sequences that are more relevant to pronunciation. The back-end uses the phoneme sequences to generate acoustic parameters and then uses a vocoder to restore the speech waveform. Therefore, the text sequences relied upon in speech synthesis are crucial, and making the text used in speech synthesis more colloquial has become a key concern for researchers.
[0004] [Summary of the invention]
[0005] The main technical problem solved by this application is to provide a spoken text generation method, a speech synthesis method and related devices, which can obtain a more colloquial spoken text.
[0006] To solve the above technical problems, the first aspect of the present application provides a spoken text generation method, which includes: obtaining a target written text and a prompt, wherein the prompt is used to instruct a spoken text generation model to perform a spoken text generation task; using the spoken text generation model to perform the spoken text generation task on the target written text according to the first instruction content of the prompt, to obtain the target spoken text.
[0007] To solve the above technical problems, the second aspect of the present application provides a speech synthesis method, which includes: obtaining a target written text and a prompt, wherein the prompt is used to instruct a spoken text generation model to perform a spoken text generation task; using the spoken text generation model to perform the spoken text generation task on the target written text according to the first instruction content of the prompt to obtain a target spoken text; and synthesizing spoken speech based on the target spoken text.
[0008] To solve the above-mentioned technical problems, the third aspect of the present application provides a spoken text generation device, which includes a first acquisition module and a first generation module. The first acquisition module is used to acquire the target written text and prompt words, wherein the prompt words are used to instruct the spoken text generation model to perform the spoken text generation task; the first generation module is used to use the spoken text generation model to perform the spoken text generation task on the target written text according to the first instruction content of the prompt words to obtain the target spoken text.
[0009] To solve the above technical problems, the fourth aspect of the present application provides a speech synthesis device, which includes: a second acquisition module, a second generation module and a speech synthesis module, the second acquisition module is used to acquire the target written text and prompt words, wherein the prompt words are used to instruct the spoken text generation model to perform the spoken text generation task; the second generation module is used to use the spoken text generation model to perform the spoken text generation task on the target written text according to the first instruction content of the prompt words to obtain the target spoken text; the speech synthesis module is used to synthesize spoken speech based on the target spoken text.
[0010] In order to solve the above technical problems, the fifth aspect of this application provides an electronic device, which includes a memory and a processor coupled to each other, and the memory stores program instructions; the processor is used to execute the program instructions stored in the memory to implement the method provided by the first or second aspect above.
[0011] In order to solve the above technical problems, the sixth aspect of the present application provides a computer-readable storage medium, which stores a program file. The program file can be executed to implement the method provided by the first aspect or the second aspect above.
[0012] The beneficial effects of the present application are as follows: unlike the prior art, the present application obtains a target written text and a prompt, wherein the prompt is used to instruct a spoken text generation model to perform a spoken text generation task; and the spoken text generation model performs the spoken text generation task on the target written text according to a first instruction of the prompt, thereby obtaining a target spoken text. By setting a personalized prompt, the spoken text generation model can generate a more colloquial target spoken text using the target written text according to the first instruction of the prompt.
Brief Description of the Drawings
[0013] FIG1 is a flow chart of an embodiment of a method for generating a spoken text provided by the present application;
[0014] FIG2 is a flow chart of an embodiment of a training process for a spoken text generation model provided by the present application;
[0015] FIG3 is a flow chart of an embodiment of a speech synthesis method provided by the present application;
[0016] FIG4 is a schematic diagram of a framework of an embodiment of a spoken text generation device provided by the present application;
[0017] FIG5 is a schematic diagram of a framework of an embodiment of a speech synthesis device provided by the present application;
[0018] FIG6 is a schematic diagram of a frame structure of an electronic device according to an embodiment of the present application;
[0019] FIG7 is a schematic diagram of a framework of an embodiment of a computer-readable storage medium provided in the present application. [Specific implementation method]
[0020] The following is a clear and complete description of the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0021] It should be noted that the terms "first," "second," and so on, used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include at least one of these features.
[0022] In the description of the application, the meaning of "multiple" is at least two, for example, two, three, etc., unless otherwise clearly and specifically limited. In the embodiments of the present application, all directional indications (such as up, down, left, right, front, back ...) are only used to explain the relative position relationship, movement situation, etc. between each component under a certain specific posture (as shown in the drawings). If this specific posture changes, this directional indication also changes accordingly. In addition, the terms "comprise" and "have" and any of their deformations are intended to cover non-exclusive inclusions. For example, the process, method, system, product or equipment comprising a series of steps or units is not limited to the steps or units listed, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or equipment.
[0023] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0024] After long-term research, researchers have discovered that existing speech synthesis methods typically use models with the same speech synthesis front-end and different speech synthesis back-ends for different speech synthesis styles and speakers. This means that existing speech synthesis systems can better meet the needs of scenarios such as navigation, news broadcasting, e-book reading, and documentary dubbing, but are difficult to meet usage requirements in more personalized human-computer interaction and virtual dialogue scenarios.
[0025] Therefore, in order to achieve more personalized human-computer interaction and virtual dialogue, this application first generates a more colloquial target spoken text, and then generates a more colloquial speech based on the target spoken text. The target spoken text can be obtained by using the spoken text generation model to perform the spoken text generation task on the target written text according to the first instruction content of the prompt.
[0026] Please refer to FIG1 , which is a flow chart of an embodiment of a method for generating spoken text provided by the present application. The method includes:
[0027] S11: Obtain the target written text and prompts.
[0028] This embodiment is used to generate a more colloquial spoken text. In one embodiment, the target written text is a formal text used to generate a spoken text. The target written text does not contain modal particles or contains fewer modal particles than the spoken text. The target written text does not contain pauses, repetitive words, onomatopoeia, etc. The target written text is, for example, a speech, an article, etc. The target written text and prompts can be input by the user. For example, the spoken text generation model provides a human-computer interaction interface, and the user can input the target written text and prompts in the human-computer interaction interface. The target written text and prompts can also be obtained from other devices by the spoken text generation model. In other embodiments, the method for obtaining the target written text and prompts is not limited to the above method, and can also be obtained using any existing data acquisition method.
[0029] The prompt is used to instruct the oral text generation model to perform the oral text generation task. Specifically, the prompt may include the first indication content, which is used to inform the oral text generation model what kind of target oral text to generate according to the target written text. In a specific embodiment, the first indication content includes at least one oral label, and the oral label is used to represent the paralinguistic phenomena to which each word in the target oral text belongs. The oral label may include a modal particle label [YQ], a discourse symbol label [HF], a repetition label [CF], a self-correction label [XZ], a sound imitation label [NS], a stuttering label [KD], etc. For example, for the target oral text "Today, I went shopping with my sister", where "呢" is a modal particle, the label [YQ] of the modal particle can be marked for "呢" in the target oral text, or it can be not marked. For the target oral text "Hello", where "喂" is a discourse symbol. A discourse symbol is a word that can appear alone, while a modal particle cannot appear alone and needs to be associated with other words. For the target oral text "It's hard to say. When I first, first went to Shanghai", where "第一次,第一次" belongs to repetition, the repetition label can be marked for the two words. For the target oral text "You'd better study culture well and don't talk about or engage in these art-related things", where "不要聊不要搞" belongs to self-correction, the self-correction label can be marked for both the words before and after the correction. For the target oral text "With a '呲儿' sound, the tight pants ripped", where "呲儿一声" is a sound imitation word. For the target oral text "When it came to 一九, it was 1988", where the first "一九" is a stutter, the stuttering label can be marked after the first "一九".
[0030] In another specific embodiment, the first prompt content further includes the application scenario of the target oral text, and the application scenario is used to indicate that the target oral text generated by the oral text generation model matches the application scenario. The application scenario may include but is not limited to the gender and personality of the speaker of the oral text, the application scenario of the oral text, oral labels, etc. For example, the prompt is "You are a very cheerful big boy. Transcribe the written text into oral text according to your personality and speaking habits. Oral labels include: modal particle [YQ], discourse symbol [HF], repetition [CF], self-correction [XZ], sound imitation [NS], stuttering [KD]".
[0031] S12: Use the oral text generation model to perform the oral text generation task on the target written text according to the first indication content of the prompt, and obtain the target oral text.
[0032] In an embodiment, the oral text generation model can directly generate the target oral text according to the first indication content.
[0033] In another embodiment, a spoken text generation model can be first used to generate at least one candidate spoken text for the target written text according to the first indication content of the prompt, and the first probability corresponding to each candidate spoken text can be obtained. In a specific embodiment, for each candidate spoken text, its corresponding first probability can be obtained using the third probability of each word in the candidate spoken text. For example, the third probabilities of each word can be summed or weighted to obtain the first probability of the candidate spoken text. Based on the first probabilities corresponding to each candidate spoken text, a candidate spoken text is selected from at least one candidate spoken text as the target spoken text. Specifically, the candidate spoken text corresponding to the largest first probability can be selected as the target spoken text.
[0034] The third probability corresponding to each word in the candidate spoken text can be a probability sequence. In one embodiment, each element in the probability sequence can represent the second probability of the word belonging to each spoken tag. In one embodiment, each element in the probability sequence can be determined by the second probability of the word belonging to each spoken tag and the prior probability of the spoken tag. Specifically, an element in the probability sequence represents the probability of the word corresponding to the spoken tag. Each element in the probability sequence can be equal to the product of the second probability of the word belonging to each spoken tag and the prior probability of the spoken tag. For example, if there are 3 spoken tags, the probability sequence contains 3 elements. Element A represents the probability of the word corresponding to the first spoken tag, which can be equal to the product of the second probability of the word belonging to the first spoken tag and the prior probability of the first spoken tag. Element B represents the probability of the word corresponding to the second spoken tag, which can be equal to the product of the second probability of the word belonging to the second spoken tag and the prior probability of the second spoken tag. Element C represents the probability of the word corresponding to the third spoken tag, which can be equal to the product of the second probability of the word belonging to the third spoken tag and the prior probability of the third spoken tag. At this time, in a specific embodiment, the third probability corresponding to each word in the candidate spoken text is used to obtain the first probability corresponding to the candidate spoken text. The maximum probability can be selected from the probability sequence corresponding to each word as the final third probability of each word. The final third probability of each word is processed to obtain the first probability corresponding to the candidate spoken text. For example, the candidate spoken text contains 3 words and the spoken label contains 2. The third probability of word 1 is [a1, a2] T , the third probability of word 2 is [b1,b2] T , the third probability of word 3 is [c1,c2] T, where a1 is greater than a2, b1 is greater than b2, and c1 is less than c2. Then a1 is taken as the final third probability of word 1, and the target spoken label of word 1 is the spoken label corresponding to probability a1; b1 is taken as the final third probability of word 2, and the target spoken label of word 2 is the spoken label corresponding to probability b1; c2 is taken as the final third probability of word 1, and the target spoken label of word 3 is the spoken label corresponding to probability c2.
[0035] In another specific embodiment, the first probability corresponding to the candidate spoken text is obtained by using the third probability corresponding to each word in the candidate spoken text. Alternatively, a probability can be selected from the probability sequence corresponding to each word each time to form multiple candidate probability sequences, and the candidate probability sequence with the largest probability is selected to obtain the first probability. For example, the candidate spoken text contains 3 words, the spoken label contains 2, and the third probability of word 1 is [a1, a2] T , the third probability of word 2 is [b1,b2] T , the third probability of word 3 is [c1,c2] T , then the candidate probability sequence may include [a1, b1, c1] T , [a1, b1, c2] T , [a1,b2,c1] T , [a1,b2,c2] T ,[a2,b1,c2]T,[a2,b1,c2] T , [a2,b2,c1] T , [a2, b2, c2] T The probability of the candidate probability sequence is obtained by summing each probability in the candidate probability sequence, and the candidate probability sequence with the largest probability is selected to obtain the first probability.
[0036] The third probability corresponding to each word in the candidate spoken text can be a probability value. For example, a spoken label in the first prompt content of the prompt is obtained as the target spoken label, and the product of the second probability that the word belongs to the target spoken label and the prior probability of the target spoken label is used as the third probability corresponding to the word.
[0037] The above scheme obtains a target written text and a prompt, wherein the prompt is used to instruct the spoken text generation model to perform a spoken text generation task. The spoken text generation model then performs the spoken text generation task on the target written text according to the first instruction of the prompt, thereby obtaining a target spoken text. By setting a personalized prompt, the spoken text generation model can generate a more colloquial target spoken text using the target written text according to the first instruction of the prompt.
[0038] In other embodiments, the prompt may further include a second prompt content. After generating the target spoken text, the spoken text generation method further includes: using the spoken text generation model to mark the target spoken text according to the second indication content of the prompt. Specifically, the second indication content of the prompt may include at least one emotional label, and the emotional label includes: [neutral], [surprise], [happiness], [sadness], [fear], [disgust] and other emotions. Then, the spoken text generation model may be used to mark each sentence in the target spoken text according to the second indication content of the prompt at a preset position of each sentence. The preset position may be the beginning of each sentence, and the preset position may be given by the prompt. For example, the emotion adopted by the sentence may be marked at the beginning of each sentence. It is understandable that in other embodiments, the emotion expressed by the sentence may also be marked. To facilitate understanding, the following is an example of a prompt: "You are a very cheerful boy. Based on your personality and speaking habits, transcribe the written language into spoken text. At the same time, judge the emotions and feelings used in each sentence and mark them at the beginning of the sentence. Spoken component labels include: modal particles [YQ], discourse symbols [HF], repetition [CF], self-correction [XZ], onomatopoeia [NS], and freeze [KD]. Emotional labels include: [neutral], [surprise], [happiness], [sadness], [fear], and [disgust]. Please transcribe in a very colloquial form."
[0039] In one embodiment, the spoken text generation method further includes a spoken text generation model training process. Referring to FIG. 2 , FIG. 2 is a flow chart illustrating an embodiment of the spoken text generation model training process provided by the present application. The spoken text generation model training process includes:
[0040] S21: Using at least one sample prompt, respectively control the spoken text generation model to perform training tasks corresponding to each sample prompt on the sample written text, and obtain model output results corresponding to each training task.
[0041] In one embodiment, during the training of a spoken text generation model, the same batch of training data can be used to perform different training tasks to improve the overall performance of the spoken text generation model. A batch of training data can include multiple sample written texts. The user can set at least one sample prompt to control the spoken text generation model to perform training tasks corresponding to each sample prompt on the sample written texts, thereby obtaining model output results corresponding to each training task. In a specific embodiment, there are multiple sample prompts, and the multiple sample prompts can be used to complete different training tasks. The training tasks can include spoken text generation tasks and text analysis tasks. The text analysis tasks can include at least one of part-of-speech tagging, text segmentation, text sentiment classification, dependency parsing, text translation, and text grammar checking. If the training task is a spoken text generation task, the model output can be the target spoken text; if the training task is a text segmentation task, the model output can be the segmentation result of the sample written text; if the training task is a text translation task, the model output can be the translation result of the sample written text. Depending on the training task, each of these tasks will not be listed here. Among them, sample prompts can be used to prompt the model what training tasks to perform and how to perform them.
[0042] It is understandable that in other implementations, the spoken text generation model may also only perform the spoken text generation task.
[0043] S22: Adjusting parameters of the spoken text generation model based on the training differences of each training task, where the training differences of the training tasks are the differences between the model output results corresponding to the training tasks and the annotation results of the sample written text corresponding to the training tasks.
[0044] In one embodiment, after obtaining the model output result, the annotation result of the sample written text corresponding to the training task is obtained, and the difference between the model output result and the annotation result of the sample written text corresponding to the training task is calculated to obtain the training difference. Based on the training difference, the parameters of the spoken text generation model are adjusted. Among them, if the model output result is the target spoken text, the annotation result is the reference spoken text; if the model output result is the word segmentation result after word segmentation of the sample written text, the annotation result is the annotated word segmentation result of the sample written text; if the model output result can be the translation result after translating the sample written text, the annotation result is the annotated translation result of the sample written text. It can be understood that the model output result and the annotation result correspond to each other, and they are not listed one by one here.
[0045] In one embodiment, to improve the spoken language generation capability of a spoken text generation model, spoken audio data is obtained before the spoken text generation model is controlled to execute training tasks corresponding to each sample prompt on sample written text using at least one sample prompt and obtaining model output results corresponding to each training task. The spoken audio data is transcribed to obtain a reference spoken text, and the reference spoken text is used as the annotation result corresponding to the spoken text generation task. Based on the spoken audio data and the reference spoken text, a sample written text corresponding to the spoken text generation task is composed. By utilizing the spoken audio data to obtain the reference spoken text, the accuracy of the reference spoken text can be improved, thereby improving the spoken text generation capability of the spoken text generation model.
[0046] By training the spoken text generation model in the above manner, the spoken text generation capability of the spoken text generation model can be improved, so that the spoken text generation model can generate more colloquial spoken text.
[0047] Please refer to FIG3 , which is a flowchart of an embodiment of a speech synthesis method provided by the present application. The method includes:
[0048] S31: Obtain the target written text and prompts.
[0049] The prompt is used to instruct the spoken text generation model to perform the spoken text generation task.
[0050] S32: Using the spoken text generation model, perform a spoken text generation task on the target written text according to the first instruction content of the prompt to obtain a target spoken text.
[0051] For the specific implementation of steps S31 and S32, please refer to the relevant description in any embodiment of the spoken text generation method provided in this application, which will not be repeated here.
[0052] S33: Synthesize spoken speech based on the target spoken text.
[0053] In one embodiment, after obtaining the target spoken text, any existing speech synthesis method can be used to synthesize spoken speech based on the target spoken text. For example, a speech synthesis model can be used to synthesize spoken speech.
[0054] Please refer to FIG4 , which is a schematic diagram of a framework of an embodiment of a spoken text generation device provided in this application.
[0055] The spoken text generation device 40 includes a first acquisition module 41 and a first generation module 42. The first acquisition module 41 is used to obtain the target written text and prompt words, wherein the prompt words are used to instruct the spoken text generation model to perform the spoken text generation task; the first generation module 42 is used to use the spoken text generation model to perform the spoken text generation task on the target written text according to the first instruction content of the prompt words to obtain the target spoken text.
[0056] In one embodiment, the first generation module 42 is further used to generate at least one candidate spoken text for the target written text according to the first indication content of the prompt using the spoken text generation model, and obtain a first probability corresponding to each candidate spoken text; based on the first probability corresponding to each candidate spoken text, select a candidate spoken text as the target spoken text from the at least one candidate spoken text.
[0057] In one embodiment, for each candidate spoken text, the first generation module 42 is further used to obtain a second probability that each word in the candidate spoken text belongs to the spoken label; for each word in the candidate spoken text, based on the second probability that the word belongs to the target spoken label and the prior probability of the target spoken label, determine a third probability corresponding to the word; and using the third probability corresponding to each word in the candidate spoken text, obtain the first probability corresponding to the candidate spoken text.
[0058] In one embodiment, the first prompt content of the prompt includes at least one spoken tag, and the first acquisition module 41 is also used to obtain at least one spoken tag in the first prompt content of the prompt as a target spoken tag; the first generation module 42 is also used to take the product of the second probability that the word belongs to the target spoken tag and the prior probability of the target spoken tag as the third probability corresponding to the word.
[0059] In one embodiment, the first prompt content of the prompt includes an application scenario of the target spoken text, where the application scenario is used to indicate that the target spoken text generated by the spoken text generation model matches the application scenario.
[0060] In one embodiment, the prompt further includes a second prompt content, and the spoken text generation device further includes a marking module, which is used to use the spoken text generation model to mark the target spoken text according to the second indication content of the prompt.
[0061] In one embodiment, the second indication content of the prompt includes at least one emotional tag; the marking module is also used to use the spoken text generation model to mark each sentence at a preset position in the target spoken text according to the second indication content of the prompt.
[0062] In one embodiment, the spoken text generation device also includes a training module, which is used to use at least one sample prompt to control the spoken text generation model to perform training tasks corresponding to each sample prompt on the sample written text, and obtain the model output results corresponding to each training task; based on the training differences of each training task, the parameters of the spoken text generation model are adjusted, and the training difference of the training task is the difference between the model output result corresponding to the training task and the annotation result of the sample written text corresponding to the training task; wherein the training task includes the spoken text generation task.
[0063] In one embodiment, there are multiple sample prompts, and the training task also includes a text analysis task, which includes at least one of a text part-of-speech tagging task, a text segmentation task, a text sentiment classification task, a dependency syntax analysis task, a text translation task, and a text grammar checking task; and / or, the sample written texts corresponding to different training tasks are different or the same.
[0064] In one embodiment, the first acquisition module 41 is also used to acquire spoken audio data; the spoken text generation device also includes a transcription module and a written text generation module, the transcription module is used to transcribe the spoken audio data to obtain a reference spoken text as a annotation result corresponding to the spoken text generation task; the written text generation module is used to obtain a sample written text corresponding to the spoken text generation task based on the spoken audio data and the reference spoken text.
[0065] In this embodiment, the spoken text generation device 40 is used to execute the spoken text generation method provided above. For the detailed implementation of the spoken text generation method, please refer to the above description and will not be repeated here.
[0066] Please refer to FIG5 , which is a schematic diagram of a framework of an embodiment of a speech synthesis device provided in the present application.
[0067] The speech synthesis device 50 includes a second acquisition module 51, a second generation module 52 and a speech synthesis module 53. The second acquisition module 51 is used to acquire the target written text and prompts, wherein the prompts are used to instruct the spoken text generation model to perform the spoken text generation task; the second generation module 52 is used to use the spoken text generation model to perform the spoken text generation task on the target written text according to the first instruction content of the prompt to obtain the target spoken text; the speech synthesis module 53 is used to synthesize spoken speech based on the target spoken text.
[0068] Please refer to FIG6 , which is a schematic diagram of the framework structure of an embodiment of an electronic device provided in this application.
[0069] The electronic device 60 includes a memory 61 and a processor 62 coupled to each other. The memory 61 stores program instructions, and the processor 62 is configured to execute the program instructions stored in the memory 61 to implement the steps of any of the above-described method implementations. In a specific implementation scenario, the electronic device 60 may include, but is not limited to, a microcomputer and a server. Furthermore, the electronic device 60 may also include a mobile device such as a laptop computer or a tablet computer, which is not limited herein.
[0070] Specifically, the processor 62 is used to control itself and the memory 61 to implement the steps of any of the above-mentioned method implementation methods. The processor 62 can also be called a CPU (Central Processing Unit). The processor 62 may be an integrated circuit chip with signal processing capabilities. The processor 62 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 62 can be implemented by an integrated circuit chip.
[0071] Please refer to FIG. 7 , which is a schematic diagram of a framework of an embodiment of a computer-readable storage medium provided in this application.
[0072] The computer-readable storage medium 70 stores program instructions 71 , which, when executed by a processor, are used to implement the steps of any of the above-mentioned method implementations.
[0073] The computer-readable storage medium 70 can specifically be a medium that can store computer programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or it can also be a server that stores the computer program. The server can send the stored computer program to other devices for execution, or it can also run the stored computer program itself.
[0074] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.
[0075] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation methods described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0076] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0077] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0078] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0079] If the technical solution of this application involves personal information, the product that applies the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves sensitive personal information, the product that applies the technical solution of this application has obtained the individual's separate consent before processing sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information; among which, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0080] The above description is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for generating spoken text, wherein: include: Obtaining a target written text and a prompt, wherein the prompt is used to instruct a spoken text generation model to perform a spoken text generation task; The spoken text generation model is used to perform the spoken text generation task on the target written text according to the first instruction content of the prompt to obtain a target spoken text.
2. The method according to claim 1, wherein The step of using the spoken text generation model to perform the spoken text generation task on the target written text according to the first instruction content of the prompt to obtain the target spoken text includes: generating at least one candidate spoken text for the target written text according to the first indication content of the prompt using a spoken text generation model, and obtaining a first probability corresponding to each candidate spoken text; Based on the first probability corresponding to each candidate spoken text, a candidate spoken text is selected from the at least one candidate spoken text as the target spoken text.
3. The method according to claim 2, wherein: Obtaining the first probability corresponding to each candidate spoken text includes: For each candidate spoken text, obtaining a second probability that each word in the candidate spoken text belongs to the spoken label; For each word in the candidate spoken text, determining a third probability corresponding to the word based on the second probability that the word belongs to the target spoken label and the prior probability of the target spoken label; The first probability corresponding to the candidate spoken text is obtained by using the third probability corresponding to each word in the candidate spoken text.
4. The method according to claim 3, wherein: The first prompt content of the prompt includes at least one spoken language tag. Before determining, for each word in the candidate spoken text, a third probability corresponding to the word based on the second probability that the word belongs to the target spoken language tag and the prior probability of the target spoken language tag, the method further includes: Obtain the at least one spoken tag in the first prompt content of the prompt as the target spoken language label; And / or, determining a third probability corresponding to the word based on the second probability that the word belongs to the target spoken language label and the prior probability of the target spoken language label includes: The product of the second probability that the word belongs to the target spoken language label and the prior probability of the target spoken language label is used as the third probability corresponding to the word.
5. The method according to claim 3, wherein: The determining, for each word in the candidate spoken text, a third probability corresponding to the word based on the second probability that the word belongs to the target spoken label and the prior probability of the target spoken label, includes: For each of the words, obtaining the product of the second probability that the word belongs to each spoken label and the prior probability of the spoken label, to obtain the probability that the word corresponds to each spoken label; From the probabilities of the word corresponding to the spoken labels, the largest probability is selected as the final third probability of the word; wherein the spoken label corresponding to the largest probability is used as the target spoken label of the word.
6. The method according to claim 1, wherein The first prompt content of the prompt includes an application scenario of the target spoken text, and the application scenario is used to indicate that the target spoken text generated by the spoken text generation model matches the application scenario.
7. The method according to claim 1, wherein The prompt also includes a second prompt content, and the method further includes: The target spoken text is marked using the spoken text generation model according to the second instruction content of the prompt.
8. The method according to claim 7, wherein: The second indication content of the prompt includes at least one emotion tag; And / or, the using the spoken text generation model to mark the target spoken text according to the second instruction content of the prompt includes: The spoken text generation model is used to mark each sentence at a preset position in the target spoken text according to the second instruction content of the prompt.
9. The method according to claim 1, wherein The method further comprises: Using at least one sample prompt, the spoken text generation model is controlled to generate the sample The written text executes the training tasks corresponding to the sample prompts to obtain the model output results corresponding to the training tasks; Based on the training differences of each of the training tasks, the parameters of the spoken text generation model are adjusted, wherein the training differences of the training tasks are the differences between the model output results corresponding to the training tasks and the annotation results of the sample written text corresponding to the training tasks; wherein the training tasks include spoken text generation tasks.
10. The method according to claim 9, wherein: There are multiple sample prompts, and the training task also includes a text analysis task, which includes at least one of a part-of-speech tagging task, a text segmentation task, a text sentiment classification task, a dependency syntax analysis task, a text translation task, and a text grammar checking task; And / or, the sample written texts corresponding to different training tasks are different or the same.
11. The method according to claim 9, wherein: Before using at least one sample prompt to control the spoken text generation model to perform training tasks corresponding to each sample prompt on the sample written text and obtaining model output results corresponding to each training task, the method further includes: Get spoken audio data; Transcribing the spoken audio data to obtain a reference spoken text as a labeling result corresponding to the spoken text generation task; Based on the spoken audio data and the reference spoken text, a sample written text corresponding to the spoken text generation task is obtained.
12. A speech synthesis method, wherein: include: Acquiring a target written text and a prompt, wherein the prompt is used to instruct the spoken text generation model to perform a spoken text generation task; Using a spoken text generation model, performing the spoken text generation task on the target written text according to the first instruction content of the prompt, to obtain a target spoken text; Based on the target spoken text, a spoken speech is synthesized.
13. A spoken text generation device, wherein: include: A first acquisition module is configured to acquire a target written text and a prompt, wherein the prompt is used to instruct the spoken text generation model to perform a spoken text generation task; The first generation module is configured to use a spoken text generation model to perform the spoken text generation task on the target written text according to the first instruction content of the prompt to obtain a target spoken text.
14. A speech synthesis device, wherein: include: A second acquisition module is configured to acquire a target written text and a prompt, wherein the prompt is used to instruct the spoken text generation model to perform a spoken text generation task; A second generation module is configured to use a spoken text generation model to perform the spoken text generation task on the target written text according to the first instruction content of the prompt to obtain a target spoken text; The speech synthesis module is used to synthesize spoken speech based on the target spoken text.
15. An electronic device, wherein: include: A memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is configured to execute the program instructions stored in the memory to perform the method according to any one of claims 1 to 12.
16. A computer-readable storage medium, wherein: The computer-readable storage medium stores a program file, and the program file can be executed to implement the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Spoken language text generation method, speech synthesis method and related devices
CN117975936A
Speech synthesis method and related device, electronic equipment and storage medium
CN114299911A
Spoken language text generation method and device, equipment and storage medium
CN115081459A
Voice generation method and device, equipment and storage medium
CN115472185A
Text-based speech generation
CN115602145A