Text-based voice generation method and device and electronic equipment
By using the mute duration between words in the speech generation system for fine-grained pronunciation labeling, combined with multimodal deep learning and ASR technology, the time-consuming and labor-intensive and labeling distortion problems of manual pronunciation in existing systems are solved, achieving a more natural and smooth speech synthesis effect.
Patent Information
- Application Number
- CN202311617843.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
The existing speech generation system requires manual annotation of pronunciation information, which is time-consuming and laborious and not suitable for large-scale data sets. In addition, the traditional pronunciation labeling scheme is separated from the pronunciation itself, making it prone to pronunciation distortion problems.
The mute duration between words is used to replace the traditional pronunciation level, and finer-grained pronunciation annotation is performed. The multimodal deep learning scheme and ASR technology are used to identify the interval duration between speech text and text, and the NLP model is combined for pinyin annotation and correction.
It realizes more refined speech synthesis control, improves the naturalness and fluency of speech synthesis, optimizes the accuracy of polyphonic characters and voice change phenomena in pinyin, and is suitable for large-scale data sets.
Smart Images

Figure CN120071890A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech synthesis technology, and particularly to a text-based speech generation method, device, storage medium, and electronic device. Background Art
[0002] The purpose of the background art provided here is to generally give the background of this application. The statements in this part only provide the background related to this application and do not necessarily constitute the prior art.
[0003] Currently, with the rise of short video software, a large number of audio and video files are generated on the Internet every day.
[0004] However, current speech generation systems often require manual annotation of prosody information, which makes this processing scheme not only time-consuming and laborious, but also not applicable to large-scale data sets.
[0005] In addition, traditional speech generation prosody schemes predict prosody based on text. For example, existing prosody annotations represent the intervals between speech, but the durations of these intervals are not specified. Since they are divorced from the speech itself, problems such as prosody annotation distortion often occur. Summary of the Invention
[0006] In view of the above problems, this application proposes a text-based speech generation method, device, storage medium, and electronic device. By using the silent duration between words (i.e., the interval duration between words) to replace the traditional prosody level, finer-grained prosody annotation is performed, thereby achieving more refined speech synthesis control.
[0007] In the first aspect of this application, a text-based speech generation method is provided. The method includes:
[0008] Determine the pinyin sequence of the target text; wherein, the pinyin of all words in the target text is saved in the pinyin sequence;
[0009] Determine the inter-word interval sequence of the target text; wherein, the interval duration between all adjacent words in the target text is saved in the inter-word interval sequence;
[0010] Generate the speech file of the target text through the trained TTS model according to the inter-word interval sequence and the pinyin sequence.
[0011] Further, the pinyin sequence and the inter-word interval sequence of the target text are determined by a preset text processing model, and the preset text processing model includes an NLP model.
[0012] Further, the pinyin of each word in the target text in the pinyin sequence is saved in a preset order;
[0013] In the inter-character interval sequence, the interval durations between all adjacent characters in the target text are saved in the preset order.
[0014] Furthermore, it further includes:
[0015] Training the TTS model with preset training data to obtain the trained TTS model; wherein, in the preset training data, there are included a pinyin sequence, an inter-character interval sequence, and audio data determined based on preset video data, and the audio data is used as a label.
[0016] Furthermore, both the pinyin sequence and the inter-character interval sequence are generated based on an annotated text, and the determination method of the annotated text includes: screening out the audio data from the video data, and determining the voice emission time and the ASR recognition result from the audio data through a preset ASR recognition model;
[0017] Determining the OCR recognition result from the video frame through a preset OCR recognition model;
[0018] Performing background denoising processing on the OCR recognition result according to the ASR recognition result to obtain a first processing result, and performing misspelling correction processing on the ASR recognition result according to the OCR recognition result to obtain a second processing result;
[0019] Determining the annotated text according to the first processing result and the second processing result.
[0020] Furthermore, the determination method of the pinyin sequence includes:
[0021] Performing pinyin annotation on the annotated text through the NLP model to obtain an initial pinyin sequence of the annotated text;
[0022] Performing pinyin annotation on the tonal change part in the initial pinyin sequence again through the NLP model to obtain a processed pinyin sequence and determining the processed pinyin sequence as the pinyin sequence.
[0023] Furthermore, the determination method of the inter-character interval sequence includes:
[0024] Identifying the voice emission time period of the annotated text through an ASR recognition model;
[0025] Within the voice emission time period, determining the interval durations between all adjacent characters in the annotated text through a VAD recognition model to determine the inter-character interval sequence of the annotated text and determining the inter-character interval sequence as the inter-character interval sequence.
[0026] In a second aspect of the present application, there is provided a text-based speech generation device, the device comprising:
[0027] a pinyin sequence determination module, configured to determine the pinyin sequence of the target text; wherein, all the pinyins of the characters in the target text are stored in the pinyin sequence;
[0028] a character interval sequence determination module, configured to determine the character interval sequence of the target text; wherein, the interval durations between all adjacent characters in the target text are stored in the character interval sequence;
[0029] a synthesis module, configured to generate a speech file of the target text through a trained TTS model according to the character interval sequence and the pinyin sequence.
[0030] In a third aspect of the present application, there is provided a computer-readable storage medium, and a computer program stored in the computer-readable storage medium can be executed by one or more processors to implement the steps of the method as described above.
[0031] In a fourth aspect of the present application, there is provided an electronic device, comprising a memory and one or more processors, a computer program is stored on the memory, and the memory and the one or more processors are communicatively connected to each other. When the computer program is executed by the one or more processors, the steps of the method as described above are implemented.
[0032] Compared with the prior art, the advantages or beneficial effects of the technical solution of the present application include:
[0033] Speech synthesis is performed through a multi-modal deep learning solution, and the existing annotation scheme is improved; the accuracy of recognizing speech text is improved by using visual and speech features; the interval duration between speech text and characters is recognized through ASR technology; the pinyin information is corrected through ASR and NLP, and the accuracy of polyphonic characters and tone-changing phenomena in pinyin is optimized, thereby improving the speech synthesis effect. Description of the Drawings
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0035] It should also be noted that, for ease of description, only the parts related to the present disclosure are shown in the drawings. The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The illustrative embodiments and descriptions in this application are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0036] Figure 1 is a flowchart of a text-based speech generation method provided by an embodiment of this application;
[0037] Figure 2 is a schematic diagram of a text annotation provided by an embodiment of this application;
[0038] Figure 3 is another schematic diagram of a text annotation provided by an embodiment of this application;
[0039] Figure 4 is a schematic diagram of determining the inter-character interval duration provided by an embodiment of this application;
[0040] Figure 5 is a schematic diagram of an overall annotation scheme provided by an embodiment of this application;
[0041] Figure 6 is another text-based speech synthesis flowchart provided by an embodiment of this application. Detailed Embodiments
[0042] The following will describe in detail the implementation manners of this application in combination with the drawings and embodiments, so as to fully understand how this application uses technical means to solve technical problems and the implementation process of achieving corresponding technical effects and implement accordingly. The embodiments of this application and each feature in the embodiments can be combined with each other without conflict, and the formed technical solutions are all within the protection scope of this application.
[0043] It should be clear that the embodiments described below are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments in this application without creative efforts belong to the protection scope of this application.
[0044] The technical solution disclosed in this application can be used to solve the annotation of speech generation data.
[0045] Under the current technology, in the field of speech generation, improving the naturalness and fluency of synthesized speech has always been the focus of research. Prosody is one of the key factors determining the fluency and naturalness of speech. Current speech generation systems often require manual annotation of prosody information. However, this method is time-consuming and laborious, and is not applicable to large-scale data sets. Therefore, an efficient automated method is needed to annotate the prosody information of speech generation data.
[0046] In view of this, the present application discloses a text-based speech generation method, which at least solves the above-mentioned defects.
[0047] The present application aims to solve the problem of generating annotated speech synthesis data from videos, aiming to use a multimodal deep learning scheme to perform speech synthesis data on videos and improve it according to existing annotations, and proposes a new prosody annotation method. Currently, the existing data prosody annotation method is as follows:
[0048] The prosody is annotated by the symbols "#0#1#2#3#4", where the meanings represented by each symbol are as follows:
[0049] #0: Word boundary;
[0050] #1: Word boundary (word segmentation result);
[0051] #2: Phrase boundary (between 1 and 3);
[0052] #3: Short sentence boundary (with an audible pause);
[0053] #4: Long sentence boundary (long pause, generally at the end of the sentence).
[0054] The existing prosody annotation represents the interval between voices. However, these intervals do not have a specified specific interval duration. In view of this, this solution selects the interval duration between words as the prosody representation to enable more refined speech synthesis control.
[0055] For example, the existing prosody annotation represents the interval between voices and is annotated by the symbols "#0#1#2#3#4":
[0056]
[0057] In the prosody annotation scheme proposed in this solution, the prosody is annotated by the interval duration between words. For example:
[0058]
[0059] As shown above, the numbers under each word represent the interval duration between the corresponding word and the next word. For example, the first "50" represents that the pronunciation interval duration between the word "person" and the next word "of" is 50 ms, and the first "100" represents that the pronunciation interval duration between the word "of" and the next word "year" is 100 ms, and so on.
[0060] The present application provides a text-based speech generation method. Figure 1 is a flowchart of a text-based speech generation method provided by an embodiment of the present application, as Figure 1As shown in the figure, the method disclosed in this application includes the following steps:
[0061] Step 110: Determine the pinyin sequence of the target text; wherein, the pinyin of all characters in the target text is stored in the pinyin sequence.
[0062] Optionally, the pinyin sequence of the target text is determined by using a preset text processing model.
[0063] As an example, the preset text processing model includes:
[0064] NLP (Natural Language Processing, hereinafter referred to as NLP) model.
[0065] Optionally, the NLP model includes Bert (full English name Bidirectional Encoder Representations from Transformers), LSTM (full English name Long Short Term Memory networks), etc. Which NLP model to specifically select can be determined according to actual needs.
[0066] Furthermore, the pinyin of each character in the target text in the pinyin sequence is stored in a preset order.
[0067] Optionally, in the pinyin sequence, the pinyin of each character in the target text is stored in sequence according to the order of the characters in the target text from left to right. That is to say, the preset order can be set as the order from left to right, or can be set according to actual needs.
[0068] Step 120: Determine the inter-character interval sequence of the target text; wherein, the interval duration between all adjacent characters in the target text is stored in the inter-character interval sequence.
[0069] Optionally, the inter-character interval sequence of the target text is determined by using a preset text processing model, and the preset text processing model includes an NLP model.
[0070] Furthermore, the interval duration between all adjacent characters in the target text in the inter-character interval sequence is stored in the preset order.
[0071] Step 130: Generate a voice file of the target text through a trained TTS (Text-To-Speech, hereinafter referred to as TTS) model according to the inter-character interval sequence and the pinyin sequence.
[0072] As an example, it further includes:
[0073] Train the TTS model with preset training data to obtain the trained TTS model; wherein, the preset training data includes a pinyin sequence, an inter-character interval sequence, and audio data determined based on preset video data, and the audio data is used as a label.
[0074] Wherein, the method for determining the audio data according to the preset video data can be selected according to actual needs, and the audio data includes the audio information corresponding to the preset video data.
[0075] This solution performs speech information annotation through video-audio-text multi-modalities to achieve more accurate data annotation. Specifically, the following takes video data as an example for illustration:
[0076] As an example, both the pinyin sequence and the inter-character interval sequence are generated based on the annotated text, and the method for determining the annotated text includes:
[0077] Determine the speech emission time and the ASR recognition result from the audio data through a preset ASR (Automatic Speech Recognition) recognition model;
[0078] Determine the video frame corresponding to the speech emission time from the video data, and determine the OCR recognition result from the video frame through a preset OCR (Optical Character Recognition) recognition model;
[0079] Perform background denoising processing on the OCR recognition result according to the ASR recognition result to obtain a first processing result, and perform misspelling correction processing on the ASR recognition result according to the OCR recognition result to obtain a second processing result;
[0080] Determine the annotated text according to the first processing result and the second processing result.
[0081] In this application, speech text annotation is performed based on video and speech dual-modalities. Simply using ASR to recognize the text content for annotation may result in annotation errors. Therefore, OCR and ASR can be used to complement each other in features. The ASR technology is used to delimit the text range and determine the image sequence time, that is, the ASR is used to determine the time when the speech appears. Therefore, the OCR recognition results at other times can be de-duplicated, and the background text of the OCR recognition text can also be removed through the ASR technology. At the same time, the OCR recognition result can be used to correct misspellings in the ASR recognition result. The specific steps may include:
[0082] 1. Duplicate removal using time: The ASR recognition model recognizes the speech recognition result to obtain the speech text and the corresponding time series. Therefore, it is only necessary to process the images within this period through the OCR recognition model, which greatly reduces the complexity of the OCR recognition model's work.
[0083] 2. Background text removal: The traditional method uses the difference and position information between the previous frame and the subtitle frame for duplicate removal, while this application proposes a method for background removal through the text similarity between the ASR recognition result and the OCR recognition result. Please refer to Figure 2 and Figure 3 , which specifically includes the following steps:
[0084] (1) Perform time alignment: Take the frame where the results before and after the OCR recognition result change as t1, and the longest continuous ASR time containing t1 as t2.
[0085] (2) Obtain the ASR recognition result text text1 within the t2 time, and obtain the OCR recognition result text text2 at the t1 moment.
[0086] (3) Divide the OCR recognition result text text2 into different text blocks according to the distance principle: text 1, text 2, text 3, text 4......
[0087] (4) Calculate the length (which can be set as the longest subsequence) of {text 1, text 2, text 3, text 4} existing in text1.
[0088] (5) Take the longest text in the longest subsequence (i.e., text 1) as the correct text.
[0089] In this solution, it can be considered that the subtitles recognized by the OCR recognition model are more accurate text recognition results, and the ASR recognition result is used as an auxiliary positioning means to locate the OCR recognition result more accurately in terms of time and physical position.
[0090] As an example, the determination method of the inter-character interval sequence includes:
[0091] Recognize the speech emission time period of the labeled text through the ASR recognition model.
[0092] Within the speech emission time period, determine the interval duration between all adjacent characters in the labeled text through the VAD (Voice Activity Detection, abbreviated as VAD) recognition model to determine the inter-character interval sequence of the labeled text and determine the inter-character interval sequence as the inter-character interval sequence.
[0093] Specifically, this application proposes a solution to improve the accuracy of ASR timestamps using a VAD recognition model, which improves the accuracy of the voice mute segment using the VAD recognition model. As shown in Figure 4 Using the ASR recognition model to recognize the text content and corresponding time period in the voice; using the VAD recognition model to recognize and detect the corresponding mute time period within this time period as the interval duration between adjacent texts.
[0094] As an example, the determination method of the pinyin sequence includes:
[0095] Performing pinyin annotation on the annotated text through the NLP model to obtain the initial pinyin sequence of the annotated text;
[0096] Performing pinyin annotation on the tone-changing part in the initial pinyin sequence again through the NLP model to obtain the processed pinyin sequence and using the processed pinyin sequence as the pinyin sequence.
[0097] This application proposes a pinyin tone-changing annotation strategy based on the ASR recognition model. Pinyin annotation is an essential part of Chinese TTS, and its role is to convert Chinese into phonemes with finer granularity and more generality in pinyin annotation. Specifically, it can include the following steps:
[0098] (1) Data preparation: Collect the TTS training data set with text annotations, ensure the data quality, including covering target vocabulary that needs tone-changing, such as "once a year", "one by one"; further perform data preprocessing, remove noise, and normalize the text to ensure the consistency and quality of the data;
[0099] (2) Extract pinyin and tone-changing annotations: Use a Chinese pinyin annotation tool (such as an NLP model) to perform pinyin annotation on the text data to obtain the corresponding text pinyin sequence; use a text annotation tool to manually annotate the tone-changing part, that is, represent the tone-changing sounds of characters such as "one" with special marks;
[0100] (3) ASR recognition model training: Train the speech recognition model, input content: speech; output content: corresponding speech pinyin with tone-changing;
[0101] (3) Perform pinyin annotation on the training data of the TTS to be annotated using the ASR recognition model.
[0102] Furthermore, during the training process of the TTS model:
[0103] It is necessary to input text, pinyin sequence, inter-character interval sequence, and speech; use the above text, pinyin sequence, and inter-character interval sequence as the input, and speech as the target, and input them into the TTS model for training and tuning to finally obtain a trained TTS model.
[0104] Optionally, the annotation scheme disclosed in this application can also refer to Figure 5 .
[0105] Furthermore, in the process of generating speech through the TTS model:
[0106] Only the input text is required, and then the corresponding pinyin sequence and inter-character interval sequence are determined according to the input text, and then the corresponding speech file is generated according to the pinyin sequence and inter-character interval sequence.
[0107] Optionally, it can also refer to Figure 6 , Figure 6 which is another text-based speech synthesis flowchart provided by the embodiments of this application. According to the Figure 6 shown speech synthesis process: First, a factor sequence (pinyin sequence of the speech text) and an interval sequence (inter-character interval sequence of the speech text) are generated according to the input speech text; then, the interval sequence is processed into corresponding vector data through the prosody module of the TTS model; finally, the factor sequence and the vector data corresponding to the interval sequence are synthesized into speech through the speech synthesis module of the TTS model and output.
[0108] In this application, through OCR and ASR recognition technologies, the content contained in the speech is recognized, and the accuracy of recognizing the speech text is improved by using visual and speech features; the interval duration between the speech text and the characters is recognized through the ASR technology; the pinyin information is corrected through ASR and NLP, optimizing the accuracy of polyphonic characters and tone-changing phenomena in the pinyin, and thus improving the speech synthesis effect.
[0109] This application also provides a text-based speech generation device. The embodiments of this device can be used to execute the method embodiments of this application. For the details not disclosed in the embodiments of this device, please refer to the method embodiments of this application. The device disclosed in this application includes:
[0110] A pinyin sequence determination module, configured to determine the pinyin sequence of the target text; wherein, all the pinyins of the characters in the target text are stored in the pinyin sequence;
[0111] An inter-character interval sequence determination module, configured to determine the inter-character interval sequence of the target text; wherein, the interval duration between all adjacent characters in the target text is stored in the inter-character interval sequence;
[0112] A synthesis module, configured to generate a speech file of the target text through a trained TTS model according to the inter-character interval sequence and the pinyin sequence.
[0113] In some embodiments, the pinyin sequence and the inter-character interval sequence of the target text are determined by a preset text processing model, and the preset text processing model includes an NLP model.
[0114] In some embodiments, in the pinyin sequence, the pinyin of each character in the target text is saved in a preset order;
[0115] In the inter-character interval sequence, the interval durations between all adjacent characters in the target text are saved in the preset order.
[0116] In some embodiments, it further includes a training module for training the TTS model with preset training data to obtain the trained TTS model; wherein, the preset training data includes labeled text, pinyin sequence, and inter-character interval sequence generated based on video data.
[0117] In some embodiments, the method for determining the labeled text includes:
[0118] Filter out the audio data from the video data, and determine the voice emission time and the ASR recognition result from the audio data through a preset ASR recognition model;
[0119] Determine the video frame corresponding to the voice emission time from the video data, and determine the OCR recognition result from the video frame through a preset OCR recognition model;
[0120] Perform background denoising processing on the OCR recognition result according to the ASR recognition result to obtain a first processing result, and perform misspelling correction processing on the ASR recognition result according to the OCR recognition result to obtain a second processing result;
[0121] Determine the labeled text according to the first processing result and the second processing result.
[0122] In some embodiments, the method for determining the pinyin sequence includes:
[0123] Perform pinyin annotation on the labeled text through the NLP model to obtain the initial pinyin sequence of the labeled text;
[0124] Perform pinyin annotation on the tone-changing part in the initial pinyin sequence again through the NLP model to obtain the processed pinyin sequence, and use the processed pinyin sequence as the pinyin sequence.
[0125] In some embodiments, the method for determining the inter-character interval sequence includes:
[0126] Identify the voice emission time period of the labeled text through an ASR recognition model;
[0127] During the voice emission time period, the VAD recognition model is used to determine the interval duration between all adjacent words in the labeled text, so as to determine the inter-word interval sequence of the labeled text and determine the inter-word interval sequence as the inter-word interval sequence.
[0128] Those skilled in the art should understand that the above-mentioned modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module for implementation.
[0129] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the modules in the text-based voice generation device can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.
[0130] The present application also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the method steps in the foregoing method embodiments can be implemented and will not be repeated here.
[0131] Among them, the computer-readable storage medium may also separately include a computer program, a data file, a data structure, etc., or include a combination thereof. The computer-readable storage medium or the computer program can be specifically designed and understood by those skilled in the computer software field, or the computer-readable storage medium may be well-known and available to those skilled in the computer software field. Examples of computer-readable storage media include: magnetic media, such as hard disks, floppy disks, and magnetic tapes; optical media, such as CD-ROM discs and DVDs; magneto-optical media, such as optical discs; and hardware devices specifically configured to store and execute computer programs, such as read-only memory (ROM), random access memory (RAM), flash memory; or servers, app application stores, etc. Examples of computer programs include machine code (e.g., code generated by a compiler) and files containing high-level code that can be executed by a computer by using an interpreter. The described hardware devices can be configured to act as one or more software modules to execute the operations and methods described above, and vice versa. In addition, the computer-readable storage medium can be distributed in a networked computer system and can store and execute program codes or computer programs in a decentralized manner.
[0132] The present application also provides a computer program product. The computer program product includes a computer program or instructions which, when executed by a processor, implement all or part of the steps of the method in the foregoing method embodiments, and details thereof will not be repeated herein.
[0133] Further, the computer program product may include one or more computer-executable components configured to execute the embodiments when the program is running; the computer program product may also include a computer program tangibly embodied on a computer-readable medium, the computer program including program code for executing any method in the embodiments of the present disclosure. In such an embodiment, the computer program may be downloaded and installed from a network through a communication part, and / or installed from a removable medium.
[0134] The present application also provides an electronic device, which may include: one or more processors, a memory, a multimedia component, an input / output (I / O) interface, and a communication component.
[0135] Wherein, the one or more processors are configured to execute all or part of the steps in the foregoing method embodiments. The memory is configured to store various types of data, which may include, for example, instructions of any application program or method in the electronic device, as well as data related to the application program.
[0136] The one or more processors may be implemented by an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components, and are configured to execute the method in the foregoing method embodiments.
[0137] The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0138] The multimedia component may include a screen and an audio component. The screen can be a touch screen. The audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal can be further stored in the memory or sent via the communication component. The audio component also includes at least one speaker for outputting audio signals.
[0139] The I / O interface provides an interface between one or more processors and other interface modules, and the other interface modules can be a keyboard, a mouse, buttons, etc. These buttons can be virtual buttons or physical buttons.
[0140] The communication component is used for wired or wireless communication between the electronic device and other devices. Wired communication includes communication via a network port, a serial port, etc.; wireless communication includes Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G, 4G, 5G, or a combination of one or more of them.
[0141] It should also be understood that the methods or devices disclosed in the embodiments provided in this application can also be implemented in other ways. The method or system embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the methods and devices according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a computer program segment, or a part of a computer program. A module, a computer program segment, or a part of a computer program contains one or more computer programs for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. In fact, they may also be executed substantially in parallel, and sometimes they may be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and a computer program.
[0142] In this application, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, device or apparatus comprising the element; if there is a description of "first", "second", etc., it is only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features or implicitly specifying the sequence of the indicated technical features; in the description of this application, unless otherwise specified, the meaning of the terms "plurality", "many" is at least two; if there is a description of a server, it should be noted that the server can be an independent physical server or terminal, or a server cluster composed of multiple physical servers, and can be a cloud server capable of providing basic cloud computing services such as cloud servers, cloud databases, cloud storage and CDN; in this application, if there is a description of a smart terminal or mobile device, it should be noted that the smart terminal or mobile device can be a mobile phone, tablet computer, smart watch, netbook, wearable electronic device, personal digital assistant (Personal Digital Assistant, abbreviated as PDA), augmented reality device (Augmented Reality, abbreviated as AR), virtual reality device (Virtual Reality, abbreviated as VR), smart TV, smart speaker, personal computer (Personal Computer, abbreviated as PC), etc., but is not limited thereto, and this application does not make special limitations on the specific form of the smart terminal or mobile device.
[0143] Finally, it should be noted that in the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "one example" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0144] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are all exemplary. The content described above is only an implementation manner adopted for the convenience of understanding the present application, and is not used to limit the present application. Any person skilled in the art within the technical field to which the present application pertains may make any modifications and changes in the form of implementation and details without departing from the spirit and scope disclosed by the present application. However, the protection scope of the present application shall still be subject to the scope defined by the appended claims.
Claims
1. A text-based speech generation method, characterized in that, the method includes: Determine the pinyin sequence of the target text; wherein, the pinyin of all characters in the target text is stored in the pinyin sequence; Determine the inter-character interval sequence of the target text; wherein, the interval duration between all adjacent characters in the target text is stored in the inter-character interval sequence; According to the inter-character interval sequence and the pinyin sequence, generate the speech file of the target text through the trained TTS model.
2. The text-based speech generation method according to claim 1, characterized in that, the pinyin sequence and the inter-character interval sequence of the target text are determined by a preset text processing model, and the preset text processing model includes an NLP model.
3. The text-based speech generation method according to claim 1, characterized in that, the pinyin of each character in the target text in the pinyin sequence is stored in a preset order; the interval duration between all adjacent characters in the target text in the inter-character interval sequence is stored in the preset order.
4. The text-based speech generation method according to claim 1, characterized in that, further includes: Train the TTS model with preset training data to obtain the trained TTS model; wherein, the preset training data includes a pinyin sequence, an inter-character interval sequence, and audio data determined based on preset video data, and the audio data is used as a label.
5. The text-based speech generation method according to claim 4, characterized in that, both the pinyin sequence and the inter-character interval sequence are generated based on an annotated text, and the determination method of the annotated text includes: Determine the speech emission time and the ASR recognition result from the audio data through a preset ASR recognition model; Determine the video frame corresponding to the speech emission time from the video data, and determine the OCR recognition result from the video frame through a preset OCR recognition model; Perform background denoising processing on the OCR recognition result according to the ASR recognition result to obtain a first processing result, and perform typo correction processing on the ASR recognition result according to the OCR recognition result to obtain a second processing result; Determine the annotated text according to the first processing result and the second processing result.
6. The text-based speech generation method according to claim 5, characterized in that, the determination method of the pinyin sequence includes: Perform pinyin annotation on the annotated text through the NLP model to obtain the initial pinyin sequence of the annotated text; Perform pinyin annotation on the tone-changing part in the initial pinyin sequence again through the NLP model to obtain the processed pinyin sequence and determine the processed pinyin sequence as the pinyin sequence.
7. The text-based speech generation method according to claim 5, characterized in that, the determination method of the inter-character interval sequence includes: Identify the speech emission time period of the annotated text through an ASR recognition model; During the voice emission time period, the VAD recognition model is used to determine the interval duration between all adjacent characters in the labeled text, so as to determine the inter-character interval sequence of the labeled text and determine the inter-character interval sequence as the inter-character interval sequence.
8. A text-based voice generation device Characterized in that Comprising: A pinyin sequence determination module, configured to determine the pinyin sequence of the target text; wherein, the pinyin of all characters in the target text is stored in the pinyin sequence; An inter-character interval sequence determination module, configured to determine the inter-character interval sequence of the target text; wherein, the interval duration between all adjacent characters in the target text is stored in the inter-character interval sequence; A synthesis module, configured to generate a voice file of the target text through a trained TTS model according to the inter-character interval sequence and the pinyin sequence.
9. A computer-readable storage medium Characterized in that The computer program stored in the computer-readable storage medium, when executed by one or more processors, implements the image quality compensation method of the display module according to any one of claims 1 to 7.
10. An electronic device Characterized in that Comprising a memory and a processor, a computer program is stored on the memory, and when the computer program is executed by the processor, the text-based voice generation method according to any one of claims 1 to 7 is implemented.