A speech synthesis method, apparatus and electronic device

By using deep learning models and speech correction technology to convert government documents into speech, the problem of low efficiency in handling business through self-service machines in government service halls has been solved, enabling efficient completion of user operation guidance and business processes.

CN116264072BActive Publication Date: 2026-04-14AISINO CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-15
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

The self-service machines in the government service hall are inefficient, and some users are unable to complete their transactions smoothly.

Method used

The government documents are converted into pinyin sequences using a deep learning model, and standard Mandarin speech is generated using a speech correction model and a vocoder to guide user operations.

Benefits of technology

It has improved the efficiency of handling business in the government service hall and helped users complete business processes smoothly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116264072B_ABST
    Figure CN116264072B_ABST
Patent Text Reader

Abstract

The application discloses a speech synthesis method and device and electronic equipment to improve the efficiency of government office business. The method comprises the following steps: inputting a first government text into a first deep learning model to obtain a first pinyin sequence corresponding to the first government text; wherein the first pinyin sequence indicates the pinyin and pause position corresponding to the first government text; inputting the first pinyin sequence into a second deep learning model to obtain a first acoustic feature sequence corresponding to the first pinyin sequence; wherein the first acoustic feature sequence indicates the acoustic features of the first pinyin sequence; and inputting the first acoustic feature sequence into a vocoder to obtain a first voice corresponding to the first government text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a speech synthesis method, apparatus and electronic device. Background Technology

[0002] Currently, in order to improve the efficiency of users handling business, the government has set up multiple self-service machines in government service halls across the country. However, due to the wide variety of services and the fact that users come from multiple age groups, a large number of users are unable to complete their business smoothly when using the self-service machines, which reduces the efficiency of business handling in government service halls. Summary of the Invention

[0003] This application provides a speech synthesis method, apparatus, and electronic device to improve the efficiency of handling business in government service halls.

[0004] In a first aspect, this application provides a speech synthesis method, the method comprising:

[0005] The first government document is input into the first deep learning model to obtain the first pinyin sequence corresponding to the first government document; wherein, the first pinyin sequence indicates the pinyin and pause position corresponding to the first government document.

[0006] The first pinyin sequence is input into the second deep learning model to obtain the first acoustic feature sequence corresponding to the first pinyin sequence; wherein, the first acoustic feature sequence indicates the acoustic features of the first pinyin sequence;

[0007] The first acoustic feature sequence is input into a vocoder to obtain the first speech corresponding to the first government document.

[0008] The above method, through a deep learning model, can efficiently convert the first government document into speech to guide users, thereby improving the efficiency of handling government service hall business.

[0009] One possible implementation involves inputting the first government document into a first deep learning model to obtain the first pinyin sequence corresponding to the first government document, which includes:

[0010] The Jieba word segmentation algorithm is used to divide the first government document into characters and words and mark the corresponding pauses to obtain the first segmented text.

[0011] The first segmented text is converted into the first pinyin sequence using a pinyin conversion program.

[0012] One possible implementation includes, before converting the first segmented text into a first pinyin sequence using a pinyin conversion program:

[0013] The second segmented text is obtained by adjusting the segmentation positions of the characters and words in the first segmented text using a text segmentation model.

[0014] The step of converting the first segmented text into a first pinyin sequence using a pinyin conversion program includes:

[0015] The second segmented text is converted into the first pinyin sequence using a pinyin conversion program.

[0016] In one possible implementation, the second deep learning model is a speech correction model;

[0017] The step of inputting the first pinyin sequence into the second deep learning model to obtain the first acoustic feature sequence includes:

[0018] Based on the speech correction model, the standard timbre corresponding to any element in the first pinyin sequence is determined; wherein, the standard timbre indicates the frequency at which the first pinyin sequence is played using Mandarin.

[0019] Construct a first feature vector corresponding to the first pinyin sequence based on the timbre;

[0020] According to preset rules, the first feature vector is transformed into a first acoustic feature sequence.

[0021] Secondly, this application provides a speech synthesis apparatus, the apparatus comprising:

[0022] Pinyin unit: used to input the first government text into the first deep learning model to obtain the first pinyin sequence corresponding to the first government text; wherein, the first pinyin sequence indicates the pinyin and pause position corresponding to the first government text;

[0023] Acoustic unit: used to input the first pinyin sequence into the second deep learning model to obtain the first acoustic feature sequence corresponding to the first pinyin sequence; wherein, the first acoustic feature sequence indicates the acoustic features of the first pinyin sequence;

[0024] Speech unit: used to input the first acoustic feature sequence into a vocoder to obtain the first speech corresponding to the first government text.

[0025] In one possible implementation, the pinyin unit is specifically used to segment the characters and words of the first government document using the Jieba word segmentation algorithm and mark the corresponding pauses to obtain the first segmented text; and to convert the first segmented text into a first pinyin sequence using a pinyin conversion program.

[0026] In one possible implementation, the device further includes an adjustment unit, specifically configured to adjust the segmentation positions of the characters and words in the first segmented text using a text segmentation model to obtain a second segmented text;

[0027] The pinyin unit is also used to convert the second word segmented text into a first pinyin sequence through a pinyin conversion program.

[0028] In one possible implementation, the second deep learning model is a speech correction model;

[0029] The acoustic unit is specifically used to determine the standard timbre corresponding to any element in the first pinyin sequence based on the speech correction model; wherein the standard timbre indicates the frequency at which the first pinyin sequence is played in Mandarin; construct a first feature vector corresponding to the first pinyin sequence based on the timbre; and convert the first feature vector into a first acoustic feature sequence according to a preset rule.

[0030] Thirdly, this application provides an electronic device, comprising:

[0031] Memory, used to store computer programs;

[0032] When a processor executes a computer program stored in the memory, it implements the method as described in the first aspect and any possible implementation.

[0033] Fourthly, this application provides a readable storage medium, including,

[0034] memory,

[0035] The memory is used to store instructions that, when executed by a processor, cause an apparatus including the readable storage medium to perform the method as described in the first aspect and any possible implementation. Attached Figure Description

[0036] Figure 1 A flowchart of a speech synthesis method provided in an embodiment of this application;

[0037] Figure 2 A schematic diagram illustrating the determination of a first pinyin sequence based on a first government document, provided as an embodiment of this application;

[0038] Figure 3 This is a schematic diagram of the structure of a speech synthesis device provided in an embodiment of this application;

[0039] Figure 4 This is a schematic diagram of the structure of a speech synthesis electronic device provided in an embodiment of this application. Detailed Implementation

[0040] To address the issue of low efficiency in handling business in government service halls, this application provides a speech synthesis method: using a deep learning model, text content that provides prompts and guidance to users is converted into speech, thereby improving the efficiency of business handling in government service halls.

[0041] To better understand the above technical solutions, the technical solutions of this application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of this application and the specific features in the embodiments are detailed descriptions of the technical solutions of this application, rather than limitations on the technical solutions of this application. In the absence of conflict, the embodiments of this application and the technical features in the embodiments can be combined with each other.

[0042] Please refer to Figure 1 This application provides a speech synthesis method that prevents de-identified data from being repeatedly recognized or ensures high availability. The method specifically includes the following implementation steps:

[0043] Step 101: Input the first government document into the first deep learning model to obtain the first pinyin sequence corresponding to the first government document.

[0044] The first pinyin sequence indicates the pinyin and pause positions corresponding to the first government document.

[0045] Specifically, firstly, all services that can be handled at the self-service machines in the government service hall are categorized according to service type, and the operation steps under each service type are converted into the first government service document.

[0046] Then, using the Jieba word segmentation algorithm, the first government document is divided into characters and words, and corresponding pause markers are applied to obtain the first segmented text. These pause markers may include a single word ending marker #0, multiple words ending marker #1, and short sentences ending marker #2.

[0047] Finally, the first segmented text is converted into a first pinyin sequence using a pinyin conversion program. In this embodiment, the pinyin conversion program is set to g2pc.

[0048] Furthermore, to make the pauses in the first segmented text more accurate, the segmentation positions of the characters and words in the first segmented text can be adjusted using a text segmentation model before determining the first pinyin sequence, resulting in the second segmented text. In this embodiment, the text segmentation model is set to the BLSTM+CRF model. Then, the second segmented text can be converted into the first pinyin sequence using a pinyin conversion program.

[0049] like Figure 2The diagram illustrates how to determine the first pinyin sequence based on a first government document, as provided in this embodiment of the application. Assume the first government document is: "How to reissue a social security card". Inputting this first government document into a first deep learning model, the Jieba word segmentation algorithm can perform pause division based on the characters and parts of speech in the first government document, resulting in the first segmented text: "How#0 reissue#0 social security card#2". Inputting this first segmented text into a BLSTM+CRF model allows for further adjustment of pause positions, resulting in more accurate pause markings in the second segmented text. The second segmented text could be: "How#0 reissue#0 issue#1 social security card#2". In other words, the second segmented text obtained through the BLSTM+CRF model further identifies "reissue" as multiple words, indicating supplementary processing, thus adjusting the pause mark after "reissue" to #1. Further, inputting this second segmented model into a pinyin conversion program (G2PC) yields the first pinyin sequence: "ruhe#0bu#0ban#1shebaoka#2".

[0050] Step 102: Input the first pinyin sequence into the second deep learning model to obtain the first acoustic feature sequence corresponding to the first pinyin sequence.

[0051] The first acoustic feature sequence indicates the acoustic features of the first pinyin sequence.

[0052] Specifically, the second deep learning model can be a speech correction model. Therefore, based on the speech correction model, the standard timbre corresponding to any element in the first pinyin sequence can be determined. This standard timbre indicates the frequency at which the first pinyin sequence is played using Mandarin. A first feature vector corresponding to the first pinyin sequence is constructed based on the timbre. After determining the first feature vector, it can be transformed into a first acoustic feature sequence according to preset rules.

[0053] In this embodiment, the speech correction model is set to Tactron2, which includes three modules: an encoding module, a decoding module, and an attention module connecting the encoding and decoding modules. The encoding module is used for character representation correction, the decoding module is used for timbre correction of all characters, and the attention module is used to align characters with speech, ensuring that when each character begins, the speech of the previous character ends.

[0054] Before using the speech correction model, to improve its effectiveness, it can be trained on a large amount of open-source data. First, after training the model parameters in the encoder, decode, and attention modules using different open-source datasets, it can be ensured that the model parameters in the encoder and attention modules remain unchanged, and only the parameters in the decode module are adjusted. This ensures that the similarity between the first acoustic feature sequence obtained through the Tactron2 model and standard Mandarin meets the set threshold.

[0055] Step 103: Input the first acoustic feature sequence into the vocoder to obtain the first speech corresponding to the first government document.

[0056] Specifically, the vocoder is set to a melgan vocoder, which includes a synthesizer and a discriminator. The discriminator is used to adjust the timbre of the first speech based on the previously synthesized speech before each playback of the first speech, so that the first speech corresponding to the first government text is more accurate and natural.

[0057] Based on the same inventive concept, this application also provides a speech synthesis device, which is similar to the aforementioned... Figure 1 The speech synthesis method shown corresponds to a specific implementation of the device, which can be found in the description of the aforementioned method embodiments. Repeated descriptions will not be repeated here. Figure 3 The device includes:

[0058] Pinyin unit 301: used to input the first government text into the first deep learning model to obtain the first pinyin sequence corresponding to the first government text.

[0059] The first pinyin sequence indicates the pinyin and pause positions corresponding to the first government document.

[0060] Specifically, the first government document is segmented into characters and words using the Jieba word segmentation algorithm and marked with corresponding pauses to obtain the first segmented text; the first segmented text is then converted into a first pinyin sequence using a pinyin conversion program.

[0061] The speech synthesis device further includes an adjustment unit, specifically used to adjust the segmentation positions of the characters and words in the first segmented text through a text segmentation model to obtain the second segmented text;

[0062] Then, the pinyin unit 301 is also used to convert the second word segmented text into the first pinyin sequence through a pinyin conversion program.

[0063] Acoustic unit 302: used to input the first pinyin sequence into the second deep learning model to obtain the first acoustic feature sequence corresponding to the first pinyin sequence.

[0064] The first acoustic feature sequence indicates the acoustic features of the first pinyin sequence.

[0065] Specifically, the second deep learning model is a speech correction model;

[0066] The acoustic unit 302 is specifically used to determine the standard timbre corresponding to any element in the first pinyin sequence based on the speech correction model; wherein the standard timbre indicates the frequency at which the first pinyin sequence is played using Mandarin; construct a first feature vector corresponding to the first pinyin sequence based on the timbre; and convert the first feature vector into a first acoustic feature sequence according to a preset rule.

[0067] Speech unit 303: used to input the first acoustic feature sequence into a vocoder to obtain the first speech corresponding to the first government document.

[0068] Based on the same inventive concept, this application also provides an electronic device that can implement the function of the aforementioned speech synthesis method. (Refer to...) Figure 4 The electronic device includes:

[0069] At least one processor 401 and a memory 402 connected to at least one processor 401. In this embodiment, the specific connection medium between the processor 401 and the memory 402 is not limited. Figure 4 The example shown is the connection between processor 401 and memory 402 via bus 400. Bus 400 is... Figure 4 The connections between other components are indicated by thick lines and are for illustrative purposes only, not as limiting information. The 400 bus can be divided into address bus, data bus, control bus, etc., for ease of representation. Figure 4 The term is represented by a single thick line, but this does not imply that there is only one bus or one type of bus. Alternatively, processor 401 can also be called a controller; there is no restriction on the name.

[0070] In this embodiment, memory 402 stores instructions executable by at least one processor 401. By executing the instructions stored in memory 402, at least one processor 401 can perform the speech synthesis method discussed above. Processor 401 can implement... Figure 3 The functions of each module in the device shown.

[0071] The processor 401 is the control center of the device. It can connect to various parts of the control device through various interfaces and lines. By running or executing instructions stored in memory 402 and calling data stored in memory 402, the processor can perform various functions and process data, thereby monitoring the device as a whole.

[0072] In one possible design, processor 401 may include one or more processing units. Processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 401. In some embodiments, processor 401 and memory 402 may be implemented on the same chip; in some embodiments, they may also be implemented separately on separate chips.

[0073] Processor 401 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor, application-specific integrated circuit, field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the speech synthesis method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules within the processor.

[0074] Memory 402, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 402 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 402 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 402 may also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.

[0075] By designing and programming the processor 401, the code corresponding to the speech synthesis method described in the foregoing embodiments can be embedded into the chip, enabling the chip to execute the code during operation. Figure 1The steps of the speech synthesis method in the illustrated embodiment are as follows. How to design and program the processor 401 is a technique well-known to those skilled in the art and will not be described further here.

[0076] Based on the same inventive concept, embodiments of this application also provide a readable storage medium, including:

[0077] memory,

[0078] The memory is used to store instructions that, when executed by a processor, cause the apparatus including the readable storage medium to perform the speech synthesis method as described above.

[0079] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0080] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0081] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0082] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0083] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: Universal Serial Bus flash disks, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.

[0084] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A speech synthesis method, characterized in that, The method includes: The first government document is input into the first deep learning model to obtain the first pinyin sequence corresponding to the first government document; wherein, the first pinyin sequence indicates the pinyin and pause position corresponding to the first government document; the pause mark corresponding to the pause position includes single word end mark, multiple word end mark, and short sentence end mark; The first pinyin sequence is input into the second deep learning model to obtain the first acoustic feature sequence corresponding to the first pinyin sequence; wherein, the first acoustic feature sequence indicates the acoustic features of the first pinyin sequence; The first acoustic feature sequence is input into a vocoder to obtain the first speech corresponding to the first government document; The second deep learning model is a speech correction model; the speech correction model includes an encoding module, a decoding module, and an attention module connecting the encoding module and the decoding module; the encoding module is used for character expression correction, the decoding module is used for timbre correction for all characters, and the attention module is used for aligning characters with speech. The step of inputting the first pinyin sequence into the second deep learning model to obtain the first acoustic feature sequence includes: The model parameters in the encoder, decode, and attention modules are trained using different open-source data to determine the model parameters in the encoder and attention modules. Adjustments are made only to the model parameters in the decode module to obtain the speech correction model. Based on the speech correction model, the standard timbre corresponding to any element in the first pinyin sequence is determined. The standard timbre indicates the frequency at which the first pinyin sequence is played in Mandarin. A first feature vector corresponding to the first pinyin sequence is constructed based on the timbre. According to a preset rule, the first feature vector is converted into the first acoustic feature sequence.

2. The method as described in claim 1, characterized in that, The step of inputting the first government document into the first deep learning model to obtain the first pinyin sequence corresponding to the first government document includes: The Jieba word segmentation algorithm is used to divide the first government document into characters and words and mark the corresponding pauses to obtain the first segmented text. The first segmented text is converted into the first pinyin sequence using a pinyin conversion program.

3. The method as described in claim 2, characterized in that, Before converting the first segmented text into the first pinyin sequence using the pinyin conversion program, the process includes: The second segmented text is obtained by adjusting the segmentation positions of the characters and words in the first segmented text using a text segmentation model. The step of converting the first segmented text into a first pinyin sequence using a pinyin conversion program includes: The second segmented text is converted into the first pinyin sequence using a pinyin conversion program.

4. A speech synthesis device, characterized in that, The device includes: Pinyin unit: used to input the first government affairs text into the first deep learning model to obtain the first pinyin sequence corresponding to the first government affairs text; wherein, the first pinyin sequence indicates the pinyin and pause position corresponding to the first government affairs text; the pause mark corresponding to the pause position includes single word end mark, multiple word end mark, and short sentence end mark; Acoustic Unit: Used to input the first pinyin sequence into the second deep learning model to obtain the first acoustic feature sequence corresponding to the first pinyin sequence; wherein, the first acoustic feature sequence indicates the acoustic features of the first pinyin sequence; the second deep learning model is a speech correction model; the speech correction model includes an encoding module, a decoding module, and an attention module connecting the encoding module and the decoding module; the encoding module is used for character expression correction, the decoding module is used for timbre correction for all characters, and the attention module is used for aligning characters with speech; then the acoustic unit inputs the first pinyin sequence into the second deep learning model to obtain the first acoustic feature sequence corresponding to the first pinyin sequence. When generating the first acoustic feature sequence corresponding to a pinyin sequence, the process specifically involves training the model parameters in the encoder, decode, and attention modules using different open-source data, determining the model parameters in the encoder and attention modules, and adjusting only the model parameters in the decode module to obtain the speech correction model. Based on the speech correction model, the standard timbre corresponding to any element in the first pinyin sequence is determined. The standard timbre indicates the frequency at which the first pinyin sequence is played using Mandarin. A first feature vector corresponding to the first pinyin sequence is constructed based on the timbre. According to a preset rule, the first feature vector is converted into the first acoustic feature sequence. Speech unit: used to input the first acoustic feature sequence into a vocoder to obtain the first speech corresponding to the first government text.

5. The apparatus as described in claim 4, characterized in that, The pinyin unit is specifically used to divide the first government document into characters and words and mark corresponding pauses using the Jieba word segmentation algorithm to obtain the first segmented text; and to convert the first segmented text into a first pinyin sequence using a pinyin conversion program.

6. The apparatus as claimed in claim 5, characterized in that, The device further includes an adjustment unit, specifically used to adjust the segmentation positions of the characters and words in the first segmented text through a text segmentation model to obtain the second segmented text; The pinyin unit is also used to convert the second word segmented text into a first pinyin sequence through a pinyin conversion program.

7. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, when executing a computer program stored in the memory, implements the method as described in any one of claims 1 to 3.

8. A readable storage medium, characterized in that, include, memory, The memory is used to store instructions that, when executed by a processor, cause a device including the readable storage medium to perform the method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Synthetic speech transmission method, cloud server and terminal device

    CN106504742A

  • Speech synthesis method and system based on multi-task acoustic model

    CN110718208A