A TTS text regularization method and device based on regular expressions and WFST

By building the basic WFST graph structure and generating regular expression replacement rules, the problem of time-consuming and error-prone for ordinary software developers to write WFST is solved, and it is easy to expand and maintain the regularization of virtual human TTS text, reducing maintenance costs.

CN116312540BActive Publication Date: 2025-07-18SHANGHAI GAUDIAN INTELLIGENT TECHNOLOGY GROUP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310276496.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-21
Publication Date
2025-07-18
Estimated Expiration
2043-03-21

AI Technical Summary

Technical Problem

In the prior art, ordinary software developers are time-consuming and error-prone to writing WFST syntax, resulting in high maintenance costs for regularization of virtual human TTS text and difficult to scale.

Method used

Provide a TTS text regularization method based on regular expressions and WFST. By building the basic WFST graph structure, writing regular expression replacement rules, and generating the first WFST through combinatorial algorithms and optimization, and finally fusing it into the second WFST, reducing the difficulty of developers writing WFST and improving scalability and maintenance.

Benefits of technology

It reduces the difficulty of writing WFST by ordinary software developers, makes the regularization of virtual human TTS text easier to expand and maintain, and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312540B_ABST
    Figure CN116312540B_ABST
Patent Text Reader

Abstract

The present invention provides a TTS text regularization method and device based on regular expressions and WFST. The method includes the steps of obtaining a target text to be recognized; determining the type of the target text; obtaining a corresponding regular expression replacement rule from a pre-constructed second WFST based on the type of the target text; converting the target text into a corresponding regular text based on the corresponding regular expression replacement rule; and converting the corresponding regular text into voice information. By adopting the TTS text regularization method and device based on regular expressions and WFST provided by the present invention, ordinary software developers can effectively fuse a large number of regular expression replacement rules into a WFST graph structure by simply writing regular expressions, thereby making the virtual human TTS text regularization method easier to expand and maintain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition, and in particular to a TTS text regularization method and device based on regular expressions and WFST. Background Art

[0002] Nowadays, virtual humans are used in more and more scenarios, including news broadcasting, service provision, chat interaction, content generation, etc. The chat interaction between virtual humans and users mainly relies on the dialogue system in the field of natural language. In order to make the pronunciation more accurate when virtual humans chat with users, a text regularization module needs to be added before TTS (text-to-speech) to convert some non-standard characters into characters that can be pronounced correctly. For example, "7+8=15" should be converted to "seven plus eight equals fifteen".

[0003] Currently, the methods for TTS text regularization in Chinese generally use regularization methods, or deep learning-based methods, or a combination of both. The regularization method is controllable and easy to expand, while the deep learning method requires a large amount of labeled training data.

[0004] The regularization method mainly uses WFST (Weighted Finite State Transducer). WFST requires manual writing of grammar, but the grammar writing is time-consuming and error-prone. It is a difficult task for ordinary software developers. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for converting text replacement rules based on regular expressions into WFST grammar on the basis of the regularization method, to provide support for the virtual human TTS text regularization method, and to reduce the maintenance cost of the virtual human TTS text regularization.

[0006] The technical solutions of a TTS text regularization method and device based on regular expressions and WFST provided by the present invention are as follows:

[0007] A TTS text regularization method based on regular expressions and WFST, comprising:

[0008] Obtaining a target text to be recognized;

[0009] Determining the type of the target text;

[0010] Based on the type of the target text, obtaining a corresponding regular expression replacement rule from a pre-constructed second WFST;

[0011] Based on the corresponding regular expression replacement rule, converting the target text into a corresponding regular text;

[0012] Converting the corresponding regular text into voice information.

[0013] In some embodiments, before obtaining the target text to be recognized, it further includes: constructing the second WFST; specifically including:

[0014] Based on the basic grammar of the WFST written in advance, construct the basic WFST graph structure;

[0015] For different types of texts, write corresponding regular expression replacement rules, and the regular expression replacement rules include regular expressions and WFST basic grammar;

[0016] Perform forward and backward stacking operations on the WFST basic grammar part in each of the regular expression replacement rules through a combination algorithm, fuse them into the first WFST, and optimize the first WFST;

[0017] Fuse all the first WFSTs into the second WFST, and optimize the second WFST.

[0018] In some embodiments, the basic WFST graph structure specifically includes: WFST graph structures for numbers, currencies, dates, telephones, and units.

[0019] In some embodiments, the optimization of the first WFST specifically includes:

[0020] Optimize the first WFST through the subset construction method, and thus realize the transformation of the first WFST from "non-deterministic" to "deterministic".

[0021] In some embodiments, the fusion of all the first WFSTs into the second WFST specifically includes:

[0022] Fuse multiple first WFSTs by means of parallel combination up and down to form the second WFST.

[0023] In some embodiments, a TTS text regularization device based on regular expressions and WFST includes:

[0024] A text acquisition module for acquiring the target text to be recognized;

[0025] A type determination module for determining the type of the target text;

[0026] A regular expression replacement rule acquisition module for obtaining the corresponding regular expression replacement rule from the pre-constructed second WFST based on the type of the target text;

[0027] A regular text conversion module for converting the target text into the corresponding regular text based on the corresponding regular expression replacement rule;

[0028] A voice information conversion module, configured to convert the corresponding regular text into voice information.

[0029] In some embodiments, the TTS text regularization device based on regular expressions and WFST further includes a second WFST construction module, and the second WFST construction module specifically includes:

[0030] A basic WFST construction sub-module, configured to construct a basic WFST graph structure based on the basic syntax of the pre-written WFST;

[0031] A regular expression replacement rule writing sub-module, configured to write corresponding regular expression replacement rules for different types of texts, where the regular expression replacement rules include regular expressions and WFST basic syntax;

[0032] A first WFST generation sub-module, configured to perform a front-back superposition operation on the WFST basic syntax part in each regular expression replacement rule through a combination algorithm, fuse them into a first WFST, and optimize the first WFST;

[0033] A second WFST generation sub-module, configured to fuse all the first WFSTs into a second WFST and optimize the second WFST.

[0034] In some embodiments, the basic WFST construction sub-module specifically includes:

[0035] A digital WFST construction unit, configured to construct a digital WFST graph structure;

[0036] A currency WFST construction unit, configured to construct a currency WFST graph structure;

[0037] A date WFST construction unit, configured to construct a date WFST graph structure;

[0038] A phone WFST construction unit, configured to construct a phone WFST graph structure;

[0039] A unit WFST construction unit, configured to construct a unit WFST graph structure.

[0040] In some embodiments, the first WFST generation sub-module specifically includes:

[0041] A fusion unit, configured to perform a front-back superposition operation on the WFST basic syntax part in each regular expression replacement rule through a combination algorithm, fuse them into a first WFST;

[0042] An optimization unit is used to optimize the first WFST by the subset construction method, thereby realizing the transformation of the first WFST from "non-deterministic" to "deterministic".

[0043] In some embodiments, the second WFST generation sub-module fuses a plurality of the first WFSTs in a vertically parallel combination manner to form the second WFST.

[0044] Adopting a TTS text regularization method and device based on regular expressions and WFST provided by the present invention has at least one of the following beneficial effects:

[0045] 1. The TTS text regularization method based on regular expressions and WFST provided by the present invention reduces the difficulty of writing WFST for ordinary software developers by writing basic WFST syntax and calling the corresponding WFST basic syntax according to the text type, enabling ordinary software developers to only write regular expressions and fuse a large number of regular expression replacement rules into a WFST graph structure.

[0046] 2. The TTS text regularization method based on regular expressions and WFST provided by the invention makes the virtual human TTS text regularization method more easily extensible and maintainable due to the existence of the basic WFST part, reduces the maintenance cost of the virtual human TTS text regularization, and makes it more in line with actual needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The above characteristics, technical features, advantages and implementation manners of a TTS text regularization method and device based on regular expressions and WFST will be further described below in a clear and understandable manner in combination with the drawings of the preferred embodiments.

[0048] Figure 1 is a schematic flowchart of an embodiment of a TTS text regularization method based on regular expressions and WFST of the present invention;

[0049] Figure 2 is a partial schematic flowchart of another embodiment of a TTS text regularization method based on regular expressions and WFST of the present invention;

[0050] Figure 3 is a digital WFST graph in another embodiment of a TTS text regularization method based on regular expressions and WFST of the present invention;

[0051] Figure 4 is another embodiment in a TTS text regularization method based on regular expressions and WFST of the present invention Figure 3 schematic diagram of the syntax of the digital WFST graph;

[0052] Figure 5 It is a block diagram of a module of an embodiment of a TTS text regularization device based on regular expressions and WFST according to the present invention;

[0053] Figure 6 A string WFST graph in another embodiment of a TTS text regularization device based on regular expressions and WFST according to the present invention. Detailed implementation manners

[0054] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are set forth in order to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present application.

[0055] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups.

[0056] It should also be further understood that the term "and / or" as used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0057] In addition, in the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the specific implementation manners of the present invention will be described below with reference to the accompanying drawings. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts, and other implementation manners can also be obtained.

[0059] In one embodiment, the present invention provides a TTS text regularization method based on regular expressions and WFST. Refer to Figure 1 , and it includes the steps:

[0060] S100, obtaining a target text to be recognized;

[0061] S200, determining the type of the target text;

[0062] S300. Obtain the corresponding regular expression replacement rule from the pre-constructed second WFST based on the type of the target text;

[0063] S400. Convert the target text into the corresponding regular text based on the corresponding regular expression replacement rule;

[0064] S500. Convert the corresponding regular text into voice information.

[0065] Specifically, according to the scenarios where voice conversion is required, such as virtual characters in scenarios like voice customer service, obtain their text content and determine the text type. For example, the text type is numbers, currency, dates, or phone numbers, etc. For these text types, there are pre-written WFSTs. You can call from them to obtain the regular expression replacement rules corresponding to the text types. For example, the regular expression replacement rule for the text "15.6%" is "([\d.]+)(%)->(2: unit - percent)(1: number)". Then, convert the corresponding text into regular text according to the corresponding regular expression replacement rule, and finally convert it into the voice information for broadcasting.

[0066] In another embodiment, based on the above embodiment, refer to Figure 2 , before obtaining the target text to be recognized, it further includes: constructing a second WFST; specifically including:

[0067] S001. Construct the basic WFST graph structure based on the basic grammar of the pre-written WFST;

[0068] S002. Write the corresponding regular expression replacement rules for different types of texts. The regular expression replacement rules include regular expressions and WFST basic grammar;

[0069] S003. Perform a front-to-back superposition operation on the WFST basic grammar part in each regular expression replacement rule through a combination algorithm, fuse it into the first WFST, and optimize the first WFST;

[0070] S004. Fuse all the first WFSTs into the second WFST, and optimize the second WFST.

[0071] Specifically, the basic WFST grammar is written, that is, the WFST graph structure of the basic text type is constructed, such as the graph structure of WFST such as numbers, currencies, dates, telephone numbers, units (such as percent signs), etc. According to different types of text, the corresponding regular expression replacement rules are written. The regular expression replacement rules have two parts, one is the regular expression of the corresponding text, and the other is the above-mentioned basic WFST grammar. When various text types are combined in a voice, such as numbers and currencies, the weighted finite state transcription machine (WFST) part in the regular expression replacement rule is superimposed using the WFST combination algorithm, and then merged into the above-mentioned first WFST, and then the first WFST is optimized from non-deterministic to deterministic to form a second WFST, and it is optimized by determinization, weight movement, minimization, etc. Through the TTS text regularization method based on regular expressions and WFST provided by the present invention, the relevant code can be deployed in the back end using docker, c++ and other technologies to improve the program execution efficiency in the production environment.

[0072] In another embodiment, based on the above embodiment, the basic WFST graph structure specifically includes: WFST graph structures of numbers, currencies, dates, telephone numbers and units.

[0073] Specifically, the construction of a digital weighted finite state transcription machine can refer to Figure 3 The figure shows a WFST that converts three-digit numbers into Chinese characters (the default weight is 1.0). This WFST can convert the number 282 into "二百八十二". Of course, when encountering four-digit and five-digit numbers, all such state machines can be combined to form a larger state machine, which can more comprehensively process the numbers appearing in the text. To construct such a state machine, you can use the relevant syntax of the open source framework Openfst to achieve this. For example, Figure 3 The syntax in Figure 4 As shown, "0 1 1 -" means that when jumping from state 0 to state 1, the character "1" is rewritten into the character "-".

[0074] In another embodiment, based on the above embodiment, the first WFST is optimized, including optimizing the first WFST from "non-deterministic" to "deterministic", specifically including:

[0075] Specifically, Figure 3The WFST in the figure is a form optimized by the self-construction method. Its initial form is NFST, i.e., a non-deterministic finite state machine, which means that the same input symbol between two nodes can have multiple outputs. This NFST is highly understandable, but inefficient, and needs to be converted into DFST, i.e., a deterministic finite state machine. The first WFST can be optimized by the subset construction method, thereby realizing "non-determinism" to "determinism". When optimization from non-determinism to determinism is required, the existing subset construction method can be used for optimization to convert the uncertain finite state machine into a deterministic finite state machine.

[0076] In another embodiment, based on the above embodiment, merging all first WFSTs into a second WFST specifically includes:

[0077] Multiple first WFSTs are combined in parallel up and down to form a second WFST.

[0078] Specifically, the fusion of the second WFST is different from the fusion of the first WFST. The first WFST is fused by front-to-back superposition, while the second WFST is fused by upper-lower parallel combination.

[0079] Another method embodiment of the present application, based on the first method embodiment, provides a specific implementation method for constructing and deploying a second WFST, including the following steps:

[0080] 1. Write the basic WFST grammar, that is, construct the graph structure of WFST such as numbers, currencies, dates, telephone numbers, units (such as percent signs), etc.

[0081] Specifically, for example, to construct a digital weighted state transcription machine, you can refer to Figure 3 To understand the structure of Figure 3 This is a WFST that converts three-digit numbers into Chinese characters. The weights are omitted (default is 1.0), for example, "282" can be converted into "二百八十二". In addition to three-digit numbers, there are four-digit, five-digit... All these state machines are combined to form a larger state machine, which can more comprehensively process the numbers that appear in the text.

[0082] To construct such a state machine, you need to use the syntax of Openfst (an open source framework that can be used to build WFST), such as Figure 3 The required syntax is Figure 4 shown.

[0083] For example, “0 1 1 一” means that when jumping from state 0 to state 1, the character “1” is rewritten as the character “一”.

[0084] 2. Write regular expression replacement rules for special texts that need to be regularized.

[0085] 3. Convert the first half of the regular expression replacement rule in 2 into a DFA (Deterministic Finite State Machine). The specific algorithm can be: first convert the regular expression into an NFA (Non-deterministic Finite State Machine), and then use the subset construction method to convert the NFA into a DFA;

[0086] For example, the regular expression replacement rule for the text "15.6%" is to replace the string in the pattern of "([\d.]+)(%)" with the pattern of "(2: unit - percent)(1: number)". The former pattern is a pattern that conforms to the regular expression syntax, and the latter pattern is a temporary pattern defined by ourselves. It is defined by the basic syntax in 1. This replacement rule will automatically create a combined WFST according to the pattern of "(unit - percent)(number)", and the first half is the unit WFST in 1, and the second half is the number WFST in 1. Similar to the following:

[0087] As Figure 6 shown, the string "123%" will be converted into "one hundred and twenty-three percent".

[0088] The meaning of the pattern "(2: unit - percent)(1: number)" is: match the matching value in the first parentheses of the regular expression "([\d.]+)(%)" with the number WFST, and match the matching value in the second parentheses with the unit WFST. Combine the unit WFST and the number WFST for combined operation, and the corresponding first WFST in Figure 6 can be generated.

[0089] The above Figure 3 and Figure 6 The first WFSTs within are all optimized by the subset construction method. Their initial form is NFST, that is, Non-deterministic Finite Transcriptor, which means that for the same input symbol between two nodes, there can be multiple outputs. This kind of NFST is highly understandable, but has low efficiency and needs to be converted into DFST, that is, Deterministic Finite State Transcriptor. There is a set of mature algorithms for converting NFST to DFST, such as the subset construction method, which is an effective method strictly proven by mathematics.

[0090] 5. Loop through steps 2 to 4, and according to the actual business requirements, write multiple regular expression replacement rules and convert them into the corresponding first WFSTs;

[0091] Generally speaking, TTS text regularization requires a lot of text conversion rules. Therefore, a large number of conversion rules similar to those in 2 written by developers or product personnel can be used in a loop to batch generate a large number of first WFSTs.

[0092] 6. Combine the multiple first WFSTs obtained in step 5 into a large second WFST, and perform optimizations such as determinization, weight shifting, and minimization on it;

[0093] The combination here is to combine multiple first WFSTs in a parallel manner, not by stacking them one after another, but by combining them in a parallel up-and-down manner.

[0094] 7. Deploy the obtained second WFST using docker and c++ for backend deployment to improve the program execution efficiency in the production environment.

[0095] In one embodiment, based on the same technical concept, the present invention also provides a TTS text regularization device based on regular expressions and WFST. Refer to Figure 5 , including:

[0096] A text acquisition module 10 for acquiring the target text to be recognized;

[0097] A type determination module 20 for determining the type of the target text;

[0098] A regular expression replacement rule acquisition module 30 for acquiring the corresponding regular expression replacement rule from the pre-constructed second WFST based on the type of the target text;

[0099] A regular text conversion module 40 for converting the target text into the corresponding regular text based on the corresponding regular expression replacement rule;

[0100] A voice information conversion module 50 for converting the corresponding regular text into voice information.

[0101] Specifically, the text acquisition module 10 acquires the target text to be converted into voice, calls the type determination module 20 to determine which type of text the target text is, and calls the regular expression replacement rule acquisition module 30 to acquire the corresponding replacement rule from the pre-constructed second WFST according to the corresponding text type. Combine the acquired replacement rule and the regular text conversion module 40 to convert the corresponding text into regular text, and finally implement TTS (text-to-speech) through the voice information conversion module 50. This device is not limited to voice conversion in Chinese and is also applicable to English.

[0102] In another embodiment, on the basis of the above embodiment, the TTS text regularization device based on regular expressions and WFST further includes a second WFST construction module, and the second WFST construction module specifically includes:

[0103] A basic WFST construction sub-module for constructing a basic WFST graph structure based on the basic grammar of the pre-written WFST;

[0104] A regular expression replacement rule writing sub-module is used to write corresponding regular expression replacement rules for different types of texts. The regular expression replacement rules include regular expressions and WFST basic syntax;

[0105] The first WFST generation sub-module is used to perform a front-back stacking operation on the WFST basic syntax part in each regular expression replacement rule through a combination algorithm, fuse it into a first WFST, and optimize the first WFST from "non-deterministic" to "deterministic";

[0106] The second WFST generation sub-module is used to fuse all the first WFSTs into a second WFST and optimize the second WFST.

[0107] Specifically, a basic WFST construction sub-module is used to determine the basic syntax of WFST, such as the WFST graph structure expression of numbers, currencies, dates, etc. Then, the regular expression replacement rule writing sub-module writes the regular expression replacement rules as needed. After that, the first WFST generation sub-module first performs a front-back stacking fusion operation on the WFST basic syntax part in the replacement rules to generate a first WFST. Then, the generated first WFST is fused by the second WFST generation sub-module to form a second WFST and is optimized such as determinization, weight shifting, minimization, etc. Refer to Figure 6 , the string "123%" will be converted to "one hundred and twenty-three percent". The meaning of the pattern (2: unit - percent)(1: number) is: match the value in the first parentheses of the regular expression "([\d.]+)(%)" with the number WFST, and match the value in the second parentheses with the unit WFST. Combine the unit WFST and the number WFST through an operation to generate the WFST in the above figure.

[0108] In another embodiment, based on the above embodiment, the basic WFST construction sub-module specifically includes:

[0109] A number WFST construction unit is used to construct a number WFST graph structure;

[0110] A currency WFST construction unit is used to construct a currency WFST graph structure;

[0111] A date WFST construction unit is used to construct a date WFST graph structure;

[0112] A phone WFST construction unit is used to construct a phone WFST graph structure;

[0113] A unit WFST construction unit is used to construct a unit WFST graph structure.

[0114] Specifically, the relevant grammar of WFST is constructed through the open-source framework Openfst, and the basic WFST graph structure is generated. The same applies to other basic construction units.

[0115] In another embodiment, based on the above embodiment, the first WFST generation sub-module specifically includes:

[0116] A fusion unit, configured to perform a front-back stacking operation on the WFST basic grammar part in each regular expression replacement rule through a combination algorithm, and fuse it into the first WFST;

[0117] An optimization unit, configured to optimize the first WFST through the subset construction method, thereby realizing the transformation from "non-deterministic" to "deterministic".

[0118] Specifically, in the first WFST generation sub-module, the WFST basic grammar part is fused through a front-back stacking combination algorithm by the fusion unit, and then the generated first WFST is optimized by the optimization unit. Here, the subset construction method is adopted to realize the optimization from a non-deterministic finite state transducer to a deterministic finite state transducer.

[0119] In another embodiment, based on the above embodiment, the second WFST generation sub-module fuses multiple first WFSTs in a way of combining them in parallel up and down to form the second WFST.

[0120] Specifically, mainly for the purpose of distinguishing from the fusion during the generation of the first WFST in the above text, here the first WFSTs are fused in a way of combining them in parallel up and down to generate the second WFST.

[0121] They can be implemented by program codes executable by a computing device. Thus, they can be stored in a storage device and executed by the computing device, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.

[0122] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0123] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed by hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0124] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0125] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may also exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0126] It should be noted that the above-mentioned embodiments can be freely combined according to needs. The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A TTS text regularization method based on regular expressions and WFST, characterized in that, Including: Obtain the target text to be recognized; Determine the type of the target text; Based on the type of the target text, obtain the corresponding regular expression replacement rule from the pre-constructed second WFST; Based on the corresponding regular expression replacement rule, convert the target text into the corresponding regular text; Convert the corresponding regular text into voice information; Before obtaining the target text to be recognized, it further includes: constructing the second WFST; specifically including: Based on the basic grammar of the WFST written in advance, construct the basic WFST graph structure; For different types of texts, write the corresponding regular expression replacement rules, and the regular expression replacement rules include regular expressions and WFST basic grammar; Perform forward and backward stacking operations on the WFST basic grammar part in each regular expression replacement rule through a combination algorithm, fuse them into the first WFST, and optimize the first WFST; Fuse all the first WFSTs into the second WFST, and optimize the second WFST.

2. The TTS text regularization method based on regular expressions and WFST according to claim 1, characterized in that The basic WFST graph structure specifically includes: the WFST graph structures of numbers, currencies, dates, telephones, and units.

3. A TTS text regularization method based on regular expressions and WFST according to claim 1, characterized in that The optimization of the first WFST specifically includes: Optimize the first WFST through the subset construction method, thereby realizing the transformation of the first WFST from "non-deterministic" to "deterministic".

4. A TTS text regularization method based on regular expressions and WFST according to claim 1, characterized in that, The fusion of all the first WFSTs into the second WFST specifically includes: Fuse multiple first WFSTs in a way of parallel combination up and down to form the second WFST.

5. A TTS text regularization device based on regular expressions and WFST, characterized in that, Including: A text acquisition module for obtaining the target text to be recognized; A type determination module for determining the type of the target text; A regular expression replacement rule acquisition module for obtaining the corresponding regular expression replacement rule from the pre-constructed second WFST based on the type of the target text; A regular text conversion module for converting the target text into the corresponding regular text based on the corresponding regular expression replacement rule; A voice information conversion module for converting the corresponding regular text into voice information; A second WFST construction module, and the second WFST construction module specifically includes: A basic WFST construction sub-module for constructing the basic WFST graph structure based on the basic grammar of the WFST written in advance; A regular replacement rule writing sub-module for writing the corresponding regular expression replacement rules for different types of texts, and the regular expression replacement rules include regular expressions and WFST basic grammar; A first WFST generation sub-module for performing forward and backward stacking operations on the WFST basic grammar part in each regular expression replacement rule through a combination algorithm, fusing them into the first WFST, and optimizing the first WFST; A second WFST generation sub-module for fusing all the first WFSTs into the second WFST and optimizing the second WFST.

6. The TTS text regularization device based on regular expressions and WFST according to claim 5, characterized in that The basic WFST construction sub-module specifically includes: A digital WFST construction unit for constructing the digital WFST graph structure; Currency WFST construction unit, used to construct the currency WFST graph structure; Date WFST construction unit, used to construct the date WFST graph structure; Telephone WFST construction unit, used to construct the telephone WFST graph structure; Unit WFST construction unit, used to construct the unit WFST graph structure.

7. The TTS text regularization device based on regular expressions and WFST according to claim 5, characterized in that The first WFST generation sub-module specifically includes: A fusion unit, used to perform a front-back superposition operation on the WFST basic syntax part in each of the regular expression replacement rules through a combination algorithm, and fuse them into a first WFST; An optimization unit, used to optimize the first WFST through the subset construction method, thereby realizing the transformation of the WFST from "non-deterministic" to "deterministic".

8. A TTS text regularization device based on regular expressions and WFST according to claim 5, characterized in that, The second WFST generation sub-module fuses a plurality of the first WFSTs by means of parallel combination up and down to form the second WFST.

Citation Information

Patent Citations

  • Text normalization method and system based on WFST

    CN108536656A

  • Text regularization method and related device, electronic equipment and storage medium

    CN114330286A