A Chinese speech synthesis normalization method, device and computing device

By dynamically calculating the priority of rules in the speech synthesis system, the problem of the inability to quantitatively resolve and the recognition accuracy of rule conflicts in the prior art is solved, and a higher accuracy of speech synthesis is achieved.

CN114428831BActive Publication Date: 2025-06-20BEIJING ZHONGKE JINDEZHU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011097297.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-14
Publication Date
2025-06-20
Estimated Expiration
2040-10-14

AI Technical Summary

Technical Problem

The existing speech synthesis system has problems such as rule conflicts that cannot be quantitatively resolved and the accuracy of recognition is not high when processing unconventional texts.

Method used

By initializing a matrix P0 of size M×N, using M rules to scan the synthetic text, updating matrix P0 to obtain matrix P1, and calculating priority Q for each column of matrix P1, selecting the rule with the highest priority for identification.

Benefits of technology

It realizes quantitative selection of the best rules for identification, solves the problem of rule conflict, and improves the accuracy of normalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114428831B_ABST
    Figure CN114428831B_ABST
Patent Text Reader

Abstract

The present application discloses a Chinese speech synthesis normalization method, apparatus and computing device. The method includes: initializing a matrix P0 with a size of M×N and initial elements of 0; scanning the text to be synthesized using M rules respectively, if a certain character in the text to be synthesized matches a certain rule, then updating the element value corresponding to the character and the rule in the matrix P0 to a non-zero value to obtain an updated matrix P1; when there are at least two non-zero elements in a certain column of the matrix P1, calculating the priorities of the rules corresponding to the elements respectively, and retaining the element value corresponding to the rule with the highest priority, and resetting the other elements to zero. The apparatus includes an initialization module, a matrix update module, a priority calculation module and a merging and processing module. The computing device includes a memory, a processor and a computer program stored in the memory and capable of running by the processor, and when the processor executes the computer program, the method described in the present application is implemented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech synthesis, and in particular to the normalization processing technology for unconventional texts in speech synthesis. Background Art

[0002] The function of a speech synthesis system is to generate synthesized speech based on the input text to be synthesized, usually referring to a TTS (text to speech) system, that is, a text-to-speech system. In a commercial speech synthesis system, the speech synthesis service needs to have the ability to process unconventional texts in the text to be synthesized, such as identifying texts like mobile phone numbers, well-known brands, date and time, etc., and be able to pronounce them correctly.

[0003] To solve the problem of correct pronunciation of the above-mentioned texts, the usual processing method is to add many rules in the form of regular expressions in the normalization module. Normalization, as a front-end processing step of the TTS system, is used to identify unconventional characters in the text and convert them into a form that the TTS system can process. For example, when the input text is "The temperature today is 35℃", the normalization module needs to identify that "35℃" represents temperature and convert it into "35 degrees Celsius". A regular expression is a logical formula for string operations. Using some predefined specific characters and combinations of these specific characters, a "rule string" is formed, and this "rule string" is used to express a filtering logic for the string. When performing speech synthesis, these rules are used to match the text to be synthesized. When a certain rule is hit, the corresponding preprocessing function (corresponding to the verbalise function in the following text) is used for preprocessing and converted into a form that the TTS system can process, such as a Chinese character text sequence or a pinyin (phoneme) sequence. The basic process of speech synthesis is as Figure 1 shown.

[0004] In the existing normalization module using rule-based technology, each rule is used to scan the text to be synthesized separately to identify all parts of the text that meet the rule. Each rule can match multiple positions in the text. When multiple rules hit the same part of the text to be synthesized, the recognition result of a certain rule is selected according to a preset strategy, such as preferentially using one of the rules.

[0005] The existing technical solutions have the following disadvantages:

[0006] 1. It is impossible to quantitatively solve the problem of rule conflicts. A rule conflict means that the same position in the text to be synthesized is hit by multiple rules simultaneously. When this phenomenon occurs, the common strategy in existing technical solutions is to select one of the multiple hit rules according to a pre-set strategy by humans. For example, for the text to be synthesized "The thing is good, but for the so-called full 199 minus 100 during 618, the price is directly increased to 128 per bottle and there are no gifts", when the rules of "number range" and "full reduction offer" both hit "199 - 100" in the text, if the pre-set strategy is to give priority to using the "number range" rule, then an incorrect pronunciation of the speech synthesis "one hundred and ninety-nine to one hundred" will be generated.

[0007] 2. The recognition accuracy is not high. As mentioned in point 1 above, when a rule conflict occurs, simply giving priority to using a certain rule according to a pre-set strategy by humans will bring problems of incorrect recognition. Summary of the Invention

[0008] The purpose of the present application is to overcome the above problems or at least partially solve or alleviate the above problems.

[0009] According to one aspect of the present application, a Chinese speech synthesis normalization method is provided, including:

[0010] Initialize a matrix P0 with a size of M×N and initial elements of 0, where M is the total number of rules and N is the length of the text to be synthesized;

[0011] Use the M rules to scan the text to be synthesized respectively. If there are t positions in the text to be synthesized that satisfy the i-th rule, where i is any integer from 1 to M, and the starting positions of the r-th position among the t positions are s r and e r , where r is any integer from 1 to t, then update the values of the elements from the s r to e r in the i-th row of the matrix P0 to any non-zero number to obtain an updated matrix P1;

[0012] Scan each column of the matrix P1. When there are at least two non-zero elements in a certain column of the matrix P1, calculate the priority Q for the rules corresponding to the at least two non-zero elements respectively. The calculation formula for the priority Q is:

[0013]

[0014] where K is the number of preset priority indicators, and q k is the priority value of the k-th preset priority indicator;

[0015] For each non-zero element in each column of the matrix P1, retain the value of the element corresponding to the rule with the highest priority, and reset the other elements to zero to obtain the merged matrix P2. Then, the recognition result of the rule corresponding to each non-zero element in the matrix P2 for the text corresponding to the non-zero element is the normalization result.

[0016] Optionally, the priority metrics include, but are not limited to, containing Chinese characters, containing symbols, containing English letters, containing numbers, and the length of the matched text.

[0017] Optionally, the priority value of the priority metric "containing Chinese characters", the priority value of the priority metric "containing symbols", the priority value of the priority metric "containing English letters", and the priority value of the priority metric "containing numbers" decrease in sequence.

[0018] Optionally, the priority value of the priority metric "length of the matched text" is the length of the matched text.

[0019] According to another aspect of the present application, there is also provided a Chinese speech synthesis normalization device, including:

[0020] An initialization module configured to initialize a matrix P0 with a size of M×N and initial elements of 0, where M is the total number of rules and N is the length of the text to be synthesized;

[0021] A matrix update module configured to scan the text to be synthesized using the M rules respectively. If there are t places in the text to be synthesized that satisfy the i-th rule, where i is any integer from 1 to M, and the starting positions of the r-th place among the t places are s r and e r , where r is any integer from 1 to t, then update the values of the elements from the s r to e r in the i-th row of the matrix P0 to any non-zero number to obtain the updated matrix P1;

[0022] A priority calculation module configured to scan each column of the matrix P1. When there are at least two non-zero elements in a certain column of the matrix P1, calculate the priority Q of the rules corresponding to the at least two non-zero elements respectively. The calculation formula for the priority Q is:

[0023]

[0024] where K is the number of preset priority metrics, q k is the priority value of the preset k-th priority metric; and

[0025] The merging processing module is configured to retain the value of the element corresponding to the rule with the highest priority for the non-zero elements in each column of the matrix P1, and reset the other elements to zero to obtain the merged matrix P2. The recognition result of the rule corresponding to each non-zero element of the matrix P2 for the text corresponding to the non-zero element is the normalized processing result.

[0026] Optionally, the priority indicators include but are not limited to containing Chinese characters, containing symbols, containing English letters, containing numbers, and matching the length of the text.

[0027] Optionally, the priority value of the priority indicator "containing Chinese characters", the priority value of the priority indicator "containing symbols", the priority value of the priority indicator "containing English letters", and the priority value of the priority indicator "containing numbers" decrease in sequence.

[0028] Optionally, the priority value of the priority indicator "length of the matched text" is the length of the matched text.

[0029] According to a third aspect of the present application, a computing device is also provided, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the method described in the present application is implemented when the processor executes the computer program.

[0030] The Chinese speech synthesis normalization method and device of the present application does not simply preset the rule priority, but dynamically calculates the priority of each rule based on the matching results of all rules. Therefore, it is possible to quantitatively select the optimal rule to recognize the text to be synthesized, solve the rule conflict problem, and improve the accuracy of normalization.

[0031] Based on the detailed description of the specific embodiments of the present application in combination with the accompanying drawings below, those skilled in the art will become more aware of the above and other objects, advantages and features of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Hereinafter, some specific embodiments of the present application will be described in detail in an exemplary and non-limiting manner with reference to the accompanying drawings. The same reference numerals in the accompanying drawings indicate the same or similar components or parts. It should be understood by those skilled in the art that these drawings are not necessarily drawn to scale. In the accompanying drawings:

[0033] Figure 1 is a schematic flow chart of speech synthesis according to the background technology of this application;

[0034] Figure 2 is a schematic flow chart of a Chinese speech synthesis normalization method according to an embodiment of the present application;

[0035] Figure 3It is a schematic flowchart of speech synthesis according to an embodiment of the present application;

[0036] Figure 4 It is a schematic structural diagram of a Chinese speech synthesis normalization device according to an embodiment of the present application;

[0037] Figure 5 It is a schematic structural diagram of a computing device according to an embodiment of the present application;

[0038] Figure 6 It is a schematic structural diagram of a computer-readable storage medium according to an embodiment of the present application. Detailed implementation manners

[0039] Figure 2 It is a schematic flowchart of a Chinese speech synthesis normalization method according to an embodiment of the present application. The Chinese speech synthesis normalization method generally may include the following steps S1 to S4:

[0040] Step S1, initialize a matrix P0 with a size of M×N and initial elements of 0 (all elements of the matrix P0 are 0), where M is the total number of rules and N is the length of the text to be synthesized (i.e., the number of characters); taking the text to be synthesized "The thing is good, but for the so-called full 199 minus 100 during 618, the price is directly increased to 128 per bottle and there are no gifts" as an example, the total number of rules M≥5, and three of the rules are the amount rule, the date rule, and the full reduction rule. Table 1 intercepts the elements of the matrix P0 related to the above three rules.

[0041] Table 1 Partial elements of the matrix P0

[0042]

[0043] Step S2, use the M rules to scan the text to be synthesized respectively. If there are t places in the text to be synthesized that satisfy the i-th rule, where i is any integer from 1 to M, and record the starting positions of the r-th place among the t places as s r and e r , where r is any integer from 1 to t, then update the values of the elements from the s r to e r th of the i-th row of the matrix P0 to any non-zero number to obtain the updated matrix P1;

[0044] In step S2 above, each rule is used to scan the text separately to identify all parts of the text that satisfy the rule. Each rule may match multiple positions in the text, and the same position may also match multiple rules. Suppose when scanning the text using a certain rule (for example, the i-th rule, 1 ≤ i ≤ M), this rule matches a total of t parts of the text. Among them, the starting point of the first part of the text matched is the s1-th character of the text, and the ending point is the e1-th character of the text. The starting point of the second part of the text matched is the s2-th character of the text, and the ending point is the e2-th character of the text, and so on. The starting point of the t-th part of the text matched is the s t -th character, and the ending point is the e t -th character. Then, the elements from the s1-th element to the e1-th element, from the s2-th element to the e2-th element, ……, and from the s t -th element to the e t -th element in the i-th row of the matrix P0 are all updated to any non-zero number. After all the M rules have scanned the text to be synthesized, the matrix P0 is updated completely to obtain the updated matrix P1, as shown in Table 2. In the updated matrix P1, non-zero elements indicate that the characters in the corresponding text match a certain (or certain) rule, and zero elements indicate that they do not match any rule.

[0045] Table 2 Partial elements of matrix P1

[0046]

[0047] Step S3: Scan each column of the matrix P1. When there are at least two non-zero elements in a certain column of the matrix P1, calculate the priority Q for the rules corresponding to the at least two non-zero elements respectively. The calculation formula for the priority Q is:

[0048]

[0049] where K is the number of preset priority indicators, and q k is the priority value of the k-th preset priority indicator;

[0050] In the above step S3, the purpose of scanning each column of the matrix P1 is to determine whether there are rule conflicts. Suppose a certain column of the matrix P1 (for example, the j-th column, 1 ≤ j ≤ N) contains two or more non-zero elements, indicating that the j-th character of the text matches two or more rules simultaneously. To solve the rule conflict problem, it is necessary to select a rule with the highest priority from the two or more rules, and use the recognition result of the rule with the highest priority for the j-th character of the text as the final recognition result of the character. The priority of the rule is calculated according to the preset priority indicators and the priority values of each priority indicator. The priority indicators of each rule can be the same or different, and the priority values of each priority indicator in different rules can be the same or different. For example, it can be set that a certain rule contains five priority indicators, namely: contains Chinese characters, contains symbols, contains English letters, contains numbers, and the length of the matched text. The corresponding priority values are shown in Table 3, where n represents the number of characters of the matched text.

[0051] In practical applications, it can be set that all rule priority values are as shown in Table 3. When calculating the priority, if there is a situation where the priorities of two rules that match certain characters are the same, the first matched rule shall prevail.

[0052] Table 3 Priority indicators and corresponding priority values

[0053] Priority index Priority value Whether it contains Chinese characters 8 Whether it contains symbols 4 Whether it contains English letters 2 Whether it contains numbers 1 Length of the matched text n

[0054] Suppose there are two non-zero elements in the j-th column of the matrix P1, that is, the j-th character in the text matches two rules (rule A and rule B) simultaneously. Then, it is necessary to calculate the priorities of these two rules respectively. Suppose further that the characters corresponding to the j - 3 to j + 3 columns (a total of 7 columns) all match rule A, but the j - 4 column and the j + 4 column do not match rule A, and the 7 characters corresponding to the j - 3 to j + 3 columns contain numbers and symbols, but do not contain Chinese characters and English letters. The priority indicators and corresponding priority values of rule A are shown in Table 3. Then, for the j-th character, the calculation method of the priority of rule A is: 4 + 1 + 7 = 12, that is, the priority of rule A is 12. The priority of rule B is calculated in the same way.

[0055] Step S4: For the non-zero elements of each column of the matrix P1, retain the values of the elements corresponding to the rule with the highest priority, and reset the other elements to zero to obtain the merged matrix P2. Then, the recognition result of the rule corresponding to each non-zero element of the matrix P2 for the text corresponding to the non-zero element is the normalization processing result.

[0056] The above step S4 performs a combined normalization process on the recognition results of multiple rules. Assume that the priority of rule B in step S3 is 9. Then, select rule A with a higher priority and discard rule B with a lower priority. Retain the value of the element corresponding to rule A in the j-th column of matrix P1, and change the value of the element corresponding to rule B in the j-th column of matrix P1 to 0. Process the non-zero elements in each column of the matrix P1 in the same way to obtain matrix P2, and its elements are shown in Table 4. At this point, there is at most one non-zero element in each column of matrix P2, that is, for each non-conventional text character in the text to be synthesized, only one rule is selected to recognize it, and the recognition result is the normalization result.

[0057] Table 4 Partial elements of matrix P2

[0058]

[0059] The Chinese speech synthesis normalization method according to the embodiments of the present invention can dynamically calculate the priority according to the recognition results of all rules, quantitatively select the optimal rule, and solve the problem of rule conflicts.

[0060] The process of the Chinese speech synthesis method based on the above Chinese speech synthesis normalization method is as Figure 3 shown.

[0061] Synthesis parameter parsing: Parse the parameters of the speech synthesis interface. In addition to conventional information such as the speaker and sampling rate, there is also a field containing the text to be synthesized;

[0062] Full-width to half-width conversion: Identify the full-width characters in the text to be synthesized and convert them into corresponding half-width characters. For example, convert "cm" to half-width "cm";

[0063] Normalization: Use the Chinese speech synthesis normalization method according to the embodiments of the present invention to recognize the text to be synthesized;

[0064] Verbalise: Process the text at the corresponding position in the text to be synthesized using the verbalise function corresponding to the selected rule. For example, use the verbalise function of the amount rule to convert "128" to "one hundred and twenty-eight", use the verbalise function of the date rule to convert "618" to "six one eight", and use the verbalise function of the full reduction rule to convert "full 199 - 100" to "full one hundred and ninety-nine minus one hundred";

[0065] Pinyin conversion: Perform pinyin conversion on the result of the normalization process to obtain a pinyin sequence;

[0066] Speech synthesis: Use a speech synthesis model to convert the pinyin sequence into an audio file;

[0067] Post - processing of speech: The synthesized audio file is post - processed in terms of volume, smoothing, etc., and then the synthesized speech is output.

[0068] Actual tests show that, compared with the prior art, using the Chinese speech synthesis normalization method of the embodiments of the present invention for Chinese speech synthesis can obtain a higher recognition accuracy rate.

[0069] Embodiments of the present invention also provide a Chinese speech synthesis normalization device, as Figure 4 shown. Generally, the device may include:

[0070] Initialization module 1, configured to initialize a matrix P0 with a size of M×N and initial elements of 0, where M is the total number of rules and N is the length of the text to be synthesized;

[0071] Matrix update module 2, configured to scan the text to be synthesized using the M rules respectively. If there are t places in the text to be synthesized that satisfy the i - th rule, where i is any integer from 1 to M, and the starting positions of the r - th place among the t places are s r and e r , where r is any integer from 1 to t, then update the values of the s r to e r elements in the i - th row of the matrix P0 to any non - zero number, obtaining an updated matrix P1;

[0072] Priority calculation module 3, configured to scan each column of the matrix P1. When there are at least two non - zero elements in a certain column of the matrix P1, calculate the priority Q for the rules corresponding to the at least two non - zero elements respectively. The calculation formula for the priority Q is:

[0073]

[0074] where K is the number of preset priority indicators, q k is the priority value of the k - th preset priority indicator; and

[0075] Merging processing module 4, configured to, for the non - zero elements in each column of the matrix P1, retain the value of the element corresponding to the rule with the highest priority, and reset the other elements to zero, obtaining a merged matrix P2. Then, the recognition result of the rule corresponding to each non - zero element in the matrix P2 for the text corresponding to the non - zero element is the normalization processing result.

[0076] The priority indicators include, but are not limited to, containing Chinese characters, containing symbols, containing English letters, containing numbers, and the length of the matched text.

[0077] The priority values of the priority metrics "including Chinese characters", "including symbols", "including English letters", and "including numbers" decrease in sequence.

[0078] The priority value of the priority metric "length of the matched text" is the length of the matched text.

[0079] The Chinese speech synthesis normalization device according to the embodiments of the present invention can execute the steps of the Chinese speech synthesis normalization method according to the embodiments of the present invention, and its principle will not be elaborated here.

[0080] Embodiments of the present application further provide a computing device. Referring to Figure 5 , the computing device includes a memory 1120, a processor 1110, and a computer program stored in the memory 1120 and capable of being run by the processor 1110. The computer program is stored in a space 1130 for program code in the memory 1120. When the computer program is executed by the processor 1110, it implements method steps 1131 for executing any one according to the present invention.

[0081] Embodiments of the present application further provide a computer-readable storage medium. Referring to Figure 6 , the computer-readable storage medium includes a storage unit for program code. The storage unit is provided with a program 1131' for executing the method steps according to the present invention, and the program is executed by the processor.

[0082] Embodiments of the present application further provide a computer program product containing instructions. When the computer program product runs on a computer, it causes the computer to execute the method steps according to the present invention.

[0083] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer loads and executes the computer program instructions, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0084] Those skilled in the art should further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0085] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above-described embodiment methods can be completed by a program instructing a processor. The program can be stored in a computer-readable storage medium, and the storage medium is a non-transitory medium, such as a random access memory, a read-only memory, a flash memory, a hard disk, a solid state drive, a magnetic tape, a floppy disk, an optical disc, and any combination thereof.

[0086] As described above, it is only the preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A Chinese speech synthesis normalization method, comprising: Initialize a matrix P0 with a size of M×N and initial elements of 0, where M is the total number of rules and N is the length of the text to be synthesized; Scan the text to be synthesized using the M rules respectively. If there are t places in the text to be synthesized that satisfy the i-th rule, where i is any integer from 1 to M, and the starting positions of the r-th place among the t places are s r and e r , where r is any integer from 1 to t, then update the values of the s r to e r -th elements in the i-th row of the matrix P0 to any non-zero number to obtain the updated matrix P1; Scan each column of the matrix P1. When there are at least two non-zero elements in a certain column of the matrix P1, calculate the priority Q for the rules corresponding to the at least two non-zero elements respectively. The calculation formula for the priority Q is: where K is the number of preset priority metrics, and q k is the priority value of the k-th preset priority metric; For the non-zero elements in each column of the matrix P1, retain the value of the element corresponding to the rule with the highest priority, and reset the other elements to zero to obtain the merged matrix P2. Then, the recognition result of the rule corresponding to each non-zero element in the matrix P2 for the text corresponding to the non-zero element is the normalization result.

2. The method according to claim 1, wherein, The priority indicators include but are not limited to containing Chinese characters, containing symbols, containing English letters, containing numbers, and the length of the matched text.

3. The method according to claim 2, wherein, The priority value of the priority indicator "containing Chinese characters", the priority value of the priority indicator "containing symbols", the priority value of the priority indicator "containing English letters", and the priority value of the priority indicator "containing numbers" decrease in sequence.

4. The method according to claim 2 or 3, wherein, The priority value of the priority indicator "length of the matched text" is the length of the matched text.

5. A Chinese speech synthesis normalization device, comprising: Initialization module, configured to initialize a matrix P0 with a size of M×N and initial elements of 0, where M is the total number of rules and N is the length of the text to be synthesized; A matrix update module, configured to scan the text to be synthesized using the M rules respectively. If there are t places in the text to be synthesized that satisfy the i-th rule, where i is any integer from 1 to M, and the starting positions of the r-th place among the t places are denoted as s r and e r , where r is any integer from 1 to t, then update the values of the s r to e r th elements in the i-th row of the matrix P0 to any non-zero number to obtain an updated matrix P1; Priority calculation module, configured to scan each column of the matrix P1. When there are at least two non-zero elements in a certain column of the matrix P1, calculate the priority Q for the rules corresponding to the at least two non-zero elements respectively. The calculation formula for the priority Q is: where K is the number of preset priority metrics, and q k is the priority value of the k-th preset priority metric; and Merging processing module, configured to, for the non-zero elements in each column of the matrix P1, retain the value of the element corresponding to the rule with the highest priority, and reset the other elements to zero to obtain the merged matrix P2. Then, the recognition result of the rule corresponding to each non-zero element in the matrix P2 for the text corresponding to the non-zero element is the normalization result.

6. The device according to claim 5, wherein, The priority indicators include but are not limited to containing Chinese characters, containing symbols, containing English letters, containing numbers, and the length of the matched text.

7. The device according to claim 6, wherein, The priority value of the priority indicator "containing Chinese characters", the priority value of the priority indicator "containing symbols", the priority value of the priority indicator "containing English letters", and the priority value of the priority indicator "containing numbers" decrease in sequence.

8. The device according to claim 6 or 7, wherein, The priority value of the priority indicator "length of the matched text" is the length of the matched text.

9. A computing device, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein, When the processor executes the computer program, it implements the method described in any one of claims 1-4.

Citation Information

Patent Citations

  • Speech recognition method and related products

    CN109003603A

  • Speech synthesis method, device and equipment and storage medium

    CN110136692A