Method, apparatus, electronic device and storage medium for predicting the pronunciation of polyphonic characters

By obtaining character features in the Chinese pronunciation synthesis system and integrating global and local information, using Conformer model and conditional random field units to predict multi-tone pronunciation, the problem of multi-tone pronunciation analysis is solved, and the accuracy and fluency of pronunciation synthesis are improved.

CN115512682BActive Publication Date: 2025-07-18BEIJING CENTURY TAL EDUCATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211138255.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-19
Publication Date
2025-07-18
Estimated Expiration
2042-09-19

AI Technical Summary

Technical Problem

The existing Chinese pronunciation synthesis technology is difficult to accurately distinguish the pronunciation of polyphonic characters, which affects the fluency of pronunciation.

Method used

By obtaining the characteristics of each character in the target text, using the BERT model to extract character encoding and position encoding, and combining the Conformer model to fuse global information and local information, and using conditional random field units to predict multitone pronunciation.

Benefits of technology

It improves the accuracy of pronunciation prediction of polyphonic characters, ensures the fluency of speech synthesis, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115512682B_ABST
    Figure CN115512682B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, electronic device, and storage medium for predicting the pronunciation of polyphonic characters, including: obtaining respective character features corresponding to each character in a target text, using a polyphonic character prediction model to perform prediction and fusion on the global information and local information of the target text according to the respective character features corresponding to each character, obtaining respective target features corresponding to each character, and performing pronunciation prediction on the polyphonic characters in the target text according to the respective target features corresponding to each character, so as to obtain a prediction result of the pronunciation of the polyphonic characters in the target text. Thereby, the present disclosure can correctly distinguish the pronunciation of polyphonic characters to improve the fluency of speech synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text recognition, and in particular, to a method, device, electronic device and storage medium for predicting the pronunciation of polyphonic characters. Background Art

[0002] The current speech synthesis technology mainly includes two parts: text analysis and speech generation. Among them, text analysis is used to provide a basis for subsequent speech synthesis to ensure the fluency of speech synthesis.

[0003] In the Chinese language system, there are a certain number of polyphonic characters. For example, "looking at the pedestrians walking in rows (hang2) on the street". Among them, the specific pronunciation of polyphonic characters depends not only on the context information but also on the specific local information. Therefore, whether the pronunciation of polyphonic characters can be correctly distinguished is of great significance for subsequent speech synthesis processing.

[0004] In view of this, there is an urgent need for a technical means to correctly distinguish the pronunciation of polyphonic characters. Summary of the Invention

[0005] In view of this, embodiments of the present disclosure provide a method, device, electronic device and storage medium for predicting the pronunciation of polyphonic characters, which can correctly distinguish the pronunciation of polyphonic characters.

[0006] According to one aspect of the embodiments of the present disclosure, a method for predicting the pronunciation of polyphonic characters is provided, including: obtaining each character feature corresponding to each character in the target text; using a polyphonic character prediction model to predict and fuse the global information and local information of the target text according to each character feature corresponding to each character, obtaining each target feature corresponding to each character, and performing pronunciation prediction on the polyphonic characters in the target text according to each target feature corresponding to each character, to obtain the pronunciation prediction result of the polyphonic characters in the target text.

[0007] According to a second aspect of the embodiments of the present disclosure, a device for predicting the pronunciation of polyphonic characters is provided, including an obtaining module for obtaining each character feature corresponding to each character in the target text; a polyphonic character prediction model for predicting the global information and local information of the target text according to each character feature corresponding to each character, fusing the global information and local information of the target text, obtaining each target feature corresponding to each character, and performing pronunciation prediction on the polyphonic characters in the target text according to each target feature corresponding to each character, to obtain the pronunciation prediction result of the polyphonic characters in the target text.

[0008] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including: a processor; and a memory storing a program, where the program includes instructions that, when executed by the processor, cause the processor to execute the polyphonic pronunciation prediction method described in the first aspect above.

[0009] According to a fourth aspect of the embodiments of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause the computer to execute the polyphonic pronunciation prediction method described in the first aspect.

[0010] The polyphonic pronunciation prediction method provided by the embodiments of the present disclosure can predict the local information and global information of the target text according to the respective character features corresponding to the characters in the target text, and by fusing the local information and global information of the target text, perform the prediction of the polyphonic pronunciation of the target text. Through the technical solution provided by the present disclosure, not only can the accuracy of the polyphonic pronunciation prediction result be improved, but also the fluency of the subsequent speech synthesis can be ensured, so as to improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In the following description of the exemplary embodiments in conjunction with the drawings, more details, features and advantages of the present disclosure are disclosed. In the drawings:

[0012] Figure 1 is a flowchart of the polyphonic pronunciation prediction method according to an exemplary embodiment of the present disclosure.

[0013] Figure 2 is a flowchart of the polyphonic pronunciation prediction method according to another exemplary embodiment of the present disclosure.

[0014] Figure 3 is Figure 2 a schematic diagram of the data generation process of the illustrated embodiment.

[0015] Figure 4 is a flowchart of the polyphonic pronunciation prediction method according to another exemplary embodiment of the present disclosure.

[0016] Figure 5 is a schematic diagram of the architecture of the polyphonic pronunciation prediction device according to an exemplary embodiment of the present disclosure.

[0017] Figure 6 is a schematic diagram of the architecture of the electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0019] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0020] The term "including" and its variations used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions executed by these devices, modules or units or their interdependent relationships.

[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more". The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0022] In a current typical Chinese speech synthesis system, the front-end module at least includes two parts, namely prosody structure prediction (PSP) and grapheme-to-phoneme conversion (G2P).

[0023] Among them, the prosody structure prediction part is mainly used to predict the distribution of prosody boundaries of prosody words, prosody phrases, and intonation phrases in the input text. For example, converting the input text "Comrade Chang Jianghe of Nanjing City" to "Nanjing#1 Mayor#2 Jianghe#1 Comrade#3". The grapheme-to-phoneme conversion part is used to convert each Chinese character in the input text into the corresponding pronunciation. For example, converting "Ancient Capital Xi'an" to "gu3 du1 xi1 an1". Therefore, the front-end module in the Chinese speech synthesis system has an important impact on the intelligibility and naturalness of the synthesized speech in the backend.

[0024] As described in the background section above, the specific pronunciation of a polyphonic character depends not only on the context information of the input text (i.e., global information), but also on the local information of the input text (i.e., the information within the scope of focus of the target text, such as the characters within a certain range around the target text). Based on the above characteristics, the present disclosure proposes a polyphonic character pronunciation prediction scheme, which captures both the global information and the local information of the input text to improve the accuracy of the polyphonic character pronunciation prediction result.

[0025] The following will describe in detail each specific embodiment of the present disclosure with reference to the accompanying drawings.

[0026] Figure 1 FIG. is a processing flowchart of the polyphonic character pronunciation prediction method according to an exemplary embodiment of the present disclosure, which mainly includes the following steps:

[0027] Step S102, obtain the character features corresponding to each character in the target text.

[0028] Optionally, a language prediction model can be used to perform feature extraction on each character in the target text to obtain the character features corresponding to each character in the target text.

[0029] In this embodiment, the language prediction model may include, but is not limited to, the BERT model (Bidirectional Encoder Representation from Transformers; a bidirectional encoding representation model based on transformers).

[0030] Specifically, the BERT model can be used to perform feature extraction on each character in the target text to obtain the character encoding and position encoding of each character in the target text.

[0031] Exemplarily, the character features corresponding to each character may include a 768-dimensional feature vector.

[0032] It should be noted that the feature dimension of the character features is not limited to 768 dimensions, and can be arbitrarily adjusted according to conditions such as the actual application scenario of the target text and the prediction accuracy requirements. The present disclosure does not limit this.

[0033] Step S104, use the polyphonic character prediction model to predict and fuse the global information and local information of the target text according to the character features corresponding to each character, obtain the target features corresponding to each character, and perform pronunciation prediction on the polyphonic characters in the target text according to the target features corresponding to each character to obtain the polyphonic character pronunciation prediction result of the target text.

[0034] Optionally, the polyphonic character prediction model may include a Convolution-augmented Transformer (Conformer), which can predict and fuse the global information and local information of the target text based on the character features corresponding to each character, so as to obtain the target features corresponding to each character (refer to Figure 3 ).

[0035] Thereby, the present disclosure uses Conformer as the encoder, which can combine the global modeling ability of the attention mechanism and the local modeling ability of the CNN (Convolutional Neural Network) to effectively fuse the global information and local information of the target text, thereby improving the accuracy of the polyphonic character pronunciation prediction result.

[0036] Optionally, the polyphonic character prediction model may include a conditional random field unit (refer to Figure 3 ).

[0037] Specifically, the conditional random field unit can be used to perform pronunciation prediction on at least one polyphonic character in the target text according to the target features corresponding to each character, and obtain the predicted pronunciations of at least one polyphonic character.

[0038] In summary, the polyphonic character pronunciation prediction method provided in this embodiment predicts and fuses the global information and local information of the target text according to the character features corresponding to each character in the target text, so as to take into account both the global information of the target text (i.e., the context information of the target text) and the local information (i.e., the information in the key attention range of the target text, such as the characters within a certain range around the target text, etc.), thereby effectively improving the accuracy of the polyphonic character pronunciation prediction result and contributing to the improvement of the fluency of speech synthesis.

[0039] Figure 2 FIG. is a processing flowchart of the polyphonic character pronunciation prediction method according to another exemplary embodiment of the present disclosure. This embodiment is a specific implementation of the above step S104. The following will be combined with Figure 3 This embodiment will be described in detail, which mainly includes the following steps:

[0040] Step S202: Perform a first feed-forward process on the character features corresponding to each character to obtain the first feed-forward features corresponding to each character.

[0041] In this embodiment, the character features may include character encoding and position encoding. Among them, the character encoding is used to convert each character (such as a Chinese character) in the target text into information that can be recognized by a computer, and the position encoding is used to represent the position information of each character (such as a Chinese character) in the target text.

[0042] Optionally, the polyphonic character prediction model may include a first feed-forward unit.

[0043] In this embodiment, the first feedforward unit of the polyphonic character prediction model can perform a first-dimensional conversion process on the character features corresponding to each character to obtain the first conversion features corresponding to each character.

[0044] Specifically, the first feedforward unit can be used to perform a first-dimensional conversion process on the character features corresponding to each character to obtain the first conversion features corresponding to each character (for example, converting the 738-dimensional character features into 1024-dimensional or 512-dimensional first conversion features), and fuse the character features and the first conversion features of the same character to obtain the first feedforward features corresponding to each character (refer to Figure 3 ).

[0045] It should be noted that the feature dimension of the first conversion feature is not limited to the above 1024 dimensions and 512 dimensions, and can be arbitrarily adjusted according to the actual application scenario, prediction accuracy, etc. of the target text, and the present disclosure does not limit this.

[0046] In this embodiment, the character features and the first conversion features of the same character can be added and fused (refer to the symbol “+” in Figure 3 ) to obtain the first feedforward features corresponding to each character.

[0047] Optionally, the first feedforward unit may at least include an activation function layer (ReLU layer) and a dropout layer (Dropout layer).

[0048] Step S204: Predict the global information and local information of the target text according to the first feedforward features corresponding to each character, and fuse the global information and the local information to obtain the fused features corresponding to each character.

[0049] Optionally, the polyphonic character prediction model may include a multi-head self-attention (Multi-Head Self Attention) unit and a convolution (Convolution) unit.

[0050] Specifically, the multi-head self-attention unit (for example, the multi-head self-attention mechanism of the transformer encoder) can be used to perform prediction according to the first feedforward features corresponding to each character to obtain the global features corresponding to each character, and fuse the first feedforward features and the global features of the same character to obtain the intermediate features corresponding to each character (refer to Figure 3 ).

[0051] In this embodiment, the first feedforward features and the global features of the same character can be added and fused (refer to the symbol “+” in Figure 3 ) to obtain the intermediate features corresponding to each character.

[0052] Specifically, a convolution unit (e.g., the convolution layer of a CNN network) can be utilized to perform a prediction (e.g., convolution processing) based on the intermediate features corresponding to each character, obtain the local features corresponding to each character, and fuse the intermediate features and local features of the same character to obtain the fused features corresponding to each character (refer to Figure 3 ).

[0053] In this embodiment, addition fusion can be performed on the intermediate features and local features of the same character (refer to the symbol “+” in Figure 3 ) to obtain the fused features corresponding to each character.

[0054] Step S206: Perform a second feed-forward process on the fused features corresponding to each character to obtain the second feed-forward features corresponding to each character.

[0055] Optionally, the polyphonic character prediction model may include a second feed-forward unit.

[0056] In this embodiment, the second feed-forward unit of the polyphonic character prediction model can be utilized to perform a second-dimensionality conversion process on the character features corresponding to each character to obtain the second conversion features corresponding to each character.

[0057] Specifically, the second feed-forward unit can be utilized to perform a second-dimensionality conversion process on the fused features corresponding to each character to obtain the second conversion features corresponding to each character (e.g., convert the 1024-dimensional or 512-dimensional fused features into 768-dimensional second conversion features), and fuse the fused features and second conversion features of the same character to obtain the second feed-forward features corresponding to each character (refer to Figure 3 ).

[0058] In this embodiment, the feature dimension of the second conversion features and the feature dimension of the character features can be the same (e.g., both the second conversion features and the character features have a feature dimension of 768), but this is not a limitation. The feature dimension of the second conversion features and the feature dimension of the character features can also be set differently, and those skilled in the art can adjust according to actual prediction accuracy requirements. The present disclosure does not limit this.

[0059] In this embodiment, addition fusion can be performed on the fused features and second conversion features of the same character (refer to the symbol “+” in Figure 3 ) to obtain the second feed-forward features corresponding to each character.

[0060] Optionally, the second feed-forward unit may at least include an activation function layer (ReLU layer) and a dropout layer (Dropout layer).

[0061] Step S208: Perform a normalization process on the second feed-forward features corresponding to each character to obtain the target features corresponding to each character.

[0062] Optionally, the polyphonic character prediction model may include a normalization processing unit (e.g., a LayerNorm layer).

[0063] In this embodiment, the normalization processing unit of the polyphonic character prediction model may be used to perform normalization processing on each second feedforward feature corresponding to each character to obtain each target feature corresponding to each character (refer to Figure 3 ).

[0064] Step S210: According to each target feature corresponding to each character, perform pronunciation prediction on the polyphonic characters in the target text to obtain the polyphonic character pronunciation prediction result of the target text.

[0065] Optionally, the polyphonic character prediction model may include a Conditional Random Field (CRF) unit.

[0066] Specifically, the CRF unit may be used to identify at least one polyphonic character in each character according to each target feature corresponding to each character in the target text, and for any current polyphonic character among the polyphonic characters, according to multiple candidate pronunciations of the current polyphonic character and the target feature of the current polyphonic character, predict each pronunciation probability value corresponding to each candidate pronunciation (e.g., Chinese pinyin) of the current polyphonic character, and determine the candidate pronunciation with the largest pronunciation probability value as the predicted pronunciation of the current polyphonic character.

[0067] In summary, the polyphonic character pronunciation prediction method of this embodiment uses the multi-head self-attention unit and the convolutional unit in the polyphonic character prediction model to learn the global information and local information in the target text, and by fusing the global information and local information of the target text at each prediction stage (that is, fusing the input data and output data of the first feedforward unit, the multi-head self-attention unit, the convolutional unit, and the second feedforward unit respectively), the present solution can focus on the key local information in the target text while taking into account the global context information of the target text, thereby effectively improving the accuracy of the polyphonic character pronunciation prediction result.

[0068] Figure 4 The flowchart of the polyphonic character pronunciation prediction method according to another exemplary embodiment of the present disclosure is shown. This embodiment mainly shows the training implementation scheme of the polyphonic character prediction model in step S104, which mainly includes the following steps:

[0069] Step S402: Use the trained language prediction model to perform feature extraction on each character in the training text to obtain each character feature corresponding to each character in the training text.

[0070] In this embodiment, the trained language prediction model may include, but is not limited to, a BERT model.

[0071] Specifically, the BERT model can be used to perform feature extraction on each character in the training text to obtain the character encoding and position encoding of each character.

[0072] Step S404: Using the polyphonic character prediction model, perform pronunciation prediction on the polyphonic characters in the training text according to the respective character features corresponding to the characters, and obtain the predicted pronunciations of the polyphonic characters in the training text.

[0073] Specifically, the polyphonic character prediction model can be used to learn the global information and local information in the training text according to the respective character features corresponding to the characters in the training text, and perform pronunciation prediction on the polyphonic characters in the training text according to the fusion result of the global information and local information in the training text, so as to obtain the predicted pronunciations of the polyphonic characters in the training text.

[0074] Regarding the specific prediction scheme of this step, reference can be made to the relevant descriptions of the above Figure 2 illustrated embodiments and will not be elaborated here.

[0075] Step S406: Compare the labeled pronunciations and the predicted pronunciations of the polyphonic characters in the training text to obtain the loss function of the polyphonic character prediction model.

[0076] Optionally, the loss function of the polyphonic character prediction model may include but is not limited to: Mean Absolute Error (MAE), Mean Square Error (MSE), etc.

[0077] Step S408: Update the polyphonic character prediction model according to the loss function.

[0078] Specifically, the model parameters (such as weight parameters) of the polyphonic character prediction model can be iteratively updated according to the current loss function.

[0079] Step S410: Determine whether the current training result of the polyphonic character prediction model meets the given training end condition. If so, execute step S412; if not, execute step S402.

[0080] Optionally, if it is determined that the loss function meets the given convergence value, obtain the judgment result that the current training result of the polyphonic character prediction model meets the training end condition.

[0081] Specifically, if it is determined that the function value of the obtained loss function meets the given convergence value, it means that the function value of the loss function tends to be stable (i.e., the change of the loss function is very small). In this case, the judgment result that the current training result of the polyphonic character prediction model meets the training end condition can be obtained.

[0082] Optionally, if the number of iterations of the polyphonic character prediction model meets the given maximum number of iterations, obtain the judgment result that the current training result of the polyphonic character prediction model meets the training end condition.

[0083] In this embodiment, if the current training result of the polyphonic character prediction model does not meet the given training end condition, return to step S402 to continue executing the training task of the polyphonic character prediction model.

[0084] Step S412, obtain the trained polyphonic character prediction model.

[0085] Specifically, if the current training result of the polyphonic character prediction model meets the training end condition, the optimization update of the polyphonic character prediction model can be stopped, and based on the current model parameters of the polyphonic character prediction model, obtain the trained polyphonic character prediction model.

[0086] In summary, this embodiment uses the trained language prediction model to cooperate in the training task of the polyphonic character prediction model, which can not only achieve a better model training effect, but also effectively reduce the training cost of the model.

[0087] Figure 5 The structural block diagram of the polyphonic character pronunciation prediction device according to an exemplary embodiment of the present disclosure is shown. As shown in the figure, the polyphonic character pronunciation prediction device of this embodiment mainly includes: an acquisition module 502 and a polyphonic character prediction model 504.

[0088] The acquisition module 502 is used to acquire the respective character features corresponding to each character in the target text.

[0089] The polyphonic character prediction model 504 is used to predict the global information and local information of the target text according to the respective character features corresponding to each character, fuse the global information and local information of the target text, obtain the respective target features corresponding to each character, and perform pronunciation prediction on the polyphonic characters in the target text according to the respective target features corresponding to each character, so as to obtain the polyphonic character pronunciation prediction result of the target text.

[0090] Optionally, the acquisition module 502 is further used to: use the language prediction model to perform feature extraction on each character in the target text to obtain the character encoding and position encoding of each character.

[0091] Optionally, the polyphone prediction model 504 is further configured to: perform a first feedforward process on each character feature corresponding to each character to obtain each first feedforward feature corresponding to each character; predict the global information and local information of the target text according to each first feedforward feature corresponding to each character, and fuse the global information and local information to obtain each fusion feature corresponding to each character; perform a second feedforward process on each fusion feature corresponding to each character to obtain each second feedforward feature corresponding to each character; perform a normalization process on each second feedforward feature corresponding to each character to obtain each target feature corresponding to each character.

[0092] Optionally, the polyphone prediction model 504 includes a first feedforward unit and a second feedforward unit.

[0093] Optionally, the polyphone prediction model 504 is further configured to: use the first feedforward unit to perform a first dimensionality conversion process on each character feature corresponding to each character to obtain each first conversion feature corresponding to each character, and fuse the character feature and the first conversion feature of the same character to obtain each first feedforward feature corresponding to each character; use the second feedforward unit to perform a second dimensionality conversion process on each fusion feature corresponding to each character to obtain each second conversion feature corresponding to each character, and fuse the fusion feature and the second conversion feature of the same character to obtain each second feedforward feature corresponding to each character.

[0094] Optionally, the first feedforward unit and the second feedforward unit each include an activation function layer and a dropout layer.

[0095] Optionally, the polyphone prediction model 504 includes a multi-head self-attention unit and a convolutional unit.

[0096] Optionally, the polyphone prediction model 504 is further configured to: use the multi-head self-attention unit to perform prediction according to each first feedforward feature corresponding to each character to obtain each global feature corresponding to each character, and fuse the first feedforward feature and the global feature of the same character to obtain each intermediate feature corresponding to each character; and use the convolutional unit to perform prediction according to each intermediate feature corresponding to each character to obtain each local feature corresponding to each character, and fuse the intermediate feature and the local feature of the same character to obtain each fusion feature corresponding to each character.

[0097] Optionally, the polyphone prediction model 504 includes a conditional random field unit.

[0098] Optionally, the polyphonic character prediction model 504 is further configured to: use the conditional random field unit to identify polyphonic characters in each character according to the target features corresponding to each character, and predict the pronunciation probability values corresponding to each candidate pronunciation of the polyphonic character according to the multiple candidate pronunciations of the polyphonic character and the target features of the polyphonic character, and determine the candidate pronunciation with the largest pronunciation probability value as the predicted pronunciation of the polyphonic character according to the pronunciation probability values corresponding to each candidate pronunciation.

[0099] Optionally, the polyphonic character pronunciation prediction device 500 further includes a training module (not shown): which is configured to use the trained language prediction model to perform feature extraction on each character in the training text to obtain the character features corresponding to each character in the training text; use the polyphonic character prediction model 504 to perform pronunciation prediction on the polyphonic characters in the training text according to the character features corresponding to each character to obtain the predicted pronunciation of the polyphonic characters in the training text; compare the labeled pronunciation and the predicted pronunciation of the polyphonic characters in the training text to obtain the loss function of the polyphonic character prediction model 504; and update the polyphonic character prediction model 504 according to the loss function until the current training result of the polyphonic character prediction model 504 meets the given training end condition to obtain the trained polyphonic character prediction model 504.

[0100] The embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause the computer to execute the polyphonic character pronunciation prediction method described in the exemplary embodiments of the present disclosure.

[0101] The exemplary embodiments of the present disclosure provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program capable of being executed by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the polyphonic character pronunciation prediction method according to the exemplary embodiments of the present disclosure.

[0102] Please refer to Figure 6 , and now the structural block diagram of the electronic device 600 that can be used as the server or client of the present disclosure will be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0103] As Figure 6 shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0104] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. The input unit 606 can be any type of device capable of inputting information into the electronic device 600. The input unit 606 can receive input digital or character information and generate key signal inputs related to user settings and / or function controls of the electronic device. The output unit 607 can be any type of device capable of presenting information and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 604 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0105] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above. For example, in some embodiments, the homophone pronunciation prediction method as described above can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 600 via the ROM 602 and / or the communication unit 609. In some embodiments, the computing unit 601 can be configured to execute the above-mentioned homophone pronunciation prediction method in any other appropriate way (e.g., by means of firmware).

[0106] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0107] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0108] As used in the present disclosure, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, and / or device (e.g., a disk, an optical disk, a memory, a programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0109] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0110] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0111] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other.

[0112] It should be noted that, according to the needs of implementation, each component / step described in the embodiments of the present disclosure can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the objectives of the embodiments of the present disclosure.

[0113] The above embodiments are only used to illustrate the embodiments of the present disclosure, rather than to limit the embodiments of the present disclosure. Those of ordinary skill in the relevant art can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present disclosure. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present disclosure. The patent protection scope of the embodiments of the present disclosure shall be defined by the claims.

Claims

1. A method for predicting the pronunciation of polyphonic characters, characterized in that, Including: Using a Transformer-based bidirectional encoding representation model to perform feature extraction on each character in the target text to obtain respective character features corresponding to each character; Using a convolutional enhanced Transformer to perform a first-dimensional conversion process on the respective character features corresponding to each character to obtain respective first conversion features, fusing the character features and the first conversion features of the same character to obtain respective first feed-forward features corresponding to each character, performing global prediction based on the respective first feed-forward features corresponding to each character to obtain respective global features corresponding to each character, fusing the first feed-forward features and the global features of the same character to obtain respective intermediate features corresponding to each character, performing local prediction based on the respective intermediate features corresponding to each character to obtain respective local features corresponding to each character, fusing the intermediate features and the local features of the same character to obtain respective fusion features corresponding to each character, and performing a second feed-forward process and a normalization process on the respective fusion features corresponding to each character to obtain respective target features corresponding to each character; Using a conditional random field unit to identify polyphonic characters in each character according to the respective target features corresponding to each character, and predicting respective pronunciation probability values corresponding to each candidate pronunciation of the polyphonic character based on the multiple candidate pronunciations of the polyphonic character and the target features of the polyphonic character, and determining the candidate pronunciation with the largest pronunciation probability value as the predicted pronunciation of the polyphonic character.

2. The method for predicting the pronunciation of polyphonic characters according to claim 1, wherein The respective character features corresponding to each character in the target text can be obtained through the following manner: Using the Transformer-based bidirectional encoding representation model to perform feature extraction on each character in the target text to obtain the character encoding and position encoding of each character.

3. The method according to claim 1, wherein The convolutional enhanced Transformer includes a first feed-forward unit, a second feed-forward unit, and a normalization processing unit; and among them, The performing a first feed-forward process on the respective character features corresponding to each character to obtain respective first feed-forward features corresponding to each character includes: Using the first feed-forward unit to perform a first-dimensional conversion process on the respective character features corresponding to each character to obtain respective first conversion features corresponding to each character, and fusing the character features and the first conversion features of the same character to obtain respective first feed-forward features corresponding to each character; The performing a second feed-forward process and a normalization process on the respective fusion features corresponding to each character to obtain respective target features corresponding to each character includes: Using the second feed-forward unit to perform a second-dimensional conversion process on the respective fusion features corresponding to each character to obtain respective second conversion features corresponding to each character, and fusing the fusion features and the second conversion features of the same character to obtain respective second feed-forward features corresponding to each character; Using the normalization processing unit to perform a normalization process on the respective second feed-forward features corresponding to each character to obtain respective target features corresponding to each character.

4. The method according to claim 3, characterized in that, The first feed-forward unit and the second feed-forward unit respectively include an activation function layer and a dropout layer.

5. The method according to claim 3, characterized in that, The convolutional enhanced Transformer includes a multi-head self-attention unit and a convolutional unit; Using the multi-head self-attention unit to perform global prediction based on the respective first feed-forward features corresponding to each character to obtain respective global features corresponding to each character, and fusing the first feed-forward features and the global features of the same character to obtain respective intermediate features corresponding to each character; Using the convolutional unit, perform local prediction according to the intermediate features corresponding to each character, obtain the local features corresponding to each character, and fuse the intermediate features and local features of the same character to obtain the fused features corresponding to each character.

6. The method according to claim 1, characterized in that, The method further includes: Using the trained bidirectional encoder representations from transformers model to perform feature extraction on each character in the training text, and obtain the character features corresponding to each character in the training text; Using a polyphonic character prediction model composed of a convolutional enhanced transformer and a conditional random field unit, perform pronunciation prediction on the polyphonic characters in the training text according to the character features corresponding to each character, and obtain the predicted pronunciations of the polyphonic characters in the training text; Compare the labeled pronunciations and the predicted pronunciations of the polyphonic characters in the training text to obtain the loss function of the polyphonic character prediction model; Update the polyphonic character prediction model according to the loss function until the current training result of the polyphonic character prediction model meets the given training end condition, so as to obtain a trained polyphonic character prediction model.

7. A polyphonic pronunciation prediction device, characterized in that, It includes: A bidirectional encoder representations from transformers model, which is used to perform feature extraction on each character in the target text and obtain the character features corresponding to each character; A convolutional enhanced transformer, which is used to perform first-dimensional conversion processing on the character features corresponding to each character to obtain the first conversion features, fuse the character features and the first conversion features of the same character to obtain the first feed-forward features corresponding to each character, perform global prediction according to the first feed-forward features corresponding to each character to obtain the global features corresponding to each character, fuse the first feed-forward features and the global features of the same character to obtain the intermediate features corresponding to each character, perform local prediction according to the intermediate features corresponding to each character to obtain the local features corresponding to each character, fuse the intermediate features and the local features of the same character to obtain the fused features corresponding to each character, perform second feed-forward processing and normalization processing on the fused features corresponding to each character to obtain the target features corresponding to each character; A conditional random field unit, which is used to identify the polyphonic characters in each character according to the target features corresponding to each character, and predict the pronunciation probability values corresponding to each candidate pronunciation of the polyphonic character according to the multiple candidate pronunciations of the polyphonic character and the target features of the polyphonic character, and determine the candidate pronunciation with the largest pronunciation probability value as the predicted pronunciation of the polyphonic character.

8. An electronic device, including: A processor; And A memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the polyphonic character pronunciation prediction method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Speech recognition method and device, computer equipment and storage medium

    CN113539273A