Character-voice conversion method and device based on polyphone context recognition
By constructing a multiphonetic word recognition and prediction model, combining a multiphonetic word dictionary and attention mechanism, the problem of low accuracy in speech and text conversion is solved, and higher conversion accuracy and user experience is achieved.
Patent Information
- Application Number
- CN202510470686.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-08-05
AI Technical Summary
In the prior art, polyphonic characters have low accuracy in the process of voice-based conversion, resulting in errors in the word content generated by the word and incorrect intonation when the word generated by the word.
Multi-task learning is used to build a multi-tone word recognition model and a multi-tone word prediction model, combining a multi-tone word dictionary, attention mechanism and user feedback mechanism to improve the accuracy of multi-tone word recognition through multi-task learning architecture and context analysis.
It improves the accuracy of voice and text conversion, saves time for manual proofreading, and improves user experience.
Smart Images

Figure CN120431931A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text-to-speech conversion, and more particularly, to a text-to-speech conversion method and device based on the recognition of polyphonic characters in context. Background Art
[0002] A speech platform is a platform for converting the text and speech of an article into each other. In the process of use, it is found that there are differences in the conversion between polyphonic characters in speech and text. When used, there will be situations where the text content generated by speech is incorrect and the intonation is incorrect when text is generated into speech; for example, for text-to-speech conversion: "de, děi, de" in different contexts, the generated text is different; for text-to-speech conversion: "xing", different pronunciations need to be made according to different situations. However, there is currently no method for accurately identifying and converting polyphonic characters, resulting in the technical problem of low accuracy in speech-to-text conversion. Summary of the Invention
[0003] In view of the deficiencies of the prior art, the present invention provides a text-to-speech conversion method and device based on the recognition of polyphonic characters in context.
[0004] According to one aspect of the present invention, there is provided a text-to-speech conversion method based on the recognition of polyphonic characters in context, including:
[0005] Constructing a polyphonic character recognition model and a polyphonic character prediction model by using multi-task learning;
[0006] Using the polyphonic character recognition model to recognize the polyphonic character information in the task to be converted, where the task to be converted is the speech to be converted or the text to be converted;
[0007] Using the polyphonic character prediction model to predict the pronunciation or font of the polyphonic characters according to the polyphonic character information and the context of the task to be converted, and determining the pronunciation or font of each polyphonic character in the polyphonic character information;
[0008] Performing polyphonic character recognition in the task to be converted according to the pronunciation or font of each polyphonic character to complete the conversion of the task to be converted.
[0009] Optionally, constructing a polyphonic character recognition model and a polyphonic character prediction model by using multi-task learning includes:
[0010] Constructing a polyphonic character dictionary including polyphonic characters and their different pronunciations;
[0011] Performing machine model recognition based on the polyphonic character dictionary to construct a polyphonic character recognition model;
[0012] Training a machine model according to the pre-constructed polyphonic character training data based on context to construct a polyphonic character prediction model, where the machine model is a machine model based on an attention mechanism.
[0013] Optionally, machine model recognition is performed based on a polyphonetic word dictionary to construct a polyphonetic word recognition model, including:
[0014] Perform feature encoding on the text in the polyphone dictionary through a text encoder to obtain text features;
[0015] The speech in the polyphone dictionary is feature-encoded by a speech feature encoder to obtain speech features;
[0016] Merge text features and speech features to obtain merged features;
[0017] Perform polyphonetic character recognition and speech recognition based on the combined features, and output the polyphonetic character recognition results and speech recognition results;
[0018] A polyphone recognition model is constructed based on the polyphone recognition results and speech recognition results.
[0019] According to another aspect of the present invention, a text-to-speech conversion device based on polyphonetic word context recognition is provided, comprising:
[0020] A construction module for constructing a polyphone recognition model and a polyphone prediction model using multi-task learning;
[0021] A recognition module, configured to use a polyphone recognition model to recognize polyphone information in a task to be converted, wherein the task to be converted is speech to be converted or text to be converted;
[0022] A prediction module is used to use a polyphone prediction model to predict the pronunciation or font of the polyphone according to the polyphone information and the context of the task to be converted, and determine the pronunciation or font of each polyphone in the polyphone information;
[0023] The conversion module is used to identify the polyphones in the task to be converted according to the pronunciation or font of each polyphone, and complete the conversion of the task to be converted.
[0024] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and the computer program is used to execute the method according to any one of the above aspects of the present invention.
[0025] According to another aspect of the present invention, an electronic device is provided, comprising: a processor; a memory for storing instructions executable by the processor; and the processor for reading the executable instructions from the memory and executing the instructions to implement the method described in any one of the above aspects of the present invention.
[0026] Therefore, the present invention constructs a polyphonetic character recognition model and a polyphonetic character prediction model through multi-task learning, thereby improving the accuracy of mutual conversion between speech and text, saving time for manual proofreading, and improving user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] A more complete understanding of exemplary embodiments of the present invention may be obtained by referring to the following drawings:
[0028] Figure 1 1 is a flow chart of a method for converting text to speech based on polyphonetic character context recognition provided by an exemplary embodiment of the present invention;
[0029] Figure 2 1 is a schematic structural diagram of a text-to-speech conversion device based on polyphonetic character context recognition provided by an exemplary embodiment of the present invention;
[0030] Figure 3 This is a structure of an electronic device provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0031] Below, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described herein.
[0032] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present invention unless specifically stated otherwise.
[0033] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of the present invention are only used to distinguish different steps, devices or modules, and neither represent any specific technical meaning nor indicate the necessary logical order between them.
[0034] It should also be understood that, in the embodiments of the present invention, “a plurality of” may refer to two or more than two, and “at least one” may refer to one, two or more than two.
[0035] It should also be understood that any component, data or structure mentioned in the embodiments of the present invention can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.
[0036] In addition, the term "and / or" in this invention merely describes an association relationship between related objects, indicating that three possible relationships exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Furthermore, the character " / " in this invention generally indicates that the related objects are in an "or" relationship.
[0037] It should also be understood that the description of the various embodiments of the present invention focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced with each other. For the sake of brevity, they will not be described one by one.
[0038] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0039] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.
[0040] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0041] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0042] Embodiments of the present invention can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate in conjunction with numerous other general-purpose or specialized computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with terminal devices, computer systems, servers, and other electronic devices include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above.
[0043] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media, including storage devices.
[0044] Exemplary Methods
[0045] Figure 1This is a flow chart of a method for converting text to speech based on polyphone context recognition provided by an exemplary embodiment of the present invention. This embodiment can be applied to electronic devices such as Figure 1 As shown, the text-to-speech conversion method 100 based on polyphone context recognition includes the following steps:
[0046] Step 101, using multi-task learning to build a polyphonetic character recognition model and a polyphonetic character prediction model;
[0047] Step 102: using a polyphone recognition model to identify polyphone information in a task to be converted, wherein the task to be converted is speech to be converted or text to be converted;
[0048] Step 103, using a polyphone prediction model, predicting the pronunciation or font of the polyphone according to the polyphone information and the context of the task to be converted, and determining the pronunciation or font of each polyphone in the polyphone information;
[0049] Step 104 , identifying the polyphonetic characters in the task to be converted based on the pronunciation or font of each polyphonetic character, and completing the conversion of the task to be converted.
[0050] Specifically, to address the technical problems existing in the background technology, the present invention comprehensively utilizes natural language processing, machine learning, big data and other technologies, combined with a rich corpus and user feedback mechanism, to continuously optimize the accuracy of the system and user experience. The specific implementation scheme is as follows:
[0051] 1. Polyphonetic Character Recognition:
[0052] 1. Build a dictionary of polyphonetic characters: First, you need to build a dictionary that contains polyphonetic characters and their different pronunciations so that the system can accurately identify polyphonetic characters.
[0053] 2. Context analysis: Design an algorithm or model to determine which pronunciation should be used for polyphones based on contextual information. Natural language processing technology can be used for part-of-speech tagging and syntactic analysis.
[0054] 1. Polyphonetic Character Recognition Model Based on Deep Learning
[0055] Substantive technical solutions:
[0056] Multi-task learning: Polyphonetic character recognition is treated as an independent task and combined with speech recognition and language modeling to form a multi-task learning framework. The model not only needs to recognize polyphonetic characters but also predict their correct pronunciation and contextual semantics.
[0057] Innovation:
[0058] A new multi-task learning architecture is proposed. Through shared representation learning, the model can optimize multiple tasks simultaneously and improve the accuracy of polyphonetic character recognition.
[0059] #Sample code: Multi-task learning model
[0060] from tensorflow.keras.layers import Input,LSTM,Dense,Concatenate
[0061] from tensorflow.keras.models import Model
[0062] #Define input layer
[0063] input_text=Input(shape=(max_length,))
[0064] input_audio=Input(shape=(audio_features_length,))
[0065] #Text Encoder
[0066] encoded_text=LSTM(128)(input_text)
[0067] #Speech feature encoder
[0068] encoded_audio=LSTM(128)(input_audio)
[0069] #Merge the encoded features
[0070] combined_features=Concatenate()([encoded_text,encoded_audio])
[0071] #Polyphonetic character recognition output
[0072] multi_phone_output=Dense(num_phones,activation='softmax')(combined_features)
[0073] #Speech recognition output
[0074] speech_output=Dense(num_speech_units,activation='softmax')(combined_features)
[0075] #Create model
[0076] model=Model(inputs=[input_text,input_audio],outputs=[multi_phone_output,speech_output])
[0077] 2. Context-sensitive polyphonetic word prediction model
[0078] Substantive technical solutions:
[0079] Attention mechanism: The attention mechanism is introduced into the model, enabling the model to focus on key contextual information related to polyphones in the text.
[0080] Innovation:
[0081] By using the attention mechanism, the model can pay more attention to the context related to polyphones, thereby improving the accuracy of polyphone recognition.
[0082] #Sample code: Attention mechanism in LSTM 2. Improved speech-to-text accuracy:
[0083] 1. Contextual inference: Use contextual information to infer the correct pronunciation of polyphones. This can be done using rule-based methods or machine learning models.
[0084] 2. Model training: Create a training set containing the speech and text conversion results of polyphonetic characters to train the system's ability to recognize polyphones.
[0085] 3. System Optimization:
[0086] 1. Corpus construction: Build a corpus containing a large number of polyphonetic word contexts for system training and optimization.
[0087] 2. User feedback mechanism: Design a user feedback mechanism to allow users to correct system recognition errors to continuously improve the accuracy of the system.
[0088] 4. Real-time testing and adjustment:
[0089] 1. System testing: Conduct system testing in actual applications to evaluate the system's performance in the context of polyphones.
[0090] 2. Parameter adjustment: Adjust and optimize system parameters based on test results to improve system accuracy and stability.
[0091] Therefore, the present invention constructs a polyphonetic character recognition model and a polyphonetic character prediction model through multi-task learning, thereby improving the accuracy of mutual conversion between speech and text, saving time for manual proofreading, and improving user experience.
[0092] Exemplary devices
[0093] Figure 2 FIG is a structural diagram of a text-to-speech conversion device based on polyphone context recognition provided by an exemplary embodiment of the present invention. Figure 2 As shown, the apparatus 200 includes:
[0094] A construction module 210 is used to construct a polyphonetic character recognition model and a polyphonetic character prediction model using multi-task learning;
[0095] A recognition module 220 is configured to use a polyphone recognition model to recognize polyphone information in a task to be converted, wherein the task to be converted is speech to be converted or text to be converted;
[0096] Prediction module 230, configured to use a polyphone prediction model to predict the pronunciation or font of the polyphone according to the polyphone information and the context of the task to be converted, and determine the pronunciation or font of each polyphone in the polyphone information;
[0097] The conversion module 240 is used to identify the polyphonetic characters in the task to be converted according to the pronunciation or font of each polyphonetic character, and complete the conversion of the task to be converted.
[0098] Optionally, the building block 210 includes:
[0099] The first construction submodule is used to construct a polyphonic word dictionary including polyphonic words and their different pronunciations;
[0100] The second construction submodule is used to perform machine model recognition based on the polyphone dictionary and build a polyphone recognition model;
[0101] The third construction submodule is used to train a machine model based on pre-constructed context-based polyphone training data to build a polyphone prediction model, wherein the machine model is a machine model based on an attention mechanism.
[0102] Optionally, the second building block includes:
[0103] A first acquiring unit is configured to perform feature encoding on text in the polyphone dictionary using a text encoder to acquire text features;
[0104] A second acquiring unit is configured to perform feature encoding on the speech in the polyphonetic word dictionary through a speech feature encoder to acquire speech features;
[0105] A merging unit, configured to merge text features and speech features to obtain merged features;
[0106] An output unit, configured to perform polyphonetic character recognition and speech recognition based on the combined features, and output the polyphonetic character recognition results and the speech recognition results;
[0107] The construction unit is used to construct a polyphone recognition model based on the polyphone recognition results and the speech recognition results.
[0108] Exemplary electronic devices
[0109] Figure 3 This is the structure of an electronic device provided by an exemplary embodiment of the present invention. Figure 3 As shown, the electronic device 30 includes one or more processors 31 and a memory 32 .
[0110] The processor 31 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0111] The memory 32 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may, for example, include read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 31 may execute the program instructions to implement the methods of the software programs of the various embodiments of the present invention described above and / or other desired functions. In one example, the electronic device may further include: an input device 33 and an output device 34, which are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0112] In addition, the input device 33 may also include, for example, a keyboard, a mouse, and the like.
[0113] The output device 34 can output various information to the outside. The output device 34 can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto.
[0114] Of course, to simplify, Figure 3 Only some of the components related to the present invention in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device may further include any other appropriate components according to specific application scenarios.
[0115] Exemplary computer program products and computer-readable storage media
[0116] In addition to the above-mentioned methods and devices, an embodiment of the present invention may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to perform the steps of the method according to various embodiments of the present invention described in the above "Exemplary Method" section of this specification.
[0117] The computer program product may be written in any combination of one or more programming languages to implement the operations of embodiments of the present invention, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0118] In addition, an embodiment of the present invention may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enables the processor to execute the steps of the method for information mining of historical change records according to various embodiments of the present invention described in the above "Exemplary Method" section of this specification.
[0119] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, system or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0120] The basic principles of the present invention have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in the present invention are merely illustrative and non-limiting, and should not be construed as necessarily possessed by each embodiment of the present invention. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, and are not intended to be limiting. These details do not necessarily limit the present invention to being implemented using these specific details.
[0121] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments are sufficient. For system embodiments, since they largely correspond to method embodiments, their description is relatively simple. For relevant parts, references to the description of the method embodiments are sufficient.
[0122] The block diagrams of the devices, systems, equipment, and systems involved in the present invention are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, systems, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "including," "comprising," "having," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0123] The method and system of the present invention may be implemented in many ways. For example, the method and system of the present invention may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above sequence of steps for the method is for illustration only, and the steps of the method of the present invention are not limited to the sequence specifically described above, unless otherwise specified. In addition, in some embodiments, the present invention may also be implemented as a program recorded in a recording medium, which includes machine-readable instructions for implementing the method according to the present invention. Thus, the present invention also covers recording media that store programs for executing the method according to the present invention.
[0124] It should also be noted that, in the system, device and method of the present invention, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent schemes of the present invention. The above description of the disclosed aspects is provided to enable any technician in this field to make or use the present invention. Various modifications to these aspects will be very obvious to those skilled in the art, and the general principles defined here can be applied to other aspects without departing from the scope of the present invention. Therefore, the present invention is not intended to be limited to the aspects shown here, but according to the widest scope consistent with the principles disclosed here and novel features.
[0125] The above description has been provided for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present invention to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A text-to-speech conversion method based on polyphone context recognition, characterized in that: include: Use multi-task learning to build polyphone recognition and prediction models; Using the polyphone recognition model to identify polyphone information in a task to be converted, wherein the task to be converted is speech to be converted or text to be converted; Using a polyphone prediction model, predicting the pronunciation or font of the polyphone according to the polyphone information and the context of the task to be converted, and determining the pronunciation or font of each polyphone in the polyphone information; The polyphonetic characters in the task to be converted are identified according to the pronunciation or font of each polyphonetic character, and the conversion of the task to be converted is completed.
2. The method according to claim 1, characterized in that Multi-task learning is used to build polyphone recognition and prediction models, including: Construct a polyphonetic dictionary that includes polyphonetic characters and their different pronunciations; Perform machine model recognition based on the polyphonetic word dictionary to construct the polyphonetic word recognition model; A machine model is trained based on pre-constructed context-based polyphone training data to construct the polyphone prediction model, wherein the machine model is a machine model based on an attention mechanism.
3. The method according to claim 2, characterized in that Performing machine model recognition based on the polyphonetic word dictionary to construct the polyphonetic word recognition model includes: Performing feature encoding on the text in the polyphonetic word dictionary by a text encoder to obtain text features; Performing feature encoding on the speech in the polyphonetic word dictionary by a speech feature encoder to obtain speech features; Merging the text features and the speech features to obtain a merged feature; Performing polyphonetic character recognition and speech recognition according to the combined features, and outputting polyphonetic character recognition results and speech recognition results; The polyphonetic character recognition model is constructed according to the polyphonetic character recognition result and the speech recognition result.
4. A text-to-speech conversion device based on polyphone context recognition, characterized in that: include: A construction module for constructing a polyphone recognition model and a polyphone prediction model using multi-task learning; a recognition module, configured to recognize polyphone information in a task to be converted using the polyphone recognition model, wherein the task to be converted is speech to be converted or text to be converted; A prediction module, configured to use a polyphone prediction model to predict the pronunciation or font of the polyphone according to the polyphone information and the context of the task to be converted, and determine the pronunciation or font of each polyphone in the polyphone information; The conversion module is used to identify the polyphones in the task to be converted according to the pronunciation or font of each polyphone, and complete the conversion of the task to be converted.
5. The device according to claim 4, characterized in that Building blocks, including: The first construction submodule is used to construct a polyphonic word dictionary including polyphonic words and their different pronunciations; A second construction submodule is configured to perform machine model recognition based on the polyphonetic word dictionary to construct the polyphonetic word recognition model; The third construction submodule is used to perform machine model training based on pre-constructed context-based polyphone training data to construct the polyphone prediction model, wherein the machine model is a machine model based on the attention mechanism.
6. The device according to claim 5, characterized in that The second building block includes: A first acquiring unit is configured to perform feature encoding on the text in the polyphonetic word dictionary through a text encoder to acquire text features; A second acquiring unit is configured to perform feature encoding on the speech in the polyphonetic word dictionary through a speech feature encoder to acquire speech features; a merging unit, configured to merge the text feature and the speech feature to obtain a merged feature; an output unit, configured to perform polyphonetic character recognition and speech recognition based on the combined features, and output the polyphonetic character recognition result and the speech recognition result; A construction unit is used to construct the polyphonetic character recognition model according to the polyphonetic character recognition result and the speech recognition result.
7. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1 to 3.
8. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of claims 1 to 3.