Phoneme prediction method and device, electronic equipment, computer readable storage medium and computer program product

By constructing a syntax tree and performing graph convolution processing during text conversion process, and fusing text and syntax tree features, the problem of insufficient phoneme prediction accuracy in the prior art is solved, and higher phoneme prediction accuracy is achieved.

CN120388557APending Publication Date: 2025-07-29MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410132986.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the process of converting phonemes to text, the prior art focuses only on the character dimension features of the text and ignores syntax tree information, resulting in insufficient accuracy of phoneme prediction.

Method used

By obtaining text features and constructing a syntax tree, performing graph convolution processing, fusing syntax tree features and text features, obtaining a fusion encoding to determine phonemes.

Benefits of technology

The accuracy of phoneme prediction is improved, and the context relevance is enhanced by combining text and syntax tree information, and the accuracy of phoneme prediction is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388557A_ABST
    Figure CN120388557A_ABST
Patent Text Reader

Abstract

The invention provides a phoneme prediction method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: obtaining a to-be-processed text, performing text coding processing on the to-be-processed text to obtain text features, constructing syntactic tree information of the to-be-processed text based on the to-be-processed text, performing graph convolution processing on the syntactic tree information to obtain syntactic tree features, performing fusion processing on the syntactic tree features and the text features to obtain fusion codes, and obtaining the fusion codes. And determining phonemes of the to-be-processed text based on the fusion code. According to the invention, the accuracy of the predicted phonemes can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to computer technology, and in particular, to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for phoneme prediction. Background Art

[0002] With the development of computer technology, there are more and more scenarios where text needs to be converted into phonemes. For example, in daily chat scenarios, shopping scenarios, and news broadcasts, etc., the technology of text-to-phoneme is required. In the related art, the way of text-to-phoneme is mainly to directly extract text features from the text, input the text features into a prediction model, and the prediction model outputs the predicted phonemes, thereby realizing the function of text-to-phoneme. Summary of the Invention

[0003] Embodiments of this application provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for phoneme prediction, which can improve the accuracy of the predicted phonemes.

[0004] The technical solution of the embodiments of this application is implemented as follows:

[0005] Embodiments of this application provide a method for phoneme prediction, including:

[0006] Obtain the text to be processed and perform text encoding processing on the text to be processed to obtain text features;

[0007] Construct syntactic tree information of the text to be processed based on the text to be processed, and perform graph convolution processing on the syntactic tree information to obtain syntactic tree features;

[0008] Perform fusion processing on the syntactic tree features and the text features to obtain a fusion code, and determine the phonemes of the text to be processed based on the fusion code.

[0009] Embodiments of this application provide a method for speech synthesis, including:

[0010] Obtain the text to be processed, and perform the method for phoneme prediction provided by the embodiments of this application on the text to be processed to obtain the phonemes of the text to be processed;

[0011] Output speech matching the text to be processed based on the phonemes of the text to be processed.

[0012] Embodiments of this application provide a device for phoneme prediction, including:

[0013] A text encoding module that obtains the text to be processed and performs text encoding processing on the text to be processed to obtain text features;

[0014] A graph convolutional module, which is used to construct syntactic tree information of the text to be processed based on the text to be processed, and perform graph convolutional processing on the syntactic tree information to obtain syntactic tree features;

[0015] A fusion processing module, which is used to perform fusion processing on the syntactic tree features and the text features to obtain a fusion encoding, and determine phonemes of the text to be processed based on the fusion encoding.

[0016] The above-mentioned graph convolutional module is further used to construct graph structure information of the text to be processed based on the syntactic tree information; perform encoding processing on the graph structure information to obtain a syntactic tree encoding; call a graph neural network to perform graph convolutional processing on the syntactic tree encoding to obtain the syntactic tree features.

[0017] The above-mentioned graph convolutional module is further used to extract N words and syntactic relationships between the N words from the syntactic tree information, where N is an integer greater than or equal to 1; use the N words as word nodes and the syntactic relationships as associated edges to construct graph structure information of the text to be processed.

[0018] The above-mentioned graph convolutional module is further used to perform word encoding processing on word nodes in the graph structure information to obtain word node encodings; perform edge encoding processing on associated edges in the graph structure information to obtain associated edge encodings; form a two-dimensional matrix with the word node encodings and the associated edge encodings, and use the two-dimensional matrix as the syntactic tree encoding.

[0019] The above-mentioned graph convolutional module is further used to perform sentence division processing on the text to be processed to obtain multiple sentences, and determine the sentence ranking of the sentences in the text to be processed; determine row position information of the word node encodings and associated edge encodings corresponding to the sentences in the two-dimensional matrix based on the sentence ranking; determine column position information of the word node encodings and the associated edge encodings in the two-dimensional matrix based on the logical order of the word node encodings and the associated edge encodings corresponding to the sentences in the sentences.

[0020] The above-mentioned graph convolutional module is further used to perform a sliding window process on the two-dimensional matrix through a convolutional kernel of the graph neural network to obtain multiple sliding window matrices of the same size, where the size of the sliding window matrix is the same as the size of the convolutional kernel; perform graph convolutional processing on the encodings in each sliding window matrix through the convolutional kernel to obtain a convolutional value of each sliding window matrix, where the encodings in the sliding window matrix include at least one of the word node encodings and the associated edge encodings; form the convolutional values of the multiple sliding window matrices into the syntactic tree features based on the position information of the sliding window matrix in the two-dimensional matrix.

[0021] The above-mentioned text encoding module is further configured to extract the letters of the words in the text to be processed; based on the letter encoding table, map the letters of the words to letter encodings, and arrange the letter encodings corresponding to the letters according to the position information of the letters in the words to obtain the word encoding of the words; perform convolutional processing on the word encoding to obtain the text feature.

[0022] The above-mentioned fusion processing module is further configured to decode the fusion encoding to obtain the digital sequence corresponding to the fusion encoding; based on the phoneme mapping table, map the digital sequence to a phoneme sequence, and use the phoneme sequence as the phoneme of the text to be processed.

[0023] The above-mentioned fusion processing module is further configured to perform mapping processing on the digital sequence based on the phoneme mapping table to obtain the original phoneme sequence; retrieve the invalid phonemes of the letters of the words in the text to be processed from the control mapping table; if the original phoneme sequence contains the invalid phonemes, remove the invalid phonemes in the original phoneme sequence to obtain the phoneme sequence.

[0024] An embodiment of the present application provides a voice synthesis device, including:

[0025] A text acquisition module, configured to acquire the text to be processed and execute the phoneme prediction method provided by the embodiment of the present application on the text to be processed to obtain the phoneme of the text to be processed;

[0026] A voice output module, configured to output voice matching the text to be processed based on the phoneme of the text to be processed.

[0027] An embodiment of the present application provides an electronic device, including:

[0028] A memory, configured to store computer-executable instructions or computer programs;

[0029] A processor, configured to implement the phoneme prediction method provided by the embodiment of the present application or implement the voice synthesis method provided by the embodiment of the present application when executing the computer-executable instructions or computer programs stored in the memory.

[0030] An embodiment of the present application provides a computer-readable storage medium, storing computer-executable instructions or computer programs, which are used to cause a processor to implement the phoneme prediction method provided by the embodiment of the present application or implement the voice synthesis method provided by the embodiment of the present application when executed.

[0031] An embodiment of the present application provides a computer program product, including computer-executable instructions, which implement the phoneme prediction method provided by the embodiment of the present application or implement the voice synthesis method provided by the embodiment of the present application when executed by a processor.

[0032] The embodiments of the present application have the following beneficial effects:

[0033] Obtain the text to be processed, construct a syntax tree corresponding to the text to be processed based on the text to be processed, perform graph convolution processing on the syntax tree to obtain syntax tree features, then perform text encoding processing on the text to be processed to obtain text features, fuse the syntax tree features and text features to obtain fused coding, and determine the phonemes of the text to be processed based on the fused coding. It can be seen that the phonemes of the text to be processed are obtained based on the text features and syntax tree features of the text to be processed. Therefore, in the process of predicting the phonemes of the text to be processed, both text features and syntax tree features are taken into account. Since the syntax tree includes the syntactic relations between words in the text to be processed, the corresponding syntax tree features also retain the syntactic relation information between words, that is, the syntax tree features contain information that represents the syntactic relations between words, so that the fused coding obtained by fusing the syntax tree features and the text features can include more information of the text to be processed, thereby improving the accuracy of the phonemes of the text to be processed obtained based on the fused coding. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 Schematic diagram of the architecture of a phoneme prediction system provided in an embodiment of the present application;

[0035] Figure 2A is a structural diagram of an electronic device provided in an embodiment of the present application;

[0036] Figure 2B is a structural diagram of another electronic device provided in an embodiment of the present application;

[0037] Figure 3 1 is a flow chart of a method for phoneme prediction provided in an embodiment of the present application;

[0038] Figure 4 is a schematic diagram of a syntax tree provided in an embodiment of the present application;

[0039] Figure 5 Schematic diagram of sliding window convolution provided in an embodiment of the present application;

[0040] Figure 6 Schematic diagram of the flow of the speech synthesis method provided in the embodiment of the present application;

[0041] Figure 7 This is a flowchart of phoneme prediction provided by an embodiment of the present application;

[0042] Figure 8 This is a flowchart of the graph convolution provided in an embodiment of the present application. DETAILED DESCRIPTION

[0043] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in detail in conjunction with the accompanying drawings. The described embodiments should not be regarded as limitations of this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0044] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0046] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are subject to the following explanations.

[0047] 1) Syntax tree: A tree diagram representing the results of syntactic analysis. It illustrates the structure, hierarchy, and functional relationships of various linguistic components in a sentence. It can be divided into binary trees and multi-way trees.

[0048] 2) Graph convolution: Graph convolution is a convolutional neural network for image and graph data. Different from traditional convolutional neural networks, graph convolution is defined on graph structures and can process irregular graph data. Its basic idea is to generalize the convolution operation to graphs and extract features by performing convolution operations on nodes and edges.

[0049] With the development of computer technology, there are more and more scenarios where text needs to be converted into phonemes. Currently, the main method for text-to-phoneme conversion is to directly input the text into a prediction model. The prediction model extracts the text features of the text and predicts the phonemes corresponding to the text based on the text features. Although this method can predict the phonemes corresponding to the text, since this method can only obtain the character-dimensional features corresponding to each character in the text, it only focuses on the information represented by each word included in the text during the prediction process and ignores the syntax tree information included in the text, thereby reducing the accuracy of the predicted phonemes.

[0050] In the embodiments of this application, on the basis of extracting the text features corresponding to the text, the syntax tree features of the text are further fused, which increases the context relevance between texts, thereby improving the accuracy of the predicted phonemes.

[0051] See Figure 1 , Figure 1 This is a schematic architecture diagram of the phoneme prediction system 100 provided by an embodiment of the present application. To support a text-to-phoneme application, the server 200 is connected to the terminal 400 via the network 300. The network 300 can be a wide area network, a local area network, or a combination of both.

[0052] The server 200 is configured to obtain the text transmitted by the terminal 400 via the network 300 as the text to be processed, perform text encoding processing on the text to be processed to obtain text features, construct syntactic tree information of the text to be processed based on the text to be processed, perform graph convolution processing on the syntactic tree information to obtain syntactic tree features, perform fusion processing on the syntactic tree features and the text features to obtain a fusion encoding, and determine the phonemes of the text to be processed based on the fusion encoding, and transmit the obtained phonemes to the terminal 400 via the network 300, so that the terminal 400 can convert the text into phonemes according to the obtained phonemes.

[0053] The terminal 400 is configured to obtain the text input by the user, transmit the text input by the user to the server 200 via the network 300, receive the phonemes returned by the server 200 via the network 300, and convert the text into speech based on the received phonemes.

[0054] It should be noted that the phoneme prediction method provided by the embodiments of the present application can be applied to various scenarios, including but not limited to scenarios such as online chatting, shopping consumption, and news reporting.

[0055] In some other embodiments, the embodiments of the present application can be implemented with the aid of cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing.

[0056] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form a resource pool, be used on demand, and be flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources.

[0057] In some embodiments, the server 200 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication means, which is not limited in the embodiments of the present application.

[0058] The following describes an exemplary application of the electronic device provided in the embodiments of the present application for implementing phoneme prediction. Next, an exemplary application will be described when the electronic device for implementing the phoneme prediction method is implemented as a server.

[0059] See Figure 2A , Figure 2A is a schematic structural diagram of the electronic device provided in the embodiments of the present application. Figure 2A The electronic device 500A shown may be Figure 1 the server 200A or the terminal 400A in Figure 2A The electronic device 500A shown includes: at least one processor 410A, a memory 430A, and at least one network interface 420A. Each component in the electronic device 500A is coupled together through a bus system 440A. It can be understood that the bus system 440A is used to realize the connection and communication between these components. In addition to the data bus, the bus system 440A also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 2A all kinds of buses are labeled as the bus system 440A.

[0060] The processor 410A may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a Digital Signal Processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.

[0061] The user interface 430A includes one or more output devices 431A that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430A also includes one or more input devices 432A, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, and other input buttons and controls.

[0062] The memory 450A can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid state memory, hard disk drives, optical disk drives, etc. The memory 450A optionally includes one or more storage devices that are physically remote from the processor 410A.

[0063] The memory 450A includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 450A described in the embodiments of the present application is intended to include any suitable type of memory.

[0064] In some embodiments, the memory 450A is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.

[0065] The operating system 451A includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0066] The network communication module 452A is used to reach other computing devices via one or more (wired or wireless) network interfaces 420A. Exemplary network interfaces 420A include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.;

[0067] The presentation module 453A is used to enable the presentation of information (such as a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431A associated with the user interface 430A (such as a display screen, speaker, etc.);

[0068] The input processing module 454A is used to detect and translate one or more user inputs or interactions from one of the one or more input devices 432A.

[0069] In some embodiments, the apparatus provided by the embodiments of the present application may be implemented in software. Figure 2A Shown is a phoneme prediction apparatus 455A stored in a memory 450A, which may be software in the form of programs and plugins, etc., including the following software modules: a text encoding module 4551, a graph convolution module 4552, and a fusion processing module 4553. These modules are logical, so they can be combined arbitrarily or further split according to the functions implemented.

[0070] See Figure 2B , Figure 2B which is another schematic structural diagram of the electronic device provided by the embodiments of the present application. Figure 2B The electronic device 500B in further includes a voice synthesis apparatus 455B, which may be software in the form of programs and plugins, etc., including the following software modules: a text acquisition module 4554 and a voice output module 4555. Figure 2B The processor 410B, network interface 420B, user interface 430B, output device 431B, input device 432B, bus system 440B, memory 450B, operating system 451B, network communication module 452B, presentation module 453B, and input processing module 454B included in all have the same structure and the same function as the corresponding modules included in . The functions of each module will be described below. Figure 2A

[0071] In other embodiments, the apparatus provided by the embodiments of the present application may be implemented in hardware. As an example, the apparatus provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the phoneme prediction method provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.

[0072] The phoneme prediction method provided by the embodiments of the present application will be described in combination with the exemplary applications and implementations of the electronic device provided by the embodiments of the present application.

[0073] See Figure 3 , Figure 3 which is a schematic flowchart of the phoneme prediction method provided by the embodiments of the present application. It will be described in combination with Figure 3 ​​The steps shown will be described. The method for phoneme prediction provided by the embodiments of this application can be implemented independently by a server or a terminal, or jointly implemented by the server and the terminal. Below, the case where the server implements it independently will be taken as an example for description.

[0074] In step 101, the text to be processed is obtained and text encoding processing is performed on the text to be processed to obtain text features.

[0075] In practical applications, the server can obtain the text to be processed uploaded by the terminal, or obtain the text to be processed from the historical database. The way to obtain the text to be processed can be selected according to the actual situation, and no specific limitation is made here.

[0076] Exemplarily, the user can upload the text to be processed from the terminal, and the terminal sends the text to be processed to the server through the network so that the server can obtain the text to be processed. Or, the server can connect to the database and obtain the text to be processed from the database. The type of the text to be processed can be news text, chat conversation text, novel fragments, etc., and no specific limitation is made on the type of the text to be processed here.

[0077] In practical applications, after obtaining the text to be processed, the server can perform encoding processing on the text to be processed, and then obtain text features. The following introduces the detailed process of text encoding for the text to be processed.

[0078] In some embodiments, "performing text encoding processing on the text to be processed to obtain text features" in step 101 above can be implemented in the following manner:

[0079] The server can extract the letters of the words in the text to be processed, map the letters of the words to letter encodings based on the letter encoding table, and arrange the letter encodings corresponding to the letters according to the position information of the letters in the words to obtain the word encodings of the words, and perform convolutional processing on the word encodings to obtain text features.

[0080] As an example, the server can determine the letter composition of the word, and map the word to the corresponding encoding according to the pre-constructed letter mapping table. For example, for the word "word", in the corresponding letter mapping table, the encoding corresponding to "w" is "22", the encoding corresponding to "o" is "14", the encoding corresponding to "r" is "17", and the encoding corresponding to "d" is "3", so the encoding corresponding to this word can be obtained as [22, 14, 17, 3].

[0081] In step 102, based on the text to be processed, the syntactic tree information of the text to be processed is constructed, and graph convolutional processing is performed on the syntactic tree information to obtain syntactic tree features.

[0082] In practical applications, after obtaining the text to be processed, the server can construct the syntactic tree information corresponding to the text to be processed. In the embodiments of the present application, the X-bar algorithm can be used in the process of constructing the syntactic tree information. Specifically, the following steps can be adopted to construct the syntactic tree:

[0083] 1. Determine the lexical items of the sentence and label them as N (noun), V (verb), P (preposition), etc.

[0084] 2. Combine the lexical items into smaller phrases, such as NP (noun phrase) and VP (verb phrase). These phrases can be single lexical items or other smaller phrases.

[0085] 3. Combine the smaller phrases into larger phrases, such as S (sentence) and PP (prepositional phrase). These phrases can be single phrases or other smaller phrases.

[0086] 4. Use the X-bar rules to determine the structure of each phrase. The X-bar rules specify the central lexical item of each phrase (for example, the central lexical item of NP is a noun), as well as other possible modifiers (such as adjectives or adverbs).

[0087] 5. Use the X-bar rules to determine the parent node of each phrase. The parent node is a larger phrase that contains the current phrase and other phrases.

[0088] 6. Repeat steps 3-5 until the entire sentence is decomposed into smaller phrases and each phrase has a parent node.

[0089] The constructed syntactic tree can participate in Figure 4 , Figure 4 which is a schematic diagram of the syntactic tree information provided by the embodiments of the present application.

[0090] In Figure 4 ,S represents a sentence, NP represents a noun phrase, VP represents a verb phrase, and Det represents a determiner.

[0091] From Figure 4 it can be seen that "this" is a determiner used to modify the noun "tree", "is" is the verb of "tree", "illustrating" is also the verb of "tree", and "the" is the determiner of the nouns "constituency" and "relation". According to the nature and correlation relationship between each word, the final syntactic tree structure is obtained.

[0092] After the server obtains the syntactic tree information of the text to be processed, it can determine the syntactic tree features based on the syntactic tree information. The following introduces the specific process of determining the syntactic tree features provided by the embodiments of the present application.

[0093] In some embodiments, the process of "performing graph convolution processing on the syntactic tree information to obtain syntactic tree features" in step 102 above can be implemented through the following process:

[0094] The server can construct graph structure information of the text to be processed based on the syntactic tree information, perform encoding processing on the graph structure information to obtain syntactic tree encoding, and call a graph neural network to perform graph convolution processing on the syntactic tree encoding to obtain syntactic tree features.

[0095] As an example, the server can first construct the graph structure information of the text to be processed, and then encode the graph structure information, and perform graph convolution processing on the syntactic tree encoding to obtain syntactic tree features.

[0096] The following introduces the detailed process of the server constructing the graph structure information.

[0097] In some embodiments, the process of "constructing graph structure information of the text to be processed based on the syntactic tree information" above can be implemented in the following manner:

[0098] The server can extract N words and the syntactic relationships between the N words from the syntactic tree information, where N is an integer greater than or equal to 1, and use the N words as word nodes and the syntactic relationships as associated edges to construct the graph structure of the text to be processed. That is, the extracted words are used as nodes in the graph structure information, and the nodes with syntactic relationships are connected by associated edges, thereby obtaining the graph structure of the syntactic tree composed of word nodes and associated edges.

[0099] As an example, the words included in the syntactic tree information and the syntactic relationships between the words are extracted from the syntactic tree information. Here, the syntactic relationship can be the dependency relationship between words. The dependency relationship is a binary symmetric relationship between a core word and its dependents. The core word is usually a verb (it can also be a noun in a sentence without a verb). Then the dependency relationship here can be the subject-predicate relationship between a verb and its subject, or the object-verb relationship between a verb and its object, etc. For example, the words extracted from the syntactic tree are "I" and "go out shopping". Since in the syntactic tree, "go out shopping" is used to describe the action of "I", there is a subject-predicate relationship between "I" and "go out shopping".

[0100] As an example, the word "abundant" is used as node A, and the word "dinner" is used as node B. Since there is an associated relationship between the word "abundant" and the word "dinner", node A and node B are connected to form the graph structure of the syntactic tree. The above example is only a simple process of constructing the graph structure information of the syntactic tree. In actual applications, the graph structure information of the syntactic tree can include multiple nodes, and no further examples will be given here.

[0101] After obtaining the graph structure information of the syntactic tree, the server can encode the graph structure information to obtain the syntactic tree encoding. The following introduces the detailed process of encoding the graph structure information.

[0102] In some embodiments, the process of "encoding the graph structure information to obtain the syntactic tree encoding" can be implemented in the following manner:

[0103] The server can perform word encoding processing on the word nodes in the graph structure information to obtain word node encodings, perform edge encoding processing on the associated edges in the graph structure information to obtain associated edge encodings, form a two-dimensional matrix with the word node encodings and the associated edge encodings, and use the two-dimensional matrix as the syntactic tree encoding.

[0104] As an example, the server can first encode the word nodes in the graph structure information. The specific encoding process can be to encode according to the words corresponding to the word nodes. The following introduces a specific encoding method provided by the embodiments of the present application.

[0105] In practical applications, the server can determine the letter composition of the word, and map the word to the corresponding encoding according to the pre-constructed letter mapping table. For example, if the word is "word", in the corresponding letter mapping table, the encoding corresponding to "w" is "22", the encoding corresponding to "o" is "14", the encoding corresponding to "r" is "17", and the encoding corresponding to "d" is "3". Therefore, the encoding corresponding to this word can be obtained as [22, 14, 17, 3]. It should be noted that the above is only a specific encoding method provided by the embodiments of the present application. In practical applications, a suitable encoding method can be selected according to the actual situation.

[0106] After the server encodes the word nodes, it can also perform edge encoding processing on the associated edges in the graph structure information to obtain associated edge encodings (that is, the process of encoding the associated edges). The following introduces a specific method for encoding the associated edges provided by the embodiments of the present application.

[0107] In practical applications, the server can pre-construct an associated edge mapping table, which records the encodings corresponding to each association relationship. For example, the encoding corresponding to the subject-predicate relationship is "1", and the encoding corresponding to the modification relationship is "2". The server can obtain the encoding of the associated edge according to the association relationship between the word nodes by referring to the associated edge mapping table.

[0108] As an example, the encoding corresponding to the subject-predicate relationship recorded in the associated edge mapping table is "1". The word node "mood" corresponds to node A, and the word node "comfortable" corresponds to node B. Since there is a subject-predicate relationship between "mood" and "comfortable", the associated edge between node A and node B can be encoded as "1".

[0109] It should be noted that the above is only a specific method for encoding associated edges provided by the embodiments of the present application. In actual applications, the method for encoding associated edges can be selected according to actual situations.

[0110] After the server encodes the word nodes and associated edges, it can form a two-dimensional matrix with the word node encoding and the associated edge encoding, and use the two-dimensional matrix as the syntactic tree encoding. That is, a two-dimensional matrix can be constructed based on the node encoding and the associated edge encoding.

[0111] The following introduces the detailed process of the server constructing the two-dimensional matrix.

[0112] In some embodiments, the process of "forming a two-dimensional matrix with the word node encoding and the associated edge encoding" can be implemented in the following manner:

[0113] Perform sentence division processing on the text to be processed to obtain multiple sentences, and determine the sentence ranking of the sentences in the text to be processed. Based on the sentence ranking, determine the row position information of the word node encoding and the associated edge encoding corresponding to the sentence in the two-dimensional matrix. Based on the logical order of the word node encoding and the associated edge encoding corresponding to the sentence in the sentence, determine the column position information of the word node encoding and the associated edge encoding in the two-dimensional matrix.

[0114] As an example, the server can perform sentence division processing on the text to be processed to obtain multiple sentences, and determine the sentence ranking of the sentences in the text to be processed. That is, divide the text to be processed into multiple sentences and rank the sentences.

[0115] As an example, the text to be processed is "You got up really early today. Do you want to rest for a while longer or go wash your face directly?" The server can divide the text to be processed into sentence A "You got up really early today", sentence B "Do you want to rest for a while longer", and sentence C "Or go wash your face directly" according to the punctuation marks in the text to be processed. At the same time, according to the content of the text to be processed, it can be known that sentence A is before sentence B, and sentence B is before sentence C. Therefore, the sentence rankings of each sentence in the text to be processed are sentence A, sentence B, and sentence C.

[0116] As an example, the server can determine the row position information of the word node encoding and the associated edge encoding corresponding to the sentence in the two-dimensional matrix based on the sentence ranking. That is, the positions of the word node encoding and the associated edge encoding can be determined according to the positions of the sentences to which the word node encoding and the associated edge encoding belong in the ranking.

[0117] As an example, the server can determine that the word node codes and associated edge codes belonging to the same statement are located in the same row of the two-dimensional matrix, and determine the row number to which each statement belongs according to the statement ranking. For example, if the statement rankings are Statement A, Statement B, and Statement C, then the word node codes and associated edge codes belonging to Statement A can be located in the first row of the two-dimensional matrix, the word node codes and associated edge codes belonging to Statement B can be located in the second row of the two-dimensional matrix, and the word node codes and associated edge codes belonging to Statement C can be located in the third row of the two-dimensional matrix.

[0118] It should be noted that the above example is an example arranged in order according to the statement ranking. In actual applications, it is also possible to use a reverse order arrangement according to the statement ranking for sorting. For example, if the statement rankings are Statement D, Statement E, and Statement F, then the word node codes and associated edge codes belonging to Statement F can be located in the first row of the two-dimensional matrix, the word node codes and associated edge codes belonging to Statement E can be located in the second row of the two-dimensional matrix, and the word node codes and associated edge codes belonging to Statement D can be located in the third row of the two-dimensional matrix.

[0119] In actual applications, the sorting method according to the statement ranking can be selected according to the actual situation and is not specifically limited here.

[0120] Next, the sorting method for each word node code and associated edge code included in each row is introduced. The server can determine the column position information of the word node code and the associated edge code in the two-dimensional matrix based on the logical order of the word node code corresponding to the statement and the associated edge code in the statement, that is, the server can determine the column position information of the word node code and the associated edge code according to the logical order of the word code and the relationship code in each statement. Furthermore, based on the row position information and column position information of the word node code, and the row position information and column position information of the associated edge code, the two-dimensional matrix of the word node code and the associated edge code is determined.

[0121] Exemplarily, the text to be processed is "The sun is bright and the starlight is dim", and the word corresponding to node encoding A is "sun", the word corresponding to word node encoding B is "bright", the word corresponding to word node encoding C is "starlight", and the word corresponding to word node encoding D is "dim". According to the subject-predicate relationship between the words, it can be known that there is a subject-predicate relationship between the word "sun" and the word "bright", and there is a subject-predicate relationship between the word "starlight" and the word "dim". Then there is an associated edge encoding E between the word node encoding A and the word node encoding B, and there is an associated encoding F between the word node encoding C and the word node encoding D. Also, since the word node encoding A, the word node encoding B, and the associated edge encoding E belong to the same sentence "The sun is bright", and the logical order of the word node encoding A is before the word node encoding B (because the sentence is "The sun is bright" and "sun" is before "bright"), therefore, the column position of the word node encoding A is before the column position of the word node encoding B. Similarly, the column position of the word node encoding C is before the column position of the word node encoding D. And for the column position of the associated edge encoding E, it can be after the word node encoding B, and the column position of the associated edge encoding F can be after the word node encoding D. The specific arrangement of the two-dimensional matrix can be seen in Table 1 below.

[0122] Word node encoding A Word node encoding B Associated edge encoding E Word node code C Word node encoding D Associated edge code F

[0123] Table 1

[0124] It should be noted that for the position of the associated edge encoding E, based on the connection relationship, the position of the associated edge encoding E can be determined to be between the word node encoding A and the word node encoding B. For the position of the associated edge encoding F, based on the connection relationship, the position of the associated edge encoding F can be determined to be between the word node encoding C and the word node encoding D. As shown in Table 2 below.

[0125] Word node encoding A Associated edge encoding E Word node encoding B Word node code C Associated edge code F Word node encoding D

[0126] Table 2

[0127] In practical applications, the positions of the associated edge encoding and the word node encoding can be selected according to the actual situation, and no specific limitation is made here.

[0128] After obtaining the syntactic tree encoding, the server can perform graph convolution processing on the syntactic tree encoding, and then obtain the corresponding syntactic tree features. The following introduces the specific process of the server performing graph convolution on the syntactic tree encoding.

[0129] In some embodiments, the process of "invoking a graph neural network to perform graph convolution processing on the syntactic tree encoding to obtain syntactic tree features" can be implemented in the following manner:

[0130] The server can perform a sliding window process on a two-dimensional matrix through the convolutional kernel of the graph neural network to obtain multiple sliding window matrices of the same size. The size of the sliding window matrix is the same as the size of the convolutional kernel. The encoding in each sliding window matrix is processed by graph convolution through the convolutional kernel to obtain the convolution value of each sliding window matrix. The encoding in the sliding window matrix includes at least one of the word node encoding and the associated edge encoding. Based on the position information of the sliding window matrix in the two-dimensional matrix, the convolution values of multiple sliding window matrices are combined into syntactic tree features.

[0131] The following combines Figure 5 to illustrate the specific graph convolution process. Figure 5 is a schematic diagram of sliding window convolution provided by an embodiment of the present application.

[0132] In Figure 5 among them, Figure 5 in (1) is a 4×4 two-dimensional matrix composed of data A to data P. Figure 5 in (2) is a convolutional kernel composed of numerical values 1 to 4. Figure 5 in (3) represents the process of the first sliding window convolution based on the convolutional kernel. Among them, the one circled by the dotted circle is the first sliding window matrix. The calculation method of the convolution value of this sliding window matrix is as shown in Figure 5 the multiplication of the two matrices shown in (4) in it. The obtained result is the convolution value obtained after the convolution process of this sliding window matrix. After calculating the convolution value of this sliding window matrix, the next sliding window matrix can be confirmed. As shown in Figure 5 in (5), the one circled by the dotted circle is the second sliding window matrix, and the convolution value of this sliding window matrix is calculated. Repeat the above steps until the convolution values of all sliding window matrices included in the two-dimensional matrix are determined. Then, according to the position of each sliding window matrix in the two-dimensional matrix, the convolution values of each sliding window matrix are combined into syntactic tree features.

[0133] As an example, the server can perform the following processing for each position in the sliding window matrix:

[0134] Obtain the convolutional kernel weight corresponding to the position from the convolutional kernel, obtain the node weight of the word node encoding corresponding to the position and the associated edge weight of the associated edge encoding corresponding to the position. Multiply the convolutional kernel weight by the node weight as the comprehensive weight of the word node encoding, and multiply the convolutional kernel weight by the associated edge weight as the comprehensive weight of the associated edge encoding. Based on the comprehensive weight, perform a fusion process on the word node encoding and the associated edge encoding at the position to obtain the convolution value at the position. Sum the convolution values of multiple positions in the sliding window matrix to obtain the convolution value of the sliding window matrix.

[0135] As an example, the server can set weights for node encoding and weights for associated edge encoding in the convolutional kernel. It should be noted that the weights for node encoding and the weights for associated edge encoding can be the same or different, and the specific situation can be set according to the actual situation, which is not specifically limited here.

[0136] Specifically, the convolution process can be performed with reference to the following formula.

[0137]

[0138] In formula (1), A represents the convolution value, m represents the number of rows of the two-dimensional matrix, and n represents the number of columns of the two-dimensional matrix. Among them, β ij represents the weight of the associated edge encoding in the i-th row and j-th column, w ij represents the value corresponding to the convolutional kernel in the i-th row and j-th column, x ij represents the value of the associated edge encoding in the i-th row and j-th column, α ij represents the weight of the word node encoding in the i-th row and j-th column, y ij represents the value of the word node encoding in the i-th row and j-th column.

[0139] As an example, according to the above formula (1), it is possible to perform a sliding window convolution on the entire two-dimensional matrix through a convolutional kernel with weights, and then obtain the convolution values at each position.

[0140] As an example, when the position covered by the convolutional kernel is the node encoding, the associated edge encoding (x ij ) at this position can be 0. Similarly, when the position covered by the convolutional kernel is the associated edge encoding, the node encoding (y ij ) at this position can be 0.

[0141] In step 103, the syntactic tree features and the text features are fused to obtain a fused encoding, and the phonemes of the text to be processed are determined based on the fused encoding.

[0142] In practical applications, after obtaining the syntactic tree features and the text features, the server can fuse the syntactic tree features and the text features to obtain a fused encoding.

[0143] In practical applications, the process of fusing the syntactic tree features and the text features can be to splice the syntactic tree features and the text features into a vector, thereby realizing the fusion of the syntactic tree features and the text features. For example, if the syntactic tree feature is A and the text feature is B, then the fused encoding is [A, B].

[0144] After obtaining the fused encoding, the server can determine the phonemes of the text to be processed based on the fused encoding. The process of determining the phonemes of the text to be processed based on the fused encoding is introduced in detail below.

[0145] In some embodiments, "determining the phonemes of the text to be processed based on the fusion encoding" in step 103 above can be implemented in the following manner:

[0146] The server can decode the fusion encoding to obtain the digital sequence corresponding to the fusion encoding, map the digital sequence to a phoneme sequence based on the phoneme mapping table, and use the phoneme sequence as the phonemes of the text to be processed.

[0147] As an example, the server can decode the fusion encoding to obtain the digital sequence corresponding to the fusion encoding. After that, it can refer to the phoneme mapping table and convert the digital sequence into a phoneme sequence according to the corresponding relationship between the phonemes and digits recorded in the phoneme mapping table.

[0148] In practical applications, the corresponding relationship between the digital sequence and the phoneme sequence may not be a one-to-one correspondence. That is, one phoneme position in the phoneme sequence may correspond to multiple digits in the digital sequence. For example, the phoneme sequence is "AA0", "A0", "B0", "BB0", "AB0", where the digits in the digital sequence corresponding to the phoneme "AA0" may include the three digits 7, 5, and 4. In the phoneme mapping table, one digit corresponds to one phoneme. Therefore, one phoneme position in the phoneme sequence may correspond to multiple phonemes. At this time, the server can screen the multiple phonemes and use the screened phonemes as the phonemes at this phoneme position, thereby obtaining the phoneme sequence. The following introduces the detailed screening process.

[0149] In some embodiments, "mapping the digital sequence to a phoneme sequence based on the phoneme mapping table" can be implemented in the following manner:

[0150] Perform mapping processing on the digital sequence based on the phoneme mapping table to obtain the original phoneme sequence. Based on the comparison mapping table, retrieve the invalid phonemes of the letters of the words in the text to be processed from the comparison mapping table. If the original phoneme sequence contains invalid phonemes, remove the invalid phonemes from the original phoneme sequence to obtain the phoneme sequence.

[0151] As an example, the server can determine the invalid phonemes included in the phonemes corresponding to the digital sequence according to the comparison mapping table and remove the invalid phonemes to obtain the final phoneme sequence. It should be noted that the comparison mapping table records the possible phonemes corresponding to each letter. For example, for the letter a, the corresponding phonemes are ["AA0", "A0"], and there are 5 phonemes corresponding to the letter a in the digital sequence, which are ["AA0", "A0", "B0", "BB0", "AB0"]. It can be seen that the three phonemes "B0", "BB0", and "AB0" in the phoneme sequence need to be removed.

[0152] The following combines with Figure 6Describe the method for speech synthesis provided by the embodiments of the present application. Figure 6 It is a schematic flowchart of the method for speech synthesis provided by the embodiments of the present application.

[0153] In step 201, obtain the text to be processed, perform text encoding processing on the text to be processed to obtain text features, construct syntactic tree information of the text to be processed based on the text to be processed, perform graph convolution processing on the syntactic tree information to obtain syntactic tree features, perform fusion processing on the syntactic tree features and the text features to obtain a fusion encoding, and determine the phonemes of the text to be processed based on the fusion encoding.

[0154] As an example, the server can first obtain the text to be processed. The way to obtain the text to be processed can be that the server obtains the text to be processed uploaded by the terminal, or the server obtains the text to be processed from the database. The way to obtain the text to be processed can be selected according to the actual situation and will not be specifically limited here.

[0155] As an example, after obtaining the text to be processed, the server can execute the method for phoneme prediction provided by the embodiments of the present application on the text to be processed, and then obtain the phonemes of the text to be processed. Specifically, the text to be processed can be subjected to text encoding processing to obtain text features, and then the syntactic tree information of the text to be processed is constructed based on the text to be processed, and graph convolution processing is performed on the syntactic tree information to obtain syntactic tree features. Then, fusion processing is performed on the syntactic tree features and the text features to obtain a fusion encoding. Finally, the phonemes of the text to be processed are determined based on the fusion encoding. Among them, the process of performing text encoding processing on the text to be processed, the process of constructing the syntactic tree information of the text to be processed, the process of performing graph convolution processing on the syntactic tree information, the process of performing fusion processing on the syntactic tree features and the text features, and the process of finally determining the phonemes of the text to be processed based on the fusion encoding are all the same as the implementation manners of the corresponding processes in the phoneme prediction method provided by the embodiments of the present application.

[0156] In step 202, output speech that matches the text to be processed based on the phonemes of the text to be processed.

[0157] As an example, the server can output speech that matches the text to be processed based on the phonemes of the text to be processed.

[0158] As an example, the server can input the phonemes of the text to be processed into a pre-trained speech synthesis model. It should be noted that the purpose of the speech synthesis model is to synthesize the input phonemes into speech. Therefore, the training method and model structure of the speech synthesis model will not be specifically limited, as long as it can realize synthesizing phonemes into speech.

[0159] Next, introduce the application process of the method for phoneme prediction provided by the embodiments of the present application in an actual scenario.

[0160] See Figure 7 , Figure 7 which is a flowchart for predicting phonemes provided by an embodiment of the present application.

[0161] In practical applications, the phoneme prediction method provided by the embodiments of the present application can predict phonemes corresponding to languages such as English and Chinese. The following takes the phoneme prediction of English as an example for illustration.

[0162] In step 301, the server preprocesses the English text information.

[0163] The front end of the text-to-speech (TTS) technology is an important module in TTS. The result obtained by the TTS front end will directly affect the overall output performance of the subsequent model. As an important part of the TTS front end, the English grapheme-to-phoneme (g2p) is an important module in the front end and has an important impact on the subsequent Chinese-English mixed synthesis result. If the English g2p prediction is inaccurate, it will directly lead to a poor final synthesis effect and it is difficult to accurately pronounce English words.

[0164] First, the server can preprocess and construct the English text information (text to be processed) with various tags. Since the seq2seq algorithm is finally used to predict the phoneme sequence, and because the seq2seq algorithm recognizes the flag <sos>Start predicting when the flag <eos>Stop predicting afterwards, so the preprocessing process can add a start flag at the beginning of the English sentence. <sos>Add the ending mark at the end <eos>, and then encode according to the obtained sequence by putting it into the mapping table.

[0165] After that, use the X-bar algorithm to construct a syntactic tree (syntactic tree information) for the obtained word sequence, and use the GCN model to extract the obtained syntactic tree information. Since GCN is mainly an algorithm for images, using it on the syntactic tree can better extract the structural information of the syntactic tree, so as to obtain the final syntactic information to assist the final model training and prediction synthesis.

[0166] The process of constructing the syntactic tree is introduced in detail below.

[0167] In step 302, the server constructs a syntactic tree.

[0168] First, the server can regard the relationship between each word as the edge of an image, and at the same time regard each word as the node of the image. The encoding method for each node (word node) of the image can be to encode according to the encoding corresponding to each letter in the alphabet. For example, for the word "word", it can be encoded according to 26 letters, and the encoding of this word is [22, 14, 17, 3]. Based on the above encoding method, the encoding corresponding to each word is obtained in turn. At the same time, each corresponding relationship is also replaced by an encoding. Specifically, it can be encoded according to the syntactic relationship corresponding between words. For example, for the subject-predicate relationship, the relationship between words forms an encoding table, and a mapping table can be formed by multiple encoding tables, and then the final encoding result is obtained through the mapping of the table.

[0169] In step 303, the server performs graph convolution on the syntactic tree.

[0170] Before the server performs graph convolution on the syntactic tree, the server can use the graph neural network (GCN) to extract the syntactic tree information. First, encode the word in the i-th node of the syntactic tree as y ij , where the subscript j of the word encoding can be the j-th sentence of the text to be processed, and y ij can represent the i-th node of the j-th sentence in the text to be processed, reflecting two-dimensional information instead of one-dimensional sequence. If there is only one sentence in the text to be processed, then the value of j is equal to 1.

[0171] During the process of the server performing graph convolution on the syntactic tree, the relationship between every two nodes included in the syntactic tree is encoded as x ij (associated edge encoding). According to the subscript of the relationship encoding, it can be known that the relationship encoding x ij is the i-th relationship encoding in the j-th sentence included in the text to be processed.

[0172] The convolution kernel used in the convolution process can be a convolution kernel with weight tendency. This convolution kernel can be a two-dimensional convolution kernel, and the data at the m-th row and n-th column of this convolution kernel can be w mn , where the weight assigned to the corresponding word encoding (word node encoding) is α ij , and the weight assigned to the corresponding edge relationship (associated edge encoding) is β ij . It should be noted that the weight can be set to an initial value at the beginning. For example, the weight for word encoding is 0.3, and the attention weight for associated edge encoding is 0.7. The weight can be changed through the learning of the model.

[0173] During the convolution process, each word encoding and the encoding information of the relationship between words (associated edge encoding) are multiplied and added as the extracted information. That is, the convolution result obtained each time the convolution kernel sweeps through the corresponding position is (The meaning of each parameter in the formula can be referred to the above formula (1)). According to this formula, the scanning value of the convolution kernel corresponding to the syntactic tree feature can be obtained. And because of the use of different weight tendencies, the convolution kernel can obtain "attention" during the training process and can focus on the information that needs to be concerned. The purpose of using a two-dimensional convolution kernel and a two-dimensional matrix is mainly to be able to focus on the information between multiple sentences while also obtaining the front-back relationship in the same sentence, and to be able to extract and abstract the information of the feature extraction layer across multiple sentences, and then further obtain the final graph convolution output tensor information, and finally obtain the tensor of the next output.

[0174] It can be seen in Figure 8 , Figure 8 which is the schematic flow chart of the graph convolution provided by the embodiment of the present application.

[0175] In Figure 8 , X1, X2, and X3 are syntactic tree structures. The syntactic tree structures are input into the hidden layer of the graph convolution network to obtain the convolution results Z1, Z2, and Z3 of the graph convolution, and the association relationships Y1 and Y2 between Z1, Z2, and Z3.

[0176] In the above graph convolution neural network, syntactic information can be extracted at the two-dimensional level, so as to obtain further syntactic information, and the structure of the whole sentence is further analyzed, greatly improving the subsequent input information dimension and information richness, which is helpful for more effective assistance in the subsequent classification task.

[0177] In step 304, the server encodes the English text information.

[0178] The server can split the input sequence (a string sequence of input English text information, which can be broken down into different sentences based on punctuation, etc.) to form a sequence group consisting of multiple sentences, and then encode the sequence group to obtain a two-dimensional matrix. The specific encoding method can be the same as the above-mentioned encoding of the syntax tree nodes in the syntax tree to obtain a two-dimensional matrix as the encoding of the English text information.

[0179] After obtaining the encoding of the English text information, the convolution encoder can be used to obtain the front-to-back order relationship between the sequences of the input sequence matrix and the parallel relationship between the sequences. It should be noted that the convolution method provided by the present application is different from the related art in which a sequence is directly used to encode the input text for information extraction. Since the convolution is performed on a two-dimensional matrix, the front-to-back order relationship and the parallel relationship of the text can be obtained. At the same time, the use of a convolutional network is not only possible to extract information from the encoding information of the English word itself, but also to extract information from an existing syntax tree. The constructed syntax information tree can be deeply extracted by the two networks, so that it can be better added to the final text tensor information (fusion features) to form a deep reflection of the text information.

[0180] The above-mentioned method of encoding word nodes can be used to encode English text information to obtain text features corresponding to the English text information.

[0181] In step 305, the server performs convolution processing on the text features.

[0182] The server can perform convolution processing on the obtained code to obtain the coding features corresponding to the code.

[0183] In step 306, the server fuses the encoding features and the syntax tree features.

[0184] The server can concatenate the encoding features and the syntax tree features. For example, if the encoding feature is A and the syntax tree feature is B, the final fusion code is [A, B].

[0185] In step 307, the server decodes the fusion code.

[0186] The phoneme prediction process provided in the embodiment of the present application can adopt an encoder-decoder structure. Steps 301 to 306 are the encoding process. After the server obtains the result of the encoder (i.e., the fused code), it can send the fused code to the decoder for decoding. Then, the digital sequence corresponding to the fused code is obtained.

[0187] In step 308, the server performs mapping through a phoneme mapping table.

[0188] The server can map the digital sequence to the corresponding phoneme sequence through a pre-built phoneme mapping table.

[0189] In step 309, the server updates the model parameters.

[0190] The phoneme prediction method provided by the embodiments of the present application can be implemented through a model. During the training process of the model, the text to be processed can be a training sample with labels. After obtaining the phoneme sequence, a loss function can be constructed through the label values and the phoneme sequence, thereby realizing the training of the model. Specifically, the Adam optimizer can be used to optimize the CTC loss function, perform backpropagation on the entire network, and optimize the parameters in the network until the final model loss function converges completely. Thus, an optimized model can be obtained, which can be used for the inference of the multi-classification model.

[0191] In the actual application process, the server can directly use the trained model, input the text to be processed into the model, and obtain the predicted phonemes output by the model.

[0192] After obtaining the phoneme sequence, the server can eliminate the invalid phonemes in the phoneme sequence by referring to the mapping table.

[0193] Specifically, for example, similar to the letter a, the corresponding phoneme pronunciations are ["AA0", "A0"], and then all the phonemes in the phoneme sequence are 5 types, ["AA0", "A0", "B0", "BB0", "AB0"], then the mapping encoding matrix corresponding to the letter a is [1, 1, 0, 0, 0]. At this time, this mapping encoding matrix can achieve the effect of masking the other pronunciations.

[0194] The above is the mapping encoding matrix for the letter a. On this basis, adding the mapping encoding intervals of each letter included in each word in the text to be processed can achieve the effect of masking the impossible pronunciations, making the result more accurate and tending to finally obtain a result closer to the real English phonemes.

[0195] Through the above embodiments provided by the present application, by using the English syntactic tree and using a cross-domain graph convolutional neural network to extract the syntactic information of the English sequence, more information can be obtained than using only BLSTM.

[0196] In the English g2p scenario, the conv-seq2seq algorithm in the translation field is used, taking advantage of the relatively powerful information and translation-related capabilities of the convolutional network to perform targeted optimization for the case where the English g2p input and output are of variable length, and thus better results can be obtained, and the sentence error rate can be better in terms of performance and effect than the original old technology model.

[0197] Next, the exemplary structure of the phoneme prediction device 455A provided in the embodiments of the present application implemented as software modules will be further described. In some embodiments, as Figure 2A shown, the software modules stored in the phoneme prediction device 455A in the memory 450A may include:

[0198] A text encoding module 4551, which acquires the text to be processed and performs text encoding processing on the text to be processed to obtain text features;

[0199] A graph convolution module 4552, which is used to construct syntactic tree information of the text to be processed based on the text to be processed, and perform graph convolution processing on the syntactic tree information to obtain syntactic tree features;

[0200] A fusion processing module 4553, which is used to perform fusion processing on the syntactic tree features and the text features to obtain a fusion encoding, and determine the phonemes of the text to be processed based on the fusion encoding.

[0201] In some embodiments, the graph convolution module 4552 is further used to construct graph structure information of the text to be processed based on the syntactic tree information; perform encoding processing on the graph structure information to obtain a syntactic tree encoding; and call a graph neural network to perform graph convolution processing on the syntactic tree encoding to obtain the syntactic tree features.

[0202] In some embodiments, the graph convolution module 4552 is further used to extract N words and the syntactic relationships between the N words from the syntactic tree information, where N is an integer greater than or equal to 1; and use the N words as word nodes and the syntactic relationships as associated edges to construct the graph structure information of the text to be processed.

[0203] In some embodiments, the graph convolution module 4552 is further used to perform word encoding processing on the word nodes in the graph structure information to obtain word node encodings; perform edge encoding processing on the associated edges in the graph structure information to obtain associated edge encodings; form a two-dimensional matrix with the word node encodings and the associated edge encodings, and use the two-dimensional matrix as the syntactic tree encoding.

[0204] In some embodiments, the graph convolution module 4552 is further used to perform sentence division processing on the text to be processed to obtain multiple sentences, and determine the sentence rankings of the sentences in the text to be processed; determine the row position information of the word node encodings and the associated edge encodings corresponding to the sentences in the two-dimensional matrix based on the sentence rankings; and determine the column position information of the word node encodings and the associated edge encodings in the two-dimensional matrix based on the logical order of the word node encodings and the associated edge encodings corresponding to the sentences in the sentences.

[0205] In some embodiments, the graph convolution module 4552 is further configured to perform a sliding window process on the two-dimensional matrix through the convolution kernel of the graph neural network to obtain a plurality of sliding window matrices with the same size, where the size of the sliding window matrix is the same as the size of the convolution kernel; perform graph convolution processing on the encoding in each sliding window matrix through the convolution kernel to obtain the convolution value of each sliding window matrix, where the encoding in the sliding window matrix includes at least one of the word node encoding and the associated edge encoding; and form the syntactic tree feature from the convolution values of the plurality of sliding window matrices based on the position information of the sliding window matrix in the two-dimensional matrix.

[0206] In some embodiments, the text encoding module 4551 is further configured to extract the letters of the words in the text to be processed; map the letters of the words to letter encodings based on the letter encoding table, and arrange the letter encodings corresponding to the letters according to the position information of the letters in the words to obtain the word encoding of the words; and perform convolution processing on the word encoding to obtain the text feature.

[0207] In some embodiments, the fusion processing module 4553 is further configured to decode the fusion encoding to obtain a digital sequence corresponding to the fusion encoding; map the digital sequence to a phoneme sequence based on the phoneme mapping table, and use the phoneme sequence as the phoneme of the text to be processed.

[0208] In some embodiments, the fusion processing module 4553 is further configured to perform a mapping process on the digital sequence based on the phoneme mapping table to obtain an original phoneme sequence; retrieve the invalid phonemes of the letters of the words in the text to be processed from the control mapping table; and if the original phoneme sequence contains the invalid phonemes, remove the invalid phonemes from the original phoneme sequence to obtain the phoneme sequence.

[0209] In some embodiments, as Figure 2B shown, the software module stored in the speech synthesis device 455B in the memory 440B may include:

[0210] A text acquisition module 4554, configured to acquire the text to be processed and perform the phoneme prediction method provided in the embodiments of the present application on the text to be processed to obtain the phoneme of the text to be processed;

[0211] A voice output module 4555, configured to output voice matching the text to be processed based on the phoneme of the text to be processed.

[0212] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the phoneme prediction method or the speech synthesis method described above in the embodiments of the present application.

[0213] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, where the executable instructions are stored. When the executable instructions are executed by a processor, the processor will be caused to execute the phoneme prediction method provided by the embodiments of the present application, or execute the speech synthesis method of the embodiments of the present application. For example, Figure 3 the phoneme prediction method shown, or Figure 6 the speech synthesis method shown.

[0214] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0215] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, and may be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0216] As an example, the executable instructions may or may not correspond to a file in a file system, and may be stored as part of a file that stores other programs or data. For example, they may be stored in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files storing one or more modules, subroutines, or code portions).

[0217] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or on multiple computing devices distributed at multiple locations and interconnected by a communication network.

[0218] In summary, the following beneficial effects can be achieved through the embodiments of the present application:

[0219] Obtain the text to be processed, construct a syntax tree corresponding to the text to be processed based on the text to be processed, perform graph convolution processing on the syntax tree, obtain syntax tree features, then perform text encoding processing on the text to be processed to obtain text features, fuse the syntax tree features and text features to obtain fusion coding, and determine the phonemes of the text to be processed based on the fusion coding. It can be seen that the phonemes of the text to be processed are obtained based on the text features and syntax tree features of the text to be processed. Therefore, in the process of predicting the phonemes of the text to be processed, on the basis of retaining the text features of the text to be processed itself, the syntax tree features of the text to be processed are further fused. Since the syntax tree includes the words in the text to be processed and the syntactic relations between the words, the corresponding syntax tree features also retain the word features and the syntactic relation features between the words in the text to be processed. That is, the syntax tree features have features of different dimensions from the text features, so that the fusion coding includes more information of the text to be processed, thereby improving the accuracy of the phonemes of the text to be processed obtained based on the fusion coding.

[0220] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.< / eos> < / sos> < / eos> is recognized <eos>Stop predicting afterwards, so the preprocessing process can add a start flag at the beginning of the English sentence. <sos>Add the ending mark at the end <eos>, and then encode according to the obtained sequence by putting it into the mapping table.

[0165] After that, use the X-bar algorithm to construct a syntactic tree (syntactic tree information) for the obtained word sequence, and use the GCN model to extract the obtained syntactic tree information. Since GCN is mainly an algorithm for images, using it on the syntactic tree can better extract the structural information of the syntactic tree, so as to obtain the final syntactic information to assist the final model training and prediction synthesis.

[0166] The process of constructing the syntactic tree is introduced in detail below.

[0167] In step 302, the server constructs a syntactic tree.

[0168] First, the server can regard the relationship between each word as the edge of an image, and at the same time regard each word as the node of the image. The encoding method for each node (word node) of the image can be to encode according to the encoding corresponding to each letter in the alphabet. For example, for the word "word", it can be encoded according to 26 letters, and the encoding of this word is [22, 14, 17, 3]. Based on the above encoding method, the encoding corresponding to each word is obtained in turn. At the same time, each corresponding relationship is also replaced by an encoding. Specifically, it can be encoded according to the syntactic relationship corresponding between words. For example, for the subject-predicate relationship, the relationship between words forms an encoding table, and a mapping table can be formed by multiple encoding tables, and then the final encoding result is obtained through the mapping of the table.

[0169] In step 303, the server performs graph convolution on the syntactic tree.

[0170] Before the server performs graph convolution on the syntactic tree, the server can use the graph neural network (GCN) to extract the syntactic tree information. First, encode the word in the i-th node of the syntactic tree as y ij , where the subscript j of the word encoding can be the j-th sentence of the text to be processed, and y ij can represent the i-th node of the j-th sentence in the text to be processed, reflecting two-dimensional information instead of one-dimensional sequence. If there is only one sentence in the text to be processed, then the value of j is equal to 1.

[0171] During the process of the server performing graph convolution on the syntactic tree, the relationship between every two nodes included in the syntactic tree is encoded as x ij (associated edge encoding). According to the subscript of the relationship encoding, it can be known that the relationship encoding x ij is the i-th relationship encoding in the j-th sentence included in the text to be processed.

[0172] The convolution kernel used in the convolution process can be a convolution kernel with weight tendency. This convolution kernel can be a two-dimensional convolution kernel, and the data at the m-th row and n-th column of this convolution kernel can be w mn , where the weight assigned to the corresponding word encoding (word node encoding) is α ij , and the weight assigned to the corresponding edge relationship (associated edge encoding) is β ij . It should be noted that the weight can be set to an initial value at the beginning. For example, the weight for word encoding is 0.3, and the attention weight for associated edge encoding is 0.7. The weight can be changed through the learning of the model.

[0173] During the convolution process, each word encoding and the encoding information of the relationship between words (associated edge encoding) are multiplied and added as the extracted information. That is, the convolution result obtained each time the convolution kernel sweeps through the corresponding position is (The meaning of each parameter in the formula can be referred to the above formula (1)). According to this formula, the scanning value of the convolution kernel corresponding to the syntactic tree feature can be obtained. And because of the use of different weight tendencies, the convolution kernel can obtain "attention" during the training process and can focus on the information that needs to be concerned. The purpose of using a two-dimensional convolution kernel and a two-dimensional matrix is mainly to be able to focus on the information between multiple sentences while also obtaining the front-back relationship in the same sentence, and to be able to extract and abstract the information of the feature extraction layer across multiple sentences, and then further obtain the final graph convolution output tensor information, and finally obtain the tensor of the next output.

[0174] It can be seen in Figure 8 , Figure 8 which is the schematic flow chart of the graph convolution provided by the embodiment of the present application.

[0175] In Figure 8 , X1, X2, and X3 are syntactic tree structures. The syntactic tree structures are input into the hidden layer of the graph convolution network to obtain the convolution results Z1, Z2, and Z3 of the graph convolution, and the association relationships Y1 and Y2 between Z1, Z2, and Z3.

[0176] In the above graph convolution neural network, syntactic information can be extracted at the two-dimensional level, so as to obtain further syntactic information, and the structure of the whole sentence is further analyzed, greatly improving the subsequent input information dimension and information richness, which is helpful for more effective assistance in the subsequent classification task.

[0177] In step 304, the server encodes the English text information.

[0178] The server can split the input sequence (a string sequence of input English text information, which can be broken down into different sentences based on punctuation, etc.) to form a sequence group consisting of multiple sentences, and then encode the sequence group to obtain a two-dimensional matrix. The specific encoding method can be the same as the above-mentioned encoding of the syntax tree nodes in the syntax tree to obtain a two-dimensional matrix as the encoding of the English text information.

[0179] After obtaining the encoding of the English text information, the convolution encoder can be used to obtain the front-to-back order relationship between the sequences of the input sequence matrix and the parallel relationship between the sequences. It should be noted that the convolution method provided by the present application is different from the related art in which a sequence is directly used to encode the input text for information extraction. Since the convolution is performed on a two-dimensional matrix, the front-to-back order relationship and the parallel relationship of the text can be obtained. At the same time, the use of a convolutional network is not only possible to extract information from the encoding information of the English word itself, but also to extract information from an existing syntax tree. The constructed syntax information tree can be deeply extracted by the two networks, so that it can be better added to the final text tensor information (fusion features) to form a deep reflection of the text information.

[0180] The above-mentioned method of encoding word nodes can be used to encode English text information to obtain text features corresponding to the English text information.

[0181] In step 305, the server performs convolution processing on the text features.

[0182] The server can perform convolution processing on the obtained code to obtain the coding features corresponding to the code.

[0183] In step 306, the server fuses the encoding features and the syntax tree features.

[0184] The server can concatenate the encoding features and the syntax tree features. For example, if the encoding feature is A and the syntax tree feature is B, the final fusion code is [A, B].

[0185] In step 307, the server decodes the fusion code.

[0186] The phoneme prediction process provided in the embodiment of the present application can adopt an encoder-decoder structure. Steps 301 to 306 are the encoding process. After the server obtains the result of the encoder (i.e., the fused code), it can send the fused code to the decoder for decoding. Then, the digital sequence corresponding to the fused code is obtained.

[0187] In step 308, the server performs mapping through a phoneme mapping table.

[0188] The server can map the digital sequence to the corresponding phoneme sequence through a pre-built phoneme mapping table.

[0189] In step 309, the server updates the model parameters.

[0190] The phoneme prediction method provided by the embodiments of the present application can be implemented through a model. During the training process of the model, the text to be processed can be a training sample with labels. After obtaining the phoneme sequence, a loss function can be constructed through the label values and the phoneme sequence, thereby realizing the training of the model. Specifically, the Adam optimizer can be used to optimize the CTC loss function, perform backpropagation on the entire network, and optimize the parameters in the network until the final model loss function converges completely. Thus, an optimized model can be obtained, which can be used for the inference of the multi-classification model.

[0191] In the actual application process, the server can directly use the trained model, input the text to be processed into the model, and obtain the predicted phonemes output by the model.

[0192] After obtaining the phoneme sequence, the server can eliminate the invalid phonemes in the phoneme sequence by referring to the mapping table.

[0193] Specifically, for example, similar to the letter a, the corresponding phoneme pronunciations are ["AA0", "A0"], and then all the phonemes in the phoneme sequence are 5 types, ["AA0", "A0", "B0", "BB0", "AB0"], then the mapping encoding matrix corresponding to the letter a is [1, 1, 0, 0, 0]. At this time, this mapping encoding matrix can achieve the effect of masking the other pronunciations.

[0194] The above is the mapping encoding matrix for the letter a. On this basis, adding the mapping encoding intervals of each letter included in each word in the text to be processed can achieve the effect of masking the impossible pronunciations, making the result more accurate and tending to finally obtain a result closer to the real English phonemes.

[0195] Through the above embodiments provided by the present application, by using the English syntactic tree and using a cross-domain graph convolutional neural network to extract the syntactic information of the English sequence, more information can be obtained than using only BLSTM.

[0196] In the English g2p scenario, the conv-seq2seq algorithm in the translation field is used, taking advantage of the relatively powerful information and translation-related capabilities of the convolutional network to perform targeted optimization for the case where the English g2p input and output are of variable length, and thus better results can be obtained, and the sentence error rate can be better in terms of performance and effect than the original old technology model.

[0197] Next, the exemplary structure of the phoneme prediction device 455A provided in the embodiments of the present application implemented as software modules will be further described. In some embodiments, as Figure 2A shown, the software modules stored in the phoneme prediction device 455A in the memory 450A may include:

[0198] A text encoding module 4551, which acquires the text to be processed and performs text encoding processing on the text to be processed to obtain text features;

[0199] A graph convolution module 4552, which is used to construct syntactic tree information of the text to be processed based on the text to be processed, and perform graph convolution processing on the syntactic tree information to obtain syntactic tree features;

[0200] A fusion processing module 4553, which is used to perform fusion processing on the syntactic tree features and the text features to obtain a fusion encoding, and determine the phonemes of the text to be processed based on the fusion encoding.

[0201] In some embodiments, the graph convolution module 4552 is further used to construct graph structure information of the text to be processed based on the syntactic tree information; perform encoding processing on the graph structure information to obtain a syntactic tree encoding; and call a graph neural network to perform graph convolution processing on the syntactic tree encoding to obtain the syntactic tree features.

[0202] In some embodiments, the graph convolution module 4552 is further used to extract N words and the syntactic relationships between the N words from the syntactic tree information, where N is an integer greater than or equal to 1; and use the N words as word nodes and the syntactic relationships as associated edges to construct the graph structure information of the text to be processed.

[0203] In some embodiments, the graph convolution module 4552 is further used to perform word encoding processing on the word nodes in the graph structure information to obtain word node encodings; perform edge encoding processing on the associated edges in the graph structure information to obtain associated edge encodings; form a two-dimensional matrix with the word node encodings and the associated edge encodings, and use the two-dimensional matrix as the syntactic tree encoding.

[0204] In some embodiments, the graph convolution module 4552 is further used to perform sentence division processing on the text to be processed to obtain multiple sentences, and determine the sentence rankings of the sentences in the text to be processed; determine the row position information of the word node encodings and the associated edge encodings corresponding to the sentences in the two-dimensional matrix based on the sentence rankings; and determine the column position information of the word node encodings and the associated edge encodings in the two-dimensional matrix based on the logical order of the word node encodings and the associated edge encodings corresponding to the sentences in the sentences.

[0205] In some embodiments, the graph convolution module 4552 is further configured to perform a sliding window process on the two-dimensional matrix through the convolution kernel of the graph neural network to obtain a plurality of sliding window matrices with the same size, where the size of the sliding window matrix is the same as the size of the convolution kernel; perform graph convolution processing on the encoding in each sliding window matrix through the convolution kernel to obtain the convolution value of each sliding window matrix, where the encoding in the sliding window matrix includes at least one of the word node encoding and the associated edge encoding; and form the syntactic tree feature from the convolution values of the plurality of sliding window matrices based on the position information of the sliding window matrix in the two-dimensional matrix.

[0206] In some embodiments, the text encoding module 4551 is further configured to extract the letters of the words in the text to be processed; map the letters of the words to letter encodings based on the letter encoding table, and arrange the letter encodings corresponding to the letters according to the position information of the letters in the words to obtain the word encoding of the words; and perform convolution processing on the word encoding to obtain the text feature.

[0207] In some embodiments, the fusion processing module 4553 is further configured to decode the fusion encoding to obtain a digital sequence corresponding to the fusion encoding; map the digital sequence to a phoneme sequence based on the phoneme mapping table, and use the phoneme sequence as the phoneme of the text to be processed.

[0208] In some embodiments, the fusion processing module 4553 is further configured to perform a mapping process on the digital sequence based on the phoneme mapping table to obtain an original phoneme sequence; retrieve the invalid phonemes of the letters of the words in the text to be processed from the control mapping table; and if the original phoneme sequence contains the invalid phonemes, remove the invalid phonemes from the original phoneme sequence to obtain the phoneme sequence.

[0209] In some embodiments, as Figure 2B shown, the software module stored in the speech synthesis device 455B in the memory 440B may include:

[0210] A text acquisition module 4554, configured to acquire the text to be processed and perform the phoneme prediction method provided in the embodiments of the present application on the text to be processed to obtain the phoneme of the text to be processed;

[0211] A voice output module 4555, configured to output voice matching the text to be processed based on the phoneme of the text to be processed.

[0212] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the phoneme prediction method or the speech synthesis method described above in the embodiments of the present application.

[0213] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, where the executable instructions are stored. When the executable instructions are executed by a processor, the processor will be caused to execute the phoneme prediction method provided by the embodiments of the present application, or execute the speech synthesis method of the embodiments of the present application. For example, Figure 3 the phoneme prediction method shown, or Figure 6 the speech synthesis method shown.

[0214] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0215] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, and may be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0216] As an example, the executable instructions may or may not correspond to a file in a file system, and may be stored as part of a file that stores other programs or data. For example, they may be stored in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files storing one or more modules, subroutines, or code portions).

[0217] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or on multiple computing devices distributed at multiple locations and interconnected by a communication network.

[0218] In summary, the following beneficial effects can be achieved through the embodiments of the present application:

[0219] Obtain the text to be processed, construct a syntax tree corresponding to the text to be processed based on the text to be processed, perform graph convolution processing on the syntax tree, obtain syntax tree features, then perform text encoding processing on the text to be processed to obtain text features, fuse the syntax tree features and text features to obtain fusion coding, and determine the phonemes of the text to be processed based on the fusion coding. It can be seen that the phonemes of the text to be processed are obtained based on the text features and syntax tree features of the text to be processed. Therefore, in the process of predicting the phonemes of the text to be processed, on the basis of retaining the text features of the text to be processed itself, the syntax tree features of the text to be processed are further fused. Since the syntax tree includes the words in the text to be processed and the syntactic relations between the words, the corresponding syntax tree features also retain the word features and the syntactic relation features between the words in the text to be processed. That is, the syntax tree features have features of different dimensions from the text features, so that the fusion coding includes more information of the text to be processed, thereby improving the accuracy of the phonemes of the text to be processed obtained based on the fusion coding.

[0220] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.< / eos> < / sos> < / eos> < / sos>

Claims

1. A method for phoneme prediction, characterized in that The method includes: Obtaining the text to be processed and performing text encoding processing on the text to be processed to obtain text features; Constructing syntactic tree information of the text to be processed based on the text to be processed, and performing graph convolution processing on the syntactic tree information to obtain syntactic tree features; Performing fusion processing on the syntactic tree features and the text features to obtain a fusion encoding, and determining the phonemes of the text to be processed based on the fusion encoding.

2. The method according to claim 1, wherein The performing graph convolution processing on the syntactic tree information to obtain syntactic tree features includes: Constructing graph structure information of the text to be processed based on the syntactic tree information; Performing encoding processing on the graph structure information to obtain syntactic tree encodings; Invoking a graph neural network to perform graph convolution processing on the syntactic tree encodings to obtain the syntactic tree features.

3. The method according to claim 2, wherein The constructing graph structure information of the text to be processed based on the syntactic tree information includes: Extracting N words and the syntactic relationships between the N words from the syntactic tree information, where N is an integer greater than or equal to 1; Using the N words as word nodes and the syntactic relationships as associated edges to construct the graph structure information of the text to be processed.

4. The method according to claim 2, wherein The performing encoding processing on the graph structure information to obtain syntactic tree encodings includes: Performing word encoding processing on the word nodes in the graph structure information to obtain word node encodings; Performing edge encoding processing on the associated edges in the graph structure information to obtain associated edge encodings; Combining the word node encodings and the associated edge encodings into a two-dimensional matrix, and using the two-dimensional matrix as the syntactic tree encodings.

5. The method according to claim 4, characterized in that, The combining the word node encodings and the associated edge encodings into a two-dimensional matrix includes: Performing sentence division processing on the text to be processed to obtain multiple sentences, and determining the sentence rankings of the sentences in the text to be processed; Determining the row position information of the word node encodings and the associated edge encodings corresponding to the sentences in the two-dimensional matrix based on the sentence rankings; Determining the column position information of the word node encodings and the associated edge encodings in the two-dimensional matrix based on the logical order of the word node encodings and the associated edge encodings corresponding to the sentences in the sentences.

6. The method according to claim 4, wherein The invoking a graph neural network to perform graph convolution processing on the syntactic tree encodings to obtain the syntactic tree features includes: Performing a sliding window process on the two-dimensional matrix through the convolution kernel of the graph neural network to obtain multiple sliding window matrices of the same size, where the size of the sliding window matrix is the same as the size of the convolution kernel; Performing graph convolution processing on the encodings in each sliding window matrix through the convolution kernel to obtain the convolution values of each sliding window matrix, where the encodings in the sliding window matrix include at least one of the word node encodings and the associated edge encodings; Combining the convolution values of the multiple sliding window matrices into the syntactic tree features based on the position information of the sliding window matrices in the two-dimensional matrix.

7. The method according to claim 1, characterized in that, The performing text encoding processing on the text to be processed to obtain text features includes: Extracting the letters of the words in the text to be processed; Based on the alphabet coding table, map the letters of the word to letter codes, and arrange the letter codes corresponding to the letters according to the position information of the letters in the word to obtain the word code of the word; Perform convolution processing on the word code to obtain the text feature.

8. The method according to claim 1, wherein Determining the phonemes of the text to be processed based on the fusion code includes: Decode the fusion code to obtain the digital sequence corresponding to the fusion code; Based on the phoneme mapping table, map the digital sequence to a phoneme sequence, and use the phoneme sequence as the phonemes of the text to be processed.

9. The method according to claim 8, characterized in that, Based on the phoneme mapping table, mapping the digital sequence to a phoneme sequence includes: Perform mapping processing on the digital sequence based on the phoneme mapping table to obtain the original phoneme sequence; Retrieve the invalid phonemes of the letters of the words in the text to be processed from the control mapping table; If the original phoneme sequence contains the invalid phoneme, remove the invalid phoneme from the original phoneme sequence to obtain the phoneme sequence.

10. A method for speech synthesis, characterized in that, The method includes: Obtain the text to be processed, and execute the method described in any one of claims 1 to 9 above on the text to be processed to obtain the phonemes of the text to be processed; Output speech matching the text to be processed based on the phonemes of the text to be processed.

11. An apparatus for phoneme prediction, characterized in that, The device includes: A text encoding module, which obtains the text to be processed and performs text encoding processing on the text to be processed to obtain text features; A graph convolution module, configured to construct syntactic tree information of the text to be processed based on the text to be processed, and perform graph convolution processing on the syntactic tree information to obtain syntactic tree features; A fusion processing module, configured to perform fusion processing on the syntactic tree features and the text features to obtain a fusion code, and determine the phonemes of the text to be processed based on the fusion code.

12. An apparatus for speech synthesis, characterized in that, The method includes: A text acquisition module, configured to obtain the text to be processed, and execute the method described in any one of claims 1 to 9 above on the text to be processed to obtain the phonemes of the text to be processed; A speech output module, configured to output speech matching the text to be processed based on the phonemes of the text to be processed.

13. An electronic device, characterized in that, The electronic device includes: A memory, configured to store computer-executable instructions or computer programs; A processor, when executing the computer-executable instructions or computer programs stored in the memory, implements the phoneme prediction method described in any one of claims 1 to 9, or implements the speech synthesis method described in claim 10.

14. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer programs are executed by the processor, the phoneme prediction method described in any one of claims 1 to 9 is implemented, or the speech synthesis method described in claim 10 is implemented.

15. A computer program product comprising computer-executable instructions, characterized in that, When the computer-executable instructions are executed by the processor, the phoneme prediction method described in any one of claims 1 to 9 is implemented, or the speech synthesis method described in claim 10 is implemented.