A method, device, equipment and storage medium for predicting the boundary of connected tone sandhi
By fusing character and tone features, the method enhances the accuracy of tone sandhi boundary prediction in speech synthesis, addressing the inefficiencies and inaccuracies of existing methods.
Patent Information
- Application Number
- CN202210158395.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-21
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-02-21
AI Technical Summary
The existing continuous reading and changing prediction method requires manual collection of large numbers of entries, with high labor costs and low prediction accuracy.
By extracting the predicted text word tone features of the predicted text, performing fusion conversion, obtaining the predicted character fusion features, and computing them to predict the boundary of continuous reading tone change.
Improve the accuracy of continuous reading and changing boundary prediction, reduce labor costs and improve the accuracy of prediction.
Smart Images

Figure CN114611583B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technologies, and particularly to a method, apparatus, device, and storage medium for predicting the boundaries of connected speech tone sandhi. Background Art
[0002] Connected speech tone sandhi refers to the phenomenon that when two or more syllables are pronounced in succession, their tones change under the influence of the preceding and following syllables, manifested as different tones of single characters. In the front-end text processing of speech synthesis, predicting the positions of the boundaries of connected speech tone sandhi can give correct tone predictions according to the corresponding positions of connected speech tone sandhi during speech synthesis, improve the accuracy rate of tones, and further improve the intelligibility of the synthesized speech.
[0003] Existing predictions of connected speech tone sandhi generally collect entries, construct a prediction model for the boundaries of connected speech tone sandhi through a word segmentation algorithm, and use a greedy algorithm to predict the input text to obtain the boundary prediction results. The existing methods require a large number of entries to be collected manually, with high labor costs and long time consumption, and the prediction accuracy rate is low. Summary of the Invention
[0004] The main technical problem to be solved by this application is to provide a method, apparatus, device, and storage medium for predicting the boundaries of connected speech tone sandhi, which can improve the accuracy rate of predicting the boundaries of connected speech tone sandhi.
[0005] To solve the above technical problem, in the first aspect of this application, a method for predicting the boundaries of connected speech tone sandhi is provided. The method includes: extracting the predicted text tone features of the predicted text; the predicted text includes multiple characters and the tone corresponding to each character, and the predicted text tone features represent the independent features of each character in the predicted text; performing fusion transformation on the predicted text tone features to obtain predicted character fusion features, and the predicted character fusion features represent the relationship features of each character in the predicted text with other characters; calculating the predicted character fusion features to obtain the predicted connected speech tone sandhi boundaries of the predicted text.
[0006] To solve the above technical problem, in the second aspect of this application, a device for predicting the boundaries of connected speech tone sandhi is provided. The prediction device includes: an extraction module for extracting the predicted text tone features of the predicted text; the predicted text includes multiple characters and the tone corresponding to each character, and the predicted text tone features represent the independent features of each character in the predicted text; a fusion module for performing fusion transformation on the predicted text tone features to obtain predicted character fusion features, and the predicted character fusion features represent the relationship features of each character in the predicted text with other characters; a calculation module for calculating the predicted character fusion features to obtain the predicted connected speech tone sandhi boundaries of the predicted text.
[0007] To solve the above technical problems, a third aspect of the present application provides a prediction device for continuous tone sandhi boundaries. The prediction device includes a memory and a processor coupled to each other. The memory stores program instructions. The processor is configured to execute the program instructions stored in the memory to implement the method described in the first aspect above.
[0008] To solve the above technical problems, a fourth aspect of the present application provides a computer-readable storage medium for storing program instructions that can be executed to implement the method described in the first aspect above.
[0009] The beneficial effects of the present application are as follows: Different from the prior art, the present application obtains the predicted character fusion feature by fusing multiple characters and the predicted text tone features corresponding to each character, and then calculates the predicted character fusion feature to obtain the predicted continuous tone sandhi boundary of the predicted text. By fusing the features of two modalities, namely characters and tones, the accuracy of the prediction result can be improved. Description of the Drawings
[0010] Figure 1 is a schematic flowchart of an embodiment of the method for predicting continuous tone sandhi boundaries of the present application;
[0011] Figure 2 is a schematic flowchart of an embodiment of the training method of the prediction model of the present application;
[0012] Figure 3 is the framework structure of the prediction model provided by the present application;
[0013] Figure 4 is a schematic flowchart of another embodiment of the training method of the prediction model of the present application;
[0014] Figure 5 is a schematic framework diagram of an embodiment of the prediction device for continuous tone sandhi boundaries provided by the present application;
[0015] Figure 6 is a schematic framework diagram of an embodiment of the prediction device for continuous tone sandhi boundaries provided by the present application;
[0016] Figure 7 is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Detailed Embodiments
[0017] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0018] It should be noted that in the embodiments of this application, there are descriptions involving "first", "second", etc. Such descriptions of "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second" may explicitly or implicitly include at least one such feature.
[0019] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an implementation manner of the method for predicting the boundary of connected-tone changes in this application. The method includes:
[0020] S110: Extract the predicted text tone features of the predicted text.
[0021] In one implementation manner, the predicted text includes multiple characters and the tone corresponding to each character. For example, the predicted text is "along6 the6 way6 didn't8 find6 the6 first7 department7 store5", where the numbers represent the tones corresponding to the characters, that is, the tone corresponding to "along" is 6.
[0022] The predicted text tone features represent the independent features of each character in the predicted text. In one implementation manner, extracting the predicted text tone features of the predicted text includes: performing one-hot encoding on the predicted text to obtain character encodings, and performing embedding on the character encodings to obtain the predicted text tone features of the predicted text.
[0023] Specifically, the predicted text tone features of the predicted text can be extracted by a prediction model. Input the predicted text into the prediction model, and encode the characters and their corresponding tones included in the predicted text. In one specific implementation manner, one-hot encoding can be performed on the characters and tones included in the predicted text to convert the predicted text into a vector. Further obtain the character features included in the predicted text. In one specific implementation manner, the character encodings of the text can be embedded through Bert to obtain the predicted text tone features of the predicted text. Among them, the predicted text tone features include the features of several characters in the predicted text and the features of the tones corresponding to the characters. It can be understood that in other implementation manners, the predicted text tone features of the predicted text can also be extracted by other models, and the extracted features are sent to the prediction model for predicting the boundary of connected-tone changes.
[0024] Before extracting the predicted text tone features of the predicted text, the category of the tone corresponding to the characters contained in the predicted text can be manually annotated. The category of the tone is the tone corresponding to the character. In a specific embodiment, the category of the tone can be obtained according to the mapping relationship between the tone category and the tone value. The mapping relationship between the tone category and the tone value can be {6:23, 5:34, 1:53, 7:55, 8:12}, etc. For example, the tone value of the character "延" is 23. According to the mapping relationship between the tone category and the tone value, it can be known that the tone corresponding to 延 is 6.
[0025] S120: Perform fusion conversion on the predicted text tone features to obtain predicted character fusion features.
[0026] The predicted text tone features obtained in step S110 are fused and transformed so that the model can obtain the predicted connected reading tone boundary based on the predicted character fusion features after the transformation. In one embodiment, the predicted text tone features are fused and transformed using the Seq2Seq+Attention method to obtain the predicted character fusion features. In other embodiments, other methods can also be used for fusion transformation, which are not limited here.
[0027] S130: Calculate the predicted character fusion features to obtain the predicted connected reading tone boundary of the predicted text.
[0028] The prediction model can be calculated based on the predicted character fusion features to obtain the predicted connected reading tone boundary of the predicted text. In one embodiment, the prediction model can calculate the probability that several character boundaries are predicted connected reading tone boundaries, several probabilities form a probability vector, and a threshold is set. If the calculated probability is greater than the threshold, it means that the character boundary is a predicted connected reading tone boundary. For example, the input predicted text is "Yan 6 tu 6 mei 8 xiao 6 dao 5", the threshold is set to 0.5, and the calculated probability vector is [0.2, 0.7, 0.6, 0.4, 0.3], then it can be known that the boundary between the character "路" and the character "没" and the boundary between the character "没" and the character "找" are predicted connected reading tone boundaries.
[0029] In another embodiment, the prediction model can also directly obtain the boundary vector of the predicted text, and determine the predicted connected reading tone boundary through the boundary vector. Among them, in the boundary vector, if the character boundary is a predicted connected reading tone boundary, it is represented by 1; if not, it is represented by 0. For example, the input predicted text is "Yan 6 tu 6 mei 8 shao 6 dao 5", and the obtained boundary vector is [0, 1, 1, 0, 0], then the characters "tu" and "mei" and the characters "mei" and "shao" are predicted connected reading tone boundaries.
[0030] The above method obtains the predicted character fusion feature by fusing multiple characters and the predicted text tone features corresponding to each character, and then calculates the predicted character fusion feature to obtain the predicted liaison tone change boundary of the predicted text. Fusing the features of two modalities, characters and tones, can improve the accuracy of the prediction result.
[0031] The above-described prediction method can be executed by a prediction model. In one embodiment, the prediction model includes an encoding layer, a decoding layer, and a target layer. The above step S110 can be performed in the encoding layer of the prediction model, and steps S120 and S130 can be performed in the decoding layer of the prediction model. Before predicting the liaison tone change boundary of the predicted text, the prediction model can be trained to make the prediction result more accurate. Please refer to Figure 2 and Figure 3 , Figure 2 which is a schematic flowchart of an embodiment of the training method of the prediction model of the present application; Figure 3 which is the framework structure of the prediction model provided by the present application. The training method includes:
[0032] S210: Extract the training text tone characteristics of the training text.
[0033] In one embodiment, the training text includes multiple characters, the tone corresponding to each character, and the standard liaison tone change boundary. The training text tone characteristics may include character features, the tone features corresponding to each character, and the standard boundary vector. Input the training text into the prediction model. As Figure 3 shown, the prediction model may include three parts: an encoding layer, a decoding layer, and a target layer. The encoding layer can be used to extract the training text tone characteristics of the training text. Specifically, the training text can be input into the prediction model to encode the characters and tones included in the training text. In a specific embodiment, the characters and tones included in the training text can be one-hot encoded to convert the training text into an encoded vector. Further, the encoded vector of the text can be embedded by Bert to obtain the training text tone characteristics of the training text. Among them, the training text tone characteristics include the characteristics of several characters in the training text and the characteristics of the tones corresponding to the characters. It can be understood that in other embodiments, the predicted text tone characteristics of the training text can also be extracted by other models, and the extracted characteristics can be sent to the prediction model for predicting the liaison tone change boundary.
[0034] The standard boundary vector can be obtained from the target layer, which can convert the training samples labeled with standard connected-tone boundary into standard boundary vectors. In one embodiment, for the pure text in the training samples, the target layer of the prediction model can convert it into 1 or 0 according to whether the character boundary is a connected-tone boundary. For example, if the pure text of the training sample is "along the way / didn't / find / the first / department store / ", where " / " represents the connected-tone boundary, the prediction model can convert it into the boundary vector [0, 1, 1, 0, 1, 0, 1, 0, 1].
[0035] S220: Perform fusion transformation on the tone characteristics of the training text to obtain the training character fusion features.
[0036] The step of performing fusion transformation on the tone characteristics of the training text can be executed by the encoding middle layer of the prediction model. Specifically, fuse the character features obtained in step S210 and the tone features corresponding to each character. In one embodiment, the Seq2Seq+Attention method is used to perform fusion transformation on the tone features of the training text to obtain the training character fusion features. Fusing the character features and the tone features can enable the prediction model to fully learn the tone-changing rules of characters and tone features in the semantic context.
[0037] S230: Calculate the training connected-tone boundary of the training text from the training character fusion features.
[0038] In one embodiment, the encoding layer of the prediction model can calculate the training text to obtain the probability vector. The probability vector is used to represent the probability that the boundaries of several characters in the training text are the training connected-tone boundaries. Based on the probability vector, the training connected-tone boundary of the training text is obtained. In one specific embodiment, a threshold can be set. If the probability value corresponding to a certain character is greater than the threshold, it is considered that the boundary of the character is the training connected-tone boundary.
[0039] S240: Calculate the loss function based on the training connected-tone boundary and the standard connected-tone boundary.
[0040] In one embodiment, the loss function can be calculated based on the standard boundary vector and the training boundary vector. Calculating the loss function can be performed in the target layer of the prediction model. The training boundary vector can be obtained by converting the training text labeled with the standard connected-tone boundary.
[0041] S250: Train the prediction model according to the loss function.
[0042] According to the result of the loss function, adjust the parameters of the prediction model until the training ends.
[0043] Please refer to Figure 4 , Figure 4It is a schematic flowchart of another embodiment of the training method of the prediction model of the present application. The training method includes:
[0044] S410: Extract the training text tone characteristics of the training text.
[0045] In one embodiment, the encoding layer of the prediction model can directly extract the training text tone characteristics of the training text. The training text is input into the prediction model, and the prediction model performs one-hot encoding on the characters and tones included in the training text to obtain the character encoding of the training text; the character encoding of the training text is embedded to obtain the text tone characteristics of the training text.
[0046] S420: Perform fusion conversion on the training text tone characteristics to obtain training character fusion features.
[0047] In one embodiment, the decoding layer of the prediction model can use the Seq2Seq+Attention method to perform fusion conversion on the training text tone features to obtain training character fusion features.
[0048] S430: Calculate the training character fusion features to obtain the training fundamental frequency sequence vector and the training connected tone change boundary of the training text.
[0049] In one embodiment, the decoding layer of the prediction model calculates the training character fusion features, and according to the calculation results, predicts the training fundamental frequency sequence vector and the training connected tone change boundary. Among them, the training fundamental frequency sequence vector is used to represent the change of the fundamental frequency values of several characters in the training text under the tone change rule. Further, before obtaining the connected tone change boundary, the prediction model can first obtain the training probability vector of the boundaries of several characters in the training text as the training connected tone change boundary, and determine the training connected tone change boundary based on the training probability vector. In a specific embodiment, a threshold can be set. If the probability value corresponding to a certain character is greater than the threshold, it is considered that the boundary of the character is the training connected tone change boundary. For example, the input prediction text is "along6way6no8find6to5", the set threshold is 0.5, and the calculated probability vector is [0.2, 0.7, 0.6, 0.4, 0.3], then it can be known that the boundaries between the character "way" and the character "no" and between the character "no" and the character "find" are the predicted connected tone change boundaries.
[0050] S440: Analyze the training audio corresponding to the training text to obtain the standard audio sequence vector of the training text.
[0051] In one embodiment, the target layer of the prediction model can extract the fundamental frequency data of the training audio to obtain the standard audio sequence vector of the training text. In a specific embodiment, the time domain method can be used to set the extraction time interval and extract the fundamental frequency of the training audio to obtain the standard audio sequence vector. Among them, the extraction time interval can be 2s, 3s, 4s, etc., which is not limited here.
[0052] S450: Calculate the first loss function based on the training liaison tone boundary and the standard liaison tone boundary.
[0053] In one embodiment, the first loss function can be calculated based on the training probability vector of the training liaison tone boundary and the training boundary vector of the training text. Among them, the training boundary vector is obtained by transforming the training text marked with the liaison tone boundary through the prediction model. In a specific embodiment, if the obtained training probability vector is P b and the training boundary vector is G b , then the calculated first loss can be Loss1 = CrossEntropyLoss(P b , G b ).
[0054] S460: Calculate the second loss function based on the training fundamental frequency sequence vector and the standard audio sequence vector.
[0055] In one embodiment, the standard audio sequence vector is G f , the training fundamental frequency sequence vector is P f . Since the lengths between sequences may be inconsistent, the second loss Loss2 = DWTLoss(P f , G f ) can be calculated through the loss function DTW of the time series.
[0056] S470: Train the prediction model according to the first loss function and the second loss function.
[0057] In one embodiment, the first loss function and the second loss function can be weighted and summed to obtain the total loss function, and the prediction model is trained based on the total loss function. In a specific embodiment, the total loss can be Loss = alpha1 * Loss1 + alpha2 * Loss2, where alpha1 and alpha2 are hyperparameters of the experiment.
[0058] In the above method, by calculating two loss functions and training the prediction model based on the two loss functions, the accuracy of predicting the liaison tone boundary of the prediction model can be improved.
[0059] Please refer to Figure 5 , Figure 5It is a schematic framework diagram of an implementation manner of the prediction device for the boundary of connected tone sandhi provided by this application.
[0060] The prediction device 50 for the boundary of connected tone sandhi includes: an extraction module 51, a fusion module 52, and a calculation module 53. The extraction module 51 is used to extract the predicted text tone feature of the predicted text; the predicted text includes multiple characters and the tone corresponding to each character, and the predicted text tone feature represents the independent features of each character in the predicted text; the fusion module 52 is used to perform fusion transformation on the predicted text tone feature to obtain the predicted character fusion feature, and the predicted character fusion feature represents the relationship features between each character and other characters in the predicted text; the calculation module 53 is used to calculate the predicted character fusion feature to obtain the predicted boundary of connected tone sandhi of the predicted text.
[0061] Among them, the above steps are executed by a prediction model, and the training steps of the prediction model include: extracting the training text tone characteristics of the training text; the training text includes multiple characters, the tone corresponding to each character, and the standard boundary of connected tone sandhi; performing fusion transformation on the training text tone characteristics to obtain the training character fusion feature; calculating the training boundary of connected tone sandhi of the training text from the training character fusion feature; calculating the loss function based on the training boundary of connected tone sandhi and the standard boundary of connected tone sandhi; training the prediction model according to the loss function.
[0062] Among them, calculating the training boundary of connected tone sandhi of the training text from the training character fusion feature includes: calculating the training fundamental frequency sequence vector and the training boundary of connected tone sandhi of the training text from the training character fusion feature; the training steps further include: analyzing the training audio corresponding to the training text to obtain the standard audio sequence vector of the training text; calculating the loss function based on the training boundary of connected tone sandhi and the standard boundary of connected tone sandhi, and training the prediction model according to the loss function, including: calculating the first loss function based on the training boundary of connected tone sandhi and the standard boundary of connected tone sandhi; calculating the second loss function based on the training fundamental frequency sequence vector and the standard audio sequence vector; training the prediction model according to the first loss function and the second loss function.
[0063] Among them, training the prediction model according to the first loss function and the second loss function includes: performing weighted summation on the first loss function and the second loss function to obtain the total loss function; training the prediction model based on the total loss function.
[0064] Among them, performing fusion transformation on the predicted text tone feature to obtain the predicted character fusion feature; performing fusion transformation on the training text tone characteristics to obtain the training character fusion feature, both include: using the Seq2Seq+Attention method to perform fusion transformation on the text tone feature to obtain the character fusion feature.
[0065] Among them, extracting the prosody feature of the predicted text and the prosody feature of the training text both include: performing one-hot encoding on the predicted text or the training text to obtain character encodings; and performing embedding on the character encodings to obtain the prosody feature of the predicted text or the prosody feature of the training text.
[0066] Among them, calculating the predicted character fusion feature to obtain the predicted liaison tone change boundary of the predicted text; calculating the training character fusion feature to obtain the training liaison tone change boundary of the training text both include: calculating the character fusion feature to obtain the liaison tone change boundary probability of each character; and obtaining the predicted liaison tone change boundary of the predicted text based on the liaison tone change boundary probability of each character.
[0067] Please refer to Figure 6 , Figure 6 which is a schematic framework diagram of an embodiment of the device for predicting the liaison tone change boundary provided by this application.
[0068] The device 60 for predicting the liaison tone change boundary includes a memory 61 and a processor 62 that are coupled to each other. Program instructions are stored in the memory 61, and the processor 62 is configured to execute the program instructions to implement the steps in any of the above method embodiments. Specifically, the device 60 for predicting the liaison tone change boundary may include, but is not limited to: desktop computers, laptop computers, servers, mobile phones, tablet computers, etc., which are not limited herein.
[0069] Specifically, the processor 62 is configured to control itself and the memory 61 to implement the steps in any of the above method embodiments. The processor 62 may also be referred to as a CPU (Central Processing Unit). The processor 62 may be an integrated circuit chip with signal processing capabilities. The processor 62 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 62 may be implemented jointly by integrated circuit chips.
[0070] Please refer to Figure 7 , Figure 7It is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 70 stores program instructions 71, which, when executed by a processor, are used to implement the steps in any of the above method embodiments.
[0071] The computer-readable storage medium 70 can specifically be a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, etc., which can store computer programs, or it can also be a server storing the computer program. The server can send the stored computer program to other devices for running, or it can also run the stored computer program by itself.
[0072] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0073] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0074] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.
[0075] The above are only the embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A method for predicting the boundary of connected tone sandhi, characterized in that, The prediction method is executed by a prediction model, and the prediction method includes: Extracting the prosodic feature of the prediction text; the prediction text includes multiple characters and the prosody corresponding to each character, and the prosodic feature of the prediction text represents the independent features of each character in the prediction text; Performing fusion transformation on the prosodic feature of the prediction text to obtain a fused prediction character feature, where the fused prediction character feature represents the relationship features between each character and other characters in the prediction text; Calculating the predicted liaison tone change boundary of the prediction text from the fused prediction character feature; Wherein, the training steps of the prediction model include: Extracting the prosodic feature of the training text; the training text includes multiple characters, the prosody corresponding to each character, and the standard liaison tone change boundary; Performing fusion transformation on the prosodic feature of the training text to obtain a fused training character feature; Calculating the training fundamental frequency sequence vector and the training liaison tone change boundary of the training text from the fused training character feature; Analyzing the training audio corresponding to the training text to obtain the standard audio sequence vector of the training text; Calculating a first loss function based on the training liaison tone change boundary and the standard liaison tone change boundary; Calculating a second loss function based on the training fundamental frequency sequence vector and the standard audio sequence vector; Training the prediction model according to the first loss function and the second loss function.
2. The method according to claim 1, wherein The training of the prediction model according to the first loss function and the second loss function includes: Performing weighted summation on the first loss function and the second loss function to obtain a total loss function; Training the prediction model based on the total loss function.
3. The method according to any one of claims 1-2, characterized in that The performing fusion transformation on the prosodic feature of the prediction text to obtain a fused prediction character feature; and the performing fusion transformation on the prosodic feature of the training text to obtain a fused training character feature both include: Using the Seq2Seq+Attention method to perform fusion transformation on the prosodic feature of the text to obtain a fused character feature.
4. The method according to any one of claims 1-2, characterized in that The extracting the prosodic feature of the prediction text and the extracting the prosodic feature of the training text both include: Performing one-hot encoding on the text to obtain a character encoding; Embedding the character encoding to obtain the prosodic feature of the text.
5. The method according to any one of claims 1-2, characterized in that, The calculating the predicted liaison tone change boundary of the prediction text from the fused prediction character feature; and the calculating the training liaison tone change boundary of the training text from the fused training character feature both include: Calculating the liaison tone change boundary probability of each character from the fused character feature; Obtaining the predicted liaison tone change boundary of the prediction text based on the liaison tone change boundary probability of each character.
6. A prediction device for the boundary of connected tone sandhi, characterized in that, The prediction device includes: An extraction module, which is used to extract the prosodic feature of the prediction text; the prediction text includes multiple characters and the prosody corresponding to each character, and the prosodic feature of the prediction text represents the independent features of each character in the prediction text; A fusion module for fusing and transforming the predicted text tone features to obtain predicted character fusion features, where the predicted character fusion features represent the relationship features between each character and other characters in the predicted text; A calculation module for calculating the predicted character fusion features to obtain the predicted connected speech tone change boundaries of the predicted text; The steps executed by the extraction module, the fusion module, and the calculation module are executed by a prediction model, and the training steps of the prediction model include: Extracting the training text tone characteristics of the training text; the training text includes multiple characters, the tone corresponding to each character, and the standard connected speech tone change boundaries; Fusing and transforming the training text tone characteristics to obtain training character fusion features; Calculating the training text fundamental frequency sequence vectors and training connected speech tone change boundaries of the training text from the training character fusion features; Analyzing the training audio corresponding to the training text to obtain the standard audio sequence vectors of the training text; Calculating a first loss function based on the training connected speech tone change boundaries and the standard connected speech tone change boundaries; Calculating a second loss function based on the training fundamental frequency sequence vectors and the standard audio sequence vectors; Training the prediction model according to the first loss function and the second loss function.
7. A prediction device for the boundary of connected tone sandhi, characterized in that The device includes a memory and a processor coupled to each other, The memory stores program instructions; The processor is configured to execute the program instructions stored in the memory to implement the method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program instructions that can be executed to implement the method according to any one of claims 1-5.
Citation Information
Patent Citations
Tibetan tone prediction method and system
CN106294310A