Voice generation method and voice generation model training method based on phonetic symbols and semantics

By fusing the semantic feature matrix of text data with the guiding information of phonetic symbol data and inputting the speech generation model, the problem of unnatural pronunciation in the prior art is solved, and a more natural and emotionally rich speech generation effect is achieved.

CN119993118AActive Publication Date: 2025-05-13BEIJING SINOVOICE TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411973758.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-13
Estimated Expiration
2044-12-30

Smart Images

  • Figure CN119993118A_ABST
    Figure CN119993118A_ABST
Patent Text Reader

Abstract

The invention provides a voice generation method based on phonetic symbols and semantics, a voice generation model training method and device based on phonetic symbols and semantics, electronic equipment and a computer readable storage medium. In the embodiment of the invention, the semantic feature matrix in the text data and the guide information with the phonetic symbol data features are fused to obtain the fused feature matrix with the semantic information and the phonetic symbol information, and the fused feature matrix is input to the voice generation model to obtain the voice data corresponding to the text data. The phonetic symbol and semantic information is fused in the input information, and the voice corresponding to the phonetic symbol is adjusted according to different semantics, so that the voice is not averaged any more and the mechanical feeling is reduced, and the phonetic symbol and semantic based voice generation method makes full use of the information contained in the input text, does not need to add additional input data, and improves the user experience. Therefore, more vivid voice data can be generated, and the voice generation effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech generation, and specifically to a speech generation method based on phonetic symbols and semantics, a speech generation model training method based on phonetic symbols and semantics, a device, an electronic device and a computer-readable storage medium. Background Art

[0002] The speech generation function is used in all aspects of daily life. For example, the text reading function requires generating corresponding speech from text data.

[0003] In the prior art, text data is usually input into a speech generation model to obtain speech corresponding to the text data.

[0004] The voice obtained through text data lacks emotional changes, the pronunciation is too even, has a heavy mechanical feel, and is not natural enough. Summary of the invention

[0005] The embodiment of the present application provides a speech generation method based on phonetic symbols and semantics to solve the problem of unnatural pronunciation of speech generated in the prior art.

[0006] Correspondingly, the embodiments of the present application also provide a speech generation model training method, device, electronic device and computer-readable storage medium based on phonetic symbols and semantics to ensure the implementation and application of the above method.

[0007] In a first aspect, an embodiment of the present application provides a method for generating speech based on phonetic symbols and semantics, the method comprising:

[0008] Inputting text data into a semantic model to obtain a semantic feature matrix of the text data;

[0009] Translating the text data into phonetic symbol data, inputting the phonetic symbol data into a bypass guidance module of a text encoding model, outputting each coupling layer in the bypass guidance module, and obtaining guidance information; the guidance information is used to be fused with the semantic feature matrix;

[0010] The semantic feature matrix is ​​input into a text encoding model, and the guidance information of each coupling layer in the bypass guidance module is sequentially input into each convolutional layer of the text encoding model, and a fusion feature matrix is ​​output from the text encoding model; the coupling layer corresponds to the convolutional layer one by one;

[0011] The fused feature matrix is ​​input into a speech generation model to obtain speech data corresponding to the text data.

[0012] In a second aspect, an embodiment of the present application provides a method for training a speech generation model based on phonetic symbols and semantics, the method comprising:

[0013] Extracting semantic sample information and phonetic sample information from the text sample data, and inputting the semantic sample information and the phonetic sample information into a text encoding model to obtain fused sample feature data;

[0014] Inputting the speech sample data corresponding to the text sample data into a speech coding model to obtain a speech sample feature matrix;

[0015] By calculating the mean and variance of the fusion sample feature matrix and the speech sample feature matrix, a text sample distribution function and a speech sample distribution function are obtained;

[0016] Using the speech sample distribution function to train an initial model, and determining a loss value according to a difference between the text sample distribution function and the speech sample distribution function;

[0017] According to the loss value and the preset loss function, the initial model is adjusted to obtain a speech generation model, and the speech generation model is used for the speech generation method based on phonetic symbols and semantics.

[0018] In a third aspect, an embodiment of the present application provides a speech generation device based on phonetic symbols and semantics, the device comprising:

[0019] A semantic module, used for inputting text data into a semantic model to obtain a semantic feature matrix of the text data;

[0020] A phonetic symbol module, used for translating the text data into phonetic symbol data, inputting the phonetic symbol data into a bypass guidance module of a text encoding model, outputting each coupling layer in the bypass guidance module, and obtaining guidance information; the guidance information is used for merging with the semantic feature matrix;

[0021] A fusion module, used for inputting the semantic feature matrix into a text encoding model, and inputting the guidance information of each coupling layer in the bypass guidance module into each convolution layer of the text encoding model in sequence, and outputting a fusion feature matrix from the text encoding model; the coupling layers correspond to the convolution layers one by one;

[0022] A generation module is used to input the fusion feature matrix into a speech generation model to obtain speech data corresponding to the text data.

[0023] In a fourth aspect, an embodiment of the present application provides a speech generation model training device based on phonetic symbols and semantics, the device comprising:

[0024] A sample fusion module is used to extract semantic sample information and phonetic sample information from text sample data, and input the semantic sample information and the phonetic sample information into a text encoding model to obtain fused sample feature data;

[0025] A sample speech module, used for inputting speech sample data corresponding to the text sample data into a speech coding model to obtain a speech sample feature matrix;

[0026] A distribution function module, used for obtaining a text sample distribution function and a speech sample distribution function by calculating the mean and variance of the fusion sample feature matrix and the speech sample feature matrix respectively;

[0027] A loss value module, used to train an initial model using the speech sample distribution function, and determine a loss value according to a difference between the text sample distribution function and the speech sample distribution function;

[0028] The training module is used to adjust the initial model according to the loss value and the preset loss function to obtain a speech generation model, and the speech generation model is used for the speech generation method based on phonetic symbols and semantics.

[0029] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of one or more methods described in the embodiments of the present application are implemented.

[0030] In a sixth aspect, an embodiment of the present application provides a readable storage medium. When the instructions in the readable storage medium are executed by a processor of an electronic device, the electronic device can execute one or more methods described in the embodiments of the present application.

[0031] In an embodiment of the present application, a fused feature matrix having semantic information and phonetic information is obtained by fusing the semantic feature matrix in the text data and the guiding information having phonetic data characteristics, and the fused feature matrix is ​​input into a speech generation model to obtain speech data corresponding to the text data. Since the input information contains information of phonetic symbols and semantics, the speech corresponding to the phonetic symbols is adjusted according to different semantics, so that the speech is no longer averaged and the mechanical feel is reduced. The speech generation method based on phonetic symbols and semantics in an embodiment of the present application makes full use of the information contained in the input text and can generate more realistic speech data without adding additional input data, thereby improving the effect of speech generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the following is a brief introduction to the drawings required for use in the embodiments or the related technical descriptions. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 It is a flowchart of the steps of a method for generating speech based on phonetic symbols and semantics provided in an embodiment of the present application;

[0034] Figure 2 It is a schematic diagram of generating a fusion feature matrix provided in an embodiment of the present application;

[0035] Figure 3 It is a flowchart of another method for generating speech based on phonetic symbols and semantics provided in an embodiment of the present application;

[0036] Figure 4 It is a schematic diagram of coupling layer partitioning provided in an embodiment of the present application;

[0037] Figure 5 It is a flowchart of the steps of a speech generation model training method based on phonetic symbols and semantics provided in an embodiment of the present application;

[0038] Figure 6 It is an architecture diagram of a speech generation model training method based on phonetic symbols and semantics provided in an embodiment of the present application;

[0039] Figure 7 It is a structural diagram of a speech generation device based on phonetic symbols and semantics provided in an embodiment of the present application;

[0040] Figure 8 It is a structural diagram of a speech generation model training device based on phonetic symbols and semantics provided in an embodiment of the present application;

[0041] Fig. 9 is a structural block diagram of an electronic device provided in an embodiment of the present application;

[0042] Fig.10 It is a structural block diagram of another electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0043] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0044] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally a class, and the number of objects is not limited. For example, the first object can be one or more. In addition, the term "and / or" in the specification and claims is used to describe the association relationship of associated objects, indicating that three kinds of relationships can exist, for example, A and / or B can be represented: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the front and back associated objects are a kind of "or" relationship. In the embodiment of the present application, the term "multiple" refers to two or more, and other quantifiers are similar.

[0045] like Figure 1 As shown, the present application embodiment proposes a flowchart of the steps of a speech generation method based on phonetic symbols and semantics, such as Figure 1 As shown, the method may include:

[0046] Step 101: input text data into a semantic model to obtain a semantic feature matrix of the text data.

[0047] In the process of speech generation, the target information is usually input into the speech generation model in the format of text data to obtain the speech data corresponding to the text data. When the text data is input into the speech generation model, in some related technologies, the phonetic information contained in the text data is usually used to generate the corresponding speech data. For Chinese, the phonetic information is the pinyin of Chinese. Phonetic information often indicates the most standard pronunciation. In daily life, people's pronunciation when speaking will make slight changes in speech speed, pitch, weight, timbre, etc. according to different semantics. Standard pronunciation is rarely used for communication. Therefore, the pronunciation of speech data generated by using phonetic information is too standardized, has a strong mechanical feel, and lacks a natural feel.

[0048] In an embodiment of the present application, a semantic feature matrix is ​​obtained by inputting text data into a semantic model. The semantic feature matrix contains the semantic features of the text data extracted by the semantic model. These semantic features enable the same text data to be refined into richer categories based on different semantic information, thereby preventing the voice data corresponding to the text data from being too single and affecting the voice generation effect.

[0049] For example, a third-party semantic model (BERT, Bidirectional Encoder Representations from Transformers) can be used to generate a semantic feature matrix, and the size of the semantic feature matrix can be set by itself.

[0050] Step 102: Translate the text data into phonetic data, input the phonetic data into the bypass guiding module of the text encoding model, and output each coupling layer in the bypass guiding module to obtain guiding information.

[0051] In the embodiment of the present application, by translating the text data into phonetic data, generating data with the initials, finals, and tone information of the text data, the information in the text data is fully utilized, and the phonetic data is input into the bypass guiding module of the text encoding model. The data of each layer in the module is used as the guiding information of the bypass and is fused with the semantic feature matrix as the input information at the same time, preparing for generating speech data with semantic and phonetic information.

[0052] In some embodiments, a third-party phonetic annotation tool, such as PyPinyin, can be used to translate the text data into phonetic data.

[0053] For example, for the text data "good person", after translation, phonetic data such as "h3 ao3 r2 en2" will be obtained, where the initials and finals are separated, and the pronunciation of the combined characters is marked with numbers at the back. Different numbers represent the corresponding tones.

[0054] Step 103: Input the semantic feature matrix into the text encoding model, and sequentially input the guiding information of each coupling layer in the bypass guiding module into each convolutional layer of the text encoding model, and output a fused feature matrix from the text encoding model.

[0055] In the embodiment of the present application, the semantic feature matrix is input into the text encoding model, and the guiding information of each coupling layer in the bypass guiding module is sequentially input into each convolutional layer of the text encoding layer, so that when the semantic feature matrix executes the mapping relationship of each convolutional layer of the text encoding model, there is corresponding guiding information of the coupling layer for fusion, making the fusion of the semantic feature matrix and the guiding information with phonetic data characteristics more stable, which is beneficial to the stability of the generation result of the speech generation model.

[0056] Figure 2The schematic diagram of generating the fusion feature matrix is ​​as follows: the phonetic symbol data is input into the bypass guidance module, and is mapped in the coupling layers in the bypass guidance in turn, and the output of each coupling layer is input into each convolution layer in the text encoding model as the guidance information. The semantic feature matrix is ​​input into the text encoding model, and before the semantic feature matrix is ​​input into each convolution layer, it is fused with the guidance data in the corresponding coupling layer. The guidance data contains the features of the phonetic symbol data, so the semantic features and the phonetic symbol features can be fused and used as the input of the convolution layer for corresponding mapping, and the output data is fused with the guidance information of the coupling layer corresponding to the next convolution layer, and the fused information is used as the input of the next convolution layer until the text encoding model is completed, and the fusion feature matrix is ​​obtained.

[0057] Among them, the coupling layers in the bypass guided module correspond one-to-one to the convolutional layers in the text encoding model. When the text encoding model has a NLAYERS layer convolutional layer, the bypass guided module has a NLAYERS layer coupling layer.

[0058] For example, when the phonetic symbol data is "h3 ao3 r2 en2", the guidance data in each coupling layer of the bypass guidance module is a 4×D matrix, where D is an internal dimension that can be set by itself, such as guidance data of size 4×192. Correspondingly, the semantic feature matrix can be set to a 4×192 matrix, and the obtained fusion feature matrix is ​​also a 4×192 matrix.

[0059] Step 104: input the fused feature matrix into a speech generation model to obtain speech data corresponding to the text data.

[0060] In an embodiment of the present application, the fusion feature matrix has the phonetic features and semantic features of the text data, and the fusion feature matrix is ​​input into the speech generation model. Since the speech generation model is also obtained by fusing the sample data through the above method to generate a sample fusion feature matrix, and using the sample fusion feature matrix and the corresponding speech data to train the speech generation model. The speech generation model can identify the fusion feature matrix with rich semantic information and phonetic information, and calculate the speech feature matrix with the highest probability corresponding to the fusion feature matrix, and obtain the speech data by decoding the speech feature matrix. Therefore, the speech data can be classified into rich semantic information and phonetic information, corresponding to the rich speech data, to achieve subtle changes in pronunciation based on semantics in daily life, so that the generated speech is more natural.

[0061] For example, when the text data is "You are such a good person", in some related technologies, the pronunciation of the voice data generated only by phonetic data is too even. When there is no pronunciation emphasis, the sentence not only sounds unnatural, but also makes the audience focus on the "good person", and think about whether "you" have done some good things to be praised by the "good person". When the actual text data adds emphasis to the word "really" in terms of semantic features, "You are such a good person" conveys a feeling of gratitude, which is not only more vivid, but also more accurately expresses the meaning of the speaker, thereby improving the effect of the generated speech.

[0062] In summary, in the embodiments of the present application, the semantic feature matrix in the text data and the guiding information with phonetic data characteristics are fused to obtain a fused feature matrix with semantic information and phonetic information, and the fused feature matrix is ​​input into the speech generation model to obtain the speech data corresponding to the text data. Since the input information incorporates the information of phonetic symbols and semantics, the speech corresponding to the phonetic symbols is adjusted according to different semantics, so that the speech is no longer averaged, reducing the mechanical feel. The speech generation method based on phonetic symbols and semantics in the embodiments of the present application makes full use of the information contained in the input text and can generate more realistic speech data without adding additional input data, thereby improving the effect of speech generation.

[0063] like Figure 3 As shown, the present application embodiment proposes a flowchart of a method for generating speech based on phonetic symbols and semantics, such as Figure 3 As shown, the method may include:

[0064] Step 201: input text data into a semantic model to obtain a semantic feature matrix of the text data.

[0065] This step may specifically refer to the above step 101, which will not be described in detail here.

[0066] Step 202, translating the text data into phonetic symbol data, and inputting the phonetic symbol data into a bypass guidance module of a text encoding model, and outputting each coupling layer in the bypass guidance module to obtain guidance information.

[0067] This step may be specifically referred to as the above step 102, which will not be described in detail here.

[0068] Optionally, step 202 may specifically include:

[0069] Sub-step 2021, segmenting the text data into phrases, and searching for phonetic symbols mapped to the segmented phrases according to a preset phonetic symbol library to obtain phonetic symbol data.

[0070] The phonetic symbol library is a data storage structure that stores corresponding phrases and phonetic symbols. The corresponding phonetic symbol data can be found by inputting the phrase into the phonetic symbol library.

[0071] In an embodiment of the present application, by segmenting text data into phrases, data in phrase units is obtained, so that the phonetic data corresponding to the polyphones appearing in the phrases are clearer, thereby reducing the situation where errors in the generated voice data occur due to the presence of polyphones in the text data.

[0072] For example, when the text data is "You are such a good person", the results obtained after segmenting the text data are "you", "really", and "good person". Due to the specific pronunciation of "good person", the mapped phonetic data obtained by searching in the phonetic database is "h3ao3 r2 en2" instead of "h4 ao4 r2 en2".

[0073] Sub-step 2022, searching a preset vector codebook to find the corresponding vector for the phonetic symbol data, and obtaining a phonetic symbol vector matrix.

[0074] Among them, the vector codebook is a data set marked with the vector corresponding to each phonetic symbol data. According to the vector corresponding to each phonetic symbol in the vector codebook, the phonetic symbol vector matrix is ​​obtained. The vectors in the vector codebook are empirical data, which aims to extract more features of the phonetic symbol data and can be customized.

[0075] For example, there are 70 initials and finals in Mandarin Chinese, and there are 5 tones, among which the light tone can be represented by 0, and the first to fourth tones are represented by 1, 2, 3, and 4 respectively, that is, there are 70×5 units in total. The dimension of the preset vector codebook is 350×D, where D is the internal dimension and can be set by yourself. For the phonetic data "h3 ao3 r2 en2", the vectors corresponding to the four units are searched respectively, and the concatenated phonetic vector matrix is ​​a 4×D-dimensional matrix.

[0076] Sub-step 2023, inputting the phonetic symbol vector matrix into the bypass guidance module.

[0077] In an embodiment of the present application, the phonetic symbol vector matrix is ​​input into a bypass guidance module so that each coupling layer of the bypass guidance module can fuse the corresponding guidance information with the input of each convolutional layer of the text encoding model.

[0078] For example, the phonetic symbol vector matrix of the phonetic symbol data "h3 ao3 r2 en2" is a 4×D-dimensional matrix, each layer of guidance information is also a 4×D-dimensional matrix, and the size of the semantic feature matrix can also be set to 4×D. The fused feature matrix obtained after fusion is a 4×D-dimensional matrix, where D is the internal dimension set by yourself.

[0079] Optionally, each coupling layer of the bypass guiding module is divided into a first area and a second area, and step 202 may specifically include:

[0080] Sub-step 2024, for any coupling layer in the bypass guidance model, and the adjacent coupling layer of any coupling layer, while the phonetic symbol data of the first area of ​​any coupling layer is kept unchanged and the phonetic symbol data of the second area is transformed, the phonetic symbol data of the first area of ​​the adjacent coupling layer is transformed and the phonetic symbol data of the second area is not transformed; or while the phonetic symbol data of the first area of ​​any coupling layer is kept unchanged and the phonetic symbol data of the second area is not transformed, the phonetic symbol data of the first area of ​​the adjacent coupling layer is not transformed and the phonetic symbol data of the second area is transformed.

[0081] In the embodiment of the present application, the size of the phonetic symbol vector matrix input in the bypass guidance module can be used as the size of the coupling layer. Figure 4 This is the partitioning method of the coupling layer, where the coupling layer has D rows. By selecting a constant d smaller than D, the coupling layer can be divided into two regions, where the first region and the second region are shown in the figure. The first region is Z 1 -Z d , the second area is Z d+1 -Z D .

[0082] Specifically, when the phonetic symbol vector matrix passes through the coupling layer of the bypass guidance module, only one of the two regions of the coupling layer undergoes mapping transformation, and the other region keeps the original data unchanged. And in the next coupling layer, Figure 2 The dimension is flipped as shown, that is, the area transformed in coupling layer 1 is no longer transformed in coupling layer 2, while the area not transformed in coupling layer 1 is transformed in coupling layer 2, and flipped in adjacent coupling layers until all coupling layers are executed.

[0083] like Figure 2 As shown, the function of the bypass guidance module is to map the input phonetic symbol vector matrix into a normal distribution matrix, and the final self-supervised loss value output by the bypass guidance module is the difference between the matrix output by the bypass guidance module and the normal distribution matrix.

[0084] Sub-step 2025, after the phonetic symbol data in each coupling layer are transformed, the guidance information of each layer is obtained.

[0085] In an embodiment of the present application, the data in each coupling layer in the bypass guidance model are output separately to obtain the guidance information of each layer of phonetic symbol data.

[0086] In one embodiment, the coupling layers in the bypass guidance module have the same number of convolutional layers as those in the text encoding model and both are even-numbered layers.

[0087] In an embodiment of the present application, by setting the number of coupling layers in the bypass guidance module to an even number of layers, the two parts of the phonetic symbol vector matrix are transformed the same number of times, thereby reducing the error caused by the different number of transformations. At the same time, the number of coupling layers is set to be the same as the number of convolution layers, so that the guidance information and the input information of the convolution layer can be fused one by one.

[0088] Step 203, when any convolutional layer of the text encoding model is about to input data, the guidance information of each group of one-to-one corresponding coupling layers and the first data of the convolutional layer are fused to obtain fused data.

[0089] Among them, the first data is the data transmitted from the upper convolution layer to the lower convolution layer, that is, the output data of the upper convolution layer.

[0090] In the embodiments of the present application, Figure 2 As shown, when the convolution layer 2 is about to input data, the one-to-one corresponding guidance information of the coupling layer 2 and the first data of the convolution layer 2 are fused to obtain fused data. The first data for the convolution layer 2 is the data output by the convolution layer 1, and the first data for the convolution layer 1 is the semantic feature matrix.

[0091] Optionally, step 203 may specifically include:

[0092] Sub-step 2031, combining the guidance information of each group of one-to-one corresponding coupling layers with the first data of the current convolutional layer in the following manner:

[0093]

[0094] Wherein O is the fused data, I is the first data of the convolutional layer, C is the guidance information of the coupling layer, mean(I) is the mean of I, std(I) is the standard deviation of I, W() and B() are two fully connected layers of the text encoding model, W(C) represents the product of the guidance information and one fully connected layer, and B(C) represents the product of the guidance information and another fully connected layer.

[0095] In an embodiment of the present application, each group of one-to-one corresponding guidance information of the coupling layer and the first data of the current convolutional layer are combined in the manner as described above, so that the semantic feature matrix and the guidance information with phonetic data characteristics are more stably integrated, which is conducive to the stability of the generation results of the speech generation model.

[0096] Step 204: Use the fused data as input data of the convolution layer, and output second data from the convolution layer.

[0097] In the embodiment of the present application, the second data is the output data after the convolution layer input fusion data. For the convolution layer NLAYERS, the second data is the feature fusion matrix.

[0098] like Figure 2 As shown, the second data output by convolution layer 1 and the guide data output by coupling layer 2 are fused to become the input data of convolution layer 2, and the second data of convolution layer 1 is the first data of convolution layer 2.

[0099] Step 205: After all convolutional layers are output in sequence, the feature fusion matrix is ​​obtained.

[0100] In an embodiment of the present application, by fusing data between the convolutional layer of the text encoding model and the coupling layer of the bypass guidance module, the semantic information in the semantic feature matrix is ​​fused with the phonetic information in the phonetic vector matrix, so that the information contained in the input text data is fully utilized, and the speech data with the highest probability corresponding to the semantic information and the phonetic information is calculated in the speech generation model. Since the speech data is mapped from data with semantic information and phonetic information, the speech data is a pronunciation with corresponding adjustments with semantic information and phonetic information, so the obtained pronunciation data is no longer averaged and is more natural.

[0101] Step 206: input the fused feature matrix into a speech generation model to obtain speech data corresponding to the text data.

[0102] This step can be specifically referred to in the above step 104, and will not be described in detail here.

[0103] In summary, in the embodiments of the present application, the semantic feature matrix in the text data and the guiding information with the characteristics of the phonetic data are fused to obtain a fused feature matrix with semantic information and phonetic information, and the fused feature matrix is ​​input into the speech generation model to obtain the speech data corresponding to the text data. Since the input information is fused with the information of the phonetic symbols and semantics, the speech corresponding to the phonetic symbols is adjusted according to different semantics, so that the speech is no longer averaged, reducing the mechanical feel, and the data between the one-to-one corresponding coupling layer and the convolution layer are fused in a specific manner, which is conducive to the stability of the generation result of the speech generation model. The speech generation method based on phonetic symbols and semantics in the embodiments of the present application makes full use of the information contained in the input text, and can stably generate more realistic speech data without adding additional input data, thereby improving the effect of speech generation.

[0104] Figure 5 is a flowchart of a method for training a speech generation model based on phonetic symbols and semantics provided in an embodiment of the present application, such as Figure 5 As shown, the method may include:

[0105] Step 301: Extract the semantic sample information and phonetic symbol sample information from the text sample data, and input the semantic sample information and the phonetic symbol sample information into a text encoding model to obtain fused sample feature data.

[0106] Among them, the semantic sample information is obtained by inputting the text sample data into a semantic model, and the phonetic symbol sample information is obtained by translating the text sample data.

[0107] For example, if the text sample data is "你好" (Hello), then the phonetic symbol sample information is "n3 i3 h3 ao3".

[0108] In the embodiments of the present application, by adding semantic sample information and phonetic symbol sample information to the input text information, subtle differences can be recognized in the input text data, and a mapping relationship is constructed between the input data with these subtle differences and the speech data, so that the speech generated by the speech generation model also has subtle differences, and the pronunciation is no longer averaged, making the generated speech more natural.

[0109] Optionally, step 301 may specifically include:

[0110] Sub-step 3011: Input the text sample data into a semantic model to obtain the sample semantic feature matrix of the text sample data.

[0111] Sub-step 3012: Translate the text sample data into phonetic symbol sample data, and input the phonetic symbol sample data into the bypass guidance module of the text encoding model, and output the sample guidance information for each coupling layer in the bypass guidance module.

[0112] Sub-step 3013: Input the sample semantic feature matrix into the text encoding model, and sequentially input the sample guidance information of each coupling layer in the bypass guidance module into each convolutional layer of the text encoding model, and output the fused sample feature matrix from the text encoding model.

[0113] For sub-steps 3011 - 3013, as Figure 6 shown, by inputting the text sample data into a semantic model, the sample semantic feature matrix is obtained, and the text sample data is used to obtain the phonetic symbol sample data by looking up a vector codebook, and the phonetic symbol sample data is input into the bypass guidance module of the text encoding model, where the bypass guidance module can implement the function of mapping the phonetic symbol sample data into a normal distribution. Finally, the sample semantic feature matrix is input into the text encoding model, and the sample guidance information of each coupling layer in the bypass guidance module is sequentially input into each convolutional layer of the text encoding model, and fusion is performed according to the fusion data generation formula in sub-step 2031, and finally the fused sample feature matrix is output from the text encoding model.

[0114] In some embodiments, the bypass guidance module will eventually output a self-supervised loss value, which is used to record the difference between the data output by the last coupling layer of the bypass guidance module and the normal distribution data.

[0115] Step 302: Input the speech sample data corresponding to the text sample data into a speech coding model to obtain a speech sample feature matrix.

[0116] The length of each sentence in the text sample data and its corresponding voice sample data does not exceed 20 seconds, because long sentences longer than 20 seconds rarely appear in daily communication. Therefore, by controlling the sample data, the impact of long sentences with less meaning on the model is reduced.

[0117] In an embodiment of the present application, a linear spectrum is extracted from speech sample data through a speech coding model to obtain a speech sample feature matrix, and the speech sample data can be restored through the linear spectrum information in the speech sample feature matrix.

[0118] In some embodiments, the linear spectrum of speech can be extracted by a third-party signal processing tool, such as the torchaudio tool in the Pytorch neural network framework.

[0119] For example, the recording length of "Good Man" is 1 second. If the setting for extracting the linear spectrum at this time is "frame shift 10 milliseconds", the size of the obtained linear spectrum is 100×D, where D is the internal dimension set by yourself.

[0120] Step 303, obtaining a text sample distribution function and a speech sample distribution function by calculating the mean and variance of the fusion sample feature matrix and the speech sample feature matrix respectively;

[0121] In the embodiment of the present application, by calculating the mean μ of the speech sample feature matrix Q and variance σ Q , get the speech sample distribution function, and calculate the mean μ of the fusion sample feature matrix θ and variance σ θ , and get the text sample distribution function.

[0122] In some embodiments, the speech sample distribution function is randomly sampled before training the initial model, and a random number rand is randomly selected to calculate the random sampling result Z of the speech sample distribution function, where Z = μ Q +rand×σ Q , when the size of the linear spectrum in the previous example is 100×D, the dimension of random sampling is also 100×D. Since speech data has strong randomness, the robustness of the trained speech generation model is enhanced by random sampling, where D is the internal dimension set by oneself.

[0123] In some embodiments, Figure 6 As shown, the randomly sampled speech sample distribution function is input into the decoder to obtain a speech waveform, and the speech waveform is compared with the speech sample data to obtain a generated loss value, which is used to record the loss value of the speech sample data after the speech sample distribution function is generated.

[0124] Step 304, using the speech sample distribution function to train an initial model, and determining a loss value according to a difference between the text sample distribution function and the speech sample distribution function;

[0125] In some embodiments, the initial model is trained using the speech sample distribution function, and the initial model has a generating function f for mapping the speech sample distribution function to a normal distribution. θ , the training result of the initial model is f θ (Z), by f θ (Z) and μ θ and σ θ To do timing alignment, that is, according to the pronunciation duration of each syllable f θ (Z), μ θ and σ θ The corresponding phonetic data is copied in the corresponding number of copies, so that it can form and f θ (Z) Same dimension, continue with the previous example, μ θ and σ θ There are 4 phonetic symbols in θ The dimension of (Z) is 100, so we maintain the same dimension by copying, for example [1,1,1,1,1,1,2,2,2,3,3,3,3…4,4,4,4,4,4], and record the number of copies. Figure 6 As shown, the fused sample feature matrix is ​​input into the duration prediction model to obtain the predicted duration, which is the number of replications given by the duration prediction model. The difference between the two replications is recorded as the duration loss.

[0126] In some embodiments, the relative distance (KL, Kullback-Leibler Divergence) between the text sample distribution function and the speech sample distribution function is used as one of the loss values, where the calculation formula is:

[0127]

[0128] Among them, D is the KL distance, σ Q is the variance of the speech sample feature matrix, μ θ is the mean of the fusion sample feature matrix, σ θ is the variance of the fusion sample feature matrix, f θ (Z) is the result generated by the initial model.

[0129] Step 305: adjusting the initial model according to the loss value and the preset loss function to obtain a speech generation model, wherein the speech generation model is used for the speech generation method based on phonetic symbols and semantics.

[0130] In some embodiments, by adding the self-supervised loss value in sub-step 3012, the generated loss value in step 303, the KL distance and the market loss value in step 304 as the result of the loss function, the model is adjusted according to the loss function result, and finally a speech generation model is obtained.

[0131] In summary, in the embodiments of the present application, the semantic feature matrix in the text data and the guiding information with phonetic data characteristics are fused to obtain a fused feature matrix with semantic information and phonetic information, and the fused feature matrix is ​​input into the speech generation model to obtain the speech data corresponding to the text data. Since the input information incorporates the information of phonetic symbols and semantics, the speech corresponding to the phonetic symbols is adjusted according to different semantics, so that the speech is no longer averaged, reducing the mechanical feel. The speech generation method based on phonetic symbols and semantics in the embodiments of the present application makes full use of the information contained in the input text and can generate more realistic speech data without adding additional input data, thereby improving the effect of speech generation.

[0132] Figure 7 A speech generation device based on phonetic symbols and semantics is provided in an embodiment of the present application, such as Figure 7 As shown, the device may include:

[0133] Semantic module 401, used for inputting text data into a semantic model to obtain a semantic feature matrix of the text data;

[0134] The phonetic symbol module 402 is used to translate the text data into phonetic symbol data, and input the phonetic symbol data into the bypass guidance module of the text encoding model, and output each coupling layer in the bypass guidance module to obtain guidance information; the guidance information is used to be merged with the semantic feature matrix;

[0135] A fusion module 403 is used to input the semantic feature matrix into a text encoding model, and input the guidance information of each coupling layer in the bypass guidance module into each convolution layer of the text encoding model in sequence, and output a fusion feature matrix from the text encoding model; the coupling layer corresponds to the convolution layer one by one;

[0136] The generation module 404 is used to input the fused feature matrix into a speech generation model to obtain speech data corresponding to the text data.

[0137] Optionally, the phonetic symbol module 402 may specifically include:

[0138] The phonetic symbol submodule is used to segment the text data into phrases, and to search for the phonetic symbols mapped to the segmented phrases according to a preset phonetic symbol library to obtain phonetic symbol data; the phonetic symbol library stores the phonetic symbols mapped to the phrases;

[0139] An initial vector submodule, used to find the corresponding vector for the phonetic symbol data by searching a preset vector codebook to obtain a phonetic symbol vector matrix; the vector codebook is marked with the vector corresponding to each phonetic symbol data;

[0140] An input submodule is used to input the phonetic symbol vector matrix into the bypass guidance module.

[0141] Optionally, the phonetic symbol module 402 may specifically include:

[0142] The dimension flipping submodule is used for, for any coupling layer in the bypass guidance model, and the adjacent coupling layer of any coupling layer, while the phonetic symbol data of the first area of ​​any coupling layer is kept unchanged and the phonetic symbol data of the second area is transformed, the phonetic symbol data of the first area of ​​the adjacent coupling layer is transformed and the phonetic symbol data of the second area is not transformed; or while the phonetic symbol data of the first area of ​​any coupling layer is kept unchanged and the phonetic symbol data of the second area is not transformed, the phonetic symbol data of the first area of ​​the adjacent coupling layer is not transformed and the phonetic symbol data of the second area is transformed.

[0143] The guide information submodule is used to obtain the guide information of each layer after the phonetic symbol data in each coupling layer is transformed.

[0144] Optionally, the fusion module 403 may specifically include:

[0145] A fusion submodule, used for fusing the guidance information of each group of one-to-one corresponding coupling layers with the first data of the convolutional layer when any convolutional layer of the text encoding model is about to input data, so as to obtain fused data; for the first convolutional layer, the first data is the semantic feature matrix;

[0146] A sequential execution submodule, used for taking the fused data as the input data of the convolution layer, and outputting second data from the convolution layer; the second data is the first data of the next convolution layer, and for the last convolution layer, the second data is the feature fusion matrix;

[0147] The output submodule is used to obtain the feature fusion matrix after all convolutional layers are output in sequence.

[0148] Optionally, the fusion submodule may specifically include:

[0149] A fusion unit is used to combine the guidance information of each group of one-to-one corresponding coupling layers with the first data of the current convolutional layer in the following manner:

[0150]

[0151] Wherein O is the fused data, I is the first data of the convolutional layer, C is the guidance information of the coupling layer, mean(I) is the mean of I, std(I) is the standard deviation of I, W() and B() are two fully connected layers of the text encoding model, W(C) represents the product of the guidance information and one fully connected layer, and B(C) represents the product of the guidance information and another fully connected layer.

[0152] In summary, in the embodiments of the present application, the semantic feature matrix in the text data and the guiding information with phonetic data characteristics are fused to obtain a fused feature matrix with semantic information and phonetic information, and the fused feature matrix is ​​input into the speech generation model to obtain the speech data corresponding to the text data. Since the input information incorporates the information of phonetic symbols and semantics, the speech corresponding to the phonetic symbols is adjusted according to different semantics, so that the speech is no longer averaged, reducing the mechanical feel. The speech generation method based on phonetic symbols and semantics in the embodiments of the present application makes full use of the information contained in the input text and can generate more realistic speech data without adding additional input data, thereby improving the effect of speech generation.

[0153] Figure 8 A speech generation model training device based on phonetic symbols and semantics is provided in an embodiment of the present application, such as Figure 8 As shown, the device may include:

[0154] The sample fusion module 501 is used to extract semantic sample information and phonetic sample information from the text sample data, and input the semantic sample information and the phonetic sample information into the text encoding model to obtain fused sample feature data;

[0155] A sample speech module 502 is used to input the speech sample data corresponding to the text sample data into a speech coding model to obtain a speech sample feature matrix;

[0156] A distribution function module 503 is used to obtain a text sample distribution function and a speech sample distribution function by calculating the mean and variance of the fusion sample feature matrix and the speech sample feature matrix respectively;

[0157] A loss value module 504, configured to train an initial model using the speech sample distribution function, and determine a loss value according to a difference between the text sample distribution function and the speech sample distribution function;

[0158] The training module 505 is used to adjust the initial model according to the loss value and the preset loss function to obtain a speech generation model, and the speech generation model is used for the speech generation method based on phonetic symbols and semantics.

[0159] Optionally, the sample fusion module 501 may specifically include:

[0160] The sample semantics submodule is used to input the text sample data into a semantic model to obtain a sample semantic feature matrix of the text sample data.

[0161] A sample guide information submodule, used for translating the text sample data into phonetic sample data, inputting the phonetic sample data into a bypass guide module of the text encoding model, and outputting each coupling layer in the bypass guide module to obtain sample guide information;

[0162] The sample fusion submodule is used to input the sample semantic feature matrix into the text encoding model, and input the sample guidance information of each coupling layer in the bypass guidance module into each convolutional layer of the text encoding model in turn, and output the fused sample feature matrix from the text encoding model.

[0163] In summary, in the embodiments of the present application, the semantic feature matrix in the text data and the guiding information with phonetic data characteristics are fused to obtain a fused feature matrix with semantic information and phonetic information, and the fused feature matrix is ​​input into the speech generation model to obtain the speech data corresponding to the text data. Since the input information incorporates the information of phonetic symbols and semantics, the speech corresponding to the phonetic symbols is adjusted according to different semantics, so that the speech is no longer averaged, reducing the mechanical feel. The speech generation method based on phonetic symbols and semantics in the embodiments of the present application makes full use of the information contained in the input text and can generate more realistic speech data without adding additional input data, thereby improving the effect of speech generation.

[0164] See also Fig. 9 , the electronic device 400 may include one or more of the following components: a processing component 402 , a memory 404 , a power component 406 , a multimedia component 408 , an audio component 410 , an input / output (I / O) interface 412 , a sensor component 414 , and a communication component 416 .

[0165] The processing component 402 generally controls the overall operation of the electronic device 400, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 402 may include one or more processors 420 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 402 may include one or more modules to facilitate the interaction between the processing component 402 and other components. For example, the processing component 402 may include a multimedia module to facilitate the interaction between the multimedia component 408 and the processing component 402.

[0166] The memory 404 is used to store various types of data to support the operation of the electronic device 400. Examples of such data include instructions for any application or method operating on the electronic device 400, contact data, phone book data, messages, pictures, multimedia, etc. The memory 404 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0167] The power supply component 406 provides power to the various components of the electronic device 400. The power supply component 406 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 400.

[0168] The multimedia component 408 includes an interface that provides an output interface between the electronic device 400 and the user. In some embodiments, the interface may include a liquid crystal display (LCD) and a touch panel (TP). If the interface includes a touch panel, the interface may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 408 includes a front camera and / or a rear camera. When the electronic device 400 is in an operating mode, such as a shooting mode or a multimedia mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0169] The audio component 410 is used to output and / or input audio signals. For example, the audio component 410 includes a microphone (MIC), and when the electronic device 400 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is used to receive an external audio signal. The received audio signal can be further stored in the memory 404 or sent via the communication component 416. In some embodiments, the audio component 410 also includes a speaker for outputting audio signals.

[0170] The input / output I / O interface 412 provides an interface between the processing component 402 and the peripheral interface module, which may be a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0171] The sensor assembly 414 includes one or more sensors for providing various aspects of status assessment for the electronic device 400. For example, the sensor assembly 414 can detect the open / closed state of the electronic device 400, the relative positioning of components, such as the display and keypad of the electronic device 400, and the sensor assembly 414 can also detect the position change of the electronic device 400 or a component of the electronic device 400, the presence or absence of user contact with the electronic device 400, the orientation or acceleration / deceleration of the electronic device 400, and the temperature change of the electronic device 400. The sensor assembly 414 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 414 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 414 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0172] The communication component 416 is used to facilitate wired or wireless communication between the electronic device 400 and other devices. The electronic device 400 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 416 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 416 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0173] In an exemplary embodiment, the electronic device 400 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to implement a method for displaying a vehicle-road collaborative scenario provided in an embodiment of the present application.

[0174] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 404 including instructions, and the instructions can be executed by a processor 420 of an electronic device 400 to perform the above method. For example, the non-transitory storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0175] Fig.10 is a block diagram of an electronic device 500 according to another embodiment of the present application. For example, the electronic device 500 may be provided as a server. Fig.10 , the electronic device 500 includes a processing component 522, which further includes one or more processors, and a memory resource represented by a memory 532 for storing instructions executable by the processing component 522, such as an application. The application stored in the memory 532 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 522 is configured to execute instructions to execute a method for displaying a vehicle-road cooperative scenario provided in an embodiment of the present application.

[0176] The electronic device 500 may also include a power supply component 526 configured to perform power management of the electronic device 500, a wired or wireless network interface 550 configured to connect the electronic device 500 to a network, and an input / output (I / O) interface 558. The electronic device 500 may operate based on an operating system stored in the memory 532, such as Windows Server TM, Mac OS X TM, Unix TM, Linux TM, Free BSD TM or the like.

[0177] In an embodiment of the present application, the memory 632 may be used to store software programs and various data. The memory 632 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory 632 may include a volatile memory or a non-volatile memory, or the memory 632 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM). The memory 632 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0178] The processor may include one or more processing units; optionally, the processor integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to the operating system, user interface, and application programs, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It is understandable that the modem processor may not be integrated into the processor.

[0179] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned image super-resolution reconstruction method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0180] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0181] An embodiment of the present application also provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the super-resolution reconstruction method embodiment of the above-mentioned image, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0182] It should be noted that the various information and data obtained in the embodiments of the present application are obtained with the authorization of the information / data holder. All actions of obtaining signals, information or data in the present application are carried out in compliance with the relevant data protection laws and policies of the country of residence and with the authorization of the corresponding device owner.

[0183] The algorithm and display provided herein are not inherently related to any particular computer, virtual system or other device. Various general purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious that the structure required for constructing such systems. In addition, the application is not directed to any specific programming language either. It should be understood that various programming languages ​​can be utilized to realize the content of the application described herein, and the description of the specific language above is for the purpose of disclosing the best mode of implementation of the application.

[0184] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.

[0185] Similarly, it should be understood that in order to streamline the present application and help understand one or more of the various application aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be interpreted as reflecting the following intention: the claimed application requires more features than the features clearly stated in each claim. More specifically, as reflected in the claims below, the application aspects are less than all the features of the single embodiment disclosed above. Therefore, the claims following the specific embodiment are hereby expressly incorporated into the specific embodiment, wherein each claim itself serves as a separate embodiment of the present application.

[0186] Those skilled in the art will appreciate that the modules in the devices in the embodiments may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments may be combined into one module or unit or component, and in addition they may be divided into a plurality of submodules or subunits or subcomponents. Except that at least some of such features and / or processes or units are mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this manner may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature providing the same, equivalent or similar purpose.

[0187] The various component embodiments of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all functions of some or all components in the sorting device according to the present application. The present application can also be implemented as a device or apparatus program for executing part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0188] It should be noted that the above embodiments illustrate the present application rather than limit the present application, and that those skilled in the art may design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets should not be constructed as a limitation to the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of multiple such elements. The present application may be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim that lists several devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order. These words may be interpreted as names.

[0189] The user information (including but not limited to the user's device information, user personal information, etc.) and related data involved in this application are all information authorized by the user or by all parties.

[0190] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0191] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.

[0192] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A speech generation method based on phonetic symbols and semantics, characterized in that: The method comprises: Inputting text data into a semantic model to obtain a semantic feature matrix of the text data; Translating the text data into phonetic symbol data, inputting the phonetic symbol data into a bypass guidance module of a text encoding model, outputting each coupling layer in the bypass guidance module, and obtaining guidance information; the guidance information is used to be fused with the semantic feature matrix; The semantic feature matrix is ​​input into a text encoding model, and the guidance information of each coupling layer in the bypass guidance module is sequentially input into each convolutional layer of the text encoding model, and a fusion feature matrix is ​​output from the text encoding model; the coupling layer corresponds to the convolutional layer one by one; The fused feature matrix is ​​input into a speech generation model to obtain speech data corresponding to the text data.

2. The method according to claim 1, characterized in that The bypass guidance module for translating the text data into phonetic symbol data and inputting the phonetic symbol data into a text encoding model includes: Segment the text data into phrases, and search for phonetic symbols mapped to the segmented phrases according to a preset phonetic symbol library to obtain phonetic symbol data; the phonetic symbol library stores the phonetic symbols mapped to the phrases; By searching a preset vector codebook, searching for a corresponding vector for the phonetic symbol data, and obtaining a phonetic symbol vector matrix; the vector codebook is marked with a vector corresponding to each phonetic symbol data; The phonetic symbol vector matrix is ​​input into the bypass guidance module.

3. The method according to claim 1, characterized in that Each coupling layer of the bypass guiding module is divided into a first area and a second area, and the bypass guiding module that inputs the phonetic symbol data into the text encoding model outputs each coupling layer in the bypass guiding module to obtain guiding information, including: For any coupling layer in the bypass guidance model, and the adjacent coupling layer of any coupling layer, while the phonetic symbol data of the first area of ​​any coupling layer is kept unchanged and the phonetic symbol data of the second area is changed, the phonetic symbol data of the first area of ​​the adjacent coupling layer is changed and the phonetic symbol data of the second area is not changed; or while the phonetic symbol data of the first area of ​​any coupling layer is kept unchanged and the phonetic symbol data of the second area is not changed, the phonetic symbol data of the first area of ​​the adjacent coupling layer is not changed and the phonetic symbol data of the second area is changed; After the phonetic symbol data in each coupling layer are transformed, the guidance information of each layer is obtained.

4. The method according to claim 1, characterized in that: The coupling layers in the bypass guidance module have the same number of convolutional layers as those in the text encoding model and are both even-numbered layers.

5. The method according to claim 1, characterized in that The step of inputting the semantic feature matrix into the text encoding model, and sequentially inputting the guidance information of each coupling layer in the bypass guidance model into each convolutional layer of the text encoding model to obtain a fused feature matrix includes: When any convolutional layer of the text encoding model is about to input data, the guidance information of each group of one-to-one corresponding coupling layers and the first data of the convolutional layer are fused to obtain fused data; for the first convolutional layer, the first data is the semantic feature matrix; The fused data is used as the input data of the convolution layer, and second data is output from the convolution layer; the second data is the first data of the next convolution layer, and for the last convolution layer, the second data is the feature fusion matrix; When all convolutional layers are output in sequence, the feature fusion matrix is ​​obtained.

6. The method according to claim 5, characterized in that The step of fusing each group of one-to-one corresponding guidance information of the coupling layer with the first data of the current convolutional layer to obtain fused data includes: The guidance information of each group of one-to-one corresponding coupling layers and the first data of the current convolutional layer are combined as follows: Wherein O is the fused data, I is the first data of the convolutional layer, C is the guidance information of the coupling layer, mean(I) is the mean of I, std(I) is the standard deviation of I, W() and B() are two fully connected layers of the text encoding model, W(C) represents the product of the guidance information and one fully connected layer, and B(C) represents the product of the guidance information and another fully connected layer.

7. A speech generation model training method based on phonetic symbols and semantics, characterized in that: The method comprises: Extracting semantic sample information and phonetic sample information from the text sample data, and inputting the semantic sample information and the phonetic sample information into a text encoding model to obtain fused sample feature data; Inputting the speech sample data corresponding to the text sample data into a speech coding model to obtain a speech sample feature matrix; By calculating the mean and variance of the fusion sample feature matrix and the speech sample feature matrix, a text sample distribution function and a speech sample distribution function are obtained; Using the speech sample distribution function to train an initial model, and determining a loss value according to a difference between the text sample distribution function and the speech sample distribution function; According to the loss value and the preset loss function, the initial model is adjusted to obtain a speech generation model, and the speech generation model is used for the speech generation method based on phonetic symbols and semantics as described in any one of claims 1 to 6.

8. The method according to claim 7, characterized in that The extracting of semantic sample information and phonetic sample information from the text sample data, and inputting the semantic sample information and phonetic sample information into a text encoding model to obtain fused sample feature data includes: Inputting the text sample data into a semantic model to obtain a sample semantic feature matrix of the text sample data; Translating the text sample data into phonetic symbol sample data, inputting the phonetic symbol sample data into a bypass guidance module of the text encoding model, and outputting each coupling layer in the bypass guidance module to obtain sample guidance information; The sample semantic feature matrix is ​​input into a text encoding model, and the sample guidance information of each coupling layer in the bypass guidance module is sequentially input into each convolutional layer of the text encoding model, and a fused sample feature matrix is ​​output from the text encoding model.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A readable storage medium, characterized in that: When the instructions in the readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method as claimed in any one of method claims 1 to 8.

Citation Information

Patent Citations

  • Image recognition method and device and storage medium

    CN110490213A

  • Text-to-voice method and device, electronic equipment and storage medium

    CN112820269A

  • Speech synthesis method and apparatus, storage medium, and electronic device

    WO2022095754A1

  • Text-to-speech conversion method and apparatus, electronic device, and storage medium

    WO2022142105A1