Phoneme alignment model training and speech synthesis method and device, equipment and medium
By generating a hard attention matrix through convolutional attention alignment and monotonic alignment search, and combining the relative entropy loss and Mel loss training model, the problem of alignment error between phonemes and speech frames in speech synthesis is solved, thereby improving the accuracy and naturalness of speech synthesis.
Patent Information
- Application Number
- CN202510882019.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-27
AI Technical Summary
In existing speech synthesis technology, the alignment error between phonemes and speech frames is particularly significant at the junction of consonants and vowels, leading to problems such as unclear pronunciation, swallowed consonants, and pronunciation deviation, affecting the naturalness and intelligibility of the speech.
By obtaining speech training data, extracting spectral feature information and text feature information, performing convolutional attention alignment, generating the first alignment matrix, and performing monotonic alignment search to generate a binary hard attention matrix, the relative entropy loss and Mel loss are calculated, and the phoneme alignment model is trained until the preset convergence conditions are met.
The accuracy of phoneme alignment is improved, the phenomenon of silent frames being incorrectly assigned to phoneme areas is reduced, and the accuracy and naturalness of speech synthesis are improved.
Smart Images

Figure CN120748366A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech synthesis technology, and in particular to a phoneme alignment model training and speech synthesis method, apparatus, computer equipment, storage medium and computer program product. Background Art
[0002] In speech synthesis technology, text information needs to be converted into a phoneme sequence, then a spectrogram is generated, and finally a speech signal is synthesized. To ensure that the synthesized speech strictly corresponds to the original text content in time, a clear alignment relationship must be established between phonemes and speech frames. Therefore, the quality of phoneme alignment directly determines the speech rhythm, pronunciation clarity, and the accuracy of semantic expression.
[0003] At present, the alignment method in related technologies mainly uses attention weights to automatically learn the relationship between speech and text, and then generates an alignment path for the two. Although it solves the alignment problem between speech and text to a certain extent, in actual applications, the methods provided by related technologies are prone to significant alignment errors at phoneme boundaries, especially at the junction of consonants and vowels. Due to the drastic energy changes at these positions and the complex characteristics of the speech frames, it is difficult for the model to accurately judge the boundaries between phonemes, which can easily lead to phenomena such as unclear pronunciation, swallowed consonants, and pronunciation deviation. In addition, at the beginning and end of the speech segment, the transition area between the silent frame and the valid pronunciation frame is often misjudged, resulting in the first phoneme being pronounced prematurely or the last phoneme being stretched to the silent segment. This misalignment has a significant impact on the naturalness and intelligibility of the speech.
[0004] It can be seen that the relevant technology has problems such as large alignment errors between consonant and vowel boundaries and misplacement of silent segments during speech synthesis. The accuracy of its speech synthesis is not enough to cope with the changing actual application scenarios. Summary of the Invention
[0005] Based on this, it is necessary to provide a phoneme alignment model training and speech synthesis method, device, computer equipment, storage medium and computer program product that can improve the accuracy of phoneme alignment and thus improve the accuracy of speech synthesis in response to the above technical problems.
[0006] In a first aspect, the present application provides a phoneme alignment model training method, comprising:
[0007] Acquire speech training data, and obtain spectral feature information and text feature information based on the speech training data; wherein the spectral feature information includes a plurality of mel-spectrogram frames, and the text feature information includes a plurality of phonemes;
[0008] Performing convolutional attention alignment on the spectral feature information and the text feature information to obtain a first alignment matrix;
[0009] Performing a monotonic alignment search based on the first alignment matrix to generate a second alignment matrix, where the second alignment matrix is a binary hard attention matrix;
[0010] calculating a relative entropy loss based on the first alignment matrix and the second alignment matrix;
[0011] Expanding the text feature information to a Mel spectrum frame length according to the second alignment matrix, and performing a linear transformation on the expanded text feature information to generate a predicted Mel spectrum;
[0012] Calculating the Mel loss according to the spectrum feature information and the predicted Mel spectrum;
[0013] The phoneme alignment model is trained according to the relative entropy loss and the Mel loss until a preset convergence condition is met.
[0014] In one embodiment, the method further comprises:
[0015] Calculating the first alignment matrix based on a logarithmic probability distribution function to obtain a phoneme probability sequence; the phoneme probability sequence is a logarithmic value of the phoneme classification probability corresponding to each mel spectrum frame;
[0016] Calculating a connection temporal classification loss based on the text feature information and the phoneme probability sequence;
[0017] The phoneme alignment model is trained based on the connection temporal classification loss, the relative entropy loss, and the Mel loss until the preset convergence condition is met.
[0018] In one embodiment, the method further comprises:
[0019] Generate a predicted Mel band probability based on the text feature information; the predicted Mel band probability includes a classification probability distribution of the Mel band corresponding to each phoneme;
[0020] Calculating the cross entropy loss between the predicted Mel band probability and the true Mel spectrum;
[0021] The phoneme alignment model is trained based on the cross entropy loss, the connection temporal classification loss, the relative entropy loss, and the Mel loss until the preset convergence condition is met.
[0022] In one embodiment, generating the predicted Mel-band probability based on the text feature information includes:
[0023] Mapping the text feature information into the predicted Mel spectrum through a linear projection head;
[0024] The predicted Mel frequency spectrum is calculated based on the logarithmic probability distribution function to obtain a predicted Mel frequency band probability.
[0025] In one embodiment, the step of extending the text feature information to a mel spectrum frame length according to the second alignment matrix, and performing a linear transformation on the extended text feature information to generate a predicted mel spectrum includes:
[0026] extracting the duration of each phoneme according to the second alignment matrix;
[0027] Expanding the text feature information to a Mel spectrum frame length according to the duration;
[0028] The expanded text feature information is mapped to the Mel spectrum frame dimension through the Mel projection head to generate a predicted Mel spectrum.
[0029] In one embodiment, performing convolutional attention alignment on the spectral feature information and the text feature information to obtain a first alignment matrix includes:
[0030] A first alignment matrix is generated based on negative Euclidean distance calculation between the text feature information and the spectrum feature information.
[0031] In a second aspect, the present application provides a speech synthesis method, comprising:
[0032] Get the text information to be processed;
[0033] Processing the text information to be processed using the phoneme alignment model described in the first aspect to obtain target spectrum information;
[0034] A target speech signal is obtained according to the target spectrum information.
[0035] In a third aspect, the present application further provides a phoneme alignment model training device, comprising:
[0036] A data acquisition module, configured to acquire speech training data and obtain spectral feature information and text feature information based on the speech training data; wherein the spectral feature information includes a plurality of mel-spectrogram frames, and the text feature information includes a plurality of phonemes;
[0037] A data alignment module, configured to perform convolutional attention alignment on the spectral feature information and the text feature information to obtain a first alignment matrix;
[0038] A first processing module is configured to perform a monotonic alignment search based on the first alignment matrix to generate a second alignment matrix, where the second alignment matrix is a binary hard attention matrix;
[0039] a first calculation module, configured to calculate a relative entropy loss according to the first alignment matrix and the second alignment matrix;
[0040] A second processing module is configured to expand the text feature information to a Mel spectrum frame length according to the second alignment matrix, and perform a linear transformation on the expanded text feature information to generate a predicted Mel spectrum;
[0041] A second calculation module is used to calculate the Mel loss according to the spectrum feature information and the predicted Mel spectrum;
[0042] The model training module is used to train the phoneme alignment model according to the relative entropy loss and the Mel loss until a preset convergence condition is met.
[0043] In a fourth aspect, the present application provides a speech synthesis device, comprising:
[0044] An information acquisition module is used to obtain text information to be processed;
[0045] An information processing module, configured to process the text information to be processed using the phoneme alignment model described in the first aspect to obtain target spectrum information;
[0046] The speech synthesis module is used to obtain a target speech signal according to the target spectrum information.
[0047] In a fifth aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps described in the first aspect or the second aspect when executing the computer program.
[0048] In a sixth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps described in the first aspect or the second aspect when executed by a processor.
[0049] In a seventh aspect, the present application further provides a computer program product, comprising a computer program, which implements the steps described in the first aspect or the second aspect when executed by a processor.
[0050] The above-described phoneme alignment model training method, apparatus, computer device, storage medium, and computer program product obtain speech training data and, based on the speech training data, obtain spectral feature information comprising multiple mel-spectrogram frames and text feature information comprising multiple phonemes. The spectral feature information and text feature information are then convolutionally aligned to obtain a first alignment matrix. The first alignment matrix can express the model's natural preference for boundary locations. Based on this, a monotonic alignment search is further performed based on the first alignment matrix to generate a second alignment matrix. The second alignment matrix is a binarized hard attention matrix. The binarization process discretizes the first alignment matrix, thereby clearly identifying the specific phoneme to which each frame should be aligned, ensuring that the alignment path is strictly monotonic in structure. Next, a relative entropy loss is calculated based on the first and second alignment matrices to measure the difference between the two. This encourages the hard alignment path to adhere as closely as possible to the model's perception of boundaries while maintaining a clear structure, thereby minimizing boundary distortion caused by alignment path offset during the binarization process. This results in more reasonable time allocation at phoneme boundaries (such as the boundary between consonants and vowels), alleviating boundary misalignment issues such as early or late alignment. Subsequently, the text feature information is expanded to the Mel spectrum frame length according to the second alignment matrix, and the expanded text feature information is linearly transformed to generate a predicted Mel spectrum. The Mel loss is then calculated based on the spectral feature information and the predicted Mel spectrum, thereby optimizing the spectral reconstruction quality and further strengthening the temporal correspondence between frames and phonemes. This inverse optimization of the spectral reconstruction loss can effectively reduce the phenomenon of silent frames being incorrectly assigned to phoneme regions, and improve the alignment accuracy of the model in silent transition regions such as the beginning and end of sentences. Finally, the phoneme alignment model is trained using relative entropy loss and Mel loss until a preset convergence condition is met, thereby improving the accuracy of phoneme alignment. Furthermore, the above-mentioned speech synthesis method, apparatus, computer device, storage medium, and computer program product obtain the text information to be processed, process the text information using the above-mentioned phoneme alignment model to obtain target spectral information, and then obtain the target speech signal based on the target spectral information. The above-mentioned phoneme alignment model can be applied to actual speech synthesis scenarios, and can achieve higher accuracy of the target speech signal under complex requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0052] Figure 1A diagram illustrating an application environment of a phoneme alignment model training method according to an embodiment;
[0053] Figure 2 1 is a flow chart of a method for training a phoneme alignment model in one embodiment;
[0054] Figure 3 A flowchart of a phoneme alignment model training method using multi-loss joint training in one embodiment;
[0055] Figure 4 A flowchart of a phoneme alignment model training method using multi-loss joint training in another embodiment;
[0056] Figure 5 is a structural block diagram of a phoneme alignment model training device in one embodiment;
[0057] Figure 6 is a diagram of the internal structure of a computer device in one embodiment;
[0058] Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in another embodiment. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0060] The phoneme alignment model training method and speech synthesis method provided in the embodiments of the present application can be applied to Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. Terminal 102 can be used to obtain speech training data and send the speech training data to server 104 for training a phoneme alignment model. Terminal 102 can also be used to obtain text information to be processed and send the text information to server 104 for speech synthesis, thereby receiving and outputting the target speech signal processed by server 104. The terminal 102 used to obtain speech training data and the terminal 102 used to obtain text data to be processed can be the same terminal device or different terminal devices. A data storage system can store data that server 104 needs to process. The data storage system can be integrated with server 104 or placed in the cloud or other network servers. The data storage system can be used to store data such as speech training data. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart car devices, etc. Portable wearable devices can include smart watches, smart bracelets, head-mounted devices, etc. Server 104 can be implemented as a standalone server or a server cluster consisting of multiple servers.
[0061] In an exemplary embodiment, Figure 2 As shown, a phoneme alignment model training method is provided, which is applied to Figure 1 The server 104 in the example is used for explanation, including the following steps S202 to S214. It is understandable that the phoneme alignment model training method can also be applied to the terminal 102. Among them:
[0062] Step S202: Acquire speech training data, and obtain spectrum feature information and text feature information based on the speech training data.
[0063] The spectral feature information includes a plurality of mel-spectrogram frames, and the text feature information includes a plurality of phonemes. Exemplarily, the speech training data includes original speech information and original text information.
[0064] For example, server 104 may first obtain training samples containing original speech and text from the speech training dataset obtained by terminal 102. The speech training data may include one or more standard speech audio files and their corresponding text content. For example, a piece of training data may be an original speech information file, and its accompanying original text message may be "Hello, today is Wednesday." Based on this data, server 104 may further extract structured feature information from the speech and text, namely, spectral feature information and text feature information, which can be used for model learning.
[0065] Among them, spectral feature information refers to the conversion of the original speech signal into a feature representation with dual dimensions of time and frequency through acoustic processing methods. The specific form can be a Mel Spectrogram. Mel Spectrogram is a spectrum transformation method that is more sensitive to human ear perception. It performs a short-time Fourier transform on the original waveform and then maps it to the Mel scale to express the frequency energy distribution of the audio at different time points. Server 104 can divide a segment of speech into frames to obtain multiple Mel Spectrogram frames. Each frame can be understood as the sound performance within a certain time window in the speech. For example, the audio "hello" may be divided into 80 frames after Mel transformation. Each frame represents a 12.5 millisecond speech signal, which lasts for about 1 second in total and constitutes the spectral feature information of the sample.
[0066] At the same time, the server 104 can pre-process the text information corresponding to the audio and extract text feature information. Text feature information mainly refers to the phoneme sequence, which is the smallest phonetic unit that constitutes the pronunciation of the text. Phonemes are the basic pronunciation units of a language. For example, the Chinese pinyin "nǐ hǎo" can be divided into the phoneme sequence [n, i, h, ao]. The server 104 can perform standardization, word segmentation, pinyin transcription and phoneme annotation on the original text through the text analysis module to convert the natural language text into a machine-processable phoneme number sequence. This process can be implemented using a Chinese phoneme annotation dictionary or a method based on preset rules, and finally a phoneme list is obtained, such as [5, 13, 42, 8], where each number corresponds to a specific phoneme meaning.
[0067] Exemplarily, the server 104 may extract audio change features from the original speech information to obtain spectral feature information; determine a phoneme number sequence based on the original text information; input the phoneme number sequence into a phoneme encoder to obtain a phoneme encoding vector, and use the phoneme encoding vector as text feature information.
[0068] Among them, audio change features may include frequency change features and rate change features. Server 104 can extract audio change features from the original speech information. These features may include two dimensions: one is the change in frequency over time itself, namely the frequency change feature; the other is the speed of frequency change, namely the rate change feature. Both serve as important signals that reveal temporal structures in speech, such as boundary changes, speech rate transitions, and silence transitions. For example, server 104 can extract the frequency change feature of the audio frequency over time from the original speech information; extract the rate of change of the audio frequency over time from the original speech information to obtain the rate change feature; and obtain spectral feature information including first-order difference features based on the frequency change feature and the rate change feature. Specifically, server 104 can perform a short-time Fourier transform (STFT) on the original speech waveform and map it to the Mel scale to obtain a basic spectral frame sequence as the basis for the frequency change feature. On this basis, the server 104 can further perform a differential calculation on these spectrum frame sequences in the frame sequence direction, that is, perform a vector difference operation on each frame spectrum with its previous frame to obtain the time derivative of the spectrum, thereby capturing the speed of change of speech in the frequency dimension over time. This differential result is the rate change feature, which can also be called the first-order difference or delta feature.
[0069] For example, if the Mel spectrum of a certain frame is [0.3, 0.4, 0.6, ...], and the next frame is [0.35, 0.38, 0.58, ...], then the subtraction of the two is [+0.05, -0.02, -0.02, ...], which represents whether the energy in these frequency bands is increasing or decreasing, and how fast the change is. This type of information changes more dramatically at the boundaries between phonemes, especially between consonants and vowels, or when switching from silent segments to voiced segments, and can significantly reveal the changing trend of the sound structure. Server 104 can combine the first-order difference feature with the original Mel spectrum as the final spectral feature information to enhance the model's sensitivity to boundaries.
[0070] On the other hand, the server 104 can perform structured processing on the original text information to determine the phoneme number sequence. This process can include standardizing the original text and converting it into pinyin, and then mapping the pinyin into a set of phoneme ID numbers according to the phoneme dictionary. Taking the phrase "hello" as an example, the server 104 can decompose it into phonemes [n, i, h, ao] and assign numbers according to the phoneme dictionary, such as [12, 4, 19, 31], to form a clear phoneme sequence. Subsequently, the server 104 can input the above phoneme number sequence into a phoneme encoder for modeling. The phoneme encoder can include an embedding layer and multiple one-dimensional convolutional layers, which map each phoneme number into a high-dimensional semantic vector through learning and capture the contextual connection between phonemes. For example, the number "12" represents the phoneme / n / , which becomes a 512-dimensional vector [0.21, -0.04, 0.38, ...] after embedding. After convolution to extract the context, the final phoneme encoding vector is formed. The server 104 may combine the encoding vectors of all phonemes into an ordered sequence as the above-mentioned text feature information.
[0071] In this step, server 104 not only extracts basic features from the spectrum and phonemes, but also incorporates dynamic information about audio variation trends (such as delta features) and the ability to express context between phonemes. This provides a more robust temporal foundation and semantic expression capabilities for subsequent attention alignment and duration prediction. Especially for speech passages with drastic changes in speech rate or blurred boundaries, the variation features introduced in this step can effectively improve the model's recognition accuracy and alignment robustness for boundary regions.
[0072] Step S204: Perform convolutional attention alignment on the spectral feature information and the text feature information to obtain a first alignment matrix.
[0073] The first alignment matrix represents the alignment between mel-spectrogram frames and phonemes. Spectral feature information can be represented as [B, T, d], and text feature information can be represented as [B, L, d]. The first alignment matrix can be represented as [B, T, L], where B is the batch size, T is the number of spectral frames (e.g., the number of time steps in each speech signal), L is the phoneme sequence length, and d is the feature dimension.
[0074] For example, server 104 can input spectral feature information as a query and text feature information as a key into two sets of one-dimensional convolutional networks with identical structures but independent parameters. Each set of convolutional networks uses multiple one-dimensional convolution kernels to slide along the time dimension to extract local contextual information of the features. After convolution, the spectral feature information [B, T, d] is enhanced to obtain the query expression [B, T, d], while the text feature information [B, L, d] is enhanced to obtain the key expression [B, L, d] through another set of convolution operations. While preserving the temporal structure of the original features, the receptive field of the convolution kernel captures the correlation patterns between local phonemes or spectra.
[0075] Next, server 104 can calculate the similarity between the processed query and key. For example, server 104 can generate a first alignment matrix based on the negative Euclidean distance calculation of the text feature information and the spectral feature information. For example, server 104 can use a broadcast mechanism to expand the query and key into a three-dimensional tensor that can be calculated item by item, and calculate the sum of squared differences between the two along the feature dimension d, resulting in a distance matrix of shape [B, T, L].
[0076] Furthermore, to convert the distance matrix into a probabilistic attention matrix, server 104 can then negate it and normalize it using a softmax function. The softmax operation can be performed on L (the phoneme dimension) so that the sum of all phoneme matching scores for each frame is 1. The normalized output is the first alignment matrix [B, T, L], in which each element represents the attention weight of a frame's spectral features on a particular phoneme, i.e., the normalized matching probability. For example, a row [0.65, 0.3, 0.05, 0] indicates that the spectral frame is primarily determined by the first phoneme, but may also be related to the second phoneme.
[0077] The final output, the first alignment matrix, not only possesses probabilistic distribution properties, facilitating subsequent loss function calculations and hard alignment path generation, but also possesses differentiability, serving as a crucial bridge for backpropagation optimization. This matrix preserves the temporal characteristics of the original input while expressing the matching tendency of each frame for each phoneme. This provides a high-resolution probabilistic matrix foundation for constructing a stable, monotonic, and semantically consistent phoneme alignment path.
[0078] In addition, the server 104 can also use the cosine distribution scheme to calculate the first alignment matrix. The server 104 can first perform L2 normalization on each Mel-spectrogram frame vector (Query) and each phoneme vector (Key) so that they are all on the unit sphere, and then calculate the dot product value between them as the similarity score. The value range of the cosine similarity is [–1,1], which indicates the degree of consistency in the direction of the two vectors. Subsequently, the server 104 can normalize these similarity scores through the softmax operation to obtain a first alignment matrix of the shape of [B, T, L], in which each row represents the probability distribution of the attribution of the frame to all phonemes. Cosine similarity pays more attention to direction rather than amplitude, so it is easier to maintain the stability of alignment when faced with samples of different pronunciation intensities or large amplitude changes. It is especially suitable for situations where the boundaries between phonemes are blurred and the energy changes drastically.
[0079] Alternatively, server 104 can employ a multi-head attention mechanism to calculate the first alignment matrix, splitting each input vector into multiple subspace representations and independently performing attention matching on each. Server 104 can split the input spectral features [B, T, d] and text features [B, L, d] into h subvectors, for example, splitting d = 512 into 8 subheads, each with 64 dimensions. Attention scores are then calculated for each subvector, resulting in multiple [B, T, L] alignment graphs. These subgraphs capture the frame-phoneme relationships in different representation subspaces, and the server then performs a weighted fusion of them to produce the final first alignment matrix. The multi-head attention mechanism allows the model to understand the alignment relationship between frames and phonemes from multiple perspectives simultaneously, thereby constructing a more robust and diverse attention graph. This is suitable for synthesis scenarios such as long sentences, complex speech rates, and multiple speaker styles, significantly reducing instances where a single attention head fails to align or misaligns.
[0080] Step S206 : performing a monotonic alignment search based on the first alignment matrix to generate a second alignment matrix.
[0081] Among them, the second alignment matrix is a binary hard attention matrix.
[0082] For example, server 104 can perform a monotonic alignment search (MAS) based on a first alignment matrix. This alignment matrix is a probability distribution tensor of shape [B, T, L], where each value represents the probability of matching a frame's spectrum across all phonemes. During the MAS process, server 104 can construct a dynamic programming table D, searching for the path with the highest cumulative probability along the time dimension T and the phoneme dimension L according to a monotonically increasing path rule. This algorithm adheres to the monotonic constraint that frames must be matched to phonemes sequentially and cannot jump back to the previous phoneme, ensuring that the generated alignment path conforms to the temporal pattern of phoneme emission frame by frame in natural speech.
[0083] After the search is complete, server 104 can backtrack from the dynamic programming table to find the path with the highest probability and generate a second alignment matrix based on it. This matrix still has the shape [B, T, L], but is strictly binary, with only one element in each row being 1 and the rest being 0, clearly indicating the unique phoneme corresponding to each frame. This binary structure is the hard attention matrix, which no longer retains the probabilistic uncertainty of the first alignment matrix. Instead, it constructs the precise distribution of phonemes along the time axis based on the clear frame-phoneme attribution relationship.
[0084] Step S208: Calculate relative entropy loss according to the first alignment matrix and the second alignment matrix.
[0085] For example, server 104 can calculate the relative entropy loss (Kullback-Leibler divergence, KL Loss) based on the first and second alignment matrices obtained in the previous steps. The first alignment matrix is a probabilistic soft attention matrix A_soft calculated by the model via the convolutional attention mechanism. Its shape is [B, T, L]. Each frame has a normalized probability representation for each phoneme, is differentiable, and can participate in gradient calculations. The second alignment matrix is a binary matrix generated via monotone alignment search (MAS), also known as the hard attention matrix A_hard. It is also [B, T, L], but only one position in each row is 1, with the rest being 0, indicating the unique phoneme corresponding to each frame. Since the second alignment matrix is inherently discrete, to ensure its participation in backpropagation, server 104 can relax this binary matrix using the Gumbel-Softmax mechanism to make it approximately differentiable. The server 104 may use the first alignment matrix as the “prediction distribution” and the relaxed second alignment matrix as the “target distribution”, and perform a frame-by-frame and phoneme-by-phoneme KL divergence calculation on the two, thereby measuring the difference between the current attention output of the model and the ideal hard alignment path.
[0086] Furthermore, the loss function can be expressed as:
[0087]
[0088] KL divergence is used to measure the information redundancy or offset of the model's current predicted distribution compared to the target distribution. The smaller the value, the closer the model is to the ideal alignment path. For example, if the distribution of a frame in the first alignment matrix is [0.7, 0.2, 0.1], and the corresponding second alignment matrix is [1, 0, 0], then KL Loss will penalize the non-maximum probability part, guiding the model to focus further on the correct phonemes in the next iteration. During the training process, the server 104 does not directly replace the output of the soft alignment matrix with the hard alignment matrix, but uses it as a learning target through KL divergence, thereby retaining the model's tolerance for fuzzy boundary areas, and at the same time guiding the model's convergence direction through the monotonic path provided by hard alignment to avoid attention drift or jump mismatches. Through relative entropy loss, the server 104 can guide the model to continuously approach the ideal path in each round of training, thereby improving the accuracy of frame-level attention distribution. Especially in passages where the boundaries between consonants and vowels are blurred, pronunciation changes dramatically, or there are silent transitions, the guidance of the soft alignment matrix combined with the structure of the hard alignment matrix enables the model to have clear and definite boundary judgment capabilities while maintaining flexibility, fundamentally reducing problems such as phoneme pronunciation misalignment, boundary sliding, and silent frame crowding.
[0089] Step S210 : Expand the text feature information to the Mel spectrum frame length according to the second alignment matrix, and perform a linear transformation on the expanded text feature information to generate a predicted Mel spectrum.
[0090] For example, server 104 can expand the text feature vector from phoneme granularity to frame-level granularity based on the binarized second alignment matrix, so that it is aligned with the spectrum in the time dimension. Specifically, server 104 can copy the corresponding phoneme vector to the position of the frame where the "1" is located in each frame in the second alignment matrix, and ultimately construct a frame-level text feature sequence with a shape of [B, T, d], that is, each frame has a high-dimensional semantic representation of the corresponding phoneme. This expansion process ensures that the phoneme expression is strictly aligned with each Mel-spectrogram frame in time, providing structural support for subsequent spectrum prediction.
[0091] After constructing the frame-level text features, server 104 can input them into a linear transformation layer, namely a fully connected projection layer, to map the frame-level phoneme vector from the encoding dimension d to the Mel-spectrogram dimension n_mels (e.g., 80). The output shape is [B, T, n_mels], which is the predicted Mel-band energy value for each frame spectrum. This output is the predicted Mel-spectrogram. Through this linear mapping, server 104 achieves the transformation from structured text semantics to specific speech energy distribution, enabling the model to have actual speech generation capabilities. Through the above steps, based on the completion of the soft-to-hard alignment transition, the phoneme information is expanded to the full speech frame length, and the spectrum corresponding to the real speech is predicted.
[0092] Exemplarily, the server 104 may extract the duration of each phoneme according to the second alignment matrix; expand the text feature information to the Mel spectrum frame length according to the duration; map the expanded text feature information to the Mel spectrum frame dimension through a Mel projection head to generate a predicted Mel spectrum.
[0093] Exemplarily, the server 104 can extract the duration of each phoneme from the second alignment matrix. The second alignment matrix is a binary attention matrix of shape [B, T, L], in which each row represents a unique phoneme position corresponding to a frame, and "1" in the matrix indicates that the frame belongs to a certain phoneme, and "0" indicates that it does not belong. The server traverses the matrix and counts the total number of values "1" in each column, that is, the number of frames assigned to each phoneme in the speech segment. The statistical result can be organized into a phoneme duration vector with a shape of [B, L], in which each element represents the frame-level duration of the corresponding phoneme. For example, if the phoneme " / a / " is marked with 10 frames in the matrix, then its duration value is 10.
[0094] After obtaining the phoneme duration vector, the server 104 can expand the text feature information in the time dimension. The original shape of the text feature information is [B, L, d], which represents the semantic vector corresponding to each phoneme. According to the duration of each phoneme, the server copies the corresponding vector a corresponding number of times in the time dimension, thereby obtaining a frame-level text vector sequence [B, T, d], where T is the total number of Mel spectrum frames. For example, if the phoneme " / t / " lasts for 4 frames, its encoding vector will be copied 4 times and filled in 4 consecutive positions of the generated sequence. Next, the server 104 can input the expanded frame-level text features into the Mel projection head. The Mel projection head is essentially a fully connected layer or a set of linear transformation modules. Its function is to map the input d-dimensional phoneme representation vector to an n_mels-dimensional Mel spectrum space (such as 80 dimensions). The output result is in the shape of [B, T, n_mels], which represents the Mel spectrum energy distribution corresponding to each frame. Through the above steps, the server 104 can obtain a complete predicted Mel spectrum, which is based on text semantic information and is expanded by phoneme duration control. It not only has language content but also is strictly aligned with the actual pronunciation rhythm.
[0095] In addition, the server 104 can also splice the expanded text feature information with the speaker's timbre vector. The timbre vector can be extracted by the server according to the speaker number, indicating the pronunciation style and voice characteristics corresponding to the voice sample. By splicing the timbre information, the server 104 can provide a more complete acoustic condition input, so that the subsequently generated spectrum results can accurately reflect the timbre characteristics of a specific speaker.
[0096] Step S212: Calculate the Mel loss based on the spectrum feature information and the predicted Mel spectrum.
[0097] For example, server 104 can calculate a mel loss based on the predicted mel spectrum and the true spectrum data to constrain the accuracy of spectrum generation. The predicted mel spectrum is the [B, T, n_mels] tensor generated by the mel projection head in the above steps, while the true spectrum is the original mel spectrum extracted from the training speech sample, with the same shape. Server 104 can square the element-by-element difference between the two and average it across all time frames and frequency band dimensions to obtain the mean squared error (MSE), which is the mel loss. This loss reflects the degree of match between the model-generated spectrum and the true speech spectrum and is an important indicator for measuring the naturalness and clarity of synthesized speech.
[0098] Furthermore, the loss function can be expressed as:
[0099]
[0100] Specifically, the squared difference of the energy values for each frame and frequency band is calculated and averaged. Smaller values indicate that the model's predicted spectrum is closer to the true speech spectrum. For example, if the true spectrum values for a frame are [0.3, 0.5, 0.4] and the model predicts [0.28, 0.52, 0.39], the error for that frame is ((0.3-0.28)^2+(0.5-0.52)^2+(0.4-0.39)^2) / 3. The errors across all frames are then averaged to obtain the overall mel loss. By guiding spectral reconstruction through hard alignment and applying mel loss constraints, the model accurately controls spectral boundaries, preventing silence from being mistakenly occupied by phonemes and preventing excessive extension of phonemes at the end, thereby improving speech naturalness and boundary clarity.
[0101] Step S214 : training the phoneme alignment model according to the relative entropy loss and the Mel loss until a preset convergence condition is met.
[0102] For example, after obtaining the two types of loss values, the server 104 can combine them according to preset weights to form a joint loss function, such as total_loss = 1*mel_loss + 1*kl_loss, and perform backpropagation training accordingly. During the training process, the server 104 can continuously iteratively update various learnable parameters in the phoneme alignment model, including the convolution kernel weights of the convolutional attention module, the embedding parameters of the phoneme encoder, the projection matrix of the Mel projection head, etc. After each round of training, the server evaluates the downward trend of the total loss value and compares it with preset convergence conditions, which may include the loss value decreasing by less than a threshold, the performance of the validation set tending to be stable, and the number of training rounds reaching a set upper limit.
[0103] During model training, server 104 can continuously monitor the changing trend of the total loss and set preset convergence conditions to determine when to terminate training. These conditions can include reaching an upper limit for the number of training rounds, the total loss decreasing by less than a threshold for several consecutive rounds, or the cessation of validation set accuracy improvement. For example, if the total loss decreases by less than 0.001 over five consecutive rounds of training, server 104 can determine that the model has stabilized and converged. At this point, the current model parameter state can be recorded as the final model parameters, and the output is the trained phoneme alignment model.
[0104] Finally, when the total loss meets the preset convergence criteria, server 104 considers model training complete, and the current model is a phoneme alignment model with high-quality alignment and spectral representation capabilities. Through this step, the server not only achieves explicit optimization of the phoneme-speech frame alignment structure, but also ensures that the predicted spectrum is highly accurate in reproducing the actual pronunciation.
[0105] In the aforementioned phoneme alignment model training method, speech training data is acquired and, based on the data, spectral feature information comprising multiple mel-spectrogram frames and textual feature information comprising multiple phonemes are obtained. Convolutional attention alignment is then performed on the spectral and textual feature information to produce a first alignment matrix. This first alignment matrix reflects the model's natural preference for boundary locations. Based on this, a monotonic alignment search is performed based on the first alignment matrix to generate a second alignment matrix. The second alignment matrix is a binarized hard attention matrix. The binarization discretizes the first alignment matrix, thereby clearly identifying the specific phoneme to which each frame should be aligned, ensuring that the alignment path is structurally strictly monotonic. Next, a relative entropy loss is calculated based on the first and second alignment matrices to measure the difference between the two. This encourages the hard alignment path to adhere as closely as possible to the model's perception of boundaries while maintaining a clear structure. This minimizes boundary distortion caused by alignment path offset during the binarization process, achieves more appropriate time allocation at phoneme boundaries (such as between consonants and vowels), and mitigates boundary misalignment issues such as early or late alignment. Subsequently, the text feature information is expanded to the Mel spectrum frame length based on the second alignment matrix, and the expanded text feature information is linearly transformed to generate a predicted Mel spectrum. The Mel loss is then calculated based on the spectral feature information and the predicted Mel spectrum, thereby optimizing the quality of spectrum reconstruction and further strengthening the temporal correspondence between frames and phonemes. This reverse optimization of the spectrum reconstruction loss can effectively reduce the phenomenon of silent frames being incorrectly assigned to phoneme regions, and improve the model's alignment accuracy in silent transition areas such as sentence beginnings and end. Finally, the phoneme alignment model is trained using relative entropy loss and Mel loss until the preset convergence conditions are met, thereby improving the accuracy of phoneme alignment. Furthermore, the above-mentioned speech synthesis method, apparatus, computer equipment, storage medium and computer program product obtain the text information to be processed, use the above-mentioned phoneme alignment model to process the text information to be processed, obtain the target spectrum information, and then obtain the target speech signal based on the target spectrum information. The above-mentioned phoneme alignment model can be applied to actual speech synthesis scenarios, and can make the target speech signal under complex requirements have higher accuracy, and then be applied to actual speech synthesis scenarios, and can make the speech synthesis results under complex requirements have higher accuracy, thereby improving the pronunciation clarity and temporal naturalness of speech synthesis.
[0106] In an exemplary embodiment, Figure 3 As shown, the above method further includes steps S302 to S306. In which:
[0107] Step S302: Calculate the first alignment matrix based on the logarithmic probability distribution function to obtain a phoneme probability sequence.
[0108] The phoneme probability sequence is the logarithmic value of the phoneme classification probability corresponding to each Mel spectrum frame.
[0109] Exemplarily, the first alignment matrix is a three-dimensional tensor of the shape [B, T, L], representing the matching probability distribution of each mel-spectrogram frame to each phoneme. To make it recognizable by the connection-time classification loss function, server 104 can perform a log_softmax operation on each row of the matrix to obtain the phoneme classification distribution in the form of logarithmic probability, forming a phoneme probability sequence. This sequence still maintains the structure of [B, T, L], where each value in each frame represents the logarithmic probability that the frame belongs to the corresponding phoneme. Compared to the ordinary softmax output, the logarithmic probability has better numerical stability and is suitable for subsequent path total probability accumulation calculation.
[0110] Step S304 : Calculate the connection temporal classification loss based on the text feature information and the phoneme probability sequence.
[0111] For example, server 104 can use the phoneme probability sequence as input to the model prediction path and the actual phoneme number sequence in the text feature information as the target label to perform the calculation of the connection temporal classification loss. The connection temporal classification loss (CTC loss) does not require explicit labeling of the boundaries between frames and phonemes. Instead, it uses a forward-backward algorithm to enumerate all possible paths that map the spectrum frame sequence to the target phoneme sequence and accumulate their total probabilities.
[0112] Furthermore, server 104 can allow special blank symbols to be inserted into the path during this step to accommodate natural speech rate variations and frame redundancy, making the alignment path time-flexible. The computational goal of the CTC loss is to minimize the distance between the predicted path distribution and the true phoneme sequence, that is, to maximize the total path probability, thereby guiding the model to automatically establish the optimal mapping structure from frames to phonemes during the learning process. The loss function can be expressed as:
[0113]
[0114] Where π is the set of all valid alignment paths, and z is the true phoneme sequence (including silence markers). The CTC loss maximizes the probability that the model generates the correct phoneme sequence, or equivalently, minimizes the negative logarithmic total probability. If the model's currently predicted path distribution deviates significantly from the true phoneme order, or if it mistakenly classifies silent frames as valid phones, the loss value will increase significantly, indicating an unreasonable alignment path. Therefore, the CTC loss has the advantages of eliminating the need for manual boundary annotation, enabling automatic alignment, and maintaining temporal monotonicity. It is particularly suitable for tasks where the length of speech sequences is significantly longer than that of phoneme sequences.
[0115] Step S306 , training the phoneme alignment model based on the connection temporal classification loss, relative entropy loss, and Mel loss until a preset convergence condition is met.
[0116] For example, server 104 can combine this connection temporal classification loss with the relative entropy loss (KL Loss) and Mel Loss calculated in the previous step as a joint training objective, forming a complete multi-task loss function. For example, this joint loss can be set as: total_loss = 1*mel_loss + 1*kl_loss + 0.1*ctc_loss. Server 104 can perform gradient backpropagation based on this joint loss to update the parameter weights in the model, particularly the learnable parameters of the convolutional attention module, phoneme encoder, and CTC path prediction head.
[0117] The introduction of CTC loss in this embodiment enhances the ability to model the monotonic path between frame sequences and phoneme sequences, compensates for the local instability of the attention mechanism in boundary processing, and improves the global rationality and stability of the phoneme alignment model in the selection of frame-phoneme alignment paths. In particular, in areas with polysyllabic words, continuous pronunciation, or varying speech rates, CTC loss can effectively suppress attention jumps, uneven frame distribution, and speech advance. By controlling the distribution of blank labels, it effectively reduces the problem of silent frames being misidentified as valid phonemes and alleviates the deviation of consonants being pronounced in advance.
[0118] In an exemplary embodiment, Figure 4 As shown, the above method further includes steps S402 to S406. In which:
[0119] Step S402: Generate predicted Mel-band probability based on text feature information.
[0120] The predicted Mel-band probability includes the classification probability distribution of the Mel-band corresponding to each phoneme.
[0121] For example, server 104 can receive text feature information output by the phoneme encoder in the shape of [B, L, d], representing a high-dimensional embedded representation of the phoneme sequence in each batch. The vector encoding of each phoneme carries abstract features such as the semantics and timbre of the phoneme in the current context. Server 104 can introduce a linear transformation module (fully connected layer) to project each phoneme vector of dimension d into a space of dimension n_mels, resulting in an output shape of [B, L, n_mels]. This output can be understood as a predicted score for the energy of each frequency band corresponding to the phoneme. For example, if the phoneme " / a / " has high values in the 4th, 17th, and 40th dimensions, the model believes that the phoneme's primary spectral energy is concentrated in these frequency bands. Server 104 can then normalize the prediction results using a log_softmax operation to obtain a logarithmic probability distribution over the dimension n_mels. This output is still [B, L, n_mels], where the position of each phoneme represents the probability of it belonging to a specific mel-band morphology category. This is the "predicted Mel band probability", which is essentially a classification distribution rather than a continuous spectral curve. It can be seen as the model's tendency to choose which type of sound template this phoneme belongs to.
[0122] Exemplarily, the server 104 can map the text feature information into a predicted Mel spectrum through a linear projection head; and calculate the predicted Mel spectrum based on the logarithmic probability distribution function to obtain the predicted Mel band probability. The server 104 can input the text feature information [B, L, d] output by the phoneme encoder into a linear projection head, which is a fully connected layer, and its function is to map the high-dimensional semantic vector of each phoneme to the spectral dimension space. Subsequently, in order to convert the prediction result into a probability distribution, the server 104 performs a log_softmax operation on the predicted Mel spectrum using the logarithmic probability distribution function. During the operation, the server 104 can perform logarithmic normalization on the spectral prediction value of each phoneme in the n_mels dimension, thereby obtaining a [B, L, n_mels] tensor with the same shape but different semantics, in which each item represents the logarithmic probability that the phoneme belongs to a certain spectral category.
[0123] Through this classification probability modeling method, server 104 can transform the original spectral modeling problem into a phoneme acoustic feature recognition problem, introduce more accurate and interpretable supervision signals into the training process, make the relationship between phoneme semantics and acoustic features clearer and more stable, and provide a more discriminative intermediate representation for subsequent speech synthesis.
[0124] Step S404 : Calculate the cross entropy loss between the predicted Mel-band probability and the true Mel-band spectrum.
[0125] For example, server 104 can extract the corresponding true mel-spectrogram [B, T, n_mels] from the training speech sample, and then segment and average-pool the spectrum along phoneme boundaries based on the previously obtained phoneme duration (derived from the hard alignment matrix). For example, if a phoneme corresponds to 10 frames, the server can average the spectra of these 10 frames to obtain an n_mels-dimensional spectrum vector, ultimately constructing the true spectrum feature matrix [B, L, n_mels].
[0126] In some embodiments, to enhance the effectiveness of classification supervision, the server may also perform K-means clustering on these true spectrum vectors to obtain spectrum category labels, converting continuous values into discrete categories so that each phoneme has a corresponding category label. For example, if the average spectrum of " / a / " falls near cluster center #4, it may be labeled "Category 4."
[0127] Next, server 104 can introduce a cross entropy loss function (Cross Entropy Loss, ce loss) to calculate the cross entropy loss between the predicted mel-band probability [B, L, n_mels] and the true label, and compare the deviation between the predicted probability distribution of each phoneme and the true label for each phoneme. Further, the cross entropy loss function is as follows:
[0128]
[0129] Where y is the one-hot label of the actual mel-band, and p is the mel-band probability predicted by the model. A smaller cross-entropy indicates a more accurate model prediction; a larger cross-entropy indicates a lower degree of match between the semantics of the phoneme and the actual spectrum.
[0130] Step S406 , training the phoneme alignment model based on cross entropy loss, connection temporal classification loss, relative entropy loss, and Mel loss until a preset convergence condition is met.
[0131] Exemplarily, the server 104 can combine the cross-entropy loss calculated in this embodiment with the other three types of losses obtained in the aforementioned steps to form a joint training target: wherein, the connection temporal classification loss (CTC Loss) is used to supervise the model to learn the global path structure between speech frames and phonemes, solving the problem of inconsistent time length; the relative entropy loss (KLLoss) is used to constrain the consistency between the soft attention output of the model and the discrete hard alignment path, thereby improving the stability of the alignment; Mel Loss: used to regress the difference between the predicted spectrum and the true spectrum, directly optimizing the speech generation quality; the cross-entropy loss (CE Loss) is used to improve the acoustic discriminability of the phoneme semantic vector, making the spectral classification ability of the phoneme stronger.
[0132] Server 104 can weight the sum of the four losses according to the set weights (for example, total_loss = 1*mel_loss + 1*kl_loss + 0.1*ctc_loss + 1*ce_loss) and update the model parameters based on this total loss. Training continues until the loss value stabilizes or the set termination criteria are met, such as no further improvement in validation set accuracy or the maximum number of training rounds is reached.
[0133] This embodiment enhances the model's ability to understand the classification boundaries of each phoneme in the acoustic space, especially between similarly pronounced phonemes and between the boundaries of silence and consonants. The model can more easily and accurately distinguish them, effectively alleviating the problems of blurred boundaries and pronunciation aliasing caused by unclear phoneme expression.
[0134] The present application also provides a speech synthesis method. This embodiment mainly uses the method applied to a computer device as an example. The computer device can be Figure 1 The terminal 102 or server 104 in the embodiment of the present invention can be used alone to perform the speech synthesis method provided in the embodiment of the present invention. The terminal 102 and server 104 can also be used together to perform the speech synthesis method provided in the embodiment of the present invention.
[0135] Exemplarily, server 104 can obtain text information to be processed; process the text information using the aforementioned phoneme alignment model to obtain target spectrum information; and obtain a target speech signal based on the target spectrum information. Specifically, server 104 can obtain the text information to be processed as the input starting point for speech synthesis. The text content can be any Chinese phrase or complete sentence, such as "The sun is shining brightly today." Server 104 can call upon a built-in text processing module to perform standardization on the sentence, including steps such as word segmentation, pinyin conversion, and phoneme annotation, ultimately converting the text into an ordered sequence of phoneme numbers, such as [n, i, j, i, n, t, i, a, n, g, sh, a, n, h, e, h, a, o]. This phoneme sequence serves as the text feature portion of the model input and enters the phoneme alignment model trained in the aforementioned embodiment to generate a semantic representation. Simultaneously, server 104 can also prepare the corresponding speaker's identity number and extract its timbre embedding, or timbre vector, from the speaker database to represent the user's desired pronunciation style. The phoneme encoder then converts the phoneme number sequence into a high-dimensional semantic vector. Using the learned attention mechanism within the alignment model, it predicts the alignment relationship between each spectrum frame and the corresponding phoneme, thereby obtaining the target spectrum information. The structural parameters of this attention module have been optimized using the aforementioned phoneme alignment model training method and will not be further described here.
[0136] Subsequently, the server 104 may input the target spectrum information into a trained vocoder, which converts the spectrum information into high-fidelity speech waveform data and outputs the data as a natural speech audio file corresponding to the original text content.
[0137] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0138] Based on the same inventive concept, the present application also provides a phoneme alignment model training device for implementing the above-mentioned phoneme alignment model training method. The implementation solution provided by the device is similar to the implementation solution described in the above-mentioned method. Therefore, the specific limitations of one or more phoneme alignment model training device embodiments provided below can be found in the above-mentioned limitations on the phoneme alignment model training method, and will not be repeated here.
[0139] In an exemplary embodiment, Figure 5 As shown, a phoneme alignment model training device is provided, comprising: a data acquisition module 502, a data alignment module 504, a first processing module 506, a first calculation module 508, a second processing module 510, a second calculation module 512 and a model training module 514, wherein:
[0140] The data acquisition module 502 is used to acquire speech training data and obtain spectral feature information and text feature information based on the speech training data; wherein the spectral feature information includes a plurality of mel-spectrogram frames, and the text feature information includes a plurality of phonemes;
[0141] A data alignment module 504 is configured to perform convolutional attention alignment on the spectral feature information and the text feature information to obtain a first alignment matrix;
[0142] A first processing module 506 is configured to perform a monotonic alignment search based on the first alignment matrix to generate a second alignment matrix, where the second alignment matrix is a binary hard attention matrix;
[0143] A first calculation module 508 is configured to calculate a relative entropy loss based on the first alignment matrix and the second alignment matrix;
[0144] The second processing module 510 is configured to expand the text feature information to a Mel spectrum frame length according to the second alignment matrix, and perform a linear transformation on the expanded text feature information to generate a predicted Mel spectrum;
[0145] A second calculation module 512 is configured to calculate the Mel loss based on the spectrum feature information and the predicted Mel spectrum;
[0146] The model training module 514 is used to train the phoneme alignment model according to the relative entropy loss and the Mel loss until a preset convergence condition is met.
[0147] In one embodiment, the device also includes: a first joint training module, which is used to calculate the phoneme probability sequence of the first alignment matrix based on the logarithmic probability distribution function; the phoneme probability sequence is the logarithmic value of the phoneme classification probability corresponding to each Mel spectrum frame; the connection temporal classification loss is calculated based on the text feature information and the phoneme probability sequence; the phoneme alignment model is trained based on the connection temporal classification loss, relative entropy loss and Mel loss until the preset convergence condition is met.
[0148] In one embodiment, the method further includes: a second joint training module for generating predicted Mel band probabilities based on text feature information; the predicted Mel band probabilities include classification probability distributions of Mel bands corresponding to each phoneme; calculating the cross entropy loss between the predicted Mel band probabilities and the true Mel spectrum; and training the phoneme alignment model based on the cross entropy loss, the connection temporal classification loss, the relative entropy loss, and the Mel loss until a preset convergence condition is met.
[0149] In one embodiment, the second joint training module is further configured to: map text feature information into a predicted Mel spectrum through a linear projection head; and calculate the predicted Mel frequency band probability based on the predicted Mel spectrum based on a logarithmic probability distribution function.
[0150] In one embodiment, the second processing module 510 is specifically used to: extract the duration of each phoneme according to the second alignment matrix; expand the text feature information to the Mel spectrum frame length according to the duration; map the expanded text feature information to the Mel spectrum frame dimension through the Mel projection head to generate a predicted Mel spectrum.
[0151] In one embodiment, the data alignment module 504 is specifically configured to generate a first alignment matrix based on negative Euclidean distance calculation of text feature information and spectrum feature information.
[0152] Based on the same inventive concept, embodiments of the present application further provide a speech synthesis device for implementing the aforementioned speech synthesis method. The implementation solution provided by this device is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations in one or more speech synthesis device embodiments provided below can be found in the above-mentioned limitations on the speech synthesis method and will not be further elaborated here.
[0153] In an exemplary embodiment, a speech synthesis device is provided, comprising: an information acquisition module, an information processing module, and a speech synthesis module, wherein:
[0154] An information acquisition module is used to obtain text information to be processed;
[0155] An information processing module is used to process the text information to be processed using the above-mentioned phoneme alignment model to obtain target spectrum information;
[0156] The speech synthesis module is used to obtain the target speech signal according to the target spectrum information.
[0157] Each module in the aforementioned phoneme alignment model training device and speech synthesis device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0158] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 6 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data used in the training phase such as voice training or data used in the use phase such as text information to be processed. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a phoneme alignment model training method is implemented.
[0159] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 7 As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be achieved via Wi-Fi, mobile cellular networks, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a phoneme alignment model training method. The display unit of the computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0160] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0161] In an exemplary embodiment, a computer device is provided, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the following steps when executing the computer program: obtaining speech training data, and obtaining spectral feature information and text feature information based on the speech training data; wherein the spectral feature information includes multiple Mel spectrum frames, and the text feature information includes multiple phonemes; performing convolutional attention alignment on the spectral feature information and the text feature information to obtain a first alignment matrix; performing a monotonic alignment search based on the first alignment matrix to generate a second alignment matrix, wherein the second alignment matrix is a binary hard attention matrix; calculating a relative entropy loss based on the first alignment matrix and the second alignment matrix; extending the text feature information to the Mel spectrum frame length based on the second alignment matrix, and performing a linear transformation on the extended text feature information to generate a predicted Mel spectrum; calculating a Mel loss based on the spectral feature information and the predicted Mel spectrum; and training a phoneme alignment model based on the relative entropy loss and the Mel loss until a preset convergence condition is met.
[0162] In one embodiment, when the processor executes the computer program, the following steps are also implemented: a phoneme probability sequence is calculated for the first alignment matrix based on a logarithmic probability distribution function; the phoneme probability sequence is the logarithmic value of the phoneme classification probability corresponding to each Mel spectrum frame; the connection temporal classification loss is calculated based on the text feature information and the phoneme probability sequence; and the phoneme alignment model is trained based on the connection temporal classification loss, relative entropy loss and Mel loss until a preset convergence condition is met.
[0163] In one embodiment, when the processor executes the computer program, the following steps are further implemented: generating predicted Mel band probabilities based on text feature information; the predicted Mel band probabilities include classification probability distributions of Mel bands corresponding to each phoneme; calculating the cross entropy loss between the predicted Mel band probabilities and the true Mel spectrum; and training the phoneme alignment model based on the cross entropy loss, the connection temporal classification loss, the relative entropy loss, and the Mel loss until a preset convergence condition is met.
[0164] In one embodiment, when the processor executes the computer program, the processor further implements the following steps: mapping the text feature information into a predicted Mel spectrum through a linear projection head; and calculating the predicted Mel frequency band probability based on the predicted Mel spectrum based on a logarithmic probability distribution function.
[0165] In one embodiment, when the processor executes the computer program, the processor further implements the following steps: extracting the duration of each phoneme according to the second alignment matrix; expanding the text feature information to the Mel spectrum frame length according to the duration; mapping the expanded text feature information to the Mel spectrum frame dimension through a Mel projection head to generate a predicted Mel spectrum.
[0166] In one embodiment, when the processor executes the computer program, the processor further implements the following steps: generating a first alignment matrix based on negative Euclidean distance calculation of the text feature information and the spectrum feature information.
[0167] In one embodiment, when the processor executes the computer program, it further implements the following steps: obtaining text information to be processed; processing the text information to be processed using the above-mentioned phoneme alignment model to obtain target spectrum information; and obtaining a target speech signal based on the target spectrum information.
[0168] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0169] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0170] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0171] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0172] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A phoneme alignment model training method, characterized in that: The method comprises: Acquire speech training data, and obtain spectral feature information and text feature information based on the speech training data; wherein the spectral feature information includes a plurality of mel-spectrogram frames, and the text feature information includes a plurality of phonemes; Performing convolutional attention alignment on the spectral feature information and the text feature information to obtain a first alignment matrix; Performing a monotonic alignment search based on the first alignment matrix to generate a second alignment matrix, where the second alignment matrix is a binary hard attention matrix; calculating a relative entropy loss based on the first alignment matrix and the second alignment matrix; Expanding the text feature information to a Mel spectrum frame length according to the second alignment matrix, and performing a linear transformation on the expanded text feature information to generate a predicted Mel spectrum; Calculating the Mel loss according to the spectrum feature information and the predicted Mel spectrum; The phoneme alignment model is trained according to the relative entropy loss and the Mel loss until a preset convergence condition is met.
2. The method according to claim 1, characterized in that The method further comprises: Calculating the first alignment matrix based on a logarithmic probability distribution function to obtain a phoneme probability sequence; the phoneme probability sequence is a logarithmic value of the phoneme classification probability corresponding to each mel spectrum frame; Calculating a connection temporal classification loss based on the text feature information and the phoneme probability sequence; The phoneme alignment model is trained based on the connection temporal classification loss, the relative entropy loss, and the Mel loss until the preset convergence condition is met.
3. The method according to claim 2, characterized in that The method further comprises: Generate a predicted Mel band probability based on the text feature information; the predicted Mel band probability includes a classification probability distribution of the Mel band corresponding to each phoneme; Calculating the cross entropy loss between the predicted Mel band probability and the true Mel spectrum; The phoneme alignment model is trained based on the cross entropy loss, the connection temporal classification loss, the relative entropy loss, and the Mel loss until the preset convergence condition is met.
4. The method according to claim 3, characterized in that The generating of predicted Mel-band probability based on the text feature information includes: Mapping the text feature information into the predicted Mel spectrum through a linear projection head; The predicted Mel frequency spectrum is calculated based on the logarithmic probability distribution function to obtain a predicted Mel frequency band probability.
5. The method according to any one of claims 1 to 4, characterized in that The step of extending the text feature information to a Mel spectrum frame length according to the second alignment matrix, and performing a linear transformation on the extended text feature information to generate a predicted Mel spectrum includes: extracting the duration of each phoneme according to the second alignment matrix; Expanding the text feature information to a Mel spectrum frame length according to the duration; The expanded text feature information is mapped to the Mel spectrum frame dimension through the Mel projection head to generate a predicted Mel spectrum.
6. The method according to any one of claims 1 to 4, characterized in that The convolutional attention alignment of the spectral feature information and the text feature information to obtain a first alignment matrix includes: A first alignment matrix is generated based on negative Euclidean distance calculation between the text feature information and the spectrum feature information.
7. A speech synthesis method, characterized in that: The method comprises: Get the text information to be processed; Processing the text information to be processed using the phoneme alignment model according to any one of claims 1 to 6 to obtain target spectrum information; A target speech signal is obtained according to the target spectrum information.
8. A phoneme alignment model training device, characterized in that: The device comprises: A data acquisition module, configured to acquire speech training data and obtain spectral feature information and text feature information based on the speech training data; wherein the spectral feature information includes a plurality of mel-spectrogram frames, and the text feature information includes a plurality of phonemes; A data alignment module, configured to perform convolutional attention alignment on the spectral feature information and the text feature information to obtain a first alignment matrix; A first processing module is configured to perform a monotonic alignment search based on the first alignment matrix to generate a second alignment matrix, where the second alignment matrix is a binary hard attention matrix; a first calculation module, configured to calculate a relative entropy loss according to the first alignment matrix and the second alignment matrix; A second processing module is configured to expand the text feature information to a Mel spectrum frame length according to the second alignment matrix, and perform a linear transformation on the expanded text feature information to generate a predicted Mel spectrum; A second calculation module is used to calculate the Mel loss according to the spectrum feature information and the predicted Mel spectrum; The model training module is used to train the phoneme alignment model according to the relative entropy loss and the Mel loss until a preset convergence condition is met.
9. A speech synthesis device, characterized in that: The device comprises: An information acquisition module is used to obtain text information to be processed; an information processing module, configured to process the text information to be processed using the phoneme alignment model according to any one of claims 1 to 6 to obtain target spectrum information; The speech synthesis module is used to obtain a target speech signal according to the target spectrum information.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Speech synthesis model training method, speech synthesis method and device thereof
CN113393828A
Speech synthesis processing method and device, equipment and medium
CN117456979A
Method and device for improving speech synthesis naturalness based on transferable monotone aligner
CN118782014A
Generation method and device of acoustic model, electronic equipment and storage medium
CN119832893A
Voice generation method and device, equipment and medium
CN120148474A
Cited By
Model training method and device, electronic equipment, computer readable storage medium and computer program product
CN121983024A
Model training method and device, voice processing method and device, electronic equipment, computer readable storage medium and computer program product
CN122067509A