Lightweight BERT-based Chinese electronic book voice broadcast method
By adopting a lightweight BERT-based method in Chinese e-book voice broadcast, using fine-tuning teacher model and student model, structure search optimizer and multi-task learning, the problems of insufficient accuracy of Chinese multi-phonetic prediction and excessive complexity are solved, and efficient and accurate multi-phonetic prediction and voice broadcast are achieved.
Patent Information
- Application Number
- CN202510115309.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-03
AI Technical Summary
The prediction accuracy of Chinese polyphonic characters in the prior art is insufficient, and the model is too complex, making it difficult to completely solve the problem of polyphonic characters in Chinese e-book voice broadcasts.
The Chinese e-book voice broadcast method based on lightweight BERT is adopted, and the BERT model is simplified by building a fine-tuning teacher model and student model, and the structure search optimizer is used to simplify the BERT model, combining multi-task learning and multi-teacher joint distillation technology to reduce the amount of model parameters and maintain high accuracy.
It significantly improves the accuracy of Chinese polyphonic prediction, reduces model complexity, improves reasoning efficiency, and has high practical application value.
Smart Images

Figure CN120089125A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice broadcasting, and particularly relates to a Chinese e-book voice broadcasting method based on lightweight BERT. Background Art
[0002] The automatic voice broadcasting function of e-books relies on advanced text-to-speech (TTS) technology to convert written content into natural and fluent voice output, enabling users to not only read visually but also seamlessly enjoy the audiobook experience in multitasking scenarios such as commuting and fitness. This technological breakthrough breaks through the limitations of traditional e-books that rely on visual input, greatly improving the flexibility and convenience of reading. At the same time, it provides important technical support for visually impaired and dyslexic people, greatly expanding the usage scenarios and audience scope of e-books.
[0003] In Chinese text-to-speech (TTS) systems, the problem of polyphonic characters is a major challenge affecting the voice broadcasting quality of e-books. Since polyphonic characters may have multiple pronunciations in different contexts, it often leads to pronunciation errors in speech synthesis. The methods for disambiguating polyphonic characters are mainly divided into two categories: rule-based and statistics-based. Rule-based methods rely on predefined rules but are difficult to cover complex language phenomena; while statistics-based methods, such as the n-gram model, predict the correct pronunciation by statistically analyzing word frequencies and co-occurrence relationships. Although the accuracy is improved in some cases, it is still insufficient when dealing with diverse text content. These technical bottlenecks make it still difficult to completely solve the problem of polyphonic characters in Chinese e-book broadcasting, and further improvement is needed. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to provide a Chinese e-book voice broadcasting method based on lightweight BERT, which solves the problems of insufficient accuracy in predicting Chinese polyphonic characters and too high model complexity in the prior art.
[0005] Technical Solution: A Chinese e-book voice broadcasting method based on lightweight BERT according to the present invention includes the following steps:
[0006] (1) Obtain a data set and perform preprocessing;
[0007] (2) Construct a data set for fine-tuning the teacher model, and use tools to extract the part of speech and pronunciation of polyphonic characters as training targets; and perform preprocessing;
[0008] (3) Mask the pinyin set to reduce the number of targets that the model needs to predict.
[0009] (4) Adopt the BERT model as the student model, and use a structure search optimizer to simplify and train the BERT model;
[0010] (5) Construct multiple teacher models, including a candidate pronunciation teacher model and a part-of-speech teacher model;
[0011] (6) Extract the outputs of each teacher model and train them through the student model on the same input to obtain a lightweight Chinese pronunciation prediction student model;
[0012] (7) Input the phoneme sequence features processed by the student model into an acoustic model and a vocoder, and finally output natural speech broadcast.
[0013] Further, in step (1), the preprocessing is as follows: construct a Chinese e-book dataset that meets the task requirements, and use the G2PC tool for text annotation; training set, validation set, and test set.
[0014] Further, step (3) includes the following steps:
[0015] (31) Consult a dictionary containing all polyphonic characters and their corresponding pronunciations, identify the polyphonic characters in the input sentence and mark them.
[0016] (32) Initialize a mask vector with the same length as the pinyin set and initialize all elements to 0; by querying the dictionary, obtain the candidate pronunciations of the polyphonic characters in the sample and their corresponding ID values, and set the values at the corresponding positions in the mask vector to 1 to indicate the legal pronunciations of the polyphonic characters;
[0017] (33) Use the posseg classifier in the jieba tool to perform part-of-speech tagging on the polyphonic characters to obtain the true part-of-speech tags; the part-of-speech tags include 24 common part-of-speech tags and 4 proper name category tags.
[0018] (34) Obtain the mask vector and the part-of-speech of the polyphonic characters for all samples and form batches.
[0019] Further, step (4) is specifically as follows: First, train the input and output sample sets of the BERT-based Chinese pronunciation prediction model to determine the encoding of network weights; through encoding mapping, determine the relationship between a set of weights and the individual bits in the hyperparameter space; then, set the model hyperparameters, perform optimization operations using the performance evaluation adjustment algorithm, and randomly generate V initial hyperparameter groups; evaluate N groups of networks according to the performance evaluation function, use the mean square error of the training samples as the measurement criterion of the performance evaluation function, and calculate the performance evaluation values of each group of hyperparameters; next, perform an optimization operation in the hyperparameter space according to the fitness, and retain the individuals with larger performance evaluations for the next cycle; use the random replacement operation to process the current hyperparameter group to generate the next group of hyperparameters; iterate by repeating the above process until the termination condition of the training objective is met, thereby obtaining a set of optimized weights; input the optimized weights as the new weights of the neural network into the BERT model to generate a Chinese polyphonic character prediction student model optimized by the hyperparameter search algorithm;
[0020] For the student model, adopt the method of multi-task fine-tuning for training, including polyphonic character pronunciation prediction and part-of-speech prediction.
[0021] Further, in step (4), the hyperparameter search algorithm is used to set the model hyperparameters.
[0022] Further, in step (4), the specific process of inputting the optimized weights as the new weights of the neural network into the BERT model is as follows:
[0023] (a) Convert the text sequence to be recognized into an input format that can be processed by the BERT model; tokenize the text and remove stop words;
[0024] (b) Extract feature words through dimensionality reduction processing, assign a unique ID to each feature word, and finally convert the text into an ID sequence to construct the features and labels of the dataset;
[0025] (c) Query the embedding layer in the BERT model according to the indexes in the text sequence to be recognized, and convert it into a vector representation;
[0026] (d) Use the BERT-base-Chinese model for parameter initialization, and map the embedding vector to a high-dimensional feature space containing semantic and position information;
[0027] (e) According to the ID value of the polyphonic character, extract the corresponding polyphonic character vector from the encoded vector as the input of the subsequent classifier;
[0028] (f) Add a pinyin classifier and a part-of-speech classifier at the output backend of the encoder, and pass the encoded vector of the polyphonic character to the pinyin classifier and the part-of-speech classifier;
[0029] (g) The pinyin classifier first calculates the distribution probability of the polyphonic character on the pinyin set, and uses the generated mask vector to set the probability of non-candidate pronunciations to a minimum value, thus retaining the probability distribution of legal candidate pronunciations;
[0030] (h) Let the output vector of the pinyin classifier be V = {v 1 , v 2 ,..., v n}, where v i represents the i-th element of V; the weighted prediction probability is obtained by adding the mask vector, and the weighted probability of each candidate pronunciation is as follows:
[0031]
[0032] Among them, the mask vector is represented as M = {m 1 , m 2 ,..., m n}, and m i is a boolean value used to indicate whether to mask the pronunciation v i ;
[0033] (i) For each polyphonic character, calculate the cross-entropy loss between the predicted probability distribution and the true pronunciation label. The loss function is as follows:
[0034]
[0035] Among them, y i,c and p i,c respectively represent the true label and the predicted probability of the i-th polyphonic character on the pinyin c.
[0036] (j) Classify the encoded vector according to the part-of-speech classifier to obtain the predicted probability distribution of the part of speech of the encoded vector of the polyphonic character, and calculate the cross-entropy loss function with the obtained true label of the part of speech:
[0037]
[0038] Among them, y i,c and p i,c respectively represent the true label and the predicted probability of the i-th polyphonic character on the pinyin s.
[0039] (k) Disambiguate the polyphonic character in the pinyin classifier and the part-of-speech classifier. The formula is as follows:
[0040] L(θ) = L 1 + L 2
[0041]
[0042] (l) Predict the phonemes of each polyphonic character in the sentence through the fine-tuned BERT model to obtain the pronunciation prediction probability of each polyphonic character. The formula is as follows;
[0043] h j = BERT(e j )
[0044]
[0045] where e j represents the embedding vector of the j-th word, and the hidden vector output at position j is obtained through encoding processing using the BERT model. f 1 and f 2 represent two fully connected layers, and the classification prediction of the hidden state vector is used to obtain the prediction result of the pronunciation of the current polyphonic character
[0046] Furthermore, step (5) is specifically as follows: The candidate pronunciation teacher model and the part-of-speech teacher model adopt a 12-layer BERT-Base architecture; during the training process, first pre-train each teacher model, and then fine-tune it on specific tasks; the candidate pronunciation teacher model is used to process the pronunciation prediction of polyphonic characters, and the part-of-speech teacher model is used for part-of-speech tagging.
[0047] Furthermore, the loss function used in step (5) is defined as follows:
[0048] Loss = KL(p T || p S ) + KL(q T || q S ) + L(θ)
[0049] where p T represents the prediction probability of the candidate pronunciation teacher model, q T represents the prediction probability of the part-of-speech teacher model, p S represents the prediction probability of the candidate pronunciation student model, and q S represents the prediction probability of the part-of-speech student model. L(θ) represents the loss between the prediction probabilities of the student model on the two tasks of part-of-speech and pronunciation and the true labels.
[0050] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: First, a large number of Chinese e-books are collected, and the G2PC tool is used to annotate the text. Based on the polyphonic character dictionary, the polyphonic characters in the text are queried and annotated, and the irrelevant pronunciations are masked, thereby narrowing the candidate range of pronunciation classification. By performing single-task learning - part-of-speech prediction and pinyin prediction respectively on two identical teacher models, the understanding degree of the model for the polyphonic character task is improved. Since the number of parameters of the BERT model is large, which limits its configuration ability in a specific environment, the present invention adopts a structure search optimizer to calculate the optimal model parameters, reducing the number of parameters of the student model while ensuring the model performance. Through the multi-teacher joint distillation technology, the knowledge of the teacher model is transferred to the student model, which not only reduces the number of parameters but also maintains the high-precision performance. By combining the multi-task learning strategy with the hyperparameter search algorithm, this method effectively compresses the model scale while ensuring the model accuracy, significantly improving the inference efficiency and having high practical application value. Brief Description of the Drawings
[0051] Figure 1 is the overall structural flowchart of the present invention;
[0052] Figure 2 is the flowchart of the BERT knowledge distillation method of the present invention;
[0053] Figure 3 is the overall flowchart of the structure search optimizer of the present invention;
[0054] Figure 4 is the schematic structural diagram of the multi-task learning based on BERT of the present invention;
[0055] Figure 5 is the schematic flowchart of the multi-teacher joint knowledge distillation of the present invention;
[0056] Figure 6 is the curve of the change of the loss value of the BERT student model of the present invention under the corresponding hyperparameter settings. Detailed Embodiment
[0057] The technical solution of the present invention will be further described below with reference to the drawings.
[0058] As Figure 1 shown, the embodiment of the present invention provides a Chinese e-book voice broadcast method based on lightweight BERT, including the following steps:
[0059] S101: Construct a Chinese e-book dataset that meets the task requirements, perform noise reduction on it to prevent model overfitting, use the G2PC tool for text annotation, and divide the dataset into training set, validation set, and test set in a ratio of 7:2:1. Next, take the sentence "The Yangtze River is China's mother river" in the training set as an example to introduce the implementation steps in detail.
[0060] S102: Input the Chinese electronic book content in order, check each character word by word whether it is in the polyphonic dictionary, and mark whether it is a polyphonic character. If yes, mark the character as a polyphonic character.
[0061] S103: Next, each polyphone in the sentence will be processed separately. It is worth noting that this paper uses batch processing to process polyphones in parallel, so it does not increase the processing time. For non-polyphones, the pinyin is directly marked; polyphones are represented by the "_" placeholder, and the resulting marking sequence is: ['_','jiang1','shi4','_','guo2','_','mu3','_','he2','. '].
[0062] S104: pre-process the data and perform masking on irrelevant pronunciations to narrow the candidate range of pronunciation classification. The purpose of obtaining the mask vector is to determine the candidate pronunciations of polyphones and assign weights to subsequent classification results. First, initialize a mask vector with the same length as the entire pinyin set and set all its elements to 0.
[0063] S105: Perform a dictionary query operation on the unique polyphone in the sample to obtain candidate pronunciations of the polyphone and its corresponding ID value. According to the ID value, the value of the corresponding position in the mask vector is set to 1 to indicate the legal pronunciation of the polyphone.
[0064] Taking the sample "The Yangtze River is the mother river of China" in the batch as an example, by querying the dictionary, it is determined that the polyphone "中" has two possible pronunciations, namely 'zhong1' (ID is 228) and 'zhong4' (ID is 229). Then, the positions corresponding to 228 and 229 in the mask vector are set to 1, and the other positions are still 0 to identify the candidate pronunciations of the polyphone.
[0065] S106: Use the possess classifier in the jieba tool to perform part-of-speech tagging on the polyphones in the e-book and generate real part-of-speech tags. Use the Paddle tagging mode to set a set of 24 part-of-speech tags and 4 proper noun category tags, as shown below:
[0066]
[0067] Use the jieba tool to tokenize the sentence, and the tokenization result is ["Yangtze River", "is", "China", "the", "mother river"]. Subsequently, use the posseg classifier to label the part-of-speech of each word, and the result is [ns, v, n, uj, n]. To ensure that the part-of-speech of each character is consistent with that of the word it belongs to, this paper sets the part-of-speech of each character in the tokenization to be the same as that of the word. Therefore, the part-of-speech of the polyphonic character "zhong" is the same as that of "China", which is "n".
[0068] S107: Divide the processed text data into pronunciation labels and part-of-speech labels, and input them into two teacher models for training respectively. At this stage, each teacher model focuses on the corresponding task and learns the features of specific labels by optimizing the loss function. During the training process, the teacher model continuously adjusts and updates its parameters according to the input data to improve the accuracy of pronunciation and part-of-speech prediction.
[0069] S108: Use the structure search optimizer to find a smaller BERT model as the student model, and its architecture definition depends on a series of hyperparameters. These hyperparameters are deeply analyzed, and the model is simplified by adjusting the hyperparameters related to the model architecture. The specific hyperparameter information is as follows:
[0070] L: The number of network layers
[0071] H: The dimension of each layer
[0072] A: The number of attention heads
[0073] D: The dimension of the feed-forward layer
[0074] V: The vocabulary size
[0075] S109: Limit the adjustable range of these hyperparameters to not exceed the corresponding values in the pre-trained model, ensuring model compression while maintaining its performance. The controllable range of the parameters is as follows:
[0076]
[0077] Use the structure search optimizer to guide the compression method, optimize and iterate the student model within the preset hyperparameter range, and calculate the optimal BERT model parameters. Taking 10MB, 25MB, 50MB, and 100MB as four size targets, through the optimization calculation of the hyperparameter search algorithm, the optimal parameters of four sizes of student models are finally calculated. The specific model information is as follows:
[0078]
[0079]
[0080] S110: The present invention adopts the method of multi-teacher joint distillation, collects the outputs of the teacher models, and uses them as the training data for the student model.
[0081] S111: During the distillation process of the student model, the KL divergence and cross-entropy loss functions are used to measure the differences between the outputs of the student model and the teacher models in the two tasks of polyphone and part-of-speech prediction.
[0082] S112: The phoneme sequence features processed by the student model are input into the acoustic model, and the model gradually converts these discrete phoneme sequences into continuous Mel spectrograms through a complex neural network structure.
[0083] S113: The vocoder synthesizes the corresponding waveform signals according to the input acoustic features using the HIFI-GAN generation algorithm. By converting the Mel spectrogram into the actual sound waveform, the vocoder can generate highly natural e-book voice broadcasts.
[0084] As Figure 3 shown, the present invention proposes a structure search optimizer, which searches for hyperparameters within a specific search space to minimize the difference between the size of the small model and the given size limit and maximize the capacity of the small model.
[0085] S201: For the original BERT model, obtain the existing model structure parameters.
[0086] S202: During the hyperparameter search process, each configuration represents a potential solution and consists of a set of parameters. In this study, each parameter is a key-value pair of a hyperparameter and its specific value. These parameters are stored in a dictionary data structure. For example, a configuration can be represented as {L:2, H:512, A:10, D:1024, V:21128}, which means a network layer with 2 512-dimensional layers, 10 attention heads, a feed-forward network with 1024 dimensions, and a vocabulary of size 21128. According to the common steps of hyperparameter optimization, the values of each key-value pair are randomly set in this paper to initialize a set of configurations. These randomly initialized configurations are called the candidate set, which provides the basis for subsequent iterations.
[0087] S203: Perform crossover operations on the individuals in the candidate set, generate new individuals by exchanging parameter segments, and simulate the recombination process.
[0088] S204: Perform parameter adjustment on some individuals, randomly modify some configurations to enhance the diversity of the candidate set and avoid falling into local optima.
[0089] S205: Calculate the floating-point representation. Calculate the floating-point numerical representation of each individual for subsequent performance evaluation.
[0090] S206: Generate a selection mechanism based on performance. Generate a selection mechanism based on the performance values of the individuals so as to select individuals with better performance.
[0091] S207: The individuals are selected to enter the candidate set for the next cycle through the selection mechanism, so that individuals with better performance have a higher probability of being selected.
[0092] S208: Check whether the termination condition is met, usually reaching the preset target model size (the default value is 20MB) or other standards. If the condition is met, the optimized model parameters are output for use by the student model; if the condition is not met, return to S203 and continue iteration.
[0093] S209: After multiple iterations, the optimized model parameters are finally obtained and input into the student model to initialize the lightweight student model and complete the structural optimization process.
[0094] like Figure 4 As shown, the present invention proposes a multi-task learning structure based on BERT for simultaneously predicting the pronunciation and part of speech of polyphones in Chinese e-books, and jointly optimizing the model through multi-task learning. Taking the sentence "The Yangtze River is China's mother river" as an example, the specific calculation steps are as follows:
[0095] S301: Using a dictionary containing all polyphones and their corresponding pronunciations, check each character in the input sentence word by word, identify and mark the polyphones therein. In the sentence "The Yangtze River is the mother river of China", the dictionary identifies five polyphones: "长", "江", "中", "的", and "亲".
[0096] S302: Perform disambiguation processing on each polyphone separately, convert the sentence into multiple samples, each sample only marks one polyphone position, for example: "_Chang_Jiang is China's mother river", "Chang_Jiang_ is China's mother river", etc. These samples will be organized into batches and processed in parallel to improve disambiguation efficiency.
[0097] S303: First, initialize a mask vector with the same length as the pinyin set, and set all elements to 0. Then, obtain the candidate pronunciations of the unique polyphone in the sample and its corresponding ID value by querying the dictionary. Taking the polyphone "中" as an example, its candidate pronunciations are 'zhong1' (ID is 228) and 'zhong4' (ID is 229). Subsequently, set the corresponding ID position in the mask vector to 1 to identify the legal pronunciation of the polyphone.
[0098] S304: Use the Posseg classifier in the Jieba tool to perform part-of-speech tagging on the polyphones in the sample and generate real part-of-speech tags. These tags provide accurate basis for subsequent part-of-speech classifiers and improve disambiguation accuracy.
[0099] S305: Through dictionary query, determine the legal candidate pronunciations for each polyphone and update the mask vector. The polyphone "中" has two candidate pronunciations (IDs are 228 and 229), and the corresponding positions in the mask vector are set to 1. After using the jieba tool to segment the words, combine the part-of-speech tagging tool to tag the part of speech of each word, and keep the part of speech of the polyphone consistent with the word to which it belongs.
[0100] S306: Use the jieba tool to segment the sentence, and get the result: ["Yangtze River", "is", "China", "of", "mother river"]. Then, use the possess classifier to perform part-of-speech tagging, and the result is: [ns,v,n,uj,n]. To ensure that the part-of-speech of each character is consistent with the word it belongs to, this paper sets the part-of-speech of each character in the segmentation to be the same as the word. Therefore, the part-of-speech of the polyphonic character "中" is consistent with "中国", that is, "n".
[0101] S307: The text sequence to be identified is a text ID sequence generated after preprocessing the text containing polyphones. In order to convert it into an input format that can be processed by BERT, the text is segmented and stop words are removed.
[0102] S308: Generate feature words through dimensionality reduction mapping, assign corresponding IDs to each feature word, and convert the text into a text ID sequence. The embedding layer in BERT generates corresponding vectors based on the index query of the text sequence, which are used as input to generate the encoded representation of the text.
[0103] S309: Encode the embedded vector using a preset encoder. The encoder is initialized using the BERT pre-trained model (BERT-base-Chinese), maps the embedded vector to a high-dimensional feature space containing semantic and position information, and generates a rich encoding vector.
[0104] S310: Extract the encoding vector corresponding to the polyphonetic character and input it into the pinyin classifier and the part-of-speech classifier. The pinyin classifier calculates the probability distribution of the polyphonetic character on the pinyin set, and the part-of-speech classifier calculates its part-of-speech prediction probability distribution.
[0105] like Figure 5 As shown, the present invention proposes a multi-teacher joint knowledge distillation structure diagram, which is used to transfer the learning knowledge of multiple teachers to the student model after structural optimization, so that the student model can strengthen the learning and understanding of the pronunciation and part of speech of polyphones. The specific steps are as follows:
[0106] S401: Input the sentence to be processed into the system and transfer it to two teacher models respectively. The teacher models are respectively for two different tasks: one for predicting the pronunciation of polyphonic characters, and the other for part-of-speech tagging. Each teacher model is based on the pre-trained BERT model and is fine-tuned for its respective specific task.
[0107] S402: The two teacher models receive the same input data and generate the output results of pronunciation prediction and part-of-speech prediction respectively. The output of the pronunciation teacher model is the pronunciation prediction result of the polyphonic characters in the corresponding sentence, and the output of the part-of-speech teacher model is the part-of-speech label of each word in the corresponding sentence.
[0108] S403: The student model also receives the same input sentence and simultaneously generates the prediction results of two tasks: pronunciation prediction and part-of-speech prediction. The structure of the student model is provided by the structure search optimizer and is optimized for specific tasks.
[0109] S404: Take the output results of the teacher model as the target and calculate the error between the student model and the teacher model. Specifically, two independent loss functions are used to measure the output difference between the student model and the teacher model:
[0110] 1. Pronunciation prediction loss: Compare the output of the student model and the pronunciation teacher model and calculate the loss of pronunciation prediction.
[0111] 2. Part-of-speech prediction loss: Compare the output of the student model and the part-of-speech teacher model and calculate the loss of part-of-speech prediction.
[0112] S405: Perform a weighted sum of the values of the two loss functions to obtain the total loss of the student model. This total loss comprehensively considers the errors of the pronunciation and part-of-speech prediction tasks to ensure that the student model can be optimized simultaneously in the two tasks.
[0113] S406: Based on the total loss value, update the parameters of the student model through the backpropagation algorithm. The student model continuously adjusts the weights through backpropagation, gradually narrowing the prediction gap with the teacher model, thereby improving its performance in the two tasks.
[0114] S407: During the training process, the structure search optimizer dynamically adjusts the architecture of the student model according to the model performance. This optimizer uses the structure of the BERT pre-trained model as the basis and continuously searches for the optimal network structure to improve the compression efficiency and inference speed of the model, ensuring the deployment of an efficient model in an environment with limited resources.
[0115] S408: This multi-teacher joint distillation process is carried out through multiple iterations until the loss value of the student model converges and reaches the set performance index. Finally, the student model can accurately predict the correct pronunciation of polyphonic characters and part-of-speech labels simultaneously.
Claims
1. A Chinese e-book voice broadcasting method based on lightweight BERT, characterized in that: The following steps are involved: (1) Obtain the data set and preprocess it; (2) Construct a dataset for fine-tuning the teacher model and use tools to extract the part of speech and pronunciation of polyphonetic characters as training targets; And pre-processing; (3) Mask irrelevant pronunciations in the training target; (4) Using the BERT model as the student model, the structure search optimizer is used to simplify and train the BERT model; (5) construct multiple teacher models, including candidate pronunciation teacher models and part-of-speech teacher models; (6) extracting the output of each teacher model and training it on the same input through a student model to obtain a lightweight Chinese pronunciation prediction student model; (7) The phoneme sequence features processed by the student model are input into the acoustic model and vocoder, and finally natural speech broadcast is output.
2. According to claim 1, a Chinese e-book voice broadcasting method based on lightweight BERT is characterized in that: In step (1), the preprocessing is as follows: construct a Chinese e-book dataset that meets the task requirements and use the G2PC tool to annotate the text; training set, validation set and test set.
3. According to claim 1, a Chinese e-book voice broadcasting method based on lightweight BERT is characterized in that: Step (3) comprises the following steps: (31) Consult a dictionary containing all polyphones and their corresponding pronunciations, identify the polyphones in the input sentence and mark them. (32) Initialize a mask vector with the same length as the pinyin set, and initialize all elements to 0; obtain candidate pronunciations of polyphones in the sample and their corresponding ID values by querying the dictionary, and set the value of the corresponding position in the mask vector to 1 to indicate the legal pronunciation of the polyphone; (33) Use the possess classifier in the jieba tool to perform part-of-speech tagging on polyphones and obtain the real part-of-speech tags; the part-of-speech tags include 24 common part-of-speech tags and 4 proper noun category tags. (34) Obtain mask vectors and parts of speech of polyphones for all samples and form batches.
4. The method for Chinese e-book voice broadcasting based on lightweight BERT according to claim 1, characterized in that: Step (4) is as follows: first, the input and output sample sets of the BERT-based Chinese pronunciation prediction model are trained to determine the encoding of the network weights; through encoding mapping, the relationship between a set of weights and individual bits in the hyperparameter space is determined; then, the model hyperparameters are set, the optimization operation is performed using the performance evaluation adjustment algorithm, and V initial hyperparameter groups are randomly generated; Evaluate N groups of networks according to the performance evaluation function, use the mean square error of the training samples as the criterion of the performance evaluation function, and calculate the performance evaluation value of each group of hyperparameters; Next, the optimization operation is performed in the hyperparameter space according to the fitness, and the individuals with larger performance evaluation are retained to the next cycle; the current hyperparameter group is processed by random replacement operation to generate the next set of hyperparameters; the above process is repeated iteratively until the termination condition of the training target is met, thereby obtaining a set of optimized weights; the optimized weights are input into the BERT model as the new weights of the neural network to generate a Chinese polyphone prediction student model optimized by the hyperparameter search algorithm; For the student model, a multi-task fine-tuning approach is used for training, including polyphonetic character pronunciation prediction and part-of-speech prediction.
5. A Chinese e-book voice broadcasting method based on lightweight BERT according to claim 4, characterized in that: In step (4), the model hyperparameters are set using a hyperparameter search algorithm.
6. A Chinese e-book voice broadcasting method based on lightweight BERT according to claim 4, characterized in that: In step (4), the optimized weights are input into the BERT model as the new weights of the neural network. The specific process is as follows: (a) Convert the text sequence to be recognized into an input format that can be processed by the BERT model; segment the text and remove stop words; (b) Extract feature words through dimensionality reduction processing and assign a unique ID to each feature word. Finally, the text is converted into an ID sequence to construct the features and labels of the dataset. (c) The embedding layer in the BERT model is queried according to the index in the text sequence to be recognized and converted into a vector representation; (d) Use the BERT-base-Chinese model to initialize parameters and map the embedding vector to a high-dimensional feature space containing semantic and position information; (e) according to the ID value of the polyphone, extract the corresponding polyphone vector from the encoding vector as the input of the subsequent classifier; (f) adding a phonetic classifier and a part-of-speech classifier to the output backend of the encoder, and passing the encoding vector of the polyphonetic character to the phonetic classifier and the part-of-speech classifier; (g) The phonetic classifier first calculates the distribution probability of polyphones on the phonetic set, and uses the generated mask vector to set the probability of non-candidate pronunciations to a minimum value, thereby retaining the probability distribution of legitimate candidate pronunciations; (h) Assume that the output vector of the pinyin classifier is V = {v1,v2,...,v n },v i represents the i-th element of V; the weighted prediction probability is obtained by adding the mask vector, where the weighted probability of each candidate pronunciation is as follows: The mask vector is represented as M = {m1,m2,...,m n },m i Is a Boolean value used to indicate whether to block the pronunciation v i ; (i) For each polyphone, the cross entropy loss between the predicted probability distribution and the true pronunciation label is calculated. The loss function is as follows: Among them, y i,c and p i,c They represent the true label and predicted probability of the i-th polyphone in pinyin c respectively. (j) Classify the encoding vector according to the part-of-speech classifier to obtain the part-of-speech prediction probability distribution of the polyphone encoding vector, and calculate the cross entropy loss function with the obtained part-of-speech true label: Among them, y i,c and p i,c They represent the true label and predicted probability of the i-th polyphone in pinyin s respectively. (k) Perform polyphone disambiguation in the pinyin classifier and part-of-speech classifier, the formula is as follows: L(θ)=L1+L2 (l) Use the fine-tuned BERT model to perform phoneme prediction on each polyphone in the sentence and obtain the pronunciation prediction probability of each polyphone. The formula is as follows: h j =BERT(e j ) Among them, e j The embedding vector of the jth word is represented by BERT model, and the hidden vector output at position j is obtained by encoding. f1 and f2 represent two fully connected layers. The hidden state vector is classified and predicted to obtain the prediction result of the current polyphone pronunciation.
7. The method for Chinese e-book voice broadcasting based on lightweight BERT according to claim 1, characterized in that: Step (5) is as follows: the candidate pronunciation teacher model and the part-of-speech teacher model use a 12-layer BERT-Base architecture; during the training process, each teacher model is first pre-trained and then fine-tuned on a specific task; The candidate pronunciation teacher model is used to process the pronunciation prediction of polyphones, and the part-of-speech teacher model is used for part-of-speech tagging.
8. The method for Chinese e-book voice broadcasting based on lightweight BERT according to claim 7, characterized in that: In step (5), the loss function is defined as follows: Loss=KL(p T ||p S )+KL(q T ||q S )+L(θ) Among them, P T represents the predicted probability of the candidate pronunciation teacher model, q T represents the predicted probability of the part-of-speech teacher model, p S represents the predicted probability of the candidate pronunciation student model, q S Represents the predicted probability of the part-of-speech student model. L(θ) represents the loss between the predicted probability of the student model and the true label on the two tasks of part-of-speech and pronunciation.