Large voice model voice reconstruction method combining Lao language characteristics
By combining adapter fine-tuning with Lao pronunciation characteristics and constructing a speech feature set, the accuracy problem of Lao speech synthesis model was solved, achieving efficient speech synthesis under low resource conditions and improving speech quality and naturalness.
Patent Information
- Application Number
- CN202511035247.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-10-31
AI Technical Summary
Lao speech synthesis models suffer from poor accuracy due to data scarcity and underutilization of language characteristics, particularly in terms of phoneme pronunciation errors.
By combining a pre-trained large speech model with the characteristics of Lao pronunciation, and fine-tuning it through an adapter module, a Lao speech feature set is constructed. Then, a codec is used to reconstruct the speech waveform, reducing the number of adjustable parameters and improving the model's adaptability.
Under low-resource conditions, the accuracy and naturalness of Lao speech synthesis were improved, with a MOS score of 3.66, significantly improving speech quality.
Smart Images

Figure CN120877700A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a speech reconstruction method based on a large speech model that incorporates the characteristics of the Lao language, and belongs to the field of natural language processing technology. Background Technology
[0002] Lao, a language spoken by a relatively small number of people, is mainly distributed in Laos and parts of northeastern Thailand. Therefore, Lao corpus resources are extremely scarce, with only a few hours or even less than an hour of clean data typically available. This scarcity of data severely restricts the training of Lao speech synthesis models, making them prone to overfitting or attention collapse, leading to a decline in the quality of generated speech. Lao characters are divided into consonants and vowels, with vowels distinguished by phoneme length, resulting in long and short vowels. Many consonants exhibit both labialized and standard phoneme contrasts. Ignoring these linguistic characteristics can lead to phoneme pronunciation errors in the generated speech, affecting its accuracy. To address these issues, this invention utilizes a pre-trained large-scale speech model combined with Lao pronunciation characteristics to fine-tune the Lao language using an adapter, better adapting it to Lao speech synthesis tasks. Summary of the Invention
[0003] This invention provides a speech reconstruction method based on a large speech model that incorporates the characteristics of the Lao language, aiming to solve the problem of poor speech synthesis accuracy caused by data scarcity and the unique features of the Lao language.
[0004] The technical solution of this invention is: a speech reconstruction method based on a large-scale speech model that incorporates the characteristics of the Lao language, the method comprising:
[0005] Step 1: Lao text to phoneme conversion: Use a Lao phoneme conversion tool to convert the text words from orthographic representation to the corresponding IPA phoneme representation, and obtain a Lao phoneme dictionary;
[0006] Step 2: Construct a Lao phonetic feature set: By analyzing the pronunciation characteristics of different Lao characters, different phonemes are mapped to corresponding pronunciation methods based on a Lao phoneme dictionary, thereby constructing a Lao phonetic feature set;
[0007] Step 3: Fine-tuning the speech model based on the adapter: Add an adapter module to the intermediate layer of the pre-trained speech model, and then fine-tune the speech model using the Lao language dataset and the Lao language speech feature set constructed in Step 2 to reduce the number of adjustable parameters and model the Lao language speech model.
[0008] Step 4: Audio codec reconstructs speech waveforms: Using a finely tuned large speech model and encoder-decoder, discrete words are reconstructed into Lao speech waveforms.
[0009] Furthermore, the specific steps of Step 1 are as follows:
[0010] Step 1.1: Preprocess the input text using the preprocessor of the Lao phoneme conversion tool to remove unnecessary symbols or punctuation marks; during the preprocessing stage, the words in the input text remain unchanged;
[0011] Step 1.2: Using predefined mapping rules, the preprocessed Lao text words are mapped to phoneme representations to obtain preliminary phoneme representations;
[0012] Step 1.3: Using the post-processor of the Lao phoneme conversion tool, the initial phoneme representation is adjusted and improved by removing vowel loss and phoneme order to generate the final phoneme representation.
[0013] Furthermore, the specific steps of Step 2 are as follows:
[0014] Step 2.1 Analyze the pronunciation of Lao from a phonological perspective. The syllable structure of Lao is consonant-vowel-consonant. The beginning and middle of the syllable are consonant phonemes; at the end of the syllable, only some consonant phonemes are allowed to appear.
[0015] Step 2.2: Based on the analyzed vowel and consonant phoneme features, construct a Lao pronunciation feature set; this pronunciation feature set contains several attributes, including pronunciation feature attributes common to vowels and consonants, pronunciation feature attributes of vowels, and pronunciation feature attributes of consonants;
[0016] Different phonemes are identified according to the format <attribute symbol: attribute value>;
[0017] The specific attribute values of vowels and consonants are represented according to the English words corresponding to the Lao pronunciation features in Step 2.1;
[0018] Step 2.3: Represent each attribute category as a unique binary vector using one-hot encoding. The length of the vector is the same as the number of attributes in each category. The value at the category position is 1, and the value at other positions is 0. This yields the pronunciation feature vector.
[0019] Next, the obtained pronunciation feature vector is concatenated with the final phoneme representation from Step 1.3, so that the pronunciation features of different phonemes in Lao can be learned during the training process, thereby improving the accuracy of synthesized speech.
[0020] Furthermore, the specific steps of Step 3 are as follows:
[0021] During training, only the parameters of the adapter's downprojection feedforward layer, upprojection feedforward layer, nonlinear layer, and two layer-normalized layers in the Transformer module are adjusted.
[0022] First, the original high-dimensional input I is projected onto the low-dimensional N to reduce the number of parameters, where the projection dimension is smaller than the input dimension features.
[0023] Then, the vector features after nonlinear transformation are restored to the original high-dimensional I through another front feedback layer; finally, the original input information is preserved through residual connections, as shown in the following formula:
[0024] x+f(xW1)W2→x
[0025] Where x represents the input of the adapter module, W1 and W2 represent the weight matrices of the two feedforward networks, which map high-dimensional features to low-dimensional features and then restore the low-dimensional features to the original features, and f represents the non-linear activation function.
[0026] Furthermore, the specific steps of Step 4 are as follows:
[0027] The encoder of the audio codec encodes the input cue audio into a discrete sequence of words. Then, the fine-tuned speech model uses its autoregressive and non-autoregressive layers to predict and generate the target sequence of words. Finally, the decoder of the audio codec decodes the generated sequence of words and reconstructs it into a continuous speech waveform.
[0028] The present invention also provides a large-scale speech model speech reconstruction system that incorporates the characteristics of the Lao language, the system comprising: a module for executing the large-scale speech model speech reconstruction method that incorporates the characteristics of the Lao language.
[0029] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the speech reconstruction method based on the large speech model incorporating the characteristics of the Lao language.
[0030] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the speech reconstruction method of the large speech model combining the characteristics of the Lao language.
[0031] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the speech reconstruction method based on the large speech model that incorporates the characteristics of the Lao language.
[0032] The beneficial effects of this invention are:
[0033] This invention proposes a large-scale speech model for speech reconstruction that incorporates the characteristics of the Lao language. By integrating the pronunciation features of Lao and adopting an adapter-based fine-tuning strategy, it achieves Lao speech synthesis modeling based on a large-scale speech model. Under low-resource conditions, the proposed model can complete effective speech synthesis tasks. The proposed method achieves a MOS score of 3.66 on the Lao speech synthesis task.
[0034] This invention can effectively improve the modeling ability of the Lao speech big model to Lao language knowledge under low resource conditions, thereby enhancing the modeling effect of the Lao speech synthesis model and making the synthesized speech more natural. Attached Figure Description
[0035] Figure 1 This is a diagram illustrating the overall model framework for large-scale speech modeling in this invention, which incorporates the characteristics of the Lao language.
[0036] Figure 2 This is a flowchart illustrating the overall process of the large-scale speech model speech reconstruction method that incorporates the characteristics of the Lao language in this invention. Detailed Implementation
[0037] Example 1: As Figures 1-2 As shown, a large-scale speech model method for speech reconstruction based on the characteristics of the Lao language is described, and the method includes:
[0038] Step 1: Lao text to phoneme conversion: Use a Lao phoneme conversion tool to convert the text words from orthographic representation to the corresponding IPA phoneme representation, and obtain a Lao phoneme dictionary;
[0039] Furthermore, the specific steps of Step 1 are as follows:
[0040] Step 1.1: Preprocess the input text using the preprocessor of the Lao phoneme conversion tool to remove unnecessary symbols or punctuation marks; during the preprocessing stage, the words in the input text remain unchanged;
[0041] Step 1.2: Using predefined mapping rules, the preprocessed Lao text words are mapped to phoneme representations to obtain preliminary phoneme representations;
[0042] Step 1.3: Using the post-processor of the Lao phoneme conversion tool, the initial phoneme representation is adjusted and improved by removing vowel loss and phoneme order to generate the final phoneme representation.
[0043] Step 2: Construct a Lao phonetic feature set: By analyzing the pronunciation characteristics of different Lao characters, different phonemes are mapped to corresponding pronunciation methods based on a Lao phoneme dictionary, thereby constructing a Lao phonetic feature set;
[0044] For example, consonant phonemes
[0045]
[0046] this means It is a consonant phoneme with the place of articulation for a labial sound and the manner of articulation for a nasal sound. Each Lao phoneme is mapped to its articulation characteristics according to this method.
[0047] Furthermore, the specific steps of Step 2 are as follows:
[0048] Step 2.1: Analyze Lao pronunciation from a phonological perspective. The syllable structure of Lao is consonant-vowel-consonant, with syllables beginning and ending with a consonant phoneme; at the end of a syllable, only specific consonants are permitted. Consonant phonemes appear;
[0049] At the end of a syllable, only specific... Consonant phonemes appear.
[0050] (1) Consonants: They are divided into five categories according to the place of articulation: labial consonants, alveolar consonants, hard palatal consonants, soft palatal consonants and glottal consonants. The pronunciation characteristics of the corresponding phonemes are shown in Table 1.
[0051] Table 1. IPA Table for Consonant Pronunciation
[0052]
[0053]
[0054] (2) Regarding vowels: Lao vowels can be divided into long vowels and short vowels. Based on tongue position, tongue height, and lip shape, Lao vowels can be further classified. First, according to tongue position, they can be divided into front vowels, central vowels, and back vowels. This classification is closely related to the position of the tongue during pronunciation. Front vowels are pronounced with the tongue close to the front of the mouth, while back vowels are pronounced with the tongue at the back of the mouth. Second, according to tongue height, they can be divided into closed vowels, closed central vowels, open central vowels, and open vowels, reflecting the differences in tongue height during pronunciation. The pronunciation characteristics of the corresponding phonemes are shown in Tables 2 and 3.
[0055] Table 2. Short Vowel Pronunciation IPA Table
[0056]
[0057] Table 3. IPA Table for Pronunciation of Long Vowels
[0058]
[0059] In addition, Lao has a distinctive lip-shape contrast, consisting of rounded and unrounded vowels. Rounded vowels include [ua], [aw], [u], and [o]. [u:a], [u:], [o:] and Rounded vowels typically have a more closed, lower-pitched sound. In contrast, unrounded vowels are pronounced with the lips remaining open rather than closed. The mouth space is more spacious during pronunciation, which helps produce different timbre. Unrounded vowels in Lao include [aj], [am]、 Changes in lip shape not only affect the pronunciation quality of vowels, but may also have a certain effect on the pronunciation of subsequent consonants.
[0060] Step 2.2: Based on the analyzed vowel and consonant phoneme features, construct a Lao pronunciation feature set; this pronunciation feature set contains several attributes, including pronunciation feature attributes common to vowels and consonants, pronunciation feature attributes of vowels, and pronunciation feature attributes of consonants;
[0061] Among them, symbol_type and vowel_consonant are common pronunciation feature attributes for vowels and consonants, vowel_frontness, vowel_openness, and vowel_roundedness are pronunciation feature attributes for vowels, and consonant_place and consonant_manner are pronunciation feature attributes for consonants. The rules for each pronunciation attribute are shown in Table 4.
[0062] Table 4 Meaning of Pronunciation Feature Set Attributes
[0063]
[0064] Different phonemes are identified according to the format <attribute symbol: attribute value>;
[0065] In the pronunciation feature set, the symbol_type attribute is always fixed as phonme for both consonants and vowels. The vowel_consonant attribute has two values: "consonant" and "vowel". "Consonant" indicates that the phoneme is a consonant, and "vowel" indicates that the phoneme is a vowel.
[0066] The specific attribute values of vowels and consonants are represented according to the English words corresponding to the Lao pronunciation features in Step 2.1;
[0067] Step 2.3: Represent each attribute category as a unique binary vector using one-hot encoding. The length of the vector is the same as the number of attributes in each category. The value at the category position is 1, and the value at other positions is 0. This yields the pronunciation feature vector.
[0068] Next, the obtained pronunciation feature vector is concatenated with the final phoneme representation from Step 1.3, so that the pronunciation features of different phonemes in Lao can be learned during the training process, thereby improving the accuracy of synthesized speech.
[0069] Step 3: Fine-tuning the speech model based on the adapter: Add an adapter module to the intermediate layer of the pre-trained speech model, and then fine-tune the speech model using the Lao language dataset and the Lao language speech feature set constructed in Step 2 to reduce the number of adjustable parameters and model the Lao language speech model.
[0070] Furthermore, the specific steps of Step 3 are as follows:
[0071] During training, only the parameters of the adapter's downprojection feedforward layer, upprojection feedforward layer, nonlinear layer, and two layer-normalized layers in the Transformer module are adjusted.
[0072] First, the original high-dimensional input I is projected onto the low-dimensional N to reduce the number of parameters, where the projection dimension is smaller than the input dimension features.
[0073] Then, the vector features after nonlinear transformation are restored to the original high-dimensional I through another front feedback layer;
[0074] Finally, the original input information is preserved through residual connection, as shown in the following formula:
[0075] x+f(xW1)W2→x
[0076] Where x represents the input of the adapter module, W1 and W2 represent the weight matrices of the two feedforward networks, which map high-dimensional features to low-dimensional features and then restore the low-dimensional features to the original features, and f represents the non-linear activation function.
[0077] Step 4: Audio codec reconstructs speech waveforms: Using a finely tuned large speech model and encoder-decoder, discrete words are reconstructed into Lao speech waveforms.
[0078] Furthermore, the specific steps of Step 4 are as follows:
[0079] The encoder of the audio codec encodes the input cue audio into a discrete sequence of words. Then, the fine-tuned speech model uses its autoregressive and non-autoregressive layers to predict and generate the target sequence of words. Finally, the decoder of the audio codec decodes the generated sequence of words and reconstructs it into a continuous speech waveform.
[0080] The present invention also provides a large-scale speech model speech reconstruction system that incorporates the characteristics of the Lao language, the system comprising: a module for executing the large-scale speech model speech reconstruction method that incorporates the characteristics of the Lao language.
[0081] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the speech reconstruction method based on the large speech model incorporating the characteristics of the Lao language.
[0082] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the speech reconstruction method of the large speech model combining the characteristics of the Lao language.
[0083] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the speech reconstruction method based on the large speech model that incorporates the characteristics of the Lao language.
[0084] To verify the effectiveness of the large-scale speech model speech reconstruction method combining the characteristics of the Lao language proposed in this invention, the following experiments were conducted:
[0085] (a) Evaluation indicators
[0086] Subjective evaluation involves inviting testers to listen to the synthesized speech. The testers provide subjective ratings based on the speech's quality and comprehensibility. Participants rate the synthesized speech quality using a scale of 0-5, and the final score is the average of these ratings to obtain the Mean Opinion Score (MOS). Generally, an average MOS of 4 or higher indicates good speech quality; while an average MOS below 3.6 indicates that most listeners are dissatisfied with the speech quality.
[0087] The objective average MCD (Mel-Cepstral Distortion) represents the difference between the Mel-frequency cepstral coefficient characteristics of the converted speech and the Mel-frequency cepstral coefficient characteristics of the standard output speech; the smaller the value, the better. The RMSE (Root Mean Square Error) is the root mean square error of the fundamental frequency; the lower the value, the closer the fundamental frequency profile is to the standard output speech, and the better the result.
[0088] (b) Comparative experiments with different corpus sizes
[0089] Due to the limited availability of Lao language corpora, a comparative experiment was first conducted on English to demonstrate the usability of the proposed model for speech synthesis tasks in low-resource environments. The experiment evaluated the MOS scores of the baseline model and the proposed model by comparing the model performance under different corpus sizes. Table 5 shows the MOS scores under different corpus sizes.
[0090] Table 5. MOS scores under different English corpus sizes
[0091]
[0092] As shown in Table 5, the MOS score of the model decreases with the continuous reduction in the size of the training corpus, and the speech synthesis performance also gradually declines. The model achieves the highest MOS score of 4.20 when the total training corpus duration is 500 hours, demonstrating optimal performance. However, when the corpus size is reduced to 100 hours, the baseline model's MOS score is only 3.46, a difference of 0.74 points compared to the model trained with 500 hours of corpus, indicating that most evaluators are dissatisfied with the synthesized speech quality. This is because the reduced training data size decreases the knowledge the model can learn, leading to a decrease in its generalization ability and a significant reduction in the naturalness of the synthesized audio. For the model improved based on the method described in this chapter, under a 100-hour English dataset, the improved model achieves a MOS score of 3.60, showing a certain improvement. This is because Lao and English phonemes partially overlap; by incorporating phoneme-based pronunciation features into the baseline model, linguistic knowledge information similar to Lao is added. Experiments demonstrate that the proposed model can effectively complete speech synthesis tasks under low-resource conditions.
[0093] (c) Ablation test
[0094] To further validate the performance of the proposed method, ablation experiments were conducted to evaluate the impact of incorporating Lao language pronunciation features and using adapter fine-tuning on the speech synthesis results. The results are shown in Table 6, where w / o stands for "without".
[0095] Table 6. Lao language evaluation scores under different modules
[0096]
[0097] As shown in the table above, on the 100-hour Lao dataset, the baseline model achieved a score of 3.41, with MCD and RMSE values of 8.11 and 43.21 respectively, indicating poor synthesized speech quality without incorporating Lao language features and fine-tuning through an adapter. The W / o adapter, which only integrates the Lao pronunciation feature map on top of the baseline model, achieved a MOS score of 3.59, an improvement of 0.18 points, indicating improved naturalness of the synthesized audio. Both the MCD and RMSE values were also higher than the baseline model. This is because reducing the fine-tuning parameters improves training stability and enhances the model's adaptability to Lao. Building upon this foundation, the integration of Lao pronunciation features further improved the synthesized speech's MOS score by 0.25 points compared to the baseline model, while reducing the MCD and RMSE scores by 0.6 and 6.14 points, respectively. This demonstrates that adding pronunciation features and integrating the adapter into the model effectively enhances the model's ability to model Lao language knowledge under low-resource conditions, thereby improving the modeling performance of the Lao speech synthesis model and making the synthesized speech more natural.
[0098] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A speech reconstruction method based on a large-scale speech model that incorporates the characteristics of the Lao language, characterized by: The method includes: Step 1: Lao text to phoneme conversion: Use a Lao phoneme conversion tool to convert the text words from orthographic representation to the corresponding IPA phoneme representation, and obtain a Lao phoneme dictionary; Step 2: Construct a Lao phonetic feature set: By analyzing the pronunciation characteristics of different Lao characters, different phonemes are mapped to corresponding pronunciation methods based on a Lao phoneme dictionary, thereby constructing a Lao phonetic feature set; Step 3: Fine-tuning the speech model based on the adapter: Add an adapter module to the intermediate layer of the pre-trained speech model, and then fine-tune the speech model using the Lao language dataset and the Lao language speech feature set constructed in Step 2 to reduce the number of adjustable parameters and model the Lao language speech model. Step 4: Audio codec reconstructs speech waveforms: Using a finely tuned large speech model and encoder-decoder, discrete words are reconstructed into Lao speech waveforms.
2. The speech reconstruction method based on a large-scale speech model incorporating the characteristics of the Lao language as described in claim 1, characterized in that: The specific steps of Step 1 are as follows: Step 1.1: Preprocess the input text using the preprocessor of the Lao phoneme conversion tool to remove unnecessary symbols or punctuation marks; during the preprocessing stage, the words in the input text remain unchanged; Step 1.2: Using predefined mapping rules, the preprocessed Lao text words are mapped to phoneme representations to obtain preliminary phoneme representations; Step 1.3: Using the post-processor of the Lao phoneme conversion tool, the initial phoneme representation is adjusted and improved by removing vowel loss and phoneme order to generate the final phoneme representation.
3. The speech reconstruction method based on a large-scale speech model incorporating the characteristics of the Lao language as described in claim 1, characterized in that: The specific steps of Step 2 are as follows: Step 2.1 Analyze the pronunciation of Lao from a phonological perspective. The syllable structure of Lao is consonant-vowel-consonant. The beginning and middle of the syllable are consonant phonemes; at the end of the syllable, only some consonant phonemes are allowed to appear. Step 2.2: Based on the analyzed vowel and consonant phoneme features, construct a Lao pronunciation feature set; this pronunciation feature set contains several attributes, including pronunciation feature attributes common to vowels and consonants, pronunciation feature attributes of vowels, and pronunciation feature attributes of consonants; Different phonemes are identified according to the format <attribute symbol: attribute value>; The specific attribute values of vowels and consonants are represented according to the English words corresponding to the Lao pronunciation features in Step 2.1; Step 2.3: Represent each attribute category as a unique binary vector using one-hot encoding. The length of the vector is the same as the number of attributes in each category. The value at the category position is 1, and the value at other positions is 0. This yields the pronunciation feature vector. Next, the obtained pronunciation feature vector is concatenated with the final phoneme representation from Step 1.3, so that the pronunciation features of different phonemes in Lao can be learned during the training process, thereby improving the accuracy of synthesized speech.
4. The speech reconstruction method based on a large-scale speech model incorporating the characteristics of the Lao language as described in claim 1, characterized in that: The specific steps of Step 3 are as follows: During training, only the parameters of the adapter's downprojection feedforward layer, upprojection feedforward layer, nonlinear layer, and two layer-normalized layers in the Transformer module are adjusted. First, the original high-dimensional input I is projected onto the low-dimensional N to reduce the number of parameters, where the projection dimension is smaller than the input dimension features. Then, the vector features after nonlinear transformation are restored to the original high-dimensional I through another front feedback layer; Finally, the original input information is preserved through residual connection, as shown in the following formula: x+f(xW1)W2→x Where x represents the input of the adapter module, W1 and W2 represent the weight matrices of the two feedforward networks, which map high-dimensional features to low-dimensional features and then restore the low-dimensional features to the original features, and f represents the non-linear activation function.
5. The speech reconstruction method based on a large-scale speech model incorporating the characteristics of the Lao language as described in claim 1, characterized in that: The specific steps of Step 4 are as follows: The encoder of the audio codec encodes the input cue audio into a discrete sequence of words. Then, the fine-tuned speech model uses its autoregressive and non-autoregressive layers to predict and generate the target sequence of words. Finally, the decoder of the audio codec decodes the generated sequence of words and reconstructs it into a continuous speech waveform.
6. A large-scale speech model speech reconstruction system that incorporates the characteristics of the Lao language, characterized in that: The system includes a module for performing the large-scale speech model speech reconstruction method that incorporates the characteristics of the Lao language as described in any one of claims 1 to 5.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the speech reconstruction method of the large speech model that combines the characteristics of the Lao language as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the speech reconstruction method of the large speech model that combines the characteristics of the Lao language as described in any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech reconstruction method of the large speech model that combines the characteristics of the Lao language as described in any one of claims 1 to 5.