Knowledge distillation-based Hainan dialect speech recognition optimization system
By adopting knowledge distillation and dynamic temperature adjustment strategies of multi-teacher models in Hainan dialect speech recognition system, the shortcomings of Hainan dialect speech recognition technology in recognition accuracy and generalization capabilities are solved, and higher recognition accuracy and lower computing complexity are achieved, which is suitable for resource-constrained environments.
Patent Information
- Application Number
- CN202510462788.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-14
AI Technical Summary
Hainan dialect speech recognition technology has low recognition accuracy in actual application, mainly due to the limited scale of training data, insufficient generalization ability of model, and insufficient adaptability in complex environments, which limits its wide application in tourism, medical care, elderly care and other fields.
The Hainan dialect speech recognition optimization system based on knowledge distillation is adopted, and the MFCC speech feature sequence and text tag are extracted through the data preprocessing module, and multiple language models (RNN, CNN, Transformer) in the teacher model module are used to generate soft labels and intermediate layer features, and knowledge distillation and parameter tuning are performed through the student model training module to optimize the training process to improve the recognition accuracy.
Through the knowledge distillation and dynamic temperature adjustment strategies of multi-teacher models, the accuracy and generalization ability of Hainan dialect speech recognition are significantly improved, and the computational complexity is reduced, making the model more suitable for deployment in resource-constrained environments.
Smart Images

Figure CN120183382A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and particularly to an optimization system for Hainan dialect speech recognition based on knowledge distillation. Background Art
[0002] The protection of dialects and the development of intelligent recognition technology have become the focus of social attention. At present, although certain achievements have been made in Hainan dialect speech recognition technology, in practical applications, its recognition accuracy still needs to be improved. The primary restrictive factors are the limited scale of training data and the insufficient generalization ability of the model. Although research teams are making efforts to construct a Hainan dialect speech database, compared with languages such as Mandarin and Cantonese that have established mature corpus systems, the construction of the Hainan dialect speech resource library lags significantly, thereby restricting the training and performance optimization of the model. At the same time, the adaptability of this technology in actual application scenarios is also somewhat insufficient. Facing complex environmental noises, diverse speaker styles, and various language expression methods, the response ability of the Hainan dialect speech recognition system needs to be improved, which also limits its wide application in fields such as tourism, medical care, and elderly care.
[0003] Relative to the scarce Hainan dialect audio-text paired data, the cost of obtaining pure text data in actual production is indeed lower, and the quantity of pure text data that can be obtained is often several or even dozens of orders of magnitude more than the audio-text paired data. Integrating an external language model into a speech recognition system is very important. Traditional methods such as shallow fusion or deep fusion are effective, but the models are relatively complex and computationally intensive during decoding. Summary of the Invention
[0004] In view of the above-mentioned prior art, the present invention aims to provide an optimization system for Hainan dialect speech recognition based on knowledge distillation, mainly solving the technical problems existing in the above background art.
[0005] To achieve the above object, the technical solution of the embodiment of the present invention is implemented as follows:
[0006] An optimization system for Hainan dialect speech recognition based on knowledge distillation includes:
[0007] A data preprocessing module, configured to process Hainan dialect speech data and Hainan dialect text data, and output an MFCC speech feature sequence, a labeled text tag, and Hainan dialect text data;
[0008] A teacher model module, including an RNN language model, a CNN language model, and a Transformer language model, inputting the dialect text data into the RNN language model, the CNN language model, and the Transformer language model respectively, and obtaining soft labels and intermediate layer features through dynamic temperature adjustment;
[0009] A student model training module that performs knowledge distillation and parameter tuning on the student model according to the MFCC speech feature sequence, the labeled text tags, the soft tags, and the intermediate layer features;
[0010] An output module that uses the student model after knowledge distillation and parameter tuning to perform speech recognition on Hainan dialect and obtains a recognition result.
[0011] Optionally, the RNN language model is a recurrent neural network language model based on LSTM, which is used to capture the sequential dependencies in the Hainan dialect text data;
[0012] The CNN language model is a language model based on a convolutional neural network, which is used to capture the local features of the Hainan dialect text data;
[0013] The Transformer language model is stacked by multiple layers of Transformer encoders and is used to extract multi-scale features of the Hainan dialect text data.
[0014] Optionally, the performing knowledge distillation and parameter tuning on the student model according to the MFCC speech feature sequence, the labeled text tags, the soft tags, and the intermediate layer features includes:
[0015] For the MFCC speech feature sequence, use the forward propagation of the student model to generate a prediction result;
[0016] For the labeled text tags, use the labeled text tags as the true labels, calculate the cross-entropy loss between the prediction result and the true labels, and obtain a classification loss;
[0017] For the soft tags, perform weighted averaging on the soft tags obtained through the RNN language model, the CNN language model, and the Transformer language model respectively, and calculate the distillation loss between the prediction result of the student model and the weighted-averaged soft tags;
[0018] For the intermediate layer features, extract the real-time intermediate layer features corresponding to the teacher model in the student model, match the intermediate layer features output by the teacher model with the real-time intermediate layer features extracted by the student model, and calculate the cosine similarity loss between the intermediate layer features of the teacher model and the real-time intermediate layer features of the student model to obtain an intermediate layer feature matching loss;
[0019] Calculate a total loss function based on the classification loss, the distillation loss, and the intermediate layer feature matching loss, and perform training according to the total loss function to obtain a trained student model.
[0020] Optionally, the calculation formula of the classification loss is as follows:
[0021]
[0022] where y is the true label, t is the temperature, and q (t) is the softened probability distribution output by the student model; is the classification loss.
[0023] Optionally, the calculation formula of the distillation loss is as follows:
[0024]
[0025] where is the distillation loss, t is the temperature, p (t) is the softened probability distribution output by the teacher model, M is the number of teacher models, is the soft label processed by temperature of the m-th teacher model, w m is the weight;
[0026] The calculation formula of the weight in the distillation loss is:
[0027]
[0028] Confidence m = 1 / CER m
[0029] where Confidence m is the confidence of the m-th teacher model.
[0030] Optionally, the calculation formula of the intermediate layer feature matching loss is:
[0031]
[0032] where L is the number of intermediate layers, is the feature of the l-th layer of the teacher model, is the feature of the l-th layer of the student model.
[0033] Optionally, the expression of the total loss function is:
[0034]
[0035] where α and β are loss weights, is the total loss, is the classification loss, is the distillation loss, is the intermediate layer feature matching loss.
[0036] The beneficial effects of the present invention are as follows: The optimized Hainan dialect speech recognition system based on knowledge distillation provided by the present invention preprocesses Hainan dialect speech data and Hainan dialect text data through a data preprocessing module, extracts MFCC speech feature sequences, obtains text labels and Hainan dialect text data, processes the Hainan dialect text data processed by the data preprocessing module using a teacher model module. The teacher model module includes multiple language models, and multiple language models are used for fusion to generate soft labels and intermediate layer features. Then, a student model training module is used for knowledge distillation and parameter tuning, calculating classification loss, distillation loss, and intermediate layer feature matching loss to obtain a total loss function. The training process is optimized through the total loss function, and finally a trained student model is obtained. The trained student model is used to perform language recognition on Hainan dialect to obtain a speech recognition result; during the training stage, soft labels are generated using multiple external language models to guide the training of the student model, and no additional external language models need to be introduced during the decoding stage, thereby reducing computational complexity; and by transferring the knowledge of the teacher model to the student model, that is, transferring the knowledge in the complex model to the simple model, with fewer parameters and lower computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic structural diagram of the optimized Hainan dialect speech recognition system based on knowledge distillation provided in the embodiment of the present invention;
[0038] Figure 2 It is a schematic diagram for constructing the loss function provided in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The technical solution of the present invention will be further elaborated in detail below in conjunction with the drawings in the specification and specific embodiments. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. In the following description, the expression "some embodiments" is mentioned, which describes a subset of all possible embodiments. However, it should be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0040] In the following description, a large number of specific details are given to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, in order to avoid confusion with the present invention, some technical features well known to the public are not described.
[0041] It should be understood that the present invention can be implemented in different forms and should not be construed as limited to the embodiments presented herein. On the contrary, providing these embodiments will make the disclosure thorough and complete, and will fully convey the scope of the present invention to those skilled in the art. And the purpose of the terms used herein is only to describe specific embodiments and is not a limitation of the present invention. When used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprising" and / or "including", when used in this specification, determine the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups. When used herein, the term "and / or" includes any and all combinations of the related listed items.
[0042] It should be noted that when an element is referred to as "fixed to" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "inner", "outer", "left", "right" and similar expressions used herein are for illustrative purposes only and do not represent the only implementation.
[0043] To thoroughly understand the present invention, detailed structures will be presented in the following description to illustrate the technical solutions proposed by the present invention. The alternative embodiments of the present invention are described in detail as follows. However, in addition to these detailed descriptions, the present invention can also have other implementations.
[0044] Embodiment
[0045] Please refer to the attached Figure 1 , this application provides an optimized system for Hainan dialect speech recognition based on knowledge distillation, including:
[0046] A data preprocessing module for processing Hainan dialect speech data and Hainan dialect text data, and outputting an MFCC speech feature sequence, a labeled text tag, and Hainan dialect text data;
[0047] A teacher model module, including an RNN language model, a CNN language model, and a Transformer language model. The dialect text data is respectively input into the RNN language model, the CNN language model, and the Transformer language model, and after dynamic temperature adjustment, soft labels and intermediate layer features are obtained;
[0048] A student model training module that performs knowledge distillation and parameter tuning on the student model according to the MFCC speech feature sequence, the labeled text tags, the soft tags, and the intermediate layer features;
[0049] An output module that uses the student model after knowledge distillation and parameter tuning to perform speech recognition on Hainan dialect and obtains a recognition result.
[0050] Specifically, Hainan dialect speech data and Hainan dialect text data corresponding to the Hainan dialect speech data are obtained, and the Hainan dialect speech data and the Hainan dialect text data are input into the data preprocessing module for processing. Specifically: the Hainan dialect speech data is segmented into short-time frames, the Fourier transform is performed on each frame of the signal, converted from the time domain to the frequency domain, its power spectrum is calculated, the power spectrum is Mel-filtered through a Mel filter bank to obtain a Mel spectrum, the Mel spectrum is logarithmically processed, the logarithmically processed Mel spectrum is subjected to discrete cosine transform, cepstral coefficients are extracted, and finally the MFCC speech feature sequence is extracted; and the Hainan dialect text data is labeled to obtain text tags; then the Hainan dialect text data processed by the data preprocessing module is input into the teacher model module, and after dynamic temperature adjustment, soft tags and intermediate layer features are generated. The teacher model module includes an RNN language model, a CNN language model, and a Transformer language model. Then, knowledge distillation and parameter tuning are performed using the student model, which can integrate multiple teacher models and enable the student model to learn richer speech structure knowledge. Finally, the trained student model is obtained, and the trained student model is used for speech recognition of the Hainan dialect to obtain an accurate Hainan dialect speech recognition result.
[0051] As an optional implementation manner, the RNN language model is a recurrent neural network language model based on LSTM, which is used to capture the sequential dependence relationship in the Hainan dialect text data;
[0052] The CNN language model is a language model based on a convolutional neural network, which is used to capture the local features of the Hainan dialect text data;
[0053] The Transformer language model is stacked by multiple layers of Transformer encoders and is used to extract multi-scale features of the Hainan dialect text data.
[0054] Specifically, the RNN language model based on LSTM can capture temporal features such as unique rhythms and intonations in Hainan dialect, helping the model better understand the semantics and grammatical structures of Hainan dialect; the CNN language model based on convolutional neural network can capture local features in Hainan dialect. Exemplarily, it can capture specific syllable combinations and vowel combinations in Hainan dialect to help the model better identify key features in Hainan dialect; the Transformer language model can extract multi-scale features in Hainan dialect. Exemplarily, it can extract features from syllables to sentence level to help the model better understand the complex structure of Hainan dialect; by combining the RNN language model, the CNN language model, and the Transformer language model, the features of Hainan dialect can be captured more comprehensively; after weighted averaging the soft labels generated by the three language models, a more comprehensive supervision signal can be generated to help the student model better learn the knowledge of multiple language models in the teacher model, so that in the subsequent training process, more accurate speech recognition results can be obtained.
[0055] Please refer to the attached Figure 2 As an optional implementation manner, the knowledge distillation and parameter tuning of the student model according to the MFCC speech feature sequence, the labeled text label, the soft label, and the intermediate layer feature include:
[0056] For the MFCC speech feature sequence, use the forward propagation of the student model to generate a prediction result;
[0057] For the labeled text label, take the labeled text label as the true label, calculate the cross-entropy loss between the prediction result and the true label, and obtain the classification loss;
[0058] For the soft label, perform weighted averaging on the soft labels obtained by the RNN language model, the CNN language model, and the Transformer language model respectively, and calculate the distillation loss between the prediction result of the student model and the weighted-averaged soft label;
[0059] For the intermediate layer feature, extract the real-time intermediate layer feature corresponding to the teacher model in the student model, match the intermediate layer feature output by the teacher model with the real-time intermediate layer feature extracted by the student model, and calculate the cosine similarity loss between the intermediate layer feature of the teacher model and the real-time intermediate layer feature of the student model to obtain the intermediate layer feature matching loss;
[0060] Based on the classification loss, the distillation loss, and the intermediate layer feature matching loss, calculate the total loss function, and perform training according to the total loss function to obtain a trained student model.
[0061] Specifically, dynamic temperature adjustment is a method of adjusting the model output probability distribution by introducing temperature parameters. The original output of the teacher model is the score for each category, representing the confidence that the input Hainan dialect speech data belongs to each category; the soft label is the probability distribution of the output of the teacher model. By dynamically adjusting the original output of the teacher model, a smoother soft label can be generated. The smooth probability distribution can provide a more fine-grained supervision signal. After weighted average fusion of the soft labels generated by the RNN language model, CNN language model, and Transformer language model through dynamic temperature adjustment, multi-dimensional language knowledge can be integrated, and the weighted average soft label is compared with the prediction result output by the student model to calculate the distillation loss. By generating soft labels through the RNN language model, CNN language model, and Transformer language model to guide the training of the student model and integrating the knowledge of multiple language models in the teacher model, the generalization ability of the student model can be further improved;
[0062] The traditional fixed temperature cannot balance "global structure learning" and "detail discrimination learning" and cannot adapt to the needs of different training stages;
[0063] In the present invention, a dynamic temperature adjustment strategy is used to obtain a soft label by dynamically adjusting the original output of the teacher model. Among them, the temperature changes dynamically with the current training round, and the expression is:
[0064] t(e) = t0 × exp(-λ × e / E)
[0065] where E is the total number of training rounds, t0 is the initial temperature, and the value range is t0 ∈ (0, 1], and λ is the sensitivity score, and the value range is ∈ represents a very small temperature value;
[0066] Exemplarily, if t0 = 1, λ = 2, and E = 100 training rounds, then the change of temperature with the training round is:
[0067] Initial temperature: t(0) = 1;
[0068] Intermediate temperature (e = 50): t(50) = 1 × exp(-2 × 50 / 100) = exp(-1) ≈ 0.368;
[0069] Final temperature (e = 100): t(100) = 1 × exp(-2) ≈ 0.135;
[0070] The temperature throughout the process is within [0, 1] and monotonically decreasing.
[0071] The improved dynamic temperature uses a large t at the initial stage to make the soft labels smoother, guiding the student model to capture the overall semantic associations of the language model in the teacher model. At the later stage, a small t is used to make the soft labels closer to the hard labels, strengthening the precise classification ability (such as the text vocabulary corresponding to specific voices). Compared with the fixed temperature, the training process better conforms to the learning law and has a higher recognition accuracy.
[0072] By calculating the cosine similarity between the intermediate layer features of the teacher model and the intermediate layer features of the student model, and thus calculating the intermediate layer feature matching loss between the intermediate layer features of the teacher model and the intermediate layer features of the student model, the underlying semantic representations of multiple language models in the teacher model can be transferred.
[0073] As an optional implementation manner, the calculation formula of the classification loss is:
[0074]
[0075] where y is the true label, t is the temperature, and q (t) is the softened probability distribution output by the student model; is the classification loss.
[0076] As an optional implementation manner, the calculation formula of the distillation loss is:
[0077]
[0078] where is the distillation loss, t is the temperature, p (t) is the softened probability distribution output by the teacher model, M is the number of teacher models, is the soft label of the m-th teacher model processed by the temperature, and w m is the weight;
[0079] The calculation formula of the weight in the distillation loss is:
[0080]
[0081] Confidence m = 1 / CER m
[0082] where Confidence m is the confidence of the m-th teacher model, calculated by the CER of the teacher model validation set.
[0083] Specifically, the temperature t is used to adjust the "softening" degree of the output; by softening the output with the temperature, the teacher model can transfer richer implicit information to the student model, rather than relying only on hard labels; by learning the softened output of the teacher model, the student model can mine the deep knowledge of the teacher model and improve its own performance.
[0084] Traditional knowledge distillation methods rely on the soft label guidance of a single teacher model (such as an RNN language model), resulting in limited knowledge sources; the present invention integrates multiple teacher models (including RNN, CNN, and Transformer language models) and designs a confidence-based adaptive weight fusion strategy, that is: calculating the distillation loss based on the soft labels fused by multiple language models, enabling the student model to learn multi-granularity language features across models and enabling the student model to learn richer language structure knowledge.
[0085] As an optional implementation manner, the calculation formula for the intermediate layer feature matching loss is:
[0086]
[0087] where L is the number of intermediate layers, is the feature of the l-th layer of the teacher model, is the feature of the l-th layer of the student model.
[0088] Specifically, traditional knowledge distillation methods only perform knowledge distillation on soft labels. In the present invention, the features of the intermediate layers of the teacher model (such as the semantic representations of the hidden layers of the language model) are also transferred, enabling the student model to learn more underlying language structure encodings, thereby helping the model learn more features.
[0089] As an optional implementation manner, the expression of the total loss function is:
[0090]
[0091] where α and β are loss weights, is the total loss, is the classification loss, is the distillation loss, is the intermediate layer feature matching loss.
[0092] Specifically, use the total loss function to train, realizing the integration of an external language model during the training process; where α and β are loss weights, preset the ranges of α and β, search for the optimal combination on the validation set, and use the OneCycle strategy to dynamically adjust the learning rate.
[0093] By introducing multi-teacher model knowledge distillation and dynamic temperature adjustment strategies, the present invention significantly improves the accuracy and generalization ability of Hainan dialect speech recognition. Compared with the traditional single-teacher model, the present invention integrates the soft labels of multiple teacher models such as RNN, CNN, and Transformer, enriching the transmission of language structure knowledge. At the same time, the application of the dynamic temperature adjustment strategy also makes the training process more in line with the learning law, helping the model to obtain good performance in different training stages and optimizing the training process. In addition, the trained student model is small in size and low in computational complexity, making it convenient to be deployed in resource-constrained environments such as embedded devices. The proposed invention not only solves the current technical problems but also provides new ideas for the future development of Hainan dialect speech recognition technology, and is expected to be widely applied in fields such as tourism, healthcare, and elderly care, making positive contributions to the informatization construction and social development of Hainan region.
[0094] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. The Hainan dialect speech recognition optimization system based on knowledge distillation is characterized by: include: The data preprocessing module is used to process Hainan dialect speech data and Hainan dialect text data, and output MFCC speech feature sequences, annotated text labels, and Hainan dialect text data; The teacher model module includes an RNN language model, a CNN language model, and a Transformer language model. The dialect text data is respectively input into the RNN language model, the CNN language model, and the Transformer language model, and soft labels and intermediate layer features are obtained through dynamic temperature adjustment. A student model training module performs knowledge distillation and parameter tuning on the student model according to the MFCC speech feature sequence, the annotated text label, the soft label, and the intermediate layer features; The output module uses the student model after knowledge distillation and parameter tuning to perform speech recognition on the Hainan dialect and obtain the recognition results.
2. The Hainan dialect speech recognition optimization system based on knowledge distillation according to claim 1 is characterized in that: The RNN language model is a LSTM-based recurrent neural network language model, which is used to capture the sequence dependency in the Hainan dialect text data; The CNN language model is a language model based on a convolutional neural network, and is used to capture the local features of the Hainan dialect text data; The Transformer language model is formed by stacking multiple layers of Transformer encoders and is used to extract multi-scale features of the Hainan dialect text data.
3. The Hainan dialect speech recognition optimization system based on knowledge distillation according to claim 1 is characterized in that: The performing knowledge distillation and parameter tuning on the student model according to the MFCC speech feature sequence, the annotated text label, the soft label, and the intermediate layer feature includes: For the MFCC speech feature sequence, using the student model forward propagation to generate a prediction result; For the annotated text label, taking the annotated text label as the true label, calculating the cross entropy loss between the prediction result and the true label to obtain the classification loss; For the soft labels, weighted average is performed on the soft labels obtained by the RNN language model, the CNN language model, and the Transformer language model, and a distillation loss between a prediction result of a student model and the soft labels after weighted average is calculated; For the intermediate layer features, extract the real-time intermediate layer features in the student model corresponding to the teacher model, match the intermediate layer features output by the teacher model with the real-time intermediate layer features extracted by the student model, and calculate the cosine similarity loss between the intermediate layer features of the teacher model and the real-time intermediate layer features of the student model to obtain the intermediate layer feature matching loss; A total loss function is calculated based on the classification loss, the distillation loss, and the intermediate layer feature matching loss, and training is performed according to the total loss function to obtain a trained student model.
4. The Hainan dialect speech recognition optimization system based on knowledge distillation according to claim 3 is characterized in that: The calculation formula of the classification loss is: Among them, y is the true label, t is the temperature, and q is (t) Output the softened probability distribution for the student model; is the classification loss.
5. The Hainan dialect speech recognition optimization system based on knowledge distillation according to claim 3 is characterized in that: The calculation formula of the distillation loss is: in, is the distillation loss, t is the temperature, p (t) is the softened probability distribution of the teacher model output, M is the number of teacher models, is the temperature-processed soft label of the mth teacher model, w m is the weight; The weight calculation formula in the distillation loss is: Confidence m =1 / CER m Among them, Confidence m is the confidence of the mth teacher model.
6. The Hainan dialect speech recognition optimization system based on knowledge distillation according to claim 3 is characterized in that: The calculation formula of the intermediate layer feature matching loss is: Where L is the number of intermediate layers, is the feature of the teacher model at layer l, is the feature of the lth layer of the student model.
7. The Hainan dialect speech recognition optimization system based on knowledge distillation according to claim 3 is characterized in that: The expression of the total loss function is: Among them, α and β are loss weights, is the total loss, is the classification loss, is the distillation loss, is the intermediate layer feature matching loss.
Citation Information
Patent Citations
End-to-end long-time speech recognition method
CN113516968A
Target detection method based on self-distillation
CN116384439A
Dialect speech recognition training method and system based on self-knowledge distillation
CN117558264A
Cross-architecture video action recognition method and device based on knowledge distillation
CN118172705A
Multi-teacher model knowledge distillation method integrating semantic weight soft labels
CN119442014A
Cited By
Knowledge distillation method and device based on AI large model and electronic equipment
CN120654774A
State evaluation system based on lightweight Chinese speech rehabilitation large model
CN121393840A
A state evaluation system based on a lightweight Chinese speech rehabilitation large model
CN121393840B