Speech recognition and synthesis aided teaching method and system
By acquiring pure speech data, fusing heterogeneous features, and mapping parameters, guided demonstration speech that matches learners' vocal habits is generated. This solves the problems of impure speech data and mismatch between teaching content in existing technologies, and improves the efficiency and relevance of teaching assistance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-04-03
AI Technical Summary
Existing speech recognition and synthesis-assisted teaching technologies suffer from problems such as impure data, inability to integrate personalized features, and mismatch between teaching content and learner speech data, resulting in poor teaching assistance effects.
By acquiring clean speech data, heterogeneous features are fused to construct a personalized feature set, which is then mapped to prosody adjustment and timbre modification parameters. This allows for fusion-based speech synthesis, generating guided demonstration speech that matches the learner's vocal habits.
It enables personalized adaptation of teaching content, improves the efficiency and relevance of speech recognition and synthesis-assisted teaching, and enhances the practical value of teaching aids.
Smart Images

Figure CN121789638A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech synthesis technology, and in particular to a speech recognition and synthesis-assisted teaching method and system. Background Technology
[0002] Existing speech recognition and synthesis-assisted teaching technologies have significant shortcomings in learner speech data processing and key knowledge point analysis. They fail to standardize the audio specifications of learners' historical speech samples and fail to accurately remove long silent and noisy segments using short-time energy features and zero-crossing rate features, making it difficult to obtain high-quality, clean speech data. At the same time, the analysis of key language knowledge points in the current teaching content of the standardized curriculum syllabus lacks a systematic approach and cannot accurately locate and select core content based on teaching stage markers. This makes the generation of subsequent teaching aids lack a reliable foundation and affects the effectiveness of teaching aids.
[0003] Existing technologies have significant shortcomings in the personalized adaptation of teaching content and speech synthesis. They fail to fuse heterogeneous features from pure speech data to construct a personalized feature set for learners, and cannot dynamically adapt and reconstruct the initial teaching narrative text based on syntactic mastery and vocabulary familiarity. This results in a mismatch between the generated teaching text and the learner's ability level. Furthermore, they fail to accurately map the vocal feature components in the personalized feature set to prosodic adjustment parameters and timbre modification parameters, making the fusion-based speech synthesis of personalized retelling texts lack specificity. The resulting guiding demonstration speech is difficult to match with learners' vocal habits and cognitive characteristics, and cannot effectively guide learners to understand and retell knowledge points, thus hindering the improvement of auxiliary teaching efficiency. Therefore, how to achieve personalized adaptation of teaching content and speech synthesis that conforms to learner characteristics has become an urgent problem to be solved. Summary of the Invention
[0004] This invention provides a speech recognition and synthesis-assisted teaching method and system to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides a speech recognition and synthesis-assisted teaching method, comprising: S1. Obtain the clean speech data of the target learner and, based on the standardized curriculum outline, analyze the key language knowledge points of the current teaching content; S2. Perform heterogeneous feature fusion on the clean speech data to obtain the personalized feature set of the target learner; S3. The key language knowledge points are transcribed in a structured manner to obtain the initial teaching narrative text of the key language knowledge points; S4. Based on the personalized feature set, perform dynamic stylistic adaptation on the initial teaching narrative text to obtain the personalized retelling text for the target learner; S5. Map the vocal feature components of the personalized feature set to prosody adjustment parameters and timbre modification parameters; S6. Based on the prosody adjustment parameters and the timbre modification parameters, perform fusion speech synthesis on the personalized recitation text to obtain the guided demonstration speech of the target learner.
[0006] In a preferred embodiment, the step of acquiring the target learner's clean speech data and, based on a standardized curriculum syllabus, analyzing the key language knowledge points of the current teaching content includes: Retrieve learner identifiers and teaching stage identifiers of target learners from the learning record database; Based on the learner identifier and the teaching stage identifier, retrieve the target learner's historical speech repetition samples; Unify the audio specifications of the historical speech retelling samples to obtain standardized speech data of the historical speech retelling samples; By analyzing the short-time energy characteristics and zero-crossing rate characteristics of the standardized speech data, long silent segments and noise segments of the target learner are removed to obtain the clean speech data of the standardized speech data. Based on the teaching stage identifiers, locate the corresponding teaching content in the standardized curriculum syllabus, and select the key language knowledge points of the corresponding teaching content.
[0007] In a preferred embodiment, the heterogeneous feature fusion of the clean speech data to obtain the personalized feature set of the target learner includes: Obtain the Mel frequency cepstral coefficient sequence, fundamental frequency trajectory, and formant distribution parameters of the clean speech data; Based on the Mel frequency cepstral coefficient sequence, the fundamental frequency trajectory, and the formant distribution parameters, an acoustic feature vector of the clean speech data is constructed. Semantic text analysis is performed on the clean speech data to obtain the semantic feature vector of the clean speech data; Cross-modal alignment is performed on the acoustic feature vector and the semantic feature vector to obtain aligned feature pairs of the acoustic feature vector and the semantic feature vector; Based on the importance of the aligned feature pairs, feature integration is performed on the aligned feature pairs to obtain the personalized feature set of the target learner.
[0008] In a preferred embodiment, the step of performing structured transcribing of the key language knowledge points to obtain the initial teaching narrative text for the key language knowledge points includes: Based on the logical components of the key language knowledge points, the vocabulary, sentence structure and grammatical rules in the key language knowledge points are analyzed. By performing syntactic structure filling on the vocabulary, sentence structure, and grammatical rules, standardized sentence fragments of the key language knowledge points are obtained. Logical semantic connections are made between the standard sentence fragments to obtain the initial teaching narrative text for the key language knowledge points.
[0009] In a preferred embodiment, the step of dynamically adapting the initial instructional narrative text to the personalized retelling text for the target learner based on the personalized feature set includes: The personalized feature set is analyzed by ability dimension to obtain the syntactic mastery parameter and vocabulary familiarity parameter of the target learner; A linguistic structural analysis was performed on the initial teaching narrative text to identify potential high-difficulty syntactic structures and target vocabulary in the initial teaching narrative text; Based on the syntactic mastery parameter and the vocabulary familiarity parameter, the contextual difficulty assessment of the potential high-difficulty syntactic structures and the target vocabulary is performed to obtain the predicted repetition difficulty of the target learner. Based on the predicted retelling difficulty, the initial teaching narrative text is reconstructed using adaptable text to obtain the personalized retelling text for the target learner.
[0010] In a preferred embodiment, the step of performing a contextualized difficulty assessment of the potentially high-difficulty syntactic structures and the target vocabulary based on the syntactic mastery parameter and the vocabulary familiarity parameter to obtain the predicted repetition difficulty of the target learner includes: Based on the syntactic mastery parameter and the vocabulary familiarity parameter, the syntactic fitness and vocabulary cognition of the target learner are determined; Based on the syntactic fitness and the lexical awareness, the predicted repetition difficulty of the target learner is calculated, wherein the formula for calculating the predicted repetition difficulty is as follows: ; In the formula, This indicates the difficulty of restating the prediction. Indicates the syntactic fitness. This indicates the level of vocabulary recognition. This represents the preset weighting coefficient. This represents the preset weighting coefficient. This represents the preset weighting coefficient. This represents the preset smallest positive number. This indicates the preset scale parameters. Represents the arctangent function. This represents the natural exponential function. This represents the square root operation.
[0011] In a preferred embodiment, the step of adapting the initial instructional narrative text to a personalized retelling text for the target learner based on the predicted retelling difficulty includes: The predicted paraphrasing difficulty is compared with a preset difficulty response threshold to filter out the syntactic structures to be simplified and the target words to be replaced. Cognitive load reduction is performed on the syntactic structure to be simplified to obtain a simplified syntactic instance of the initial teaching narrative text; The vocabulary to be replaced is benchmarked against the acquisition stage to obtain the appropriate vocabulary for the initial teaching narrative text; The simplified syntax examples and the adapted vocabulary are integrated into the initial teaching narrative text, and the coherence of the integrated initial teaching narrative text is checked. Based on the verification results, a personalized restatement text for the target learner is generated.
[0012] In a preferred embodiment, mapping the vocal feature components of the personalized feature set to prosody adjustment parameters and timbre modification parameters includes: The vocal feature components of the personalized feature set are deconstructed at multiple levels to obtain the primary prosodic features, secondary prosodic features, steady-state timbre features, and transient timbre features of the vocal feature components; The primary prosodic features and the secondary prosodic features are reduced by spectral features to obtain the prosodic control vector of the vocal feature component; Based on a preset timbre transfer strategy, the steady-state timbre features and the transient timbre features are mapped to derive a timbre influence factor; By performing cross-modal parameter coordination on the prosody control vector and the timbre influence factor, the prosody adjustment parameters and timbre modification parameters of the vocal feature components are obtained.
[0013] In a preferred embodiment, the step of performing fusion speech synthesis on the personalized recitation text based on the prosodic adjustment parameters and the timbre modification parameters to obtain the guided demonstration speech for the target learner includes: The personalized paraphrased text is subjected to text acoustic analysis to obtain the phoneme sequence and basic prosodic boundaries of the personalized paraphrased text; Based on the prosodic adjustment parameters, the phoneme sequence and the basic prosodic boundary are parametrically marked to obtain the target synthesis specification of the personalized paraphrased text. Acoustic parameter prediction is performed on the target synthesis specification to obtain the initial acoustic parameter sequence of the target synthesis specification; Based on the timbre modification parameters, the initial acoustic parameter sequence is modulated in the timbre dimension to obtain the acoustic parameter sequence of the initial acoustic parameter sequence; Semantic waveform synthesis is performed on the acoustic parameter sequence to obtain the guided demonstration speech of the target learner.
[0014] To address the above problems, the present invention also provides a speech recognition and synthesis-assisted teaching system, the system comprising: The data acquisition and parsing module is used to acquire clean speech data from the target learners and, based on the standardized curriculum outline, to parse out the key language knowledge points of the current teaching content. The feature fusion module is used to perform heterogeneous feature fusion on the clean speech data to obtain the personalized feature set of the target learner; The text generation module is used to perform structured transcription of the key language knowledge points to obtain the initial teaching narrative text of the key language knowledge points; The style adaptation module is used to dynamically adapt the initial teaching narrative text to the style based on the personalized feature set, so as to obtain the personalized retelling text for the target learner. The parameter mapping module is used to map the vocal feature components of the personalized feature set into prosody adjustment parameters and timbre modification parameters; The speech synthesis module is used to perform fusion speech synthesis on the personalized recitation text based on the prosody adjustment parameters and the timbre modification parameters to obtain the guided demonstration speech of the target learner.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention accurately acquires the pure speech data of the target learners, ensuring the accuracy of speech feature extraction. At the same time, it analyzes the key language knowledge points of the current teaching content based on the standardized curriculum outline, providing a reliable foundation for the generation of teaching aids. Then, it performs heterogeneous feature fusion on the pure speech data to construct a personalized feature set for learners. Based on this feature set, it dynamically adapts the initial teaching narrative text to generate personalized retelling texts that match the learners' syntactic mastery and vocabulary familiarity, effectively adapting to the learners' ability level and helping them better understand key language knowledge points.
[0016] 2. This invention maps the vocal feature components of a personalized feature set to prosodic adjustment parameters and timbre modification parameters. Based on these parameters, it performs fusion-based speech synthesis on personalized repetition text. The generated guiding demonstration speech matches the learner's vocal habits, improves the adaptability and guiding effect of the demonstration speech, and ultimately effectively improves the efficiency of speech recognition and synthesis-assisted teaching, enhances the relevance of assisted teaching to learners, and further strengthens the practical value of teaching assistance. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a speech recognition and synthesis-assisted teaching method according to an embodiment of the present invention; Figure 2 This is a functional module diagram of a speech recognition and synthesis-assisted teaching system provided in an embodiment of the present invention; The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] This application provides a speech recognition and synthesis-assisted teaching method. The executing entity of this speech recognition and synthesis-assisted teaching method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application embodiment: a server, a terminal, etc. In other words, the speech recognition and synthesis-assisted teaching method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0020] Reference Figure 1 The diagram shown is a flowchart illustrating a speech recognition and synthesis-assisted teaching method according to an embodiment of the present invention. In this embodiment, the speech recognition and synthesis-assisted teaching method includes: S1. Obtain the clean speech data of the target learner and, based on the standardized curriculum outline, analyze the key language knowledge points of the current teaching content; In this embodiment of the invention, the step of acquiring the target learner's clean speech data and, based on a standardized curriculum syllabus, analyzing the key language knowledge points of the current teaching content includes: Retrieve learner identifiers and teaching stage identifiers of target learners from the learning record database; Based on the learner identifier and the teaching stage identifier, retrieve the target learner's historical speech repetition samples; Unify the audio specifications of the historical speech retelling samples to obtain standardized speech data of the historical speech retelling samples; By analyzing the short-time energy characteristics and zero-crossing rate characteristics of the standardized speech data, long silent segments and noise segments of the target learner are removed to obtain the clean speech data of the standardized speech data. Based on the teaching stage identifiers, locate the corresponding teaching content in the standardized curriculum syllabus, and select the key language knowledge points of the corresponding teaching content.
[0021] The learning record database is structured according to learner information categories. Each learner's information entry includes basic information such as name and registration account, along with a binding relationship between learner identifier and teaching stage identifier. During the retrieval process, staff or the system inputs the target learner's name or registration account into the database. The database then traverses its internally stored learner information entries based on the input information, first locating the learner's main entry that matches the basic information, and then extracting the learner identifier and teaching stage identifier corresponding to the learner's current learning stage from the learning stage record sub-module under the main entry.
[0022] The system employs a hierarchical storage architecture for historical speech repetition samples. This architecture uses learner identifiers as the primary category, with each primary category further divided into subcategories based on teaching stage identifiers. Different subcategories store historical speech repetition samples corresponding to the learning stage, and the sample formats include common types such as WAV and MP3. When retrieving samples, the system first navigates to the corresponding primary category based on the learner identifier, then precisely locates the subcategories based on the teaching stage identifier. Subsequently, all historical speech repetition samples stored in that subcategory are completely extracted and transferred to a dedicated buffer module for temporary data processing, awaiting further operations.
[0023] The system pre-defines clear and unified audio specifications, specifically a 16kHz sampling rate, 16-bit bit depth, and mono. When processing historical speech retelling samples, the system invokes the audio format conversion module to process each sample in the buffer module: for samples with an original sampling rate other than 16kHz, the frequency is adjusted to 16kHz through resampling; for samples with an original bit depth other than 16bit, quantization adjustment technology is used to convert the bit depth to 16bit; for multi-channel samples, a channel merging algorithm is used to combine them into mono. After processing each sample, the system performs specification verification to confirm compliance with the set standards before processing the next sample. The audio data compiled after processing all samples constitutes the standardized speech data of the historical speech retelling samples.
[0024] The system imports standardized speech data into the signal analysis module. This module first divides the standardized speech data into multiple consecutive short frames with a fixed time length of 20ms, and slides between adjacent short frames with a step size of 10ms to ensure no data is missed. Then, the short-time energy and zero-crossing rate are calculated for each short frame: When calculating the short-time energy, the amplitude of all audio signals in the frame is first obtained, each amplitude is squared, and then all squared results are summed to obtain the short-time energy value of the short frame; when calculating the zero-crossing rate, the amplitude signs of two adjacent sampling points in the frame are compared one by one, and the number of times the sign changes from positive to negative or from negative to positive is the zero-crossing rate of the short frame. The system pre-sets short-time energy thresholds and zero-crossing rate thresholds, and compares the parameters of short-time frames with the thresholds frame by frame. Frames with short-time energy and zero-crossing rate below the thresholds are identified as long-duration silent segments, while frames with zero-crossing rate above the thresholds and short-time energy in the middle range are identified as noise segments. These identified segments are then deleted from the standardized speech data, and the remaining audio data is spliced together to obtain the clean speech data of the standardized speech data.
[0025] The standardized curriculum syllabus is divided into multiple teaching stages according to learning progress and knowledge difficulty. Each teaching stage is marked with a unique corresponding teaching stage identifier, and the teaching content within each teaching stage is divided into units. Each unit clearly indicates the core knowledge point categories that support the learning objectives of that unit. The system compares the target learner's teaching stage identifier with the identifiers of each stage in the curriculum syllabus. Once a match is found, it locates all the unit content under that teaching stage. Then, it analyzes the content of each unit sentence by sentence, identifying the core vocabulary, basic sentence structures, and key grammar rules that play a crucial role in understanding the core content of the unit. These elements are then organized into a clear list, which represents the key language knowledge points of the corresponding teaching content.
[0026] The beneficial effects include the ability to obtain clean speech data of target learners that meets unified standards and is free from long periods of silence and noise through a structured retrieval and processing process, ensuring the accuracy of subsequent data processing. At the same time, it can accurately locate the course outline content based on the teaching stage identifier and select the core key language knowledge points, ensuring that subsequent teaching support content revolves around the core points, laying a solid foundation for improving the effectiveness of subsequent personalized teaching support.
[0027] S2. Perform heterogeneous feature fusion on the clean speech data to obtain the personalized feature set of the target learner; In this embodiment of the invention, the step of performing heterogeneous feature fusion on the clean speech data to obtain the personalized feature set of the target learner includes: Obtain the Mel frequency cepstral coefficient sequence, fundamental frequency trajectory, and formant distribution parameters of the clean speech data; Based on the Mel frequency cepstral coefficient sequence, the fundamental frequency trajectory, and the formant distribution parameters, an acoustic feature vector of the clean speech data is constructed. Semantic text analysis is performed on the clean speech data to obtain the semantic feature vector of the clean speech data; Cross-modal alignment is performed on the acoustic feature vector and the semantic feature vector to obtain aligned feature pairs of the acoustic feature vector and the semantic feature vector; Based on the importance of the aligned feature pairs, feature integration is performed on the aligned feature pairs to obtain the personalized feature set of the target learner.
[0028] The clean speech data is divided into consecutive short frames in a fixed 20ms time window, with a 10ms overlap between adjacent frames to avoid information loss. A Fourier transform is performed on each frame to convert the time-domain signal into a frequency-domain spectrum. This spectrum is then input into a pre-defined Mel filter bank, designed based on the nonlinear frequency perception characteristics of the human ear, with frequency intervals varying according to the Mel frequency scale. The energy value output by each filter is calculated and the Mel spectrum is constructed. The natural logarithm of the Mel spectrum is taken, followed by a discrete cosine transform. The first 13 coefficients are extracted from the transform result and arranged chronologically across all short frames to obtain the Mel frequency cepstral coefficient sequence of the clean speech data. Simultaneously, the fundamental frequency of each frame is calculated using the autocorrelation method. The correlation coefficient between this frame and signals with different delay lengths is calculated, and the delay length corresponding to the maximum correlation coefficient is found. Combined with the sampling rate of the speech data, the fundamental frequency value of this frame is calculated. The fundamental frequency values of all short frames are then concatenated chronologically to form the fundamental frequency trajectory of the clean speech data. In addition, peak detection is performed on the frequency domain spectrum of each frame of speech signal to identify the frequencies corresponding to the three highest energy peaks in the spectrum, namely the formant frequencies. The frequency difference when the energy on both sides of each formant frequency drops to half of the peak value is calculated, namely the formant bandwidth. The formant frequencies and bandwidths of all short frames are sorted in frame order to obtain the formant distribution parameters of the clean speech data.
[0029] The dimensionality information of the Mel-frequency cepstral coefficient sequence, fundamental frequency trajectory, and formant distribution parameters was statistically analyzed. Each frame of the Mel-frequency cepstral coefficient sequence contains 13 dimensions of data, each frame of the fundamental frequency trajectory contains 1 dimension of data, and each frame of the formant distribution parameters contains 6 dimensions of data, including the frequencies and bandwidths of the first 3 formants. Using short-time frames as units, the 13 Mel-frequency cepstral coefficients, 1 fundamental frequency value, and 6 formant parameters corresponding to the same frame were concatenated sequentially to form a single-frame feature vector. Then, the single-frame feature vectors of all short-time frames were stacked in chronological order to construct a two-dimensional data matrix. This two-dimensional data matrix is the acoustic feature vector of the clean speech data.
[0030] Speech recognition technology is used to convert clean speech data into corresponding text content, ensuring that the converted text is completely consistent with the semantics expressed by the speech. The converted text is then segmented into words based on grammatical rules and collocation habits. Each word or phrase is then tagged with its part-of-speech (POS) to clarify its grammatical attributes in the sentence, such as nouns, verbs, and adjectives. Based on the POS tagging results, key elements expressing the core semantics of the text are extracted, including core nouns, verbs, and logical relationships between words, such as subject-verb and verb-object relationships. These semantic key elements are converted into fixed-length numerical values using predefined encoding rules. All numerical values are then arranged in semantic logical order to form a semantic feature vector of the clean speech data.
[0031] This study analyzes the correspondence between the temporal dimension of acoustic feature vectors and the semantic unit dimension of semantic feature vectors. By detecting pauses in speech corresponding to segments with zero short-term energy in the acoustic feature vectors, natural segmentation of the speech is determined. Simultaneously, based on punctuation marks in the text corresponding to the semantic feature vectors, semantic units such as phrases and clauses are segmented. The natural speech segments are matched with the text semantic units to ensure that the temporal range of each speech segment completely overlaps with the representational range of its corresponding semantic unit. Feature segments corresponding to the speech segments are extracted from the acoustic feature vectors, and features corresponding to the semantic units are extracted from the semantic feature vectors. These two features are combined into a data pair. The aggregation of all such data pairs constitutes the aligned feature pair between the acoustic and semantic feature vectors.
[0032] For each alignment feature pair, content analysis is performed to determine the importance of its acoustic information (such as unique fundamental frequency variations and formant distribution) and semantic information (such as commonly used expressions and vocabulary selection preferences) in reflecting the individual characteristics of the target learner. Alignment feature pairs containing key personalized information are then selected. These selected alignment feature pairs are categorized by information type; for example, acoustic feature pairs related to vocalization habits are grouped into one category, and semantic feature pairs related to language expression habits are grouped into another. The information in each category is summarized, removing duplicate information and retaining core personalized information. All the summarized core information is then integrated into a structured dataset, which represents the personalized feature set of the target learner.
[0033] The beneficial effects include the ability to comprehensively extract acoustic and semantic features from pure speech data and complete cross-modal alignment. The resulting personalized feature set fully preserves the vocal characteristics and language expression habits of the target learner, providing accurate personalized basis for subsequent dynamic text adaptation and integrated speech synthesis. This ensures that subsequent teaching support can closely match the individual characteristics of learners, laying the foundation for improving the effectiveness of teaching support.
[0034] S3. The key language knowledge points are transcribed in a structured manner to obtain the initial teaching narrative text of the key language knowledge points; In this embodiment of the invention, the step of performing structured transcription of the key language knowledge points to obtain the initial teaching narrative text of the key language knowledge points includes: Based on the logical components of the key language knowledge points, the vocabulary, sentence structure and grammatical rules in the key language knowledge points are analyzed. By performing syntactic structure filling on the vocabulary, sentence structure, and grammatical rules, standardized sentence fragments of the key language knowledge points are obtained. Logical semantic connections are made between the standard sentence fragments to obtain the initial teaching narrative text for the key language knowledge points.
[0035] First, the key language knowledge points are logically broken down, divided into logical components according to a complete logical chain of "basic definition - core components - usage scenarios - constraints". Each logical component is further subdivided into several sub-modules. For example, the basic definition module is subdivided into the "concept name - essential attributes - category definition" sub-module, and the core components module is subdivided into the "main elements - relationships between elements" sub-module. For each sub-module, the original expression of the key language knowledge points is studied sentence by sentence. First, the vocabulary carrying the core meaning is extracted, distinguishing between basic general vocabulary and subject-specific terminology to ensure that no core vocabulary is omitted. Then, the sentence structure of each sentence is analyzed to determine whether it is a subject-verb-object structure stating facts, a subject-linking verb-complement structure explaining attributes, an imperative sentence structure making a request, or a causal complex sentence structure explaining reasons, and the composition of each sentence structure is recorded. Finally, grammatical rules are extracted, including vocabulary collocation rules, sentence tense rules, and sentence pattern norms. Through module-by-module and sentence-by-sentence analysis and verification, the vocabulary, sentence structure, and grammatical rules in the key language knowledge points are fully analyzed.
[0036] For each sentence structure identified, its grammatical framework is first broken down sentence by sentence to clarify the type and number of blanks that need to be filled in. For example, the subject-verb-object sentence structure is broken down into "subject blank - verb blank - object blank", the subject-linking verb-predicate sentence structure is broken down into "subject blank - linking verb blank - predicate blank", and the causal complex sentence structure is broken down into "cause clause blank - conjunction blank - result clause blank". Based on the grammatical rules established in this way, the vocabulary type and grammatical requirements corresponding to each blank are determined. For example, the subject blank needs to be filled with a noun referring to a specific thing or concept, the verb blank needs to be filled with a verb expressing an action or state, and the predicate blank needs to be filled with an adjective or noun phrase describing the attributes of the subject. The system selects words that meet the type and grammatical requirements from the parsed vocabulary database. At the same time, it performs adaptability verification by combining the semantic logic of key language knowledge points. For example, in the subject-verb-complement structure of "subject blank + predicate blank + complement blank", words such as "grammatical rules", "is", and "standards of language use" are selected and filled into the corresponding blanks. After filling them in, the word forms are checked again to see if they meet the grammatical rules and whether their semantics are consistent with the core meaning of the knowledge points. After confirming that there are no errors, a complete and standard sentence is formed. All sentences that have been filled and verified are summarized to form the standard sentence fragments of key language knowledge points.
[0037] First, all standardized statement fragments are categorized and semantically analyzed. The core semantics of each fragment are annotated; for example, some fragments are annotated as "basic definition explanation," some as "core component breakdown," some as "practical usage examples," and some as "usage prohibitions." Based on the annotation results, the logical relationships between fragments are analyzed: if fragments correspond to "definition-component-example," they are considered to have a sequential relationship; if fragments are explanations of different dimensions of "component elements," they are considered to have a parallel relationship; if one fragment explains "usage conditions" and another explains "the result after meeting the conditions," they are considered to have a causal relationship. Precise logical connectors are selected based on the type of logical relationship: for parallel relationships, "simultaneously," "in addition," and "on the other hand" are used; for sequential relationships, "firstly," "furthermore," and "on this basis" are used; and for causal relationships, "because," "therefore," and "thus it is evident" are used. Rearrange the standard sentence segments according to the cognitive logical order of the key language knowledge points, insert matching logical connectors between adjacent segments, read through the whole content after insertion, check whether the segments are connected naturally, whether the meaning is coherent, and whether the logic is closed loop. Adjust the connectors or the order of segments for parts that are not connected smoothly, and finally form a text with complete structure, clear logic, and coherent meaning. This text is the initial teaching narrative text for the key language knowledge points.
[0038] The beneficial effects include: through multi-dimensional logical decomposition and sentence-by-sentence analysis, the core language elements of key language knowledge points can be extracted comprehensively and accurately, avoiding omissions; through meticulous sentence structure matching, vocabulary selection, and grammar verification, the generated standardized sentence fragments are ensured to be both grammatically correct and semantically relevant to the knowledge points; and through semantic analysis, logical connection determination, and conjunction adaptation, a logically closed-loop and naturally connected initial teaching narrative text is constructed. The entire process ensures that the initial teaching narrative text comprehensively covers key knowledge points and possesses a clear structure and fluent expression, providing a high-quality, highly usable foundational text for subsequent dynamic stylistic adaptation based on learners' personalized characteristics. This helps learners quickly establish a holistic cognitive framework for key language knowledge points, laying a solid foundation for in-depth understanding.
[0039] S4. Based on the personalized feature set, perform dynamic stylistic adaptation on the initial teaching narrative text to obtain the personalized retelling text for the target learner; In this embodiment of the invention, the step of dynamically adapting the initial instructional narrative text to the personalized retelling text for the target learner based on the personalized feature set includes: The personalized feature set is analyzed by ability dimension to obtain the syntactic mastery parameter and vocabulary familiarity parameter of the target learner; A linguistic structural analysis was performed on the initial teaching narrative text to identify potential high-difficulty syntactic structures and target vocabulary in the initial teaching narrative text; Based on the syntactic mastery parameter and the vocabulary familiarity parameter, the contextual difficulty assessment of the potential high-difficulty syntactic structures and the target vocabulary is performed to obtain the predicted repetition difficulty of the target learner. Based on the predicted retelling difficulty, the initial teaching narrative text is reconstructed using adaptable text to obtain the personalized retelling text for the target learner.
[0040] The method of assessing the contextual difficulty of the potentially high-difficulty syntactic structures and the target vocabulary based on the syntactic mastery parameter and the vocabulary familiarity parameter, to obtain the predicted repetition difficulty of the target learner, includes: Based on the syntactic mastery parameter and the vocabulary familiarity parameter, the syntactic fitness and vocabulary cognition of the target learner are determined; Based on the syntactic fitness and the lexical awareness, the predicted repetition difficulty of the target learner is calculated, wherein the formula for calculating the predicted repetition difficulty is as follows: ; In the formula, This indicates the difficulty of restating the prediction. Indicates the syntactic fitness. This indicates the level of vocabulary recognition. This represents the preset weighting coefficient. This represents the preset weighting coefficient. This represents the preset weighting coefficient. This represents the preset smallest positive number. This indicates the preset scale parameters. Represents the arctangent function. This represents the natural exponential function. This represents the square root operation.
[0041] The process of reconstructing the initial instructional narrative text based on the predicted retelling difficulty to obtain the personalized retelling text for the target learner includes: The predicted paraphrasing difficulty is compared with a preset difficulty response threshold to filter out the syntactic structures to be simplified and the target words to be replaced. Cognitive load reduction is performed on the syntactic structure to be simplified to obtain a simplified syntactic instance of the initial teaching narrative text; The vocabulary to be replaced is benchmarked against the acquisition stage to obtain the appropriate vocabulary for the initial teaching narrative text; The simplified syntax examples and the adapted vocabulary are integrated into the initial teaching narrative text, and the coherence of the integrated initial teaching narrative text is checked. Based on the verification results, a personalized restatement text for the target learner is generated.
[0042] When analyzing the ability dimensions of the personalized feature set, relevant information from the target learner's historical speech repetition samples is first extracted. For the syntactic ability dimension, the total number of times learners correctly used different syntactic structures and the total number of errors when using these syntactic structures are counted in the sample. The ratio of correct times to correct times plus errors is calculated, and this ratio is used as the quantitative result to obtain the target learner's syntactic mastery parameter. For the vocabulary ability dimension, the total number of times learners correctly used various types of vocabulary and the total number of hesitations and pauses when using these vocabulary are counted in the sample. The ratio of correct times to correct times plus hesitations and pauses is calculated, and this ratio is used as the quantitative result to obtain the target learner's vocabulary familiarity parameter.
[0043] When conducting linguistic structural analysis on the initial teaching narrative text, the syntactic structure of the text is broken down sentence by sentence. The syntactic structure of each sentence is compared with a pre-set high-difficulty syntactic structure database. This database contains complex syntactic types that need to be mastered at each stage of the teaching syllabus, such as multi-layered nested compound sentences and inverted sentences. If the syntactic structure of a sentence matches a type in the database, it is marked as a potential high-difficulty syntactic structure in the initial teaching narrative text. At the same time, the vocabulary in the text is analyzed word by word. By comparing it with the vocabulary difficulty level table of the current teaching stage, words that are beyond the learner's current mastery or marked as high-difficulty are selected and marked as target words in the initial teaching narrative text.
[0044] When determining syntactic fitness and vocabulary recognition based on the obtained syntactic mastery and vocabulary familiarity parameters, a pre-set correspondence table of syntactic mastery and syntactic fitness is first retrieved. This table has fixed syntactic fitness values corresponding to different syntactic mastery parameter values based on a large amount of teaching data. By searching for entries in the table that perfectly match the current syntactic mastery parameter, the syntactic fitness of the target learner is determined. Then, a pre-set correspondence table of vocabulary familiarity and vocabulary recognition is retrieved. This table also has fixed vocabulary recognition values corresponding to different vocabulary familiarity parameter values. By searching for entries in the table that perfectly match the current vocabulary familiarity parameter, the vocabulary recognition of the target learner is determined.
[0045] When calculating the predicted repetition difficulty based on syntactic fitness and lexical cognition, the system first obtains three preset fixed weight coefficients (first, second, and third), as well as preset minimum normal numbers and preset scale parameters. The first part of the result is calculated by dividing the first weight coefficient by the square root of the syntactic fitness plus the preset minimum normal number. The second part of the result is calculated by first subtracting the lexical cognition from 1, then dividing this difference by the preset scale parameter, performing an arctangent operation on the result, and multiplying the result by the second weight coefficient. The third part of the result is calculated by first subtracting the syntactic fitness from 1, then subtracting the lexical cognition from 1, performing a natural exponentiation operation on this difference, multiplying the two results together, and then multiplying by the third weight coefficient. The sum of the first, second, and third parts of the result is the predicted repetition difficulty for the target learner.
[0046] When comparing the predicted repetition difficulty with the preset difficulty response threshold, if the predicted repetition difficulty is higher than the difficulty response threshold, it is determined that there is content in the initial teaching narrative text that the learner cannot repetite. From the previously identified potential high-difficulty syntactic structures, structures whose complexity exceeds the learner's ability are selected as syntactic structures to be simplified, and from the identified target vocabulary, words with low learner cognition are selected as words to be replaced. If the predicted repetition difficulty is lower than or equal to the difficulty response threshold, both the syntactic structures to be simplified and the words to be replaced are empty sets, and finally, the syntactic structures to be simplified and the words to be replaced corresponding to the target learner are obtained.
[0047] When reducing the cognitive load on syntactic structures to be simplified, for each syntactic structure to be simplified, if it is a multi-layered nested complex sentence, it is split into two or more simple sentences, each containing only one subject-predicate structure. If it is an inverted sentence, it is adjusted to a sentence with normal word order. At the same time, redundant modifiers that do not affect the core semantics in the syntactic structure are removed to ensure that the simplified syntactic structure retains the key semantics and is in line with the learner's syntactic mastery level. After each syntactic structure to be simplified is processed, a corresponding simplified syntactic instance is formed, and finally, a simplified syntactic instance of the initial teaching narrative text is obtained.
[0048] When benchmarking the vocabulary to be replaced against the learning stage, first retrieve the vocabulary acquisition list of the target learner at the current teaching stage. This list contains all the vocabulary that the learner has mastered at the current stage and their corresponding semantics. For each vocabulary to be replaced, search the list for vocabulary that has the same or similar semantics and that the learner has mastered. If multiple words exist, select the word with the closest semantics and the highest frequency of use as the matching word. If there is no mastered word that completely matches the semantics of the vocabulary to be replaced, select the basic word with the closest semantic relationship as the matching word. Finally, the matching words for the initial teaching narrative text are obtained.
[0049] When integrating simplified syntax examples and matching vocabulary into the initial teaching narrative text, the corresponding syntax structures to be simplified are replaced with simplified syntax examples, and the corresponding words to be replaced are replaced with matching vocabulary, according to the semantic logic and word order of the initial text. Then, the coherence of the integrated initial teaching narrative text is checked. The text is read sentence by sentence to check whether the logical connection between sentences is natural, whether the semantics are complete and coherent, and whether there are semantic contradictions or expression gaps caused by the replacement. If incoherence is found, the connecting words between sentences are adjusted or the word order is slightly adjusted to ensure that the integrated text is logically fluent and semantically accurate.
[0050] When generating personalized paraphrased text based on the coherence verification results, if the integrated text has no logical contradictions, is semantically coherent, and matches the target learner's ability level, then the text is directly identified as the target learner's personalized paraphrased text; if all the problems found in the verification have been corrected and the verification is passed again, the corrected text is also identified as the target learner's personalized paraphrased text, and finally the target learner's personalized paraphrased text is obtained.
[0051] The beneficial effects are as follows: By analyzing the ability dimensions of personalized feature sets, the learner can accurately obtain syntactic mastery parameters and vocabulary familiarity parameters that reflect their syntactic and vocabulary abilities, providing a reliable basis for subsequent dynamic text adaptation; by performing linguistic structural analysis on the initial teaching narrative text, the learner can accurately locate potentially difficult syntactic structures and target vocabulary that they may not be able to master, avoiding adaptation deviations; by comparing with a pre-set correspondence table, the learner's syntactic fitness and vocabulary cognition are determined, and then combined with pre-set fixed weight coefficients, minimum normal numbers, and scale parameters, the first weight coefficient is calculated by dividing the sum of syntactic fitness and minimum normal numbers by the square root; the second weight coefficient is calculated by multiplying by 1 and subtracting the difference in vocabulary cognition by the scale parameter, then the arctangent of the result is calculated; finally, the third weight coefficient is calculated by multiplying by the scale parameter. The accurate predicted repetition difficulty is obtained by multiplying the natural index of the difference between syntactic fitness (1 minus the difference between lexical cognition and 1 minus the difference between vocabulary cognition) and summing the three results. This calculation process balances the impact of syntactic and lexical dimensions on repetition difficulty while avoiding calculation anomalies, ensuring that the difficulty assessment fully matches the learner's actual ability. Based on the predicted repetition difficulty, content to be adjusted is selected, simplified syntactic examples are obtained through cognitive load reduction, and suitable vocabulary is obtained through learning stage benchmarking. After integration and coherence verification, personalized repetition text is generated. The entire process ensures that the personalized repetition text is highly matched with the learner's ability, effectively reducing the learner's repetition difficulty, helping learners better understand key language knowledge points, and ensuring the pertinence and effectiveness of auxiliary teaching, thereby significantly improving the actual effect of speech recognition and synthesis-assisted teaching.
[0052] S5. Map the vocal feature components of the personalized feature set to prosody adjustment parameters and timbre modification parameters; In this embodiment of the invention, the vocal feature components of the personalized feature set are deconstructed at multiple levels to obtain the primary prosodic features, secondary prosodic features, steady-state timbre features and transient timbre features of the vocal feature components; The primary prosodic features and the secondary prosodic features are reduced by spectral features to obtain the prosodic control vector of the vocal feature component; Based on a preset timbre transfer strategy, the steady-state timbre features and the transient timbre features are mapped to derive a timbre influence factor; By performing cross-modal parameter coordination on the prosody control vector and the timbre influence factor, the prosody adjustment parameters and timbre modification parameters of the vocal feature components are obtained.
[0053] First, the vocal feature components in the personalized feature set are decomposed into layers, and basic data related to speech rhythm and pauses are selected. Specifically, the number of syllables per unit time and the interval between speech segments are used as primary prosodic features. The core quantization value is determined by statistically analyzing the average number of syllables in 10 consecutive speech segments, and the auxiliary quantization value is determined by calculating the median of all pause intervals. The range of fundamental frequency variation within a sentence and the short-time energy peak corresponding to stressed syllables are extracted as secondary prosodic features. The difference between the maximum and minimum fundamental frequency values of each sentence is calculated to obtain the intonation fluctuation quantization value, and the ratio of short-time energy of stressed syllables to ordinary syllables is statistically analyzed to obtain the stress intensity quantization value. The average formant frequency and average morphology of the spectral envelope during continuous vowel pronunciation are separated as steady-state timbre features. The average formant frequency of five different vowels is calculated to determine the core parameters, and the average energy distribution of the spectral envelope in the 1kHz-4kHz frequency band is analyzed to determine auxiliary parameters. The formant abrupt value during consonant pronunciation and the instantaneous peak energy of plosives are captured as transient timbre features. The difference between the formant frequencies before and after consonant pronunciation is calculated to obtain the abrupt quantization value. The maximum energy value and occurrence time of plosive pronunciation are recorded to obtain the instantaneous parameters. Finally, the primary prosodic features, secondary prosodic features, steady-state timbre features, and transient timbre features are obtained in their entirety.
[0054] Spectral feature reduction was performed on primary and secondary prosodic features. First, an association mapping table was established, revealing a negative correlation between the average number of syllables in primary prosodic features and the quantized value of intonation fluctuation in secondary prosodic features (specifically, more syllables usually result in smaller intonation fluctuations). Conversely, a positive correlation was found between the median pause interval of primary prosodic features and the quantized value of stress intensity in secondary prosodic features (specifically, longer pause intervals usually result in higher stress intensity). The quantized values of both types of features were standardized to a range of 0-1 and arranged in the order of "primary prosodic feature parameters → secondary prosodic feature parameters," forming a four-dimensional numerical sequence containing the average number of syllables, median pause interval, quantized value of intonation fluctuation, and quantized value of stress intensity. Logical verification was performed on the sequence; for example, when the average number of syllables is 0.8, the quantized value of intonation fluctuation does not exceed 0.5. After successful verification, the prosodic control vector of the vocal feature components was obtained.
[0055] Based on a pre-defined timbre transfer strategy, this strategy aims to preserve the learner's core timbre characteristics while optimizing speech intelligibility to suit the teaching scenario, handling both steady-state and transient timbre features. If the learner's average formant frequency is more than 10% lower than the formant frequency of the standard teaching speech, 30% of the difference is converted into a high-frequency enhancement coefficient; if it is more than 10%, 20% of the difference is converted into a low-frequency enhancement coefficient; if the energy proportion in the 1kHz-2kHz frequency band is less than 30%, 50% of the difference is converted into an energy gain coefficient for that frequency band; if the formant abrupt change value is less than 50Hz, 20% of the difference between 50Hz and the abrupt change value is converted into an abrupt change enhancement coefficient; if the peak energy of a plosive is less than 80% of the standard value, 40% of the difference between the standard value and the peak value is converted into a plosive energy compensation coefficient. All the converted coefficients are organized in the order of "steady-state timbre feature coefficient → transient timbre feature coefficient" to form a timbre influence factor.
[0056] A parameter-coordinated correlation matrix is constructed, defining the correlation weights between each dimension of the prosodic control vector and each coefficient of the timbre influence factor. For example, the weight of the average number of syllables and the spectral high-frequency boost coefficient is -0.2, and the weight of the stress intensity quantization value and the plosive energy compensation coefficient is 0.3. The values of each dimension of the prosodic control vector are adjusted according to these weights. For instance, when the average number of syllables is 0.8 combined with the spectral high-frequency boost coefficient of 0.2, the adjusted value is 0.768. The coefficients of the timbre influence factor are adjusted inversely. For instance, when the stress intensity quantization value is 0.6 combined with the plosive energy compensation coefficient of 0.3, the adjusted coefficient is 0.354. After adjustment, the four dimensions of the prosodic control vector are defined as the speech rate adjustment value, pause duration adjustment value, intonation fluctuation adjustment amplitude, and stress intensity adjustment coefficient, which together constitute the prosodic adjustment parameters. The coefficients of the timbre influence factor are defined as the spectral envelope adjustment coefficient, frequency band energy gain coefficient, formant abrupt change adjustment coefficient, and plosive energy modification coefficient, which together constitute the timbre modification parameters.
[0057] The beneficial effects are that by decomposing the vocal features in layers, different dimensions are accurately distinguished. Through spectral reduction and timbre mapping, basic parameters that fit the learner's characteristics are formed. Then, through cross-modal collaboration, the parameters are optimized. The resulting prosodic adjustment parameters and timbre modification parameters can be highly matched with the learner's vocal habits, providing a precise basis for subsequent integrated speech synthesis. This ensures that the guided demonstration speech is suitable for the learner in terms of both prosody and timbre, effectively improving the naturalness and guidance effect of the demonstration speech and enhancing the personalization level of auxiliary teaching.
[0058] S6. Based on the prosody adjustment parameters and the timbre modification parameters, perform fusion speech synthesis on the personalized recitation text to obtain the guided demonstration speech of the target learner.
[0059] In this embodiment of the invention, the step of performing fusion speech synthesis on the personalized recitation text based on the prosody adjustment parameters and the timbre modification parameters to obtain the guided demonstration speech for the target learner includes: The personalized paraphrased text is subjected to text acoustic analysis to obtain the phoneme sequence and basic prosodic boundaries of the personalized paraphrased text; Based on the prosodic adjustment parameters, the phoneme sequence and the basic prosodic boundary are parametrically marked to obtain the target synthesis specification of the personalized paraphrased text. Acoustic parameter prediction is performed on the target synthesis specification to obtain the initial acoustic parameter sequence of the target synthesis specification; Based on the timbre modification parameters, the initial acoustic parameter sequence is modulated in the timbre dimension to obtain the acoustic parameter sequence of the initial acoustic parameter sequence; Semantic waveform synthesis is performed on the acoustic parameter sequence to obtain the guided demonstration speech of the target learner.
[0060] First, the personalized paraphrased text undergoes text preprocessing. Based on the lexical segmentation rules of the language, the text is broken down into independent words. Then, grammatical attributes are labeled for each word according to part-of-speech tagging standards. Simultaneously, potential word order inversions, omissions, and other expression deviations are corrected to ensure accurate and linguistically compliant text. After preprocessing, a pre-defined phoneme mapping dictionary is used to convert each word into its corresponding phonetic symbol. These phonetic symbols are then broken down into the smallest phonetic units, i.e., phonemes. All phonemes are concatenated according to the order of words in the text to form the phoneme sequence of the personalized paraphrased text. Subsequently, punctuation marks and semantic pauses in the text are analyzed. Combining this with standard pause durations in language expression, pause positions and basic pause durations are marked in the phoneme sequence. These markers collectively constitute the basic prosodic boundaries of the personalized paraphrased text.
[0061] First, clarify the prosodic adjustment parameters, including speech rate adjustment values, intonation rise and fall ranges, and pause duration correction values. For the phoneme sequence, calculate the standard duration of each phoneme based on the speech rate adjustment value. If the speech rate needs to be increased, shorten the duration of each phoneme; if the speech rate needs to be decreased, extend the duration of each phoneme. Mark the specific adjusted duration value next to each phoneme. For the basic prosodic boundaries, adjust the pause duration at each marked point based on the pause duration correction value. Simultaneously, mark the intonation change trend next to the corresponding phoneme sequence segment based on the intonation rise and fall range, such as the rising interval from low to high and the falling interval from high to low. Integrate the phoneme sequence with marked phoneme durations, intonation change trends, and adjusted pause durations to form the target synthesis specification of a personalized retelling text that meets the complete synthesis requirements.
[0062] A pre-defined acoustic parameter prediction model, trained on a large number of speech samples, is invoked to predict corresponding acoustic parameters based on the prosodic features of speech. Information such as the phoneme sequence, phoneme duration, intonation trends, and pause duration from the target synthesis specification is input into the model. The model first performs frame-by-frame analysis of the phoneme sequence, with each frame taking 20 milliseconds, predicting the fundamental frequency, the first three formant frequencies, and the short-time energy value for each frame. During prediction, the model smooths the parameters based on the prosodic relationships between adjacent frames, avoiding abrupt changes in acoustic parameters between adjacent frames and ensuring continuous and natural parameter variations. The fundamental frequency, formant frequencies, and short-time energy values of all frames are arranged chronologically to form an ordered parameter set, which constitutes the initial acoustic parameter sequence for the target synthesis specification.
[0063] First, the timbre modification parameters include the spectral envelope adjustment coefficient, harmonic component ratio, and noise component control value. For each frame of parameters in the initial acoustic parameter sequence, the formant frequency distribution of that frame is adjusted according to the spectral envelope adjustment coefficient. If the coefficient is positive, the high-frequency formant frequency is appropriately increased to enhance the brightness of the speech; if the coefficient is negative, the high-frequency formant frequency is appropriately decreased to increase the fullness of the speech. Next, the energy ratio of the fundamental frequency to each harmonic is adjusted according to the harmonic component ratio. To highlight the learner's vocal characteristics, the energy proportion of higher harmonics is increased; to enhance the softness of the speech, the energy proportion of the fundamental frequency is increased. Finally, the short-time energy value is fine-tuned according to the noise component control value to suppress any possible redundant noise energy, ensuring that each frame of parameters meets the timbre modification requirements. All the adjusted acoustic parameters of all frames are recombine in their original chronological order to obtain the acoustic parameter sequence of the adjusted initial acoustic parameter sequence.
[0064] A waveform splicing and synthesis algorithm is employed. First, basic speech waveform segments corresponding to each phoneme in the acoustic parameter sequence are retrieved from a pre-defined speech waveform library. Based on the fundamental frequency, formant frequency, and short-time energy value of each frame in the acoustic parameter sequence, the basic waveform segment that best matches the parameters of that frame is selected. The selected waveform segments are spliced sequentially in time. During the splicing process, a smooth transition is performed on the 5-millisecond overlap region between adjacent segments. By adjusting the amplitude variation curve of the overlap region, sudden changes in volume or timbre discontinuities are avoided at the splicing points. Simultaneously, silent waveform segments are inserted at the corresponding splicing positions according to the pause duration in the acoustic parameter sequence to ensure that the pause effect meets the target synthesis specifications. After all segments are spliced and smoothed, a complete time-domain speech waveform is generated. This waveform is then converted to standard audio formats such as WAV, and the final audio file serves as the guided demonstration speech for the target learner.
[0065] The beneficial effects are as follows: the text acoustic analysis ensures accurate phoneme and prosodic boundaries; the prosodic adjustment parameters form a clear target synthesis specification; the acoustic parameters are matched with the learner's personalized characteristics through acoustic parameter prediction and timbre modification parameter modulation; and finally, the semantic waveform synthesis generates a coherent, natural, and learner-appropriate guided demonstration speech. This speech can accurately match the learner's vocal habits and cognitive acceptance, effectively improving the guiding effect of the demonstration speech, helping learners to more efficiently imitate and understand personalized repetition texts, and further ensuring the practical application value of speech recognition and synthesis-assisted teaching.
[0066] like Figure 2 The diagram shown is a functional block diagram of a speech recognition and synthesis-assisted teaching system provided in an embodiment of the present invention.
[0067] The speech recognition and synthesis-assisted teaching system 100 described in this invention can be installed in an electronic device. Depending on the functions implemented, the speech recognition and synthesis-assisted teaching system 100 may include a data acquisition and parsing module 101, a feature fusion module 102, a text generation module 103, a style adaptation module 104, a parameter mapping module 105, and a speech synthesis module 106. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, stored in the memory of the electronic device.
[0068] In this embodiment, the functions of each module / unit are as follows: The data acquisition and parsing module 101 is used to acquire the pure speech data of the target learner and, according to the standardized curriculum outline, parse out the key language knowledge points of the current teaching content. The feature fusion module 102 is used to perform heterogeneous feature fusion on the clean speech data to obtain the personalized feature set of the target learner. The text generation module 103 is used to perform structured transcription of the key language knowledge points to obtain the initial teaching narrative text of the key language knowledge points. The style adaptation module 104 is used to dynamically adapt the initial teaching narrative text to the style based on the personalized feature set, so as to obtain the personalized retelling text of the target learner. The parameter mapping module 105 is used to map the vocal feature components of the personalized feature set into prosody adjustment parameters and timbre modification parameters. The speech synthesis module 106 is used to perform fusion speech synthesis on the personalized recitation text based on the prosody adjustment parameters and the timbre modification parameters to obtain the guided demonstration speech of the target learner.
[0069] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0070] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0071] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0072] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0073] This application embodiment can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A speech recognition and synthesis-assisted teaching method, characterized in that, The method includes: S1. Obtain the clean speech data of the target learner and, based on the standardized curriculum outline, analyze the key language knowledge points of the current teaching content; S2. Perform heterogeneous feature fusion on the clean speech data to obtain the personalized feature set of the target learner; S3. The key language knowledge points are transcribed in a structured manner to obtain the initial teaching narrative text of the key language knowledge points; S4. Based on the personalized feature set, perform dynamic stylistic adaptation on the initial teaching narrative text to obtain the personalized retelling text for the target learner; S5. Map the vocal feature components of the personalized feature set to prosody adjustment parameters and timbre modification parameters; S6. Based on the prosody adjustment parameters and the timbre modification parameters, perform fusion speech synthesis on the personalized recitation text to obtain the guided demonstration speech of the target learner.
2. The speech recognition and synthesis-assisted teaching method as described in claim 1, characterized in that, The process of acquiring the target learner's clean speech data and, based on a standardized curriculum syllabus, analyzing the key language knowledge points of the current teaching content includes: Retrieve learner identifiers and teaching stage identifiers of target learners from the learning record database; Based on the learner identifier and the teaching stage identifier, retrieve the target learner's historical speech repetition samples; Unify the audio specifications of the historical speech retelling samples to obtain standardized speech data of the historical speech retelling samples; By analyzing the short-time energy characteristics and zero-crossing rate characteristics of the standardized speech data, long silent segments and noise segments of the target learner are removed to obtain the clean speech data of the standardized speech data. Based on the teaching stage identifiers, locate the corresponding teaching content in the standardized curriculum syllabus, and select the key language knowledge points of the corresponding teaching content.
3. The speech recognition and synthesis-assisted teaching method as described in claim 1, characterized in that, The heterogeneous feature fusion of the clean speech data to obtain the personalized feature set of the target learner includes: Obtain the Mel frequency cepstral coefficient sequence, fundamental frequency trajectory, and formant distribution parameters of the clean speech data; Based on the Mel frequency cepstral coefficient sequence, the fundamental frequency trajectory, and the formant distribution parameters, an acoustic feature vector of the clean speech data is constructed. Semantic text analysis is performed on the clean speech data to obtain the semantic feature vector of the clean speech data; Cross-modal alignment is performed on the acoustic feature vector and the semantic feature vector to obtain aligned feature pairs of the acoustic feature vector and the semantic feature vector; Based on the importance of the aligned feature pairs, feature integration is performed on the aligned feature pairs to obtain the personalized feature set of the target learner.
4. The speech recognition and synthesis-assisted teaching method as described in claim 1, characterized in that, The process of structurally transcribing the key language knowledge points to obtain the initial teaching narrative text for the key language knowledge points includes: Based on the logical components of the key language knowledge points, the vocabulary, sentence structure and grammatical rules in the key language knowledge points are analyzed. By filling in the syntactic structure of the vocabulary, the sentence structure, and the grammatical rules, a standardized sentence fragment of the key language knowledge points is obtained. Logical semantic connections are made between the standard sentence fragments to obtain the initial teaching narrative text for the key language knowledge points.
5. The speech recognition and synthesis-assisted teaching method as described in claim 1, characterized in that, The process of dynamically adapting the initial instructional narrative text to the personalized retelling text for the target learner based on the personalized feature set includes: The personalized feature set is analyzed by ability dimension to obtain the syntactic mastery parameter and vocabulary familiarity parameter of the target learner; A linguistic structural analysis was performed on the initial teaching narrative text to identify potential high-difficulty syntactic structures and target vocabulary in the initial teaching narrative text; Based on the syntactic mastery parameter and the vocabulary familiarity parameter, the contextual difficulty assessment of the potential high-difficulty syntactic structures and the target vocabulary is performed to obtain the predicted repetition difficulty of the target learner. Based on the predicted retelling difficulty, the initial teaching narrative text is reconstructed using adaptable text to obtain the personalized retelling text for the target learner.
6. The speech recognition and synthesis-assisted teaching method as described in claim 5, characterized in that, The method of assessing the contextual difficulty of the potentially high-difficulty syntactic structures and the target vocabulary based on the syntactic mastery parameter and the vocabulary familiarity parameter, to obtain the predicted repetition difficulty of the target learner, includes: Based on the syntactic mastery parameter and the vocabulary familiarity parameter, the syntactic fitness and vocabulary cognition of the target learner are determined; Based on the syntactic fitness and the lexical awareness, the predicted repetition difficulty of the target learner is calculated, wherein the formula for calculating the predicted repetition difficulty is as follows: ; In the formula, This indicates the difficulty of restating the prediction. Indicates the syntactic fitness. This indicates the level of vocabulary recognition. This represents the preset weighting coefficient. This represents the preset weighting coefficient. This represents the preset weighting coefficient. This represents the preset smallest positive number. This indicates the preset scale parameters. Represents the arctangent function. This represents the natural exponential function. This represents the square root operation.
7. The speech recognition and synthesis-assisted teaching method as described in claim 5, characterized in that, The process of reconstructing the initial instructional narrative text based on the predicted retelling difficulty to obtain the personalized retelling text for the target learner includes: The predicted paraphrasing difficulty is compared with a preset difficulty response threshold to filter out the syntactic structures to be simplified and the target words to be replaced. Cognitive load reduction is performed on the syntactic structure to be simplified to obtain a simplified syntactic instance of the initial teaching narrative text; The vocabulary to be replaced is benchmarked against the acquisition stage to obtain the appropriate vocabulary for the initial teaching narrative text; The simplified syntax examples and the adapted vocabulary are integrated into the initial teaching narrative text, and the coherence of the integrated initial teaching narrative text is checked. Based on the verification results, a personalized restatement text for the target learner is generated.
8. The speech recognition and synthesis-assisted teaching method as described in claim 1, characterized in that, The step of mapping the vocal feature components of the personalized feature set to prosody adjustment parameters and timbre modification parameters includes: The vocal feature components of the personalized feature set are deconstructed at multiple levels to obtain the primary prosodic features, secondary prosodic features, steady-state timbre features, and transient timbre features of the vocal feature components; The primary prosodic features and the secondary prosodic features are reduced by spectral features to obtain the prosodic control vector of the vocal feature component; Based on a preset timbre transfer strategy, the steady-state timbre features and the transient timbre features are mapped to derive a timbre influence factor; By performing cross-modal parameter coordination on the prosody control vector and the timbre influence factor, the prosody adjustment parameters and timbre modification parameters of the vocal feature components are obtained.
9. The speech recognition and synthesis-assisted teaching method as described in claim 1, characterized in that, The process of performing fusion speech synthesis on the personalized recitation text based on the prosodic adjustment parameters and the timbre modification parameters to obtain the guided demonstration speech for the target learner includes: The personalized paraphrased text is subjected to text acoustic analysis to obtain the phoneme sequence and basic prosodic boundaries of the personalized paraphrased text; Based on the prosodic adjustment parameters, the phoneme sequence and the basic prosodic boundary are parametrically marked to obtain the target synthesis specification of the personalized paraphrased text. Acoustic parameter prediction is performed on the target synthesis specification to obtain the initial acoustic parameter sequence of the target synthesis specification; Based on the timbre modification parameters, the initial acoustic parameter sequence is modulated in the timbre dimension to obtain the acoustic parameter sequence of the initial acoustic parameter sequence; Semantic waveform synthesis is performed on the acoustic parameter sequence to obtain the guided demonstration speech of the target learner.
10. A speech recognition and synthesis-assisted teaching system for implementing the speech recognition and synthesis-assisted teaching method of claim 1, the system comprising: The data acquisition and parsing module is used to acquire clean speech data from the target learners and, based on the standardized curriculum outline, to parse out the key language knowledge points of the current teaching content. The feature fusion module is used to perform heterogeneous feature fusion on the clean speech data to obtain the personalized feature set of the target learner; The text generation module is used to perform structured transcription of the key language knowledge points to obtain the initial teaching narrative text of the key language knowledge points; The style adaptation module is used to dynamically adapt the initial teaching narrative text to the style based on the personalized feature set, so as to obtain the personalized retelling text for the target learner. The parameter mapping module is used to map the vocal feature components of the personalized feature set into prosody adjustment parameters and timbre modification parameters; The speech synthesis module is used to perform fusion speech synthesis on the personalized recitation text based on the prosody adjustment parameters and the timbre modification parameters to obtain the guided demonstration speech of the target learner.