A target language acquisition fostering system based on staged immersive context generation
By generating phased ability labels through multi-source data collection and hierarchical Bayesian inference, learning resource selection and interactive training are driven, forming a closed-loop mechanism. This solves the problems of insufficient recognition of individual learner differences and imprecise feedback in existing systems, and realizes a personalized and efficient language learning path.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QINGDAO LANYI TIANQING EDUCATION TECHNOLOGY CO LTD
- Filing Date
- 2025-10-15
- Publication Date
- 2026-05-08
AI Technical Summary
Existing target language learning systems lack dynamic identification of individual learners' differences, cannot effectively adjust according to cognitive stages, and the feedback process fails to combine multi-source data for refined analysis, resulting in a lack of continuity in the learning process.
It employs multi-source learning data acquisition and processing, hierarchical Bayesian inference, contextualized learning resource selection, and multimodal interactive training feedback. Through phased ability tags, it drives learning resource selection, interactive training, and phased updates, forming a closed-loop mechanism.
It achieves a high degree of personalization and dynamically adjustable learning paths, improves language acquisition efficiency, ensures the scientific nature and repeatability of stage assessment results, and provides real-time and continuous feedback information.
Smart Images

Figure CN121235120B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent language learning technology, and in particular to a target language acquisition and development system based on staged immersive context generation. Background Technology
[0002] Existing target language learning systems generally employ fixed courses and uniform difficulty levels, typically providing vocabulary, grammar, listening, reading, speaking, and writing content through online platforms or mobile applications. While these methods cover basic skills, they lack dynamic identification of individual learner differences and struggle to effectively adjust to learners' cognitive stages. Some systems attempt to introduce adaptive mechanisms, but these often rely solely on test accuracy, failing to comprehensively model multi-dimensional features and resulting in limited personalized effects.
[0003] Existing technologies have attempted to analyze learning behavior using statistical models or machine learning methods, but they lack a clear hierarchical structure in stage division and progression determination, resulting in somewhat one-sided outputs that fail to reflect the dependencies between different language skills. Furthermore, learning resource recommendations often rely on manual rules or shallow feature matching, and resource selection and context generation fail to form a closed loop with stage determination results, leading to a lack of coherence in the learning process.
[0004] Furthermore, the feedback mechanisms in existing systems are mostly simple prompts or direct error corrections, failing to integrate multi-source data for refined analysis, and the prompts are sometimes excessive or insufficient. The assessment of transfer learning also generally lacks mathematical and repeatable criteria, relying more on empirical thresholds, which makes it difficult to guarantee objectivity.
[0005] Therefore, how to provide a target language acquisition and development system based on phased immersive context generation is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] One objective of this invention is to propose a target language acquisition and training system based on staged immersive context generation. This invention utilizes multi-source learning data acquisition and processing, hierarchical Bayesian inference, contextualized learning resource selection, and multimodal interactive training feedback. It describes in detail the closed-loop mechanism of learning resource selection, interactive training, and stage updates driven by staged ability tags, which has the advantages of high personalization, dynamically adjustable learning paths, and improved language acquisition efficiency.
[0007] A target language acquisition and training system based on staged immersive context generation according to an embodiment of the present invention includes:
[0008] The learning data acquisition module is used to collect and process multi-source data and output a standardized learning dataset.
[0009] The cognitive diagnosis module receives a standardized learning dataset and outputs stage-based ability labels based on a hierarchical Bayesian cognitive diagnosis model.
[0010] The Stage Context Resource Module is used to store and manage contextualized learning resources divided into stages, and to call up a set of learning resources that match the learner's current stage based on the stage's ability tags.
[0011] The interactive training and feedback module is used to conduct multimodal interactive training, collect performance data in real time, and generate error correction information and minimize prompts.
[0012] The phase update module is used to receive real-time performance data, perform posterior probability updates of latent variables, generate new phase capability labels, and generate migration signals when advanced conditions are met.
[0013] The migration control and path scheduling module is used to receive migration signals, drive the stage context resource module to call learning resources of higher stages, control learning path scheduling, and record learner stage changes.
[0014] Optionally, modules can be integrated using the following methods:
[0015] Collect and process multi-source data from learners during language training to form a standardized learning dataset;
[0016] Input the standardized learning dataset into the hierarchical Bayesian cognitive diagnostic model and output stage-based ability labels;
[0017] Based on the stage-based ability tags, learning resources that match the learner's current stage are selected from the stage-based context resource library to obtain a combined learning resource set;
[0018] Multimodal interactive training is conducted for learners in an immersive context. During the interactive training process, real-time performance data of learners is collected, and error correction information and minimize prompts are generated.
[0019] The real-time performance dataset is input into the hierarchical Bayesian cognitive diagnostic model, the posterior probability of the latent variables is updated, new stage ability labels are generated, and a transfer signal is generated when the advancement conditions are met.
[0020] Based on the new stage-based competency labels, the stage-based contextual resource selection, immersive contextual interactive training, and stage updates are re-executed until stage transfer is triggered, forming a dynamic closed-loop language acquisition path.
[0021] Optionally, the step of collecting and processing multi-source data from learners during language training to form a standardized learning dataset specifically includes:
[0022] The study collects raw data from learners across six dimensions: vocabulary, grammar, listening, reading, speaking, and writing. This raw data includes vocabulary reaction time, vocabulary accuracy, grammar structure recognition accuracy, number of grammar errors, listening keyword recognition rate, listening comprehension response delay time, reading accuracy (intensive reading), reading accuracy (extensive reading), speaking speech recognition similarity, speaking pronunciation score, writing grammar accuracy, and writing vocabulary diversity index.
[0023] Perform data cleaning operations on the original data to remove missing and outlier values, and obtain the cleaned original dataset.
[0024] Missing value imputation, timestamp alignment, and feature standardization are performed on the cleaned original dataset. Feature standardization unifies data of different dimensions to the range of zero mean and unit standard deviation by subtracting the mean from each feature value and dividing by the standard deviation.
[0025] The standardized features are combined in the following order: vocabulary reaction time, vocabulary accuracy, grammatical structure recognition accuracy, number of grammatical errors, listening keyword recognition rate, listening answer delay time, reading intensive reading accuracy, reading extensive reading accuracy, spoken speech recognition similarity, spoken pronunciation score, writing grammar accuracy, and writing vocabulary diversity index to construct a standardized learning dataset.
[0026] Optionally, the step of inputting the standardized learning dataset into the hierarchical Bayesian cognitive diagnostic model and outputting stage-based ability labels specifically includes:
[0027] The standardized learning dataset is input into the hierarchical Bayesian cognitive diagnostic model, in which a latent variable layer and an observation layer are established.
[0028] The latent variable layer includes vocabulary latent variables, grammar latent variables, listening latent variables, reading latent variables, speaking latent variables, and writing latent variables;
[0029] The observation layer includes vocabulary reaction time, vocabulary accuracy, grammatical structure recognition accuracy, number of grammatical errors, listening keyword recognition rate, listening answer delay time, reading intensive reading accuracy, reading extensive reading accuracy, spoken language speech recognition similarity, spoken language pronunciation score, writing grammar accuracy and writing vocabulary diversity index.
[0030] In the latent variable layer, a hierarchical dependency structure is set up, defining lexical latent variables and grammatical latent variables as low-level latent variables, listening latent variables and speaking latent variables as mid-level latent variables, and reading latent variables and writing latent variables as high-level latent variables.
[0031] The prior conditions of the middle-level latent variables depend on the posterior results of the lower-level latent variables, and the prior conditions of the higher-level latent variables depend on the posterior results of the middle-level latent variables.
[0032] A prior distribution is set for each latent variable, where the vocabulary latent variable, grammar latent variable, listening latent variable, reading latent variable, speaking latent variable and writing latent variable correspond to the learner’s mastery probability in the six dimensions, respectively. Each latent variable is defined by a set of hyperparameters, which are determined by the number of correct and incorrect performances in the historical learning samples of that dimension.
[0033] Define a likelihood function between the observation layer and the latent variable layer, where each latent variable establishes a conditional probability relationship with its corresponding observed feature. The likelihood function represents the probability of observed data occurring given the value of the latent variable.
[0034] Specifically, each observation feature in the standardized learning dataset is modeled one by one. A joint probability distribution is established between vocabulary reaction time, vocabulary accuracy, grammatical structure recognition accuracy, number of grammatical errors, listening keyword recognition rate, listening answer delay, reading intensive reading accuracy, reading extensive reading accuracy, spoken speech recognition similarity, spoken pronunciation score, writing grammar accuracy, and writing vocabulary diversity index and the corresponding latent variables. This forms the overall probability of learners generating all observation data under the latent variable state.
[0035] The posterior probability of each latent variable is calculated using the Bayesian inference formula to obtain the set of posterior probabilities of the latent variables, and then the updates are performed sequentially under the predefined hierarchical dependency structure.
[0036] That is, firstly, the posterior probability of the lower-level latent variables is calculated using the observation layer data, then the posterior result of the lower-level latent variables is used as the prior condition of the middle-level latent variables to update the posterior probability of the middle-level latent variables, and then the posterior result of the middle-level latent variables is used as the prior condition of the higher-level latent variables to update the posterior probability of the higher-level latent variables, thus forming a bottom-up layer-by-layer inference and update process.
[0037] Based on the updated posterior probabilities of each latent variable, the learner's mastery probability in six dimensions—vocabulary, grammar, listening, reading, speaking, and writing—is output. Stage-based ability labels are generated according to preset thresholds, and these stage-based ability labels serve as input conditions for selecting learning resources in the stage-based contextual resource library.
[0038] Optionally, the step of selecting learning resources that match the learner's current stage from the stage-specific contextual resource library based on stage-specific ability tags to obtain a combined learning resource set specifically includes:
[0039] Establish a phased contextual resource library, indexing and dividing learning resources according to different phases, with each phase index set corresponding to a set of learning resources;
[0040] Based on the stage-based ability tags, the vocabulary stage-based ability tags, grammar stage-based ability tags, listening stage-based ability tags, reading stage-based ability tags, speaking stage-based ability tags, and writing stage-based ability tags are matched with the corresponding stage index sets in the stage context resource library to determine the learning resources associated with the stage index set corresponding to each dimension.
[0041] During the matching process, the stage ability tags are mapped to the stage index set in the stage context resource library. Specifically, based on the stage ability tags of six dimensions, namely vocabulary, grammar, listening, reading, speaking and writing, the learning resource set corresponding to each dimension tag is called respectively. This mapping relationship is defined as a calling function from stage ability tag to learning resource set, so that each dimension stage ability tag can be matched with a set of learning resources.
[0042] The sets of vocabulary learning resources, grammar learning resources, listening learning resources, reading learning resources, speaking learning resources, and writing learning resources obtained by filtering according to the six-dimensional stage ability tags will be merged to form a combined learning resource set that includes learning resources from all dimensions.
[0043] Optionally, the step of conducting multimodal interactive training for learners in an immersive context, collecting real-time performance data of learners during the interactive training process, and generating error correction information and minimization prompts specifically includes:
[0044] Select contextualized learning resources that correspond to the learner's current stage of ability tags from the combined learning resource set, and construct immersive interactive scenarios;
[0045] Perform interactive training in six dimensions: vocabulary, grammar, listening, reading, speaking, and writing in an immersive interactive scenario;
[0046] Among them, vocabulary training triggers vocabulary responses through images and audio, grammar training triggers grammatical structure recognition through sentence highlighting, listening training triggers keyword recognition through audio segments, reading training triggers semantic understanding through text annotation, speaking training triggers pronunciation scoring through voice input, and writing training triggers grammatical accuracy and vocabulary diversity calculation through text input.
[0047] During the immersive interactive training process, learners' real-time performance data is collected, including vocabulary reaction time, vocabulary accuracy, grammatical structure recognition accuracy, number of grammatical errors, listening keyword recognition rate, listening answer delay time, reading intensive reading accuracy, reading extensive reading accuracy, spoken language speech recognition similarity, spoken language pronunciation score, writing grammar accuracy, and writing vocabulary diversity index, to construct a real-time performance dataset.
[0048] The real-time performance dataset is input into the speech recognition module, the natural language understanding module, and the speech evaluation module. The speech recognition module receives the learner's spoken speech signal, performs acoustic feature extraction and text transcription, outputs the spoken recognition text, and calculates the similarity with the standard answer.
[0049] The natural language understanding module receives the recognized text output by the speech recognition module and the reading and writing input text from the learner, performs syntactic structure parsing and semantic component analysis, and outputs grammatical error features and semantic matching results;
[0050] The speech evaluation module receives the learner's acoustic parameters, which include the speech signal, extracts the fundamental frequency, formants, energy distribution, and pronunciation duration, calculates the pronunciation accuracy score, and outputs the pronunciation score result. After completing the three types of processing, the speech recognition result, grammatical error features, semantic matching result, and pronunciation score result are integrated to generate an error correction information set.
[0051] Based on the error correction information set, the minimum prompt required for the learner's input is determined. That is, the most critical prompt for the current input is selected from all error correction information as the output prompt, thereby generating the minimum prompt result and using it together with the error correction information as feedback content.
[0052] Optionally, the step of inputting the real-time performance dataset into the hierarchical Bayesian cognitive diagnostic model, performing latent variable posterior probability updates, generating new stage-specific ability labels, and generating transfer signals when advancement conditions are met specifically includes:
[0053] Real-time performance data is input into the hierarchical Bayesian cognitive diagnostic model. Based on the hierarchical dependency structure, the latent variable layer is updated using real-time performance data to maintain the correspondence between vocabulary latent variables, grammar latent variables, listening latent variables, reading latent variables, speaking latent variables and writing latent variables.
[0054] The posterior probability of each latent variable is updated. Specifically, given the real-time performance dataset, the likelihood function of the real-time performance data under the latent variable state is jointly calculated with the prior distribution of the original latent variable. The posterior probability of the latent variable is then solved using the Bayesian inference formula, thereby obtaining the learner's mastery probability in each ability dimension under the real-time performance data condition.
[0055] In the hierarchical dependency structure, the posterior probability of the lower-level latent variables is used as the prior input of the middle-level latent variables, and the posterior probability of the middle-level latent variables is used as the prior input of the higher-level latent variables, thus completing the hierarchical update from bottom to top.
[0056] Based on the updated posterior probabilities of the latent variables, new stage-specific capability labels are generated according to a preset threshold mapping rule;
[0057] The advancement conditions for each dimension's stage-based capability tags are determined, and these advancement conditions include two aspects of constraints:
[0058] First, the updated posterior probability of the latent variable must be greater than the corresponding preset threshold.
[0059] Second, it requires that the aforementioned probability determination conditions be met in several consecutive updates;
[0060] When both of the above constraints are met simultaneously, a migration signal is generated and a stage migration is triggered. The migration signal is used to control the stage context resource library to call learning resources of higher stages and update the learning path scheduling; otherwise, the original stage state remains unchanged.
[0061] Optionally, the process of re-executing stage-based contextual resource selection, immersive contextual interaction training, and stage updates based on new stage-based ability labels, until stage transfer is triggered, to form a dynamic closed-loop language acquisition path, specifically includes:
[0062] Based on the new stage-based ability tags, the stage-based context resource library is invoked to determine the matching set of learning resources for each stage-based ability tag in each dimension. Through mapping relationships, a set of learning resources in six dimensions corresponding to the updated stage-based ability tags is generated to form a new combined learning resource set.
[0063] The new combined learning resource set is input into the immersive context interaction training and real-time feedback steps to generate a new real-time performance dataset.
[0064] The new real-time performance dataset is input into the hierarchical Bayesian cognitive diagnostic model to calculate the updated posterior probabilities of latent variables and generate new stage-specific ability labels.
[0065] Determine whether the new stage ability label meets the advancement conditions. If it does, generate a transfer signal and trigger stage transfer, while recording the learner's stage change.
[0066] If the requirements are not met, learning resources will continue to be selected based on the new stage-based ability labels, a new set of combined learning resources will be generated, new immersive contextual interactive training will be carried out and new real-time performance data will be collected. The data will then be input into the hierarchical Bayesian cognitive diagnostic model for updating, forming a dynamic closed-loop language acquisition path.
[0067] The beneficial effects of this invention are:
[0068] This invention utilizes a hierarchical Bayesian cognitive diagnostic model to model learning performance across six dimensions: vocabulary, grammar, listening, reading, speaking, and writing. It achieves layer-by-layer updates of posterior probabilities through hierarchical dependencies between low-, mid-, and high-level latent variables. This mechanism establishes a rigorous mathematical mapping between observed data and latent abilities, generating stage-based ability labels. Compared to existing schemes that rely on answer accuracy or single-dimensional assessments, this more comprehensively reflects learners' multi-dimensional cognitive states, ensuring the scientific validity and repeatability of stage-based assessment results.
[0069] Regarding the access to learning resources, this invention establishes a phased contextual resource library, which maps contextualized learning materials for different phases to phased ability tags and uses a mapping function to complete resource filtering. This approach enables learners to obtain resources that match their current cognitive state when entering a new learning phase, avoiding the problem of resource difficulty being out of sync with learners' levels in existing technologies, while ensuring coverage of both daily and academic contexts.
[0070] In the interactive training phase, this invention transforms speech recognition, natural language understanding, and speech evaluation results into error correction information and generates feedback by combining them with a minimized prompt function, thereby reducing redundant information while ensuring the effectiveness of the feedback. This mechanism improves the continuity of the interactive process, enabling feedback to remain real-time without interrupting the immersive training experience.
[0071] In terms of learning path scheduling, this invention determines whether a learner meets the transfer requirements by setting advancement thresholds and continuous stability conditions. When the conditions are met, a transfer signal is generated, driving the system to access learning resources at a higher stage. This control mechanism transforms the probabilistic determination results into executable system operations, achieving automated stage transfer and dynamic closed-loop management of the learning path. Attached Figure Description
[0072] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0073] Figure 1 This is a flowchart of a method for a target language acquisition and development system based on staged immersive context generation proposed in this invention.
[0074] Figure 2 This is a schematic diagram of a target language acquisition and development system based on staged immersive context generation proposed in this invention.
[0075] Figure 3This is a schematic diagram of the structure of a hierarchical Bayesian cognitive diagnostic model in a target language acquisition and development system based on staged immersive context generation proposed in this invention. Detailed Implementation
[0076] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0077] refer to Figure 1-3 A target language acquisition and development system based on phased immersive context generation includes:
[0078] The learning data acquisition module is used to collect and process data from multiple sources and output a standardized learning dataset.
[0079] The cognitive diagnosis module receives a standardized learning dataset and outputs stage-based ability labels based on a hierarchical Bayesian cognitive diagnosis model.
[0080] The Stage Context Resource Module is used to store and manage contextualized learning resources divided into stages, and to call up a set of learning resources that match the learner's current stage based on the stage's ability tags.
[0081] The interactive training and feedback module is used to conduct multimodal interactive training, collect performance data in real time, and generate error correction information and minimize prompts.
[0082] The phase update module is used to receive real-time performance data, perform posterior probability updates of latent variables, generate new phase capability labels, and generate migration signals when advanced conditions are met.
[0083] The migration control and path scheduling module is used to receive migration signals, drive the stage context resource module to call learning resources of higher stages, control learning path scheduling, and record learner stage changes.
[0084] In this embodiment, the modules are connected through the following method:
[0085] Collect and process multi-source data from learners during language training to form a standardized learning dataset;
[0086] Input the standardized learning dataset into the hierarchical Bayesian cognitive diagnostic model and output stage-based ability labels;
[0087] Based on the stage-based ability tags, learning resources that match the learner's current stage are selected from the stage-based context resource library to obtain a combined learning resource set;
[0088] Multimodal interactive training is conducted for learners in an immersive context. During the interactive training process, real-time performance data of learners is collected, and error correction information and minimize prompts are generated.
[0089] The real-time performance dataset is input into the hierarchical Bayesian cognitive diagnostic model, the posterior probability of the latent variables is updated, new stage ability labels are generated, and a transfer signal is generated when the advancement conditions are met.
[0090] Based on the new stage-based competency labels, the stage-based contextual resource selection, immersive contextual interactive training, and stage updates are re-executed until stage transfer is triggered, forming a dynamic closed-loop language acquisition path.
[0091] In this embodiment, the step of collecting and processing multi-source data from learners during language training to form a standardized learning dataset specifically includes:
[0092] The study collects raw data from learners across six dimensions: vocabulary, grammar, listening, reading, speaking, and writing. This raw data includes vocabulary reaction time, vocabulary accuracy, grammar structure recognition accuracy, number of grammar errors, listening keyword recognition rate, listening comprehension response delay time, reading accuracy (intensive reading), reading accuracy (extensive reading), speaking speech recognition similarity, speaking pronunciation score, writing grammar accuracy, and writing vocabulary diversity index.
[0093] Perform data cleaning operations on the original data to remove missing and outlier values, and obtain the cleaned original dataset.
[0094] Missing value imputation, timestamp alignment, and feature standardization are performed on the cleaned original dataset. Missing value imputation is used to fill in data gaps to ensure data integrity. Timestamp alignment is used to ensure that data of different dimensions correspond consistently under the same time reference. Feature standardization is performed by subtracting the mean of each feature value and dividing by the standard deviation to unify data of different dimensions to the range of zero mean and one standard deviation, so as to eliminate the impact of dimensional differences on subsequent analysis.
[0095] The standardized features are combined in the following order: vocabulary reaction time, vocabulary accuracy, grammatical structure recognition accuracy, number of grammatical errors, listening keyword recognition rate, listening answer delay time, reading intensive reading accuracy, reading extensive reading accuracy, spoken speech recognition similarity, spoken pronunciation score, writing grammar accuracy, and writing vocabulary diversity index to construct a standardized learning dataset.
[0096] In this embodiment, the step of inputting the standardized learning dataset into the hierarchical Bayesian cognitive diagnostic model and outputting stage-based ability labels specifically includes:
[0097] The standardized learning dataset is input into the hierarchical Bayesian cognitive diagnostic model, in which a latent variable layer and an observation layer are established.
[0098] The latent variable layer includes vocabulary latent variables, grammar latent variables, listening latent variables, reading latent variables, speaking latent variables, and writing latent variables;
[0099] The observation layer includes vocabulary reaction time, vocabulary accuracy, grammatical structure recognition accuracy, number of grammatical errors, listening keyword recognition rate, listening answer delay time, reading intensive reading accuracy, reading extensive reading accuracy, spoken language speech recognition similarity, spoken language pronunciation score, writing grammar accuracy and writing vocabulary diversity index.
[0100] In the latent variable layer, a hierarchical dependency structure is set up, defining lexical latent variables and grammatical latent variables as low-level latent variables, listening latent variables and speaking latent variables as mid-level latent variables, and reading latent variables and writing latent variables as high-level latent variables.
[0101] The prior conditions of the middle-level latent variables depend on the posterior results of the lower-level latent variables, and the prior conditions of the higher-level latent variables depend on the posterior results of the middle-level latent variables.
[0102] A prior distribution is set for each latent variable, where the vocabulary latent variable, grammar latent variable, listening latent variable, reading latent variable, speaking latent variable and writing latent variable correspond to the learner’s mastery probability in six dimensions, respectively. Each latent variable is defined by a set of hyperparameters, which are determined by the number of correct and incorrect performances in the historical learning samples of that dimension, and are used to form the initial mastery probability when no observation data is input.
[0103] Define a likelihood function between the observation layer and the latent variable layer, where each latent variable establishes a conditional probability relationship with its corresponding observed feature. The likelihood function represents the probability of observed data occurring given the value of the latent variable.
[0104] Specifically, each observation feature in the standardized learning dataset is modeled one by one. A joint probability distribution is established between vocabulary reaction time, vocabulary accuracy, grammatical structure recognition accuracy, number of grammatical errors, listening keyword recognition rate, listening answer delay, reading intensive reading accuracy, reading extensive reading accuracy, spoken speech recognition similarity, spoken pronunciation score, writing grammar accuracy, and writing vocabulary diversity index and the corresponding latent variables. This forms the overall probability of learners generating all observation data under the latent variable state.
[0105] The posterior probability of each latent variable is calculated using the Bayesian inference formula to obtain the set of posterior probabilities of the latent variables, and then the updates are performed sequentially under the predefined hierarchical dependency structure.
[0106] That is, firstly, the posterior probability of the lower-level latent variables is calculated using the observation layer data, then the posterior result of the lower-level latent variables is used as the prior condition of the middle-level latent variables to update the posterior probability of the middle-level latent variables, and then the posterior result of the middle-level latent variables is used as the prior condition of the higher-level latent variables to update the posterior probability of the higher-level latent variables, thus forming a bottom-up layer-by-layer inference and update process.
[0107] Based on the updated posterior probabilities of each latent variable, the learner's mastery probability in six dimensions—vocabulary, grammar, listening, reading, speaking, and writing—is output. Stage-based ability labels are generated according to preset thresholds, and these stage-based ability labels serve as input conditions for selecting learning resources in the stage-based contextual resource library.
[0108] In this embodiment, the step of selecting learning resources that match the learner's current stage from the stage-specific contextual resource library based on stage-specific ability tags to obtain a combined learning resource set specifically includes:
[0109] Establish a phased contextual resource library, index and divide the learning resources according to different phases, with each phase index set corresponding to a set of learning resources. The learning resources include resources for daily communication scenarios and resources for academic communication scenarios, so as to ensure that matching immersive contextual content can be provided at each phase.
[0110] Based on the stage-based ability tags, the vocabulary stage-based ability tags, grammar stage-based ability tags, listening stage-based ability tags, reading stage-based ability tags, speaking stage-based ability tags, and writing stage-based ability tags are matched with the corresponding stage index sets in the stage context resource library to determine the learning resources associated with the stage index set corresponding to each dimension.
[0111] During the matching process, the stage ability tags are mapped to the stage index set in the stage context resource library. Specifically, based on the stage ability tags of six dimensions, namely vocabulary, grammar, listening, reading, speaking and writing, the learning resource set corresponding to each dimension tag is called respectively. This mapping relationship is defined as a calling function from stage ability tag to learning resource set, so that each dimension stage ability tag can be matched with a set of learning resources.
[0112] The sets of vocabulary learning resources, grammar learning resources, listening learning resources, reading learning resources, speaking learning resources, and writing learning resources obtained by filtering according to the six-dimensional stage ability tags will be merged to form a combined learning resource set that includes learning resources from all dimensions.
[0113] In this embodiment, the step of conducting multimodal interactive training for learners in an immersive context, collecting real-time performance data of learners during the interactive training process, and generating error correction information and minimization prompts specifically includes:
[0114] Select contextualized learning resources that correspond to the learner's current stage of ability tags from the combined learning resource set, and construct immersive interactive scenarios;
[0115] Perform interactive training in six dimensions: vocabulary, grammar, listening, reading, speaking, and writing in an immersive interactive scenario;
[0116] Among them, vocabulary training triggers vocabulary responses through images and audio, grammar training triggers grammatical structure recognition through sentence highlighting, listening training triggers keyword recognition through audio segments, reading training triggers semantic understanding through text annotation, speaking training triggers pronunciation scoring through voice input, and writing training triggers grammatical accuracy and vocabulary diversity calculation through text input.
[0117] During the immersive interactive training process, learners' real-time performance data is collected, including vocabulary reaction time, vocabulary accuracy, grammatical structure recognition accuracy, number of grammatical errors, listening keyword recognition rate, listening answer delay time, reading intensive reading accuracy, reading extensive reading accuracy, spoken language speech recognition similarity, spoken language pronunciation score, writing grammar accuracy, and writing vocabulary diversity index, to construct a real-time performance dataset.
[0118] The real-time performance dataset is input into the speech recognition module, the natural language understanding module, and the speech evaluation module. The speech recognition module receives the learner's spoken speech signal, performs acoustic feature extraction and text transcription, outputs the spoken recognition text, and calculates the similarity with the standard answer.
[0119] The natural language understanding module receives the recognized text output by the speech recognition module and the reading and writing input text from the learner, performs syntactic structure parsing and semantic component analysis, and outputs grammatical error features and semantic matching results;
[0120] The speech evaluation module receives the learner's acoustic parameters, which include the speech signal, extracts the fundamental frequency, formants, energy distribution, and pronunciation duration, calculates the pronunciation accuracy score, and outputs the pronunciation score result. After completing the three types of processing, the speech recognition result, grammatical error features, semantic matching result, and pronunciation score result are integrated to generate an error correction information set for error correction.
[0121] Based on the error correction information set, the minimum prompt required for the learner's input is determined. That is, the most critical prompt for the current input is selected from all error correction information as the output prompt, thereby generating the minimum prompt result and using it together with the error correction information as feedback content.
[0122] In this embodiment, the step of inputting the real-time performance dataset into the hierarchical Bayesian cognitive diagnostic model, performing latent variable posterior probability updates, generating new stage-specific ability labels, and generating a transfer signal when the advancement conditions are met specifically includes:
[0123] Real-time performance data is input into the hierarchical Bayesian cognitive diagnostic model. Based on the hierarchical dependency structure, the latent variable layer is updated using real-time performance data to maintain the correspondence between vocabulary latent variables, grammar latent variables, listening latent variables, reading latent variables, speaking latent variables and writing latent variables.
[0124] The posterior probability of each latent variable is updated. Specifically, given the real-time performance dataset, the likelihood function of the real-time performance data under the latent variable state is jointly calculated with the prior distribution of the original latent variable. The posterior probability of the latent variable is then solved using the Bayesian inference formula, thereby obtaining the learner's mastery probability in each ability dimension under the real-time performance data condition.
[0125] In the hierarchical dependency structure, the posterior probability of the lower-level latent variables is used as the prior input of the middle-level latent variables, and the posterior probability of the middle-level latent variables is used as the prior input of the higher-level latent variables, thus completing the hierarchical update from bottom to top.
[0126] Based on the updated posterior probabilities of the latent variables, new stage-specific capability labels are generated according to a preset threshold mapping rule;
[0127] The advancement conditions for each dimension's stage-based capability tags are determined, and these advancement conditions include two aspects of constraints:
[0128] First, the updated posterior probability of the latent variable must be greater than the corresponding preset threshold.
[0129] Second, it requires that the aforementioned probability determination conditions be met in several consecutive updates;
[0130] When both of the above constraints are met simultaneously, a migration signal is generated and a stage migration is triggered. The migration signal is used to control the stage context resource library to call learning resources of higher stages and update the learning path scheduling; otherwise, the original stage state remains unchanged.
[0131] In this embodiment, the process of re-executing stage-based contextual resource selection, immersive contextual interaction training, and stage updates based on new stage-based ability labels, until stage transfer is triggered, to form a dynamic closed-loop language acquisition path, specifically includes:
[0132] Based on the new stage-based ability tags, the stage-based context resource library is invoked to determine the matching set of learning resources for each stage-based ability tag in each dimension. Through mapping relationships, a set of learning resources in six dimensions corresponding to the updated stage-based ability tags is generated to form a new combined learning resource set.
[0133] The new combined learning resource set is input into the immersive context interaction training and real-time feedback steps to generate a new real-time performance dataset.
[0134] The new real-time performance dataset is input into the hierarchical Bayesian cognitive diagnostic model to calculate the updated posterior probabilities of latent variables and generate new stage-specific ability labels.
[0135] Determine whether the new stage ability label meets the advancement conditions. If it does, generate a transfer signal and trigger stage transfer, while recording the learner's stage change.
[0136] If the requirements are not met, learning resources will continue to be selected based on the new stage-based ability labels, a new set of combined learning resources will be generated, new immersive contextual interactive training will be carried out and new real-time performance data will be collected. The data will then be input into the hierarchical Bayesian cognitive diagnostic model for updating, forming a dynamic closed-loop language acquisition path.
[0137] Example 1:
[0138] To verify the feasibility of this invention in practice, it was applied to an intermediate English course training scenario at a language learning institution. Learners at this institution generally suffer from insufficient matching between learning resources and their individual skill levels, a simplistic feedback mechanism, and a lack of dynamic adjustment to their learning paths. This results in some learners stagnating in listening and speaking skills for extended periods, making it difficult for them to progress to higher levels. This invention establishes a hierarchical Bayesian cognitive diagnostic model to model learners' performance across six dimensions: vocabulary, grammar, listening, reading, speaking, and writing. This model generates stage-specific ability labels, which drive a closed-loop learning mechanism encompassing contextualized resource selection, immersive interactive training, real-time feedback, stage updates, and transfer control.
[0139] In the application process, multi-source data from learners' classroom and online practice is first collected, including vocabulary reaction time, number of grammatical errors, listening comprehension response delay, oral pronunciation score, and writing vocabulary diversity. This data is preprocessed to form a standardized dataset. The system then inputs this data into a hierarchical Bayesian cognitive diagnostic model, outputting mastery probabilities across six dimensions and corresponding stage-specific ability labels. For example, a learner with a mastery probability of 0.72 in vocabulary, 0.65 in grammar, 0.48 in listening, 0.44 in speaking, 0.70 in reading, and 0.61 in writing is classified as "Vocabulary Stage S2, Grammar Stage S2, Listening Stage S1, Speaking Stage S1, Reading Stage S2, Writing Stage S2" after threshold mapping. Based on these labels, the system automatically accesses the stage-specific contextual resource library and pushes real-world scenario materials matching the stage, such as asking for directions at an airport, academic discussions, and excerpts from news reports.
[0140] During interactive training, learners complete oral tasks via voice input. The system utilizes speech recognition and speech evaluation modules to output pronunciation similarity and intonation curve deviation in real time. Simultaneously, it combines a natural language understanding module to identify grammatical errors and generate minimal prompts. The learner's performance data is then input into the cognitive diagnostic model for updates. In three consecutive tests, the posterior probabilities of listening and speaking gradually increased from 0.48 and 0.44 to 0.71 and 0.69, respectively, meeting the set advancement threshold of 0.65. This triggers a transfer signal, and the system automatically allocates resources to higher-level tasks, such as academic lecture listening and debate-style speaking tasks.
[0141] After four weeks of continuous application, system statistics show that among the 40 participants, the average vocabulary reaction time decreased by 18%, the grammar error rate decreased by 22%, the listening keyword recognition rate increased by 19%, the average oral pronunciation score increased by 1.3 points (out of 5), and the writing vocabulary diversity index increased by 16%. These data indicate that this invention can achieve dynamic diagnosis and updating of learners' multi-dimensional cognitive levels, and on this basis, drive the adaptive scheduling of learning resources and paths, thus effectively solving the problems of resource matching difficulties, inaccurate feedback, and the inability to dynamically optimize learning paths in existing technologies.
[0142] Table 1 Comparison of Implementation Results
[0143]
[0144] As can be seen from the table above, this invention significantly improves several key performance indicators of target language acquisition compared to existing conventional learning models. Firstly, in vocabulary learning, learners' average reaction time is reduced from 3.9 seconds to 3.2 seconds, a reduction of 18%, indicating that this invention effectively reduces learners' time cost in vocabulary retrieval. Simultaneously, grammar learning performance also shows significant improvement, with the average grammar error rate decreasing from 24.6% to 19.2%, an overall reduction of 22%. This demonstrates that through phased cognitive diagnosis and targeted resource selection, learners' accuracy in understanding and applying syntactic structures has improved.
[0145] Secondly, in terms of listening comprehension, the system of this invention improved learners' keyword recognition rate from 62.4% to 74.3%, an increase of 19%, while reducing the response delay from 6.1 seconds to 4.8 seconds, a reduction of 21%. This demonstrates that by providing listening materials appropriate to learners' learning stages in an immersive context, combined with a real-time feedback mechanism, learners' auditory capture ability and information processing speed can be significantly enhanced, thereby improving their response efficiency in listening tasks.
[0146] In terms of reading, learners' accuracy in intensive reading increased from 68.5% to 78.9%, and their accuracy in extensive reading increased from 71.3% to 81.2%, representing increases of 15% and 14% respectively. This result demonstrates that the present invention, through dynamic resource selection, guides learners to reading materials that match their ability level, enabling them to gradually improve their abilities in both detailed comprehension and overall understanding, thus forming a more balanced reading skill structure.
[0147] Furthermore, in the speaking and writing sections, learners' pronunciation scores improved from 2.9 to 4.2, and similarity increased from 66.1% to 79.5%, representing improvements of 1.3 points and 20% respectively. This indicates that, supported by the speech recognition and evaluation module, the present invention can provide more accurate error correction and minimize prompts, thereby helping learners gradually optimize their pronunciation and expression. In writing, grammatical accuracy increased from 72.6% to 83.1%, and the vocabulary diversity index increased from 0.58 to 0.67, representing improvements of 14% and 16% respectively. This demonstrates that the phased ability labels and transfer mechanism can drive learners to continuously expand their vocabulary and reduce grammatical errors. Overall, the present invention, through the combination of phased cognitive diagnosis, contextual resource selection, immersive interactive training, and dynamic path scheduling, effectively solves the problems of resource matching difficulties, inaccurate feedback, and lack of dynamic adjustment of learning paths in existing technologies, achieving a simultaneous improvement in the efficiency and quality of language acquisition.
[0148] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A target language acquisition and development system based on staged immersive context generation, characterized in that, include: The learning data acquisition module is used to collect and process multi-source data and output a standardized learning dataset. The cognitive diagnosis module receives a standardized learning dataset and outputs stage-based ability labels based on a hierarchical Bayesian cognitive diagnosis model. The cognitive diagnostic module specifically includes: The standardized learning dataset is input into the hierarchical Bayesian cognitive diagnostic model, in which a latent variable layer and an observation layer are established. In the latent variable layer, a hierarchical dependency structure is set up, defining lexical latent variables and grammatical latent variables as low-level latent variables, listening latent variables and speaking latent variables as mid-level latent variables, and reading latent variables and writing latent variables as high-level latent variables. The prior conditions of the middle-level latent variables depend on the posterior results of the lower-level latent variables, and the prior conditions of the higher-level latent variables depend on the posterior results of the middle-level latent variables. A prior distribution is set for each latent variable, and an initial distribution for each latent variable is defined by a set of hyperparameters; Define the likelihood function between the observation layer and the latent variable layer, and model each observation feature in the standardized learning dataset one by one to form the overall probability of the learner generating all observation data under the latent variable state. The posterior probability of each latent variable is calculated using the Bayesian inference formula to obtain the set of posterior probabilities of the latent variables. The posterior distributions of the middle-level latent variables and the high-level latent variables are then updated sequentially based on the hierarchical dependency structure. The learner's mastery probability is output based on the updated posterior probabilities of each latent variable, and a stage-based ability label is generated according to a preset threshold. The Stage Context Resource Module is used to store and manage contextualized learning resources divided into stages, and to call up a set of learning resources that match the learner's current stage based on the stage's ability tags. The interactive training and feedback module is used to conduct multimodal interactive training, collect performance data in real time, and generate error correction information and minimize prompts. The interactive training and feedback module specifically includes: Select contextualized learning resources that correspond to the learner's current stage of ability tags from the combined learning resource set, and construct immersive interactive scenarios; Perform interactive training in six dimensions: vocabulary, grammar, listening, reading, speaking, and writing in an immersive interactive scenario; Real-time performance data of learners is collected during immersive interactive training to construct a real-time performance dataset; The real-time performance dataset is input into the speech recognition module, natural language understanding module, and speech evaluation module to generate an error correction information set. The learner determines the minimum prompt required for the input based on the error correction information set, that is, selects the most critical prompt for the current input from all error correction information as the output prompt, and generates the minimum prompt result; The phase update module is used to receive real-time performance data, perform posterior probability updates of latent variables, generate new phase capability labels, and generate migration signals when advanced conditions are met. The phase update module specifically includes: Real-time performance data is input into the hierarchical Bayesian cognitive diagnostic model. Based on the hierarchical dependency structure, the latent variable layer is updated using real-time performance data to maintain the correspondence between vocabulary latent variables, grammar latent variables, listening latent variables, reading latent variables, speaking latent variables and writing latent variables. By performing posterior probability updates on each latent variable and obtaining the real-time performance dataset, the likelihood function of the real-time performance data under the latent variable state is jointly calculated with the prior distribution of the original latent variable. The posterior probability of the latent variable is then solved using the Bayesian inference formula to obtain the learner's mastery probability in each ability dimension under the real-time performance data condition. In the hierarchical dependency structure, the posterior probability of the lower-level latent variables is used as the prior input of the middle-level latent variables, and the posterior probability of the middle-level latent variables is used as the prior input of the higher-level latent variables, thus completing the hierarchical update from bottom to top. Based on the updated posterior probabilities of the latent variables, new stage-specific capability labels are generated according to a preset threshold mapping rule; For each dimension of stage capability label, the advancement conditions are determined. If the advancement conditions are met, a migration signal is generated and stage migration is triggered. Otherwise, the original stage state remains unchanged. The migration control and path scheduling module is used to receive migration signals, drive the stage context resource module to call learning resources of higher stages, control learning path scheduling, and record learner stage changes.
2. The target language acquisition and training system based on staged immersive context generation according to claim 1, characterized in that, The modules are connected in the following way: Collect and process multi-source data from learners during language training to form a standardized learning dataset; Input the standardized learning dataset into the hierarchical Bayesian cognitive diagnostic model and output stage-based ability labels; Based on the stage-based ability tags, learning resources that match the learner's current stage are selected from the stage-based context resource library to obtain a combined learning resource set; Multimodal interactive training is conducted for learners in an immersive context. During the interactive training process, real-time performance data of learners is collected, and error correction information and minimize prompts are generated. The real-time performance dataset is input into the hierarchical Bayesian cognitive diagnostic model, the posterior probability of the latent variables is updated, new stage ability labels are generated, and a transfer signal is generated when the advancement conditions are met. Based on the new stage-based competency labels, the stage-based contextual resource selection, immersive contextual interactive training, and stage updates are re-executed until stage transfer is triggered, forming a dynamic closed-loop language acquisition path.
3. The target language acquisition and training system based on staged immersive context generation according to claim 2, characterized in that, The standardized learning dataset includes preprocessed vocabulary reaction time, vocabulary accuracy, grammatical structure recognition accuracy, number of grammatical errors, listening keyword recognition rate, listening answer delay time, reading intensive reading accuracy, reading extensive reading accuracy, spoken language speech recognition similarity, spoken language pronunciation score, writing grammar accuracy, and writing vocabulary diversity index.
4. The target language acquisition and training system based on staged immersive context generation according to claim 2, characterized in that, The step of selecting learning resources from the stage-specific context resource library that match the learner's current stage based on stage-specific ability tags to obtain a combined learning resource set specifically includes: Establish a phased contextual resource library, indexing and dividing learning resources according to different phases, with each phase index set corresponding to a set of learning resources; Match the stage-specific capability tags with the stage index set in the stage context resource library; During the matching process, the stage-specific ability tags are mapped to the stage index set in the stage context resource library to obtain the learning resource set for each dimension. The learning resource sets obtained by filtering according to the six-dimensional stage ability tags will be merged to form a combined learning resource set.
5. A target language acquisition and training system based on staged immersive context generation according to claim 2, characterized in that, The language acquisition path, which is based on new stage-specific ability labels, and involves re-executing stage-specific contextual resource selection, immersive contextual interactive training, and stage updates until stage transfer is triggered, forming a dynamic closed loop, specifically includes: Based on the new stage-based ability tags, the stage-based context resource library is invoked, and the matching set of learning resources is determined according to the stage-based ability tags of each dimension, forming a new set of combined learning resources. The new combined learning resource set is input into the immersive context interaction training and real-time feedback steps to generate a new real-time performance dataset. The new real-time performance dataset is input into the hierarchical Bayesian cognitive diagnostic model to calculate the updated posterior probabilities of latent variables and generate new stage-specific ability labels. Determine whether the new stage capability label meets the advancement conditions. If it does, generate a migration signal and trigger stage migration. If the requirements are not met, learning resources will continue to be selected based on the new stage-based ability labels, a new set of combined learning resources will be generated, new immersive contextual interactive training will be carried out and new real-time performance data will be collected. The data will then be input into the hierarchical Bayesian cognitive diagnostic model for updating, forming a dynamic closed-loop language acquisition path.
Citation Information
Patent Citations
Multilingual training management system and method
CN120495032A