A real-time interactive online Chinese language learning platform

By constructing a multi-dimensional cultural element matrix and multi-level language error analysis, combined with user multimodal data, personalized feedback and interactive optimization of the online Chinese language learning platform are achieved, improving users' Chinese learning results and experience.

CN120315598BActive Publication Date: 2025-09-23闽南科技学院
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510813618.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-23
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

Existing online Chinese language learning platforms lack a deep understanding of user learning behavior and intelligent response mechanisms, and fail to provide an intelligent and personalized learning experience, especially in terms of language error analysis and personalized feedback.

Method used

Construct a multi-dimensional cultural element matrix, establish a virtual cultural interaction scene through the scene construction module, use the interaction collection module to collect continuous dialogue fragments in real time, the correction feedback module performs multi-level language error analysis, the state evaluation module collects user multimodal data to generate multi-dimensional state feature vectors, and the interaction optimization module performs interaction control strategy reasoning to achieve personalized feedback and interaction optimization.

Benefits of technology

It provides an immersive language practice environment, accurately identifies vocabulary, grammar, and pragmatic errors, formulates progressive correction suggestions, dynamically optimizes the interactive mode, improves users' Chinese language proficiency and learning experience, and realizes personalized and efficient online Chinese learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120315598B_ABST
    Figure CN120315598B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of online learning and discloses a real-time interactive online Chinese language learning platform. The platform comprises the following steps: constructing a virtual cultural interaction scene; collecting continuous dialogue segments in real time during the process of dialogue interaction; performing multi-level language error analysis on the continuous dialogue segments to identify language errors, formulating corresponding language correction suggestions for the language errors, and feeding back to users through a progressive feedback mechanism; collecting user multimodal data, performing nonlinear feature mapping, generating multidimensional state feature vectors, and dynamically evaluating cognitive load and learning status; fusing cognitive load and learning status, performing interaction regulation strategy reasoning, and dynamically optimizing interaction modes and interaction difficulty in the dialogue interaction according to the interaction regulation strategy. The platform achieves precise control and dynamic regulation of the entire process of user Chinese language learning, provides a personalized online Chinese learning experience, and promotes the comprehensive improvement of user language ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of online learning, and more particularly to a real-time interactive online Chinese language learning platform. Background Art

[0002] As China's international influence continues to increase, the demand for Chinese language learning continues to grow worldwide, and Chinese is gradually becoming a language with extensive international influence. However, current Chinese language teaching still faces many challenges, such as uneven distribution of teaching resources, obvious geographical restrictions, and the lack of interactivity in traditional teaching methods, making it difficult to meet diverse and personalized learning needs. In recent years, the rapid development of real-time audio and video communications, artificial intelligence, and mobile Internet has provided new technical paths for language learning. Web-based online teaching methods no longer restrict learners to offline classrooms, and they can achieve real-time interaction through the Internet, obtaining a more flexible and immersive learning experience. Therefore, developing an online Chinese language learning platform with real-time interactive capabilities can effectively improve language learning efficiency, promote the transformation of Chinese teaching from indoctrination to interaction, and inject new vitality into international Chinese education.

[0003] The patent with publication number CN115292499A discloses a real-time interactive online Chinese language learning platform; it includes: a cloud storage reservation module, a user operation module, an online interaction processing module and a data execution monitoring module; the cloud storage reservation module is used to enter Chinese language learning content and conveniently classify and store the entered Chinese language learning content according to keywords; the user operation module is used for user login and synchronization with the cloud; this invention, through the coordinated design of the cloud storage reservation module, the user operation module, the online interaction processing module and the data execution monitoring module, makes the platform convenient for real-time interactive online Chinese language learning applications according to usage needs, and facilitates convenient synchronization of learning terminals, live broadcast terminals and chat terminals according to usage needs, thereby forming keyword-directed learning, live broadcast interaction and classmate interaction, and facilitates monitoring and differentiated storage of data in the learning platform.

[0004] However, although the above technologies can achieve real-time interaction and online learning, they focus on the functional stacking of modular structures. Although they cover storage, user operations, interaction processing and data monitoring, they lack a deep understanding of user learning behavior and an intelligent response mechanism. They do not involve key links such as language error analysis and personalized feedback mechanisms, making it difficult to provide learners with an intelligent and personalized learning experience.

[0005] In view of this, the present invention proposes a real-time interactive online Chinese language learning platform to solve the above problems. Summary of the Invention

[0006] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned objectives, the present invention provides the following technical solutions: a real-time interactive online Chinese language learning platform, comprising:

[0007] A scenario construction module is used to establish a multi-dimensional cultural element matrix and build a virtual cultural interaction scenario based on the multi-dimensional cultural element matrix;

[0008] An interaction collection module is used to collect continuous dialogue segments in real time during the process of dialogue interaction between users and virtual characters in the virtual cultural interaction scene;

[0009] The correction feedback module is used to perform multi-level language error analysis on continuous dialogue segments, identify language errors in continuous dialogue segments, and formulate corresponding language correction suggestions for language errors, which are fed back to users through a progressive feedback mechanism;

[0010] The state assessment module is used to collect user multimodal data, perform nonlinear feature mapping on the user multimodal data, generate a multidimensional state feature vector, and dynamically assess cognitive load and learning status based on the multidimensional state feature vector;

[0011] The interaction optimization module is used to integrate cognitive load and learning status, and to infer interaction regulation strategies, and dynamically optimize the interaction mode and interaction difficulty in dialogue interaction based on the interaction regulation strategies.

[0012] Furthermore, the method for collecting continuous conversation segments in real time includes:

[0013] Continuously collect conversation interaction information between users and avatars, including conversation content and corresponding conversation time. Conversation content is the natural language text output by the user or avatar during the conversation interaction, and conversation time is the output time of each conversation content.

[0014] Obtain historical duration, average the historical duration, and obtain a preliminary response threshold; the historical duration is the response duration of the user during the dialogue interaction with the virtual character at the historical moment; obtain the interaction mode and interaction difficulty of the dialogue interaction, dynamically adjust the preliminary response threshold based on the interaction mode and interaction difficulty, and obtain a dynamic response threshold; after each round of the virtual character outputting natural language text, monitor whether the user outputs natural language text within the dynamic response threshold; if so, determine that the dialogue interaction continues and continue to collect dialogue interaction information; if not, determine that the dialogue interaction ends, and use the collected dialogue content as the preliminary dialogue segment;

[0015] Calculate the semantic similarity between each two adjacent conversation contents in the preliminary conversation segment, and compare each semantic similarity with a preset similarity threshold; mark the conversation contents with semantic similarity less than the similarity threshold as irrelevant content, and mark the conversation contents with semantic similarity greater than the similarity threshold as continuous content; if the conversation content is marked as both irrelevant content and continuous content, unmark it as continuous content; mark the earliest conversation time as the earliest time and the latest time as the latest time among the conversation times corresponding to the irrelevant content;

[0016] The continuous content whose conversation time is earlier than the earliest time and the irrelevant content corresponding to the earliest time are regarded as a continuous conversation segment; the continuous content whose conversation time is later than the latest time and the irrelevant content corresponding to the latest time are regarded as a continuous conversation segment; it is judged whether there is continuous content between any two irrelevant contents; if so, the corresponding irrelevant content and the continuous content between them are regarded as a continuous conversation segment; if not, no operation is performed.

[0017] Furthermore, the method for obtaining the dynamic response threshold includes:

[0018] A preset coefficient set includes a mode set and a difficulty set; the mode set includes adjustment coefficients corresponding to different interaction modes, and the difficulty set includes adjustment coefficients corresponding to different interaction difficulties; according to the interaction mode and interaction difficulty of the dialogue interaction, the corresponding adjustment coefficient is obtained from the coefficient set; the product of the adjustment coefficient of the interaction mode and the adjustment coefficient of the interaction difficulty is used as the overall adjustment coefficient; the product of the overall adjustment coefficient and the preliminary response threshold is used as the dynamic response threshold;

[0019] Methods for calculating the semantic similarity between two adjacent conversation contents include:

[0020] The SimCSE model is used to convert each conversation content into a corresponding conversation vector. The cosine similarity between the conversation vectors corresponding to each two adjacent conversation contents is calculated in sequence and used as the semantic similarity.

[0021] Furthermore, the step of identifying language errors in the continuous dialogue segments includes:

[0022] Step S101: performing a lexical error analysis on each continuous dialogue segment to identify lexical errors;

[0023] Each continuous segment is input into the trained context-aware language model, and the masked prediction probability of each phrase in each continuous segment is calculated. A probability threshold is preset, and each masked prediction probability is compared with the probability threshold. Phrases with a masked prediction probability less than the probability threshold are marked as incorrect phrases, while phrases with a masked prediction probability greater than or equal to the probability threshold are not marked. The same incorrect phrases in corresponding continuous dialogue segments are merged and regarded as lexical errors in the corresponding continuous dialogue segments.

[0024] Step S102: performing grammatical error analysis on each continuous dialogue segment to identify grammatical errors;

[0025] Input each continuous dialogue segment into the trained grammatical error correction model to identify grammatical errors in each continuous dialogue segment;

[0026] Step S103: performing pragmatic error analysis on each continuous dialogue segment to identify pragmatic errors;

[0027] Each continuous dialogue segment is input into the trained context analysis model to identify pragmatic errors in each continuous dialogue segment;

[0028] Step S104: Combining vocabulary errors, grammatical errors, and pragmatic errors to determine language errors in each continuous dialogue segment;

[0029] The same lexical errors, grammatical errors and pragmatic errors in the corresponding continuous dialogue segments are combined as the language errors in the corresponding continuous dialogue segments.

[0030] Furthermore, the method for formulating corresponding language correction suggestions for language errors includes:

[0031] Each continuous dialogue segment and the corresponding language error are taken as a set of segment sets, each set of segment sets is input into the trained language correction model, and the correction set corresponding to each set of segment sets is output; the correction set includes the suggestion set corresponding to each language error, and the suggestion set includes the corresponding language error. Corrective suggestions, is an integer greater than 1;

[0032] From each suggestion set corresponding to each continuous dialogue segment, a correction suggestion is randomly selected and a set of suggestion subsets is constructed. Group suggestion subsets, is an integer greater than 1; the SimCSE model is used to convert the correction suggestions in each suggestion subset into corresponding suggestion vectors; the dialogue vector corresponding to each continuous dialogue segment and the corresponding set of suggestion subsets are used as a set of analysis sets, and each set of analysis sets is input into the trained correction analysis model to predict the corresponding indicator set; the indicator set includes language fluency and cultural fit, and the correction analysis model is a deep neural network model;

[0033] A preset ratio set includes proportional coefficients corresponding to language fluency and cultural fit; based on the ratio set, the language fluency and cultural fit in each indicator set are weighted and summed to obtain a comprehensive score for each set of suggestion subsets; the same comprehensive scores of corresponding continuous dialogue segments are compared, and the suggestion subset with the largest comprehensive score is used as the language correction suggestion for the corresponding continuous dialogue segment.

[0034] Furthermore, the step of providing feedback to the user through a progressive feedback mechanism includes:

[0035] Step S201: Feedback the language errors corresponding to each continuous dialogue segment to the user and obtain the modified content;

[0036] Step S202: Identify language errors in the modified content. If language errors still exist, proceed to step S203. If no language errors exist, the feedback ends.

[0037] Step S203: Based on the pre-built language knowledge base, obtain the error cause corresponding to the language error identified in step S202;

[0038] Step S204: Feedback the error reason to the user and obtain the secondary modification content;

[0039] Step S205: Identify language errors in the second modified content. If language errors still exist, proceed to step S206. If no language errors exist, the feedback ends.

[0040] Step S206: Feedback the correction set of all continuous dialogue segments and language correction suggestions to the user.

[0041] Furthermore, the user multimodal data includes facial image sequences and eye movement trajectories; the facial image sequences include the images collected during the user's conversation and interaction with the virtual character. facial images, is an integer greater than 1; the eye movement trajectory includes the gaze point, gaze duration, and saccade frequency;

[0042] The method for generating a multidimensional state feature vector comprises:

[0043] Using the trained emotion recognition model, each facial image in the facial image sequence is identified in turn, and the recognition results of each facial image are output. The recognition results include category labels and emotion intensity; category labels are numerical labels corresponding to emotion categories, and different emotion categories have different numerical labels; the range of emotion intensity is , According to the category label in the recognition result, the corresponding emotion category is obtained; the number of each emotion category is counted and marked as the emotion quantity; all emotion quantities are compared, and the emotion category with the largest emotion quantity is regarded as the user's emotional state; all emotion intensities corresponding to the emotional state are averaged to obtain the average emotion intensity; the emotion recognition model is a convolutional neural network model;

[0044] A clustering algorithm is used to cluster all fixation points in the eye movement trajectory to obtain multiple fixation areas; the number of fixation areas is counted and marked as the number of fixations; all gaze durations in the eye movement trajectory are averaged to obtain the average gaze duration; a weight set is preset, which includes weight coefficients corresponding to the inverse of the number of fixations, the average gaze duration, and the inverse of the skip rate; the number of fixations, the average gaze duration, and the skip rate are all standardized, and the inverse of the standardized number of fixations, the average gaze duration, and the inverse of the skip rate are weighted summed according to the weight set to obtain the attention concentration;

[0045] Count the total number of vocabulary errors, grammatical errors, and pragmatic errors in all consecutive dialogue segments and mark them as the number of language errors; count the number of times the progressive feedback mechanism provides feedback to users and mark them as the number of feedbacks;

[0046] A multidimensional state feature vector is generated based on the emotional state, average emotional intensity, attention concentration, number of language errors, and number of feedbacks.

[0047] Furthermore, methods for dynamically assessing cognitive load and learning status include:

[0048] The category label, average emotion intensity, and attention concentration corresponding to the emotional state in the multidimensional state feature vector are used as first analysis data, and the number of language errors and the number of feedback times in the multidimensional state feature vector are used as second analysis data; the learning state is dynamically evaluated based on the first analysis data, and the cognitive load is dynamically evaluated based on the second analysis data, and the method for dynamically evaluating the learning state based on the first analysis data is consistent with the method for dynamically evaluating the cognitive load based on the second analysis data;

[0049] The method for dynamically evaluating the learning status according to the first analysis data includes:

[0050] A plurality of fuzzy sets are constructed for each data in the first analysis data; each data in the first analysis data is converted into the membership of each corresponding fuzzy set through fuzzification technology; fuzzy rules are defined, the fuzzified first analysis data are matched with the fuzzy rules, and fuzzy reasoning methods are used to perform fuzzy reasoning to obtain fuzzy reasoning results, which are the membership of each learning state; each membership is compared, and the learning state with the largest membership is used as the learning state dynamically evaluated.

[0051] Furthermore, the method for performing interactive control strategy reasoning includes:

[0052] Obtain the corresponding load label based on the cognitive load, and obtain the corresponding state label based on the learning state; wherein the load label is the numerical label corresponding to the cognitive load, and the state label is the numerical label corresponding to the learning state; input the load label and the state label into the trained direction judgment model to predict the corresponding direction label, and the direction judgment model is a deep neural network model; wherein the direction label is the numerical label corresponding to the control direction, and the control direction includes positive control and negative control;

[0053] Set the corresponding guidance intensity for each interaction mode, sort all interaction modes from high to low according to the corresponding guidance intensity to generate a mode sequence; sort all interaction difficulties from low to high to generate a difficulty sequence;

[0054] If the control direction is positive control, the interaction mode that follows the current mode is obtained from the mode sequence, and the interaction difficulty that follows the current difficulty is obtained from the difficulty sequence; if the control direction is negative control, the interaction mode that precedes the current mode is obtained from the mode sequence; the interaction difficulty that precedes the current difficulty is obtained from the difficulty sequence; wherein, the current mode is the interaction mode corresponding to the user's current dialogue interaction, and the current difficulty is the interaction difficulty corresponding to the user's current dialogue interaction; the obtained interaction modes are all marked as control modes, and the obtained interaction difficulties are all marked as control difficulties;

[0055] Randomly select a control mode and a control difficulty, construct a control strategy, and build A regulatory strategy, is an integer greater than 0; calculate the control effect of each group of control strategies in turn, compare all the control effects, and take the control strategy with the greatest control effect as the interactive control strategy.

[0056] Furthermore, the method of sequentially calculating the control effect of each group of control strategies includes:

[0057] Different numerical labels are set for different interaction modes and marked as mode labels; different numerical labels are set for different interaction difficulties and marked as difficulty labels; the multidimensional state feature vector, the mode label of the current mode, the difficulty label of the current difficulty, and the mode label of the control mode and the difficulty label of the control difficulty in a control strategy are used as a set of control data; each set of control data is input into the trained effect prediction model to predict the corresponding control effect; the effect prediction model is a deep neural network model.

[0058] The technical effects and advantages of the real-time interactive online Chinese language learning platform of the present invention are as follows:

[0059] By constructing a virtual cultural interactive scene and integrating multi-dimensional historical and cultural elements, we provide users with an immersive language practice environment, enhance their understanding of traditional Chinese culture, and make Chinese language learning more interesting and engaging. During real-time conversations, we collect continuous conversation fragments and conduct multi-level language error analysis to accurately identify various types of errors in vocabulary, grammar, pragmatics, etc., and formulate targeted and progressive language correction suggestions to effectively promote the systematic improvement of users' Chinese language proficiency. We collect users' facial expressions and eye movement data, dynamically analyze cognitive load and learning status, and combine intelligent reasoning to optimize interaction modes and difficulty in real time, so that the learning process is more in line with users' learning needs and cognitive characteristics. We achieve precise control and dynamic regulation of the entire process of users' Chinese language learning, provide users with a personalized and efficient online Chinese learning experience, and promote the comprehensive improvement of users' language ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 Schematic diagram of a real-time interactive online Chinese language learning platform according to Example 1 of the present invention.

[0061] Figure 2 This is a flow chart of a real-time interactive online Chinese language learning platform according to embodiment 1 of the present invention. DETAILED DESCRIPTION

[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0063] Example 1

[0064] See also Figure 1 and Figure 2As shown, the real-time interactive online Chinese language learning platform described in this embodiment includes a scenario construction module, an interaction collection module, a correction feedback module, a status evaluation module and an interaction optimization module; each module is connected by wired and / or wireless means to realize data transmission between modules.

[0065] The scenario construction module is used to establish a multi-dimensional cultural element matrix and construct a virtual cultural interaction scenario based on the multi-dimensional cultural element matrix.

[0066] The multidimensional cultural element matrix is ​​established by those skilled in the art based on Chinese history and culture. Each element in the multidimensional cultural element matrix includes but is not limited to dynasties, regions, language expressions, social scenes, etc.

[0067] Dynasties are used to reflect the timeline and political and cultural characteristics of China's historical evolution; for example, political systems (such as the Qin Dynasty's system of prefectures and counties, the Tang Dynasty's system of three provinces and six ministries), cultural features (such as the flourishing of Confucianism in the Han Dynasty, the openness and diversity of the Tang Dynasty, and the development of Neo-Confucianism in the Song Dynasty), and social trends (such as the refined conversation and abstruse style of the Wei and Jin Dynasties, the boldness and extravagance of the Tang Dynasty, and the refinement and subtlety of the Song Dynasty).

[0068] Regions are used to reflect the spatial distribution and regional characteristics of Chinese culture; for example, natural environments (such as the arid plains in the north, the water towns and marshes in the south of the Yangtze River, and the grasslands and Gobi deserts in the northwest), cultural circles (such as Guanzhong culture, Bashu culture, and Lingnan culture), and economic characteristics (such as handicrafts and commerce in the south of the Yangtze River, agriculture in the north, and animal husbandry in the northwest).

[0069] Linguistic expressions are used to reflect the language usage characteristics of different historical periods and regions; for example, usage habits (such as titles, greetings, honorifics), literary styles (such as the boldness and elegance of Tang poetry, the elegance and subtlety of Song poetry, and the colloquial expressions of Ming and Qing novels);

[0070] Social scenes are used to reflect the specific environment and interaction methods of social role activities; for example, market scenes (such as teahouses, wine shops, streets, etc.), palace scenes (such as court meetings, hunting, etc.), academy scenes (such as academy lectures, imperial examinations, etc.), etc.

[0071] The methods for constructing virtual cultural interaction scenarios based on a multidimensional cultural element matrix include:

[0072] For each element in the multidimensional cultural element matrix, a corresponding role system, interaction rule system, and intelligent dialogue system are designed, and virtual cultural interaction scenes are constructed by comprehensively using technologies such as mixed reality and natural language processing. Mixed reality and natural language processing are both existing technologies, and the specific process will not be elaborated on here.

[0073] The method for designing a role system is as follows: those skilled in the art obtain a set of social roles corresponding to each element by consulting historical documents and cultural materials, and construct a corresponding role portrait for each social role in the social role set; the role portrait includes but is not limited to basic attributes (such as age, gender, occupation, etc.), personality traits (such as elegant and reserved, bold and confident, etc.), and behavior patterns (such as daily activities, social habits, etc.); based on the role portraits of different social roles in the social role set, different virtual characters are formed, and all virtual characters are integrated to form a role system;

[0074] The method for designing an interaction rule system is as follows: technical personnel in this field, based on historical research and cultural studies, obtain the social norms and interaction patterns corresponding to each element and construct an interaction rule system; social norms include but are not limited to etiquette norms (such as the specific etiquette for visiting elders and paying homage to the emperor), social class behavioral norms (such as the behavioral boundaries between different social roles such as commoners, scholars, and officials), and scene-specific behavioral norms (such as the behavioral norms for specific scenes such as drinking tea in teahouses, lecturing in academies, and trading in the market); interaction patterns include but are not limited to social class interaction patterns (such as the communication methods between commoners and officials, the interaction methods between scholars and merchants, etc.), scene-specific interaction patterns (bargaining methods in market transactions, debate and discussion patterns in academy lectures, etc.);

[0075] The method for designing an intelligent dialogue system is: based on the language expression corresponding to each element in the multi-dimensional cultural element matrix, a corresponding dialogue generation model is constructed. The dialogue generation model is a large language model based on the Transformer architecture (such as BERT, GPT, etc.); an intelligent dialogue system with historical language sense, cultural depth and contextual awareness is realized, supporting low-latency, high-semantic-fidelity real-time natural interaction between users and virtual characters.

[0076] The interaction collection module is used to collect continuous dialogue segments in real time during the process of dialogue interaction between users and virtual characters in the virtual cultural interaction scene.

[0077] Methods for collecting continuous conversation snippets in real time include:

[0078] Continuously collect conversation interaction information between users and avatars, including conversation content and corresponding conversation time. Conversation content is the natural language text output by the user or avatar during the conversation interaction; conversation time is the output time of each conversation content, which is used to record the timing information of the conversation;

[0079] Obtain historical duration, average the historical duration, and obtain the initial response threshold; historical duration is the response duration of the user in the process of dialogue interaction with the virtual character at the historical moment, which is obtained through the built-in database of the Chinese language learning platform; obtain the interaction mode and interaction difficulty of the dialogue interaction, dynamically adjust the initial response threshold based on the interaction mode and interaction difficulty, and obtain the dynamic response threshold; among them, the interaction mode is the structured way of interaction between the user and the virtual person, such as question-and-answer interaction (for example, the user asks a Song Dynasty academy teacher about the concept of Neo-Confucianism, and the teacher gives a systematic answer) Answers), casual interactions (e.g., users chatting with Ming Dynasty teahouse owners about weather, current affairs, and neighborhood gossip), situational task-based interactions (e.g., users completing the task of purchasing specific goods in a Song Dynasty market and needing to bargain with different vendors), role-playing interactions (e.g., users playing the role of a Tang Dynasty Jinshi who had just entered the officialdom and participated in court meetings), etc.; interaction difficulty refers to the complexity of the knowledge reserve, language skills, and cultural understanding depth required for users to participate in the interaction mode, such as beginner, intermediate, and advanced difficulty; interaction mode and interaction difficulty are directly obtained through virtual cultural interaction scenarios;

[0080] After each round of the virtual character outputting natural language text, the user is monitored within the dynamic response threshold to determine whether the user outputs natural language text. If so, the conversation interaction is determined to continue, and the conversation interaction information is continued to be collected. If not, the conversation interaction is determined to end, and the collected conversation content is used as a preliminary conversation segment; the semantic similarity between each two adjacent conversation contents in the preliminary conversation segment is calculated, and each semantic similarity is compared with a preset similarity threshold. The similarity threshold is pre-set by those skilled in the art according to actual conditions; the conversation content with a semantic similarity less than the similarity threshold is marked as irrelevant content, and the conversation content with a semantic similarity greater than the similarity threshold is marked as continuous content; if the conversation content is marked as both irrelevant content and continuous content, the mark as continuous content is cancelled; among the conversation times corresponding to the irrelevant content, the earliest conversation time is marked as the earliest time, and the latest time is marked as the latest time;

[0081] The continuous content whose conversation time is earlier than the earliest time and the irrelevant content corresponding to the earliest time are regarded as a continuous conversation segment; the continuous content whose conversation time is later than the latest time and the irrelevant content corresponding to the latest time are regarded as a continuous conversation segment; it is judged whether there is continuous content between any two irrelevant contents; if so, the corresponding irrelevant content and the continuous content between them are regarded as a continuous conversation segment; if not, no operation is performed.

[0082] Methods for obtaining dynamic response thresholds include:

[0083] A preset coefficient set includes a mode set and a difficulty set. The coefficient set is pre-set by those skilled in the art based on actual conditions. The mode set includes adjustment coefficients corresponding to different interaction modes, and the difficulty set includes adjustment coefficients corresponding to different interaction difficulties. In different interaction modes, users have different thinking and response speeds, and therefore corresponding adjustment coefficients are also different. For example, the adjustment coefficient for situational task interaction should be greater than the adjustment coefficient for casual chat interaction, which in turn should be greater than the adjustment coefficient for question-and-answer interaction. Complex or highly professional conversation content usually requires users to spend more time understanding and formulating responses. That is, the greater the interaction difficulty, the longer the user's response time, and therefore corresponding adjustment coefficients are different. For example, the adjustment coefficient for advanced difficulty should be greater than the adjustment coefficient for intermediate difficulty, which in turn should be greater than the adjustment coefficient for elementary difficulty. According to the interaction mode and interaction difficulty of the conversation interaction, the corresponding adjustment coefficient is obtained from the coefficient set. The product of the interaction mode adjustment coefficient and the interaction difficulty adjustment coefficient is used as the overall adjustment coefficient. The product of the overall adjustment coefficient and the preliminary response threshold is used as the dynamic response threshold.

[0084] Methods for calculating the semantic similarity between two adjacent conversation contents include:

[0085] Using the SimCSE model, each conversation is converted into a corresponding conversation vector. The cosine similarity between the conversation vectors of two adjacent conversations is calculated and used as the semantic similarity. It should be noted that the construction method of the SimCSE model and the calculation method of cosine similarity are both existing technologies, and the specific process will not be elaborated here. The training process of the SimCSE model includes:

[0086] Pre-collection Sentences ,in For the Each sentence is input into the SimCSE model. For each sentence, the same Transformer encoder (such as BERT) is used, and two different but semantically similar vector representations are obtained through two dropout encodings. and ;in, For sentences The vector representation after the first encoding is, For sentences The vector representation after the second encoding is, For the A sentence, ; , Where, represents the Transformer encoder, represents the randomly generated Dropout mask during the first pass through the encoder, Represents the randomly generated Dropout mask during the second pass through the encoder; dropout is a regularization method in neural networks that randomly discards some neuron connections each time it propagates forward;

[0087] Based on the two vector representations corresponding to each sentence, the vector similarity of each sentence is calculated; the expression of vector similarity is: Where, For sentences The vector similarity of express and The dot product of express Length of the module;

[0088] According to the vector similarity of each sentence, the contrast loss of each sentence is calculated; the expression of contrast loss is: Where, For sentences The contrast loss, is the temperature parameter (usually 0.05), is an indicator function, indicating that only when The value is 1 only when (avoiding comparison of the same sentence); add the contrast loss of each sentence in turn to obtain the overall loss;

[0089] Based on the overall loss, the back-propagation algorithm is used to calculate the gradient of the first model parameters, and the optimizer (such as Adam) is used to update the first model parameters. The training goal is to minimize the sum of all overall losses. The above training process is repeated until the sum of the overall losses reaches convergence and training is stopped.

[0090] The correction feedback module is used to perform multi-level language error analysis on continuous dialogue segments, identify language errors in continuous dialogue segments, and formulate corresponding language correction suggestions for language errors, which are fed back to users through a progressive feedback mechanism.

[0091] The steps for identifying language errors in continuous dialogue segments include:

[0092] Step S101: performing a lexical error analysis on each continuous dialogue segment to identify lexical errors;

[0093] Each continuous segment is input into the trained context-aware language model, and the masked prediction probability of each phrase in each continuous segment is calculated; a probability threshold is preset, and the probability threshold is preset by those skilled in the art based on actual conditions; each masked prediction probability is compared with the probability threshold, and phrases with a masked prediction probability less than the probability threshold are marked as incorrect phrases, while phrases with a masked prediction probability greater than or equal to the probability threshold are not marked; the same incorrect phrases in corresponding continuous dialogue segments are merged and regarded as lexical errors in the corresponding continuous dialogue segments; for example, in a Song Dynasty teahouse scenario, the user says: "Boss, please give me a cup of good tea." Because the unit of measurement for tea in the Song Dynasty should be "cup" rather than "cup", the "cup" is identified as an incorrect phrase;

[0094] In this embodiment, the preferred context-aware language model is the RoBERTa model. The training process of the RoBERTa model includes:

[0095] Collect large-scale text corpus in advance to build a training data set; build a vocabulary based on the training data set; take a text in the training data set as an input text sequence, and record each input text sequence as ,in, is the length of the input text sequence;

[0096] For the input text sequence, first randomly mask some of the words to obtain the masked word position set After the input text sequence passes through the word embedding layer and position encoding, it is input into the same Transformer encoder to obtain the hidden vector representation:

[0097] ;

[0098] ;

[0099] Where, is the vector representation input to the first layer of the Transformer encoder, The word vector matrix is ​​obtained after the input text sequence passes through the word embedding layer. is the position encoding vector obtained after position encoding of the input text sequence, is the input Transformer encoder The output representation of the layer, i.e. the hidden vector, , is the number of layers of the Transformer encoder, Represents the single-layer operation of the Transformer encoder, including the multi-head self-attention mechanism and feedforward neural network;

[0100] For each masked word position, calculate the corresponding conditional probability based on the hidden vector representation. The expression for the conditional probability is: ; where represents the probability that the RoBERTa model predicts the -th word as a specific word in the vocabulary when the second model parameter is and given the context (all words in the input text sequence except the -th word). is a function that converts a set of numerical values into a probability distribution, is the weight matrix of the output layer, is the bias vector of the output layer;

[0101] Define the loss function as: ; where is the value of the loss function;

[0102] According to the value of the loss function, calculate the gradient of the second model parameter using the backpropagation algorithm, and update the second model parameter using an optimizer; take minimizing the sum of all loss function values as the training objective, and repeat the above training process until the sum of the overall losses reaches convergence and stop training.

[0103] Step S102: Perform grammar error analysis on each continuous dialogue segment to identify grammar errors;

[0104] Input each continuous dialogue segment into the trained grammar error correction model to identify the grammar errors in each continuous dialogue segment; Exemplarily, in the Tang Dynasty court scene, the user says to the emperor: "I have already submitted the memorial to Your Majesty". Since "submitted to Your Majesty" is a complete verb-object structure and should not be split by the word "le", it is identified as a grammar error;

[0105] In this embodiment, the preferred grammar error correction model is the GECToR model. The training process of the GECToR model includes:

[0106] Pre-construct sentence pairs, denoted as ; where is the original sentence (containing grammar errors), is the correct sentence, is the number of training samples; Align each original sentence with the corresponding correct sentence, generate the editing operation label corresponding to each word in the original sentence, and mark it as the true label; Editing operation labels such as retain word, delete word, replace with a new word, etc.;

[0107] Each original sentence is input into the GECToR model in turn, and the BPE algorithm is used to segment each original sentence in turn to obtain the word sequence corresponding to each original sentence; each word sequence is passed through the embedding layer to obtain an embedding vector; each embedding vector is input into the Transformer encoder to obtain the context encoding representation of each word in each word sequence; the context encoding representation of each word is linearly transformed, and the expression of the linear transformation is: Where, For the The context encoding of a word represents the vector obtained after linear transformation. For the The context encoding representation of each word after linear transformation is used Function, obtains the probability of each word belonging to each editing operation label;

[0108] The cross entropy loss function is defined as: Where, is the cross entropy loss function value, For the The true label of the word, The GECToR model predicts The probability that a word belongs to the true label, is the original sentence length;

[0109] According to the cross entropy loss function value, the back propagation algorithm is used to calculate the gradient of the third model parameters, and the optimizer is used to update the third model parameters; the training goal is to minimize the sum of all cross entropy loss function values, and the above training process is repeated until the sum of the cross entropy loss function values ​​reaches convergence and the training is stopped.

[0110] Step S103: performing pragmatic error analysis on each continuous dialogue segment to identify pragmatic errors;

[0111] Each continuous dialogue segment is fed into the trained context analysis model to identify pragmatic errors in each continuous dialogue segment. For example, in a Tang Dynasty court scenario, a user says to the emperor, "Your Majesty, I have written this memorial. Please take a look at it, and don't forget to approve it." The tone of "Please take a look at it" is too colloquial and casual, which does not conform to the courtesy and politeness that an official should show to the emperor. "Don't forget to approve it" carries a sense of urging, appears disrespectful, and even has an imperative tone, which does not conform to court etiquette. Therefore, "Please take a look at it" and "Don't forget to approve it" are identified as pragmatic errors.

[0112] In this embodiment, the preferred context analysis model is a fine-tuned RoBERTa model. The fine-tuning process of the RoBERTa model includes:

[0113] Collect multiple pragmatically incorrect sentences in advance, use the BPE algorithm to perform word segmentation on each pragmatically incorrect sentence in turn, and pre-set a corresponding judgment label for each word in each pragmatically incorrect sentence; the judgment label includes a correct label and an incorrect label, the correct label is set to 0, and the incorrect label is set to 1; input each pragmatically incorrect sentence into the RoBERTa model, and after processing by a multi-layer Transformer encoder, output the predicted judgment label sequence corresponding to each pragmatically incorrect sentence, the predicted judgment label sequence includes the probability of each word in the pragmatically incorrect sentence output by the RoBERTa model belonging to each judgment label; based on the predicted judgment label sequence and the actual judgment label sequence of each pragmatically incorrect sentence, calculate the fine-tuned cross entropy loss function value corresponding to each pragmatically incorrect sentence, the actual judgment label sequence is the pre-set judgment label corresponding to each word in the pragmatically incorrect sentence, and the calculation method of the fine-tuned cross entropy loss function value is consistent with the calculation method of the above-mentioned cross entropy loss function value;

[0114] According to the fine-tuning cross entropy loss function value, the back propagation algorithm is used to calculate the gradient of the fourth model parameter, and the optimizer is used to update the fourth model parameter; the training goal is to minimize the sum of all fine-tuning cross entropy loss function values, and the above training process is repeated until the sum of the cross entropy loss function values ​​reaches convergence and the training is stopped.

[0115] Step S104: Combining vocabulary errors, grammatical errors, and pragmatic errors to determine language errors in each continuous dialogue segment;

[0116] The same lexical errors, grammatical errors and pragmatic errors in the corresponding continuous dialogue segments are combined as the language errors in the corresponding continuous dialogue segments.

[0117] Methods for formulating appropriate language correction suggestions for language errors include:

[0118] Each continuous dialogue segment and the corresponding language error are taken as a set of segment sets, and the segment sets correspond to the continuous dialogue segments one by one; each set of segment sets is input into the trained language correction model, and the correction set corresponding to each set of segment sets is output; the correction set includes the suggestion set corresponding to each language error, and the suggestion set includes the corresponding language error. Corrective suggestions, is an integer greater than 1;

[0119] In this embodiment, the preferred language correction model is the T5 model. The training process of the T5 model includes:

[0120] Collect multiple training samples in advance, each of which includes linguistically incorrect sentences 、Language error corresponding to the language error in the sentence And a target suggestion sequence composed of multiple correction suggestions; construct a corresponding input sequence based on the language error sentences and language errors in each training sample; use the BPE algorithm to segment each input sequence in turn to obtain the word sequence corresponding to each input sequence; pass each word sequence through the embedding layer to obtain the embedding vector; input each embedding vector into the Transformer encoder to obtain the context encoding representation of each word in each word sequence; use the Transformer decoder to use the target suggestion sequence as a supervision signal to gradually predict the probability of each word in the target suggestion sequence , thus generating correction suggestions word by word; where, The first words, Indicates that the generated words, is the fifth model parameter;

[0121] Calculate the prediction bias based on the probability of each word;

[0122] The expression for prediction deviation is: Where, is the prediction bias, Suggest sequence length for the target;

[0123] According to the prediction deviation, the back propagation algorithm is used to calculate the gradient of the fifth model parameters, and the optimizer is used to update the fifth model parameters; minimizing the sum of all prediction deviations is used as the training goal, and the above training process is repeated until the sum of the prediction deviations reaches convergence and the training is stopped.

[0124] From each suggestion set corresponding to each continuous dialogue segment, a correction suggestion is randomly selected and a set of suggestion subsets is constructed. Group suggestion subsets, is an integer greater than 1; the SimCSE model is used to convert the correction suggestions in each set of suggestion subsets into corresponding suggestion vectors; the dialogue vector corresponding to each continuous dialogue segment and the corresponding set of suggestion subsets are used as a set of analysis sets, and the analysis sets correspond to the continuous dialogue segments one-to-one; each set of analysis sets is input into the trained correction analysis model to predict the corresponding indicator set; the indicator set includes language fluency and cultural fit, language fluency refers to the naturalness and smoothness of the expression of the continuous dialogue segment, and cultural fit refers to the compliance of the continuous dialogue segment with the requirements of a specific cultural background, historical period and communication etiquette; the correction analysis model is a deep neural network model, which includes an input layer, a hidden layer and an output layer; each hidden layer includes multiple neurons, each neuron is connected to the neurons in the next layer, and the connection contains weights that determine the importance and influence of data transmission in the neural network; an activation function is applied to each neuron between the hidden layer and the output layer, and the activation function introduces nonlinearity, allowing the network to learn more complex patterns and features; the training process of the correction analysis model includes:

[0125] Pre-collection Set different analysis sets and set corresponding indicator sets for each analysis set. is an integer greater than 1; the analysis set and the corresponding indicator set are converted into a corresponding set of feature vectors; the indicator set corresponding to the analysis set is collected by those skilled in the art in the process of formulating language correction suggestions in history. Different analysis sets are set, and each analysis set is analyzed in turn to evaluate the corresponding language fluency and cultural fit, and used as the corresponding indicator set. Different analysis sets are set up in sequence with corresponding indicator sets;

[0126] Each set of feature vectors is used as the input of the correction analysis model. The correction analysis model takes a set of prediction indicators corresponding to each analysis set as the output, and the actual indicator set corresponding to each analysis set as the prediction target. The actual indicator set is the pre-set indicator set corresponding to the analysis set; the training goal is to minimize the sum of the prediction errors of all analysis sets; the calculation formula of the prediction error is: ,in is the prediction error, is the group number of the corresponding eigenvector of the analysis set, For the The set of prediction indicators corresponding to the group analysis set, For the The actual indicator set corresponding to the group analysis set is trained on the correction analysis model until the sum of the prediction errors reaches convergence and the training is stopped.

[0127] A preset ratio set includes proportional coefficients corresponding to language fluency and cultural fit, and the ratio set is pre-set by technical personnel in this field based on actual conditions. Based on the ratio set, the language fluency and cultural fit in each set of indicators are weighted and summed to obtain a comprehensive score for each set of suggestion subsets. The same comprehensive scores of corresponding continuous dialogue segments are compared, and the suggestion subset with the largest comprehensive score is used as the language correction suggestion for the corresponding continuous dialogue segment.

[0128] The steps to provide feedback to users through a progressive feedback mechanism include:

[0129] Step S201: Feedback the language errors corresponding to each continuous conversation segment to the user, and obtain a primary modification content; wherein the primary modification content is the natural language text output after the user modifies the corresponding conversation content based on the feedback language errors;

[0130] Step S202: Identify language errors in the modified content. If language errors still exist, proceed to step S203. If no language errors exist, the feedback ends.

[0131] Step S203: Based on a pre-built language knowledge base, the error cause corresponding to the language error identified in step S202 is obtained. The language knowledge base includes, but is not limited to, a historical and cultural knowledge base (e.g., tea was called "zhan" instead of "bei" in the Song Dynasty), a grammatical structure rule base (e.g., the verb-object structure cannot be separated by "le"), and a pragmatic norm base (e.g., court language requirements and social status discourse norms). These are pre-built by those skilled in the art based on language knowledge and cultural common sense such as historical and cultural norms, grammatical rule systems, and pragmatic etiquette guidelines, and are used to support the structured expression of background knowledge for error explanation.

[0132] Step S204: Feedback the error cause to the user and obtain secondary modified content, which is the natural language text output after the user modifies the corresponding conversation content based on the feedback error cause;

[0133] Step S205: Identify language errors in the second revised content. If language errors still exist, proceed to step S206. If no language errors exist, the feedback ends. The steps for identifying language errors in the first and second revised content are consistent with the steps for identifying language errors in the continuous dialogue segment.

[0134] Step S206: Feedback the correction set of all continuous dialogue segments and language correction suggestions to the user.

[0135] It should be noted that the beneficial effect of adopting a progressive feedback mechanism is that it provides a structured, in-depth and educational language learning process; by gradually identifying errors, providing specific explanations of the reasons and guiding users to make multiple revisions, it not only corrects superficial language errors, but also helps users understand the cultural, grammatical or pragmatic knowledge behind the errors, thereby promoting deeper language ability improvement; the final correction set and language correction suggestions reinforce learning outcomes, enabling users to systematically master correct usage, form long-term memory, and effectively improve the accuracy and cultural appropriateness of language expression.

[0136] The state evaluation module is used to collect user multimodal data, perform nonlinear feature mapping on the user multimodal data, generate a multidimensional state feature vector, and dynamically evaluate cognitive load and learning status based on the multidimensional state feature vector.

[0137] User multimodal data includes facial image sequences and eye movement trajectories; facial image sequences include those collected during the user's conversation and interaction with the virtual character. facial images, is an integer greater than 1, and the facial image sequence is obtained in real time through the built-in camera of the user terminal; the eye movement trajectory includes the gaze point, gaze duration and jump rate, which are used to characterize the user's visual attention state; among them, the gaze point is the spatial position where the user's eyes stay relatively stably, reflecting the specific area where the user's visual attention is concentrated; the gaze duration is the length of time the user's eyes stay at a single gaze point, indicating the depth of the user's processing of information in this area; the jump rate is the number of jumps per unit time, and jumps are the process of the user's eyes quickly moving from one gaze point to another; a higher jump rate indicates that the user's attention is distracted, while a lower jump rate indicates that the user is in a state of deep thinking; the eye detection algorithm and the eye tracking algorithm are comprehensively used to analyze the facial image sequence and extract the user's eye movement trajectory; the eye detection algorithm and the eye tracking algorithm are both existing technologies, and the specific process will not be described in detail here.

[0138] Methods for generating multi-dimensional state feature vectors include:

[0139] Using the trained emotion recognition model, each facial image in the facial image sequence is recognized in turn, and the recognition results of each facial image are output. The recognition results include category labels and emotion intensity; category labels are numerical labels corresponding to emotion categories, and different emotion categories have different numerical labels. Emotion categories include happiness, confusion, frustration, etc.; the range of emotion intensity is , 、 The specific value of is set by those skilled in the art according to actual conditions. ;According to the category label in the recognition result, the corresponding emotion category is obtained;The number of each emotion category is counted and marked as the emotion number;All emotion numbers are compared, and the emotion category with the largest emotion number is taken as the user's emotional state;All emotion intensities corresponding to the emotional state are averaged to obtain the average emotion intensity;The emotion recognition model is a convolutional neural network model, which is an extended form of deep neural network. It mainly includes an input layer, multiple convolution layers, a pooling layer (downsampling layer), a fully connected layer and an output layer;Among them, the convolution layer is composed of multiple convolution kernels (filters), each convolution kernel slides with the input data to extract local features, and the convolution operation contains weight parameters for learning the importance of features;The convolution layer is usually followed by a pooling layer, which is used to reduce the size of the feature map, retain the main information while reducing the computational complexity;After the convolution layer and the pooling layer, a fully connected layer is connected to map the high-dimensional features to the output space;In the convolution layer, the fully connected layer and other layers, each neuron usually applies an activation function to introduce nonlinearity to enhance the model's expression and generalization capabilities for complex image patterns;The specific training process of the emotion recognition model includes:

[0140] Collect multiple facial images in advance, and annotate the patient's face in each facial image, including the category label and emotion intensity. Divide the annotated facial images into a training set and a test set, with 70% of the facial images used as the training set and 30% of the facial images used as the test set. Use the training set to train the emotion recognition model, and use the test set to test the emotion recognition model. Preset an error threshold, and output the emotion recognition model when the mean prediction error of all facial images in the test set is less than the error threshold. The calculation formula for the mean prediction error is: ,in is the prediction error, is the number of the facial image, For the The predicted category labels corresponding to the group of facial images, For the The actual category labels corresponding to the group of facial images, For the The predicted emotion intensity corresponding to the group of facial images, For the The actual emotional intensity corresponding to the group of facial images, is the number of facial images in the test set; the error threshold is preset according to the accuracy required by the emotion recognition model.

[0141] A clustering algorithm (such as a K-means clustering algorithm or a density-based spatial clustering algorithm) is used to cluster all fixation points in the eye movement trajectory to obtain multiple fixation areas; the number of fixation areas is counted and marked as the number of fixations; all gaze durations in the eye movement trajectory are averaged to obtain the average gaze duration; a weight set is preset, which includes weight coefficients corresponding to the inverse of the number of fixations, the average gaze duration, and the inverse of the skip rate. The weight set is preset by a person skilled in the art according to actual conditions; the number of fixations, the average gaze duration, and the skip rate are all standardized (such as by Z-score standardization or Min-Max standardization), and the inverse of the number of fixations, the average gaze duration, and the inverse of the skip rate after the normalization are weighted and summed according to the weight set to obtain the degree of attention concentration;

[0142] Count the total number of vocabulary errors, grammatical errors, and pragmatic errors in all consecutive dialogue segments and mark them as the number of language errors; count the number of times the progressive feedback mechanism provides feedback to users and mark them as the number of feedbacks;

[0143] A multidimensional state feature vector is generated based on the emotional state, average emotional intensity, attention concentration, number of language errors, and number of feedbacks.

[0144] Methods for dynamically assessing cognitive load and learning status include:

[0145] The category label, average emotion intensity, and attention concentration corresponding to the emotional state in the multidimensional state feature vector are used as first analysis data, and the number of language errors and the number of feedback times in the multidimensional state feature vector are used as second analysis data; the learning state is dynamically evaluated based on the first analysis data, and the cognitive load is dynamically evaluated based on the second analysis data, and the method for dynamically evaluating the learning state based on the first analysis data is consistent with the method for dynamically evaluating the cognitive load based on the second analysis data;

[0146] The method for dynamically evaluating the learning status according to the first analysis data includes:

[0147] Construct multiple fuzzy sets for each data in the first analysis data; for example, the fuzzy sets corresponding to the category labels are positive, neutral, negative, etc.; the fuzzy sets corresponding to the average emotional intensity are high intensity, medium intensity, low intensity, etc.; the fuzzy sets corresponding to the attention concentration are high concentration, medium concentration, low concentration, etc.; each data in the first analysis data is converted into the membership of each corresponding fuzzy set through fuzzification technology; fuzzification technology is the process of converting precise numerical values ​​into the membership corresponding to the fuzzy set, and fuzzification technologies include triangular membership function, trapezoidal membership function, etc.; for example, if the value of the average emotional intensity is low, it is inferred that the membership of the high intensity is 0.1, the membership of the medium intensity is 0.3, and the membership of the low intensity is 0.8;

[0148] Define fuzzy rules, which are defined based on expert knowledge or relevant literature; for example, if the category label is positive, high intensity, and high concentration, then it is inferred that the learning state has a high positive membership; if the category label is negative, high intensity, and low concentration, then it is inferred that the learning state has a high negative membership; match the fuzzified first analysis data with the fuzzy rules respectively, and use fuzzy reasoning methods (such as Mamdani fuzzy reasoning model, Sugeno fuzzy reasoning model, etc.) to perform fuzzy reasoning to obtain fuzzy reasoning results, which are the membership of each learning state, including but not limited to positive, stable, negative, etc.; for example, the positive membership is 0.6, the neutral membership is 0.8, and the negative membership is 0.1; compare each membership separately, and take the learning state with the largest membership as the dynamically evaluated learning state.

[0149] The interaction optimization module is used to integrate cognitive load and learning status, and to infer interaction regulation strategies, and dynamically optimize the interaction mode and interaction difficulty in dialogue interaction based on the interaction regulation strategies.

[0150] Methods for reasoning about interactive regulatory strategies include:

[0151] The corresponding load label is obtained according to the cognitive load, and the corresponding state label is obtained according to the learning state; wherein, the load label is a numerical label corresponding to the cognitive load, and different cognitive loads correspond to different load labels; the state label is a numerical label corresponding to the learning state, and different cognitive loads correspond to different numerical labels; the load label and the state label are input into the trained direction judgment model to predict the corresponding direction label. The direction judgment model is a deep neural network model. The specific training process of the direction judgment model is consistent with the training process of the above-mentioned correction analysis model, except that the input of the direction judgment model is the load label and the state label, and the output is the direction label;

[0152] The direction label is a numerical label corresponding to the control direction. Different control directions have different direction labels. The control directions include positive control and negative control. Positive control indicates that the user's current learning state is good and the cognitive load is low. Increasing the difficulty and deepening the interaction can further stimulate the user's learning motivation and enhance learning results. Negative control indicates that the user's current learning state is poor and the cognitive load is high. It is necessary to reduce the difficulty and increase guidance to help the user recover and prevent cognitive overload and learning frustration.

[0153] A corresponding guidance intensity is set for each interaction mode, and all interaction modes are sorted from high to low according to the corresponding guidance intensity to generate a mode sequence. For example, question-and-answer interaction is led by a virtual character and has clear content, so the guidance intensity is higher; role-playing interaction has greater user autonomy, flexible interaction methods, and high explorability, so the guidance intensity is weaker. The guidance intensity is set by those skilled in the art based on actual conditions. All interaction difficulties are sorted from low to high to generate a difficulty sequence.

[0154] If the control direction is positive control, the interaction mode that follows the current mode is obtained from the mode sequence, and the interaction difficulty that follows the current difficulty is obtained from the difficulty sequence; if the control direction is negative control, the interaction mode that precedes the current mode is obtained from the mode sequence; the interaction difficulty that precedes the current difficulty is obtained from the difficulty sequence; wherein, the current mode is the interaction mode corresponding to the user's current dialogue interaction, and the current difficulty is the interaction difficulty corresponding to the user's current dialogue interaction; the obtained interaction modes are all marked as control modes, and the obtained interaction difficulties are all marked as control difficulties;

[0155] Randomly select a control mode and a control difficulty, construct a control strategy, and build A regulatory strategy, is an integer greater than 0; calculate the control effect of each group of control strategies in turn, compare all the control effects, and take the control strategy with the greatest control effect as the interactive control strategy.

[0156] Methods for calculating the control effect of each group of control strategies in turn include:

[0157] Different numerical labels are set for different interaction modes and marked as mode labels; different numerical labels are set for different interaction difficulties and marked as difficulty labels; the multidimensional state feature vector, the mode label of the current mode, the difficulty label of the current difficulty, and the mode label of the control mode and the difficulty label of the control difficulty in a control strategy are used as a set of control data; each set of control data is input into the trained effect prediction model to predict the corresponding control effect; the control effect reflects the improvement of the user's learning status and cognitive load under the current control strategy; the effect prediction model is a deep neural network model, and the specific training process of the effect prediction model is consistent with the training process of the above-mentioned correction analysis model, the difference is that the input of the effect prediction model is the control data, and the output is the control effect.

[0158] This embodiment constructs a virtual cultural interaction scene and integrates multi-dimensional historical and cultural elements to provide users with an immersive language practice environment, enhance their understanding of traditional Chinese culture, and improve the fun and participation of Chinese language learning. It collects continuous dialogue fragments during real-time dialogue and conducts multi-level language error analysis to accurately identify various types of errors such as vocabulary, grammar, and pragmatics, and formulates targeted and progressive language correction suggestions to effectively promote the systematic improvement of users' Chinese language proficiency. It collects users' facial expressions and eye movement data, dynamically analyzes cognitive load and learning status, and combines intelligent reasoning to optimize the interaction mode and difficulty in real time, so that the learning process is more in line with the users' learning needs and cognitive characteristics. It achieves precise control and dynamic regulation of the entire process of users' Chinese language learning, provides users with a personalized and efficient online Chinese learning experience, and promotes the comprehensive improvement of users' language ability.

[0159] Example 2

[0160] The present application also provides an electronic device. The electronic device may include one or more processors and one or more memories. The memories may store computer-readable code that, when executed by the one or more processors, may implement a real-time interactive online Chinese language learning platform as described above.

[0161] The method or system according to the embodiment of the present application can also be implemented with the help of the architecture of the electronic device shown in this application. The electronic device may include a bus, one or more CPUs, ROM, RAM, a communication port connected to a network, input / output, a hard disk, etc. The storage device in the electronic device, such as ROM or hard disk, can store a real-time interactive online Chinese language learning platform provided by this application. Furthermore, the electronic device may also include a user interface. Of course, the architecture shown in this application is only exemplary. When implementing different devices, one or more components in the electronic device shown in this application can be omitted according to actual needs.

[0162] Example 3

[0163] One embodiment of the present application discloses a computer-readable storage medium. The computer-readable storage medium stores computer-readable instructions. When the computer-readable instructions are executed by a processor, a real-time interactive online Chinese language learning platform according to an embodiment of the present application described with reference to the above figures can be executed. The storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc.

[0164] In addition, according to embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the present application provides a non-transitory machine-readable storage medium storing machine-readable instructions that can be executed by a processor to perform instructions corresponding to the steps of the method provided in the present application, such as a real-time interactive online Chinese language learning platform. When the computer program is executed by a central processing unit (CPU), the above-mentioned functions defined in the method of the present application are performed.

[0165] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art will be able to modify the technical solutions described in the foregoing embodiments or to substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

[0166] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0167] In the description of the present invention, it should be understood that the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0168] In the description of the present invention, unless otherwise specified, "plurality" means two or more.

[0169] In the description of the present invention, “several” means one or more, and “a large number” means two or more.

[0170] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0171] The formulas in this manual are all dimensionless and calculated using numerical values. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field based on actual conditions.

[0172] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

Claims

1. A real-time interactive online Chinese language learning platform, characterized by: include: A scenario construction module is used to establish a multi-dimensional cultural element matrix and build a virtual cultural interaction scenario based on the multi-dimensional cultural element matrix; An interaction collection module is used to collect continuous dialogue segments in real time during the process of dialogue interaction between users and virtual characters in the virtual cultural interaction scene; The method for collecting continuous conversation segments in real time comprises: Continuously collect conversation interaction information between users and avatars, including conversation content and corresponding conversation time. Conversation content is the natural language text output by the user or avatar during the conversation interaction, and conversation time is the output time of each conversation content. Obtain historical duration, average the historical duration, and obtain a preliminary response threshold; the historical duration is the response duration of the user during the dialogue interaction with the virtual character at the historical moment; obtain the interaction mode and interaction difficulty of the dialogue interaction, dynamically adjust the preliminary response threshold based on the interaction mode and interaction difficulty, and obtain a dynamic response threshold; after each round of the virtual character outputting natural language text, monitor whether the user outputs natural language text within the dynamic response threshold; if so, determine that the dialogue interaction continues and continue to collect dialogue interaction information; if not, determine that the dialogue interaction ends, and use the collected dialogue content as the preliminary dialogue segment; Calculate the semantic similarity between each two adjacent conversation contents in the preliminary conversation segment, and compare each semantic similarity with a preset similarity threshold; mark the conversation contents with semantic similarity less than the similarity threshold as irrelevant content, and mark the conversation contents with semantic similarity greater than the similarity threshold as continuous content; if the conversation content is marked as both irrelevant content and continuous content, unmark it as continuous content; mark the earliest conversation time as the earliest time and the latest time as the latest time among the conversation times corresponding to the irrelevant content; The continuous content whose conversation time is earlier than the earliest time and the irrelevant content corresponding to the earliest time are regarded as a continuous conversation segment; the continuous content whose conversation time is later than the latest time and the irrelevant content corresponding to the latest time are regarded as a continuous conversation segment; whether there is continuous content between any two irrelevant contents is determined; if so, the corresponding irrelevant content and the continuous content between them are regarded as a continuous conversation segment; if not, no operation is performed; The correction feedback module is used to perform multi-level language error analysis on continuous dialogue segments, identify language errors in continuous dialogue segments, and formulate corresponding language correction suggestions for language errors, which are fed back to users through a progressive feedback mechanism; The state assessment module is used to collect user multimodal data, perform nonlinear feature mapping on the user multimodal data, generate a multidimensional state feature vector, and dynamically assess cognitive load and learning status based on the multidimensional state feature vector; The interaction optimization module is used to integrate cognitive load and learning status, and to infer interaction regulation strategies, and dynamically optimize the interaction mode and interaction difficulty in dialogue interaction based on the interaction regulation strategies.

2. A real-time interactive online Chinese language learning platform according to claim 1, characterized in that: The method for obtaining the dynamic response threshold comprises: A preset coefficient set includes a mode set and a difficulty set; the mode set includes adjustment coefficients corresponding to different interaction modes, and the difficulty set includes adjustment coefficients corresponding to different interaction difficulties; according to the interaction mode and interaction difficulty of the dialogue interaction, the corresponding adjustment coefficient is obtained from the coefficient set; the product of the adjustment coefficient of the interaction mode and the adjustment coefficient of the interaction difficulty is used as the overall adjustment coefficient; the product of the overall adjustment coefficient and the preliminary response threshold is used as the dynamic response threshold; Methods for calculating the semantic similarity between two adjacent conversation contents include: The SimCSE model is used to convert each conversation content into a corresponding conversation vector. The cosine similarity between the conversation vectors corresponding to each two adjacent conversation contents is calculated in sequence and used as the semantic similarity.

3. A real-time interactive online Chinese language learning platform according to claim 2, characterized in that: The step of identifying language errors in the continuous dialogue segment comprises: Step S101: performing a lexical error analysis on each continuous dialogue segment to identify lexical errors; Each continuous segment is input into the trained context-aware language model, and the masked prediction probability of each phrase in each continuous segment is calculated. A probability threshold is preset, and each masked prediction probability is compared with the probability threshold. Phrases with a masked prediction probability less than the probability threshold are marked as incorrect phrases, while phrases with a masked prediction probability greater than or equal to the probability threshold are not marked. The same incorrect phrases in corresponding continuous dialogue segments are merged and regarded as lexical errors in the corresponding continuous dialogue segments. Step S102: performing grammatical error analysis on each continuous dialogue segment to identify grammatical errors; Input each continuous dialogue segment into the trained grammatical error correction model to identify grammatical errors in each continuous dialogue segment; Step S103: performing pragmatic error analysis on each continuous dialogue segment to identify pragmatic errors; Each continuous dialogue segment is input into the trained context analysis model to identify pragmatic errors in each continuous dialogue segment; Step S104: Combining vocabulary errors, grammatical errors, and pragmatic errors to determine language errors in each continuous dialogue segment; The same lexical errors, grammatical errors and pragmatic errors in the corresponding continuous dialogue segments are combined as the language errors in the corresponding continuous dialogue segments.

4. A real-time interactive online Chinese language learning platform according to claim 3, characterized in that: The method for formulating corresponding language correction suggestions for language errors includes: Each continuous conversation segment and the corresponding language error are treated as a set of segment sets, each set of segment sets is input into the trained language correction model, and a correction set corresponding to each set of segment sets is output; the correction set includes a suggestion set corresponding to each language error, and the suggestion set includes a correction suggestions corresponding to the language error, where a is an integer greater than 1; From each suggestion set corresponding to each continuous dialogue segment, a correction suggestion is randomly selected to construct a set of suggestion subsets. For each continuous dialogue segment, b suggestion subsets are constructed, where b is an integer greater than 1. The correction suggestions in each suggestion subset are converted into corresponding suggestion vectors using the SimCSE model. The dialogue vector corresponding to each continuous dialogue segment and the corresponding set of suggestion subsets are used as an analysis set. Each analysis set is input into a trained correction analysis model to predict the corresponding indicator set. The indicator set includes language fluency and cultural fit. The correction analysis model is a deep neural network model. A preset ratio set includes proportional coefficients corresponding to language fluency and cultural fit; based on the ratio set, the language fluency and cultural fit in each indicator set are weighted and summed to obtain a comprehensive score for each set of suggestion subsets; the same comprehensive scores of corresponding continuous dialogue segments are compared, and the suggestion subset with the largest comprehensive score is used as the language correction suggestion for the corresponding continuous dialogue segment.

5. A real-time interactive online Chinese language learning platform according to claim 4, characterized in that: The step of providing feedback to the user through the progressive feedback mechanism includes: Step S201: Feedback the language errors corresponding to each continuous dialogue segment to the user and obtain the modified content; Step S202: Identify language errors in the modified content. If language errors still exist, proceed to step S203. If no language errors exist, the feedback ends. Step S203: Based on the pre-built language knowledge base, obtain the error cause corresponding to the language error identified in step S202; Step S204: Feedback the error reason to the user and obtain the secondary modification content; Step S205: Identify language errors in the second modified content. If language errors still exist, proceed to step S206. If no language errors exist, the feedback ends. Step S206: Feedback the correction set of all continuous dialogue segments and language correction suggestions to the user.

6. A real-time interactive online Chinese language learning platform according to claim 5, characterized in that: The user multimodal data includes a facial image sequence and an eye movement trajectory; the facial image sequence includes c facial images collected during the user's conversation and interaction with the virtual character, where c is an integer greater than 1; the eye movement trajectory includes the gaze point, gaze duration, and skip rate; The method for generating a multidimensional state feature vector comprises: Using a trained emotion recognition model, each facial image in the facial image sequence is identified in turn, and the recognition results of each facial image are output. The recognition results include category labels and emotion intensity. The category labels are numerical labels corresponding to the emotion categories, and different emotion categories have different numerical labels. The range of emotion intensity is [d1, d2], 0<d1<d2. According to the category labels in the recognition results, the corresponding emotion category is obtained. The number of each emotion category is counted and marked as the number of emotions. All emotion numbers are compared, and the emotion category with the largest number of emotions is regarded as the user's emotional state. All emotion intensities corresponding to the emotional state are averaged to obtain the average emotion intensity. The emotion recognition model is a convolutional neural network model. A clustering algorithm is used to cluster all fixation points in the eye movement trajectory to obtain multiple fixation areas; the number of fixation areas is counted and marked as the number of fixations; all gaze durations in the eye movement trajectory are averaged to obtain the average gaze duration; a weight set is preset, which includes weight coefficients corresponding to the inverse of the number of fixations, the average gaze duration, and the inverse of the skip rate; the number of fixations, the average gaze duration, and the skip rate are all standardized, and the inverse of the standardized number of fixations, the average gaze duration, and the inverse of the skip rate are weighted summed according to the weight set to obtain the attention concentration; Count the total number of vocabulary errors, grammatical errors, and pragmatic errors in all consecutive dialogue segments and mark them as the number of language errors; count the number of times the progressive feedback mechanism provides feedback to users and mark them as the number of feedbacks; A multidimensional state feature vector is generated based on the emotional state, average emotional intensity, attention concentration, number of language errors, and number of feedbacks.

7. A real-time interactive online Chinese language learning platform according to claim 6, characterized in that: Methods for dynamically assessing cognitive load and learning status include: The category label, average emotion intensity, and attention concentration corresponding to the emotional state in the multidimensional state feature vector are used as first analysis data, and the number of language errors and the number of feedback times in the multidimensional state feature vector are used as second analysis data; the learning state is dynamically evaluated based on the first analysis data, and the cognitive load is dynamically evaluated based on the second analysis data, and the method for dynamically evaluating the learning state based on the first analysis data is consistent with the method for dynamically evaluating the cognitive load based on the second analysis data; The method for dynamically evaluating the learning status according to the first analysis data includes: A plurality of fuzzy sets are constructed for each data in the first analysis data; each data in the first analysis data is converted into the membership of each corresponding fuzzy set through fuzzification technology; fuzzy rules are defined, the fuzzified first analysis data are matched with the fuzzy rules, and fuzzy reasoning methods are used to perform fuzzy reasoning to obtain fuzzy reasoning results, which are the membership of each learning state; each membership is compared, and the learning state with the largest membership is used as the learning state dynamically evaluated.

8. A real-time interactive online Chinese language learning platform according to claim 7, characterized in that: The method for performing interactive regulation strategy reasoning includes: Obtain the corresponding load label based on the cognitive load, and obtain the corresponding state label based on the learning state; wherein the load label is the numerical label corresponding to the cognitive load, and the state label is the numerical label corresponding to the learning state; input the load label and the state label into the trained direction judgment model to predict the corresponding direction label, and the direction judgment model is a deep neural network model; wherein the direction label is the numerical label corresponding to the control direction, and the control direction includes positive control and negative control; Set the corresponding guidance intensity for each interaction mode, sort all interaction modes from high to low according to the corresponding guidance intensity to generate a mode sequence; sort all interaction difficulties from low to high to generate a difficulty sequence; If the control direction is positive control, the interaction mode that follows the current mode is obtained from the mode sequence, and the interaction difficulty that follows the current difficulty is obtained from the difficulty sequence; if the control direction is negative control, the interaction mode that precedes the current mode is obtained from the mode sequence; the interaction difficulty that precedes the current difficulty is obtained from the difficulty sequence; wherein, the current mode is the interaction mode corresponding to the user's current dialogue interaction, and the current difficulty is the interaction difficulty corresponding to the user's current dialogue interaction; the obtained interaction modes are all marked as control modes, and the obtained interaction difficulties are all marked as control difficulties; Randomly select a regulation mode and a regulation difficulty, construct a regulation strategy, and construct k regulation strategies in total, where k is an integer greater than 0; calculate the regulation effect of each group of regulation strategies in turn, compare all regulation effects, and take the regulation strategy with the greatest regulation effect as the interactive regulation strategy.

9. A real-time interactive online Chinese language learning platform according to claim 8, characterized in that: The method of sequentially calculating the control effect of each group of control strategies includes: Different numerical labels are set for different interaction modes and marked as mode labels; different numerical labels are set for different interaction difficulties and marked as difficulty labels; the multidimensional state feature vector, the mode label of the current mode, the difficulty label of the current difficulty, and the mode label of the control mode and the difficulty label of the control difficulty in a control strategy are used as a set of control data; each set of control data is input into the trained effect prediction model to predict the corresponding control effect; the effect prediction model is a deep neural network model.

Citation Information

Patent Citations

  • Real-time interactive online Chinese language learning platform

    CN115292499A

  • Conversational interaction method, device and system based on large model

    CN119311808A

  • Middle and primary school multi-person foreign language situational teaching method and system based on VR

    CN120031684A

  • Conversation learning system using artificial intelligence avatar tutor, and method therefor

    WO2022182064A1