Primary and secondary school english ai study companion system and method based on multi-modal education large model

By collecting and analyzing multimodal data, a quantitative parameter system was constructed to generate personalized learning paths and provide psychological intervention. This solved the problems of insufficient adaptability and psychological attention in the existing system, and improved the effectiveness of English learning and cross-cultural competence.

CN120997002BActive Publication Date: 2026-04-17SHENZHEN TONGMOJI ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN TONGMOJI ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
Filing Date
2025-08-05
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

The existing English learning systems for primary and secondary schools fail to accurately adapt to the new national English curriculum standards, lack real-time attention and intervention to students' psychological state, and have a one-dimensional design of cultural education content, failing to effectively improve students' ability to express Chinese culture in English and their comprehensive intercultural literacy.

Method used

By collecting students' multimodal data, a multi-dimensional quantitative parameter system is constructed to assess their English proficiency and anxiety levels in real time, generate personalized learning paths, and provide psychological intervention through virtual interactive tutors. This optimizes cross-cultural learning programs and combines federated learning algorithms to optimize model parameters in real time.

Benefits of technology

It enables accurate assessment and feedback of English proficiency, significantly alleviates student anxiety, enhances cross-cultural literacy, and improves the fit between learning content and domestic curriculum goals, as well as the effect of cultural integration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997002B_ABST
    Figure CN120997002B_ABST
Patent Text Reader

Abstract

The application discloses a primary and secondary school English AI learning companion system and method based on a multi-modal education large model, comprising the following steps: collecting multi-modal learning state data such as facial expressions, voices and writing pressure of students in real time; constructing multi-dimensional quantitative parameters of English oral speed, vocabulary usage frequency and cultural knowledge proportion in line with the requirements of new curriculum standards; calculating the difference value of the English ability state of students and the quantitative index of the anxiety state; generating a personalized learning path and a psychological intervention scheme; dynamically optimizing the cross-cultural learning scheme by using the Monte Carlo tree search algorithm; and updating the learning model parameters in real time through federated learning. The application realizes accurate adaptation to the requirements of the new national English curriculum standards, and significantly improves the English ability, psychological health level and cross-cultural integration ability of students.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart education technology, and in particular to a primary and secondary school English AI learning companion system and method based on a multimodal education model. Background Technology

[0002] With the deep integration of artificial intelligence technology and education, AI-assisted language learning systems have become an important development direction for primary and secondary school English teaching. Currently, most intelligent English learning systems are designed and content-matched using internationally recognized language learning grading standards. These standards are designed to serve the globalization of English as a second language teaching, but they do not fully consider the specific ability requirements and cultural and educational goals of students at different grade levels according to my country's new primary and secondary school English curriculum standards. This leads to a disconnect between these systems and domestic teaching objectives in practical applications.

[0003] Most existing language learning systems use students' academic performance as the sole evaluation indicator, focusing on the training and assessment of language skills themselves. They generally lack real-time attention and intervention measures for students' psychological state during the learning process. In particular, they fail to effectively consider the common learning anxiety problem, making it difficult to detect and guide students' psychological state and emotional changes in a timely manner, which affects students' overall learning experience and learning outcomes.

[0004] In addition, the existing language learning systems mostly design cultural education content as a one-way input of Western cultural content, lacking an effective training mechanism to integrate Chinese culture and achieve two-way integration of Chinese and foreign cultural content. As a result, students fail to effectively improve their English expression ability of Chinese culture during the English learning process, making it difficult for them to reflect their own cultural characteristics and advantages in cross-cultural communication, which is not conducive to the comprehensive improvement of students' cross-cultural literacy.

[0005] It is evident that current AI-powered English learning systems for primary and secondary schools still have significant shortcomings in areas such as accurately adapting teaching content to the new curriculum standards, providing real-time attention and intervention for learning psychology, and integrating cultural training in both directions. There is an urgent need to propose an intelligent learning solution that can comprehensively and accurately adapt to my country's new curriculum standards and combine multi-dimensional psychological guidance with cultural training in both directions.

[0006] Therefore, how to provide an AI-powered English learning companion system and methods for primary and secondary schools based on a multimodal education model is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0007] One objective of this invention is to propose an AI-powered English learning companion system and method for primary and secondary schools based on a multimodal education model. Addressing the shortcomings of existing technologies in effectively adapting to the new national English curriculum standards and lacking psychological counseling and cultural integration training, this invention proposes a technical solution that includes real-time collection of students' multimodal data, precise quantification of English proficiency and anxiety levels, dynamic adjustment of personalized learning paths and psychological interventions, and optimization of cross-cultural integration learning programs. This invention possesses the advantages of accurately adapting to the new curriculum standards, effectively alleviating learning anxiety, and enhancing cross-cultural comprehensive literacy.

[0008] According to an embodiment of the present invention, a method for providing AI-powered English learning companions for primary and secondary schools based on a multimodal education model includes:

[0009] The system collects students' facial expression data through cameras, voice data through microphones, and writing trajectory and writing pressure data through smart styluses, and integrates them into multimodal learning state feature data.

[0010] Based on the new national compulsory education English curriculum standards, a multi-dimensional quantitative parameter system was constructed for English speaking speed, vocabulary usage frequency, and the proportion of cultural knowledge.

[0011] By matching multimodal learning state feature data with a multidimensional quantitative parameter system, the difference between students' English proficiency status and the English proficiency status required by the new curriculum standards can be determined.

[0012] By integrating facial micro-expression recognition algorithm, voice tremor frequency detection algorithm, and writing pressure sensor, a quantitative index of student anxiety state is generated.

[0013] Based on the differences in English proficiency and the quantitative index of students' anxiety, a personalized learning path is generated. When the anxiety index exceeds the threshold, a psychological guidance dialogue is triggered by a virtual interactive tutor and mindfulness audio is pushed for psychological intervention.

[0014] Using data on students’ interest in Chinese culture and their cross-cultural adaptability as input, and based on a bidirectional cultural knowledge graph, the Monte Carlo tree search algorithm is used to optimize the ratio of Western cultural knowledge input to Chinese cultural knowledge output, and dynamically generate cross-cultural learning programs.

[0015] By transmitting data across terminals in real time, feedback is provided on differences in English proficiency, a quantitative index of anxiety, and data on the effectiveness of cross-cultural learning programs. Based on federated learning algorithms, learning paths and psychological intervention programs are optimized and updated in real time, serving as the basis for the next analysis of multimodal learning status characteristics and path generation.

[0016] Optionally, the step of collecting students' facial expression data via camera, voice data via microphone, and writing trajectory and pressure data via smart stylus, and fusing them into multimodal learning state feature data, specifically:

[0017] The student's facial expression data collected in real time by the camera is processed frame by frame for noise reduction, and the positions of key facial points and the amplitude of movements are extracted.

[0018] Time-domain framing and spectral feature analysis were performed on the student speech data collected in real time by the microphone to extract the fundamental frequency, harmonic frequency and formant position parameters of the speech.

[0019] The writing trajectory data of students collected in real time by the smart stylus is processed by two-dimensional coordinate mapping to obtain the writing trajectory coordinate parameters, and the pressure value parameters of the writing pressure data are recorded simultaneously.

[0020] Using facial key point location and movement amplitude parameters, fundamental frequency, harmonic frequency and formant location parameters of speech, writing trajectory coordinate parameters and pressure numerical parameters as initial inputs, principal component analysis algorithm is used to determine the weighting factors of each parameter respectively;

[0021] The weight factors of each parameter are multiplied by the quantized values ​​of the corresponding parameters and then summed to obtain the fused multimodal learning state feature data.

[0022] Optionally, the construction of a multi-dimensional quantitative parameter system for English speaking speed, vocabulary usage frequency, and the proportion of cultural knowledge specifically includes:

[0023] Based on the English oral proficiency requirements for grades 3 to 9 as stipulated in the new national compulsory education English curriculum standards, the upper and lower limits of the number of words pronounced per minute in English oral expression for each grade are determined and converted into quantitative parameters of English oral speaking speed.

[0024] Based on the specific requirements for the amount of English vocabulary mastery in each grade in the new curriculum standards, the total amount of vocabulary is clearly divided into three independent categories: basic vocabulary, extended vocabulary, and cultural vocabulary, and the frequency parameters of each type of vocabulary in the teaching content are determined.

[0025] Based on the requirement of integrating Chinese and Western cultural knowledge teaching proposed in the new curriculum standards, the teaching time of Chinese traditional cultural knowledge and Western cultural knowledge involved in classroom teaching is quantified to form parameters for the proportion of Chinese traditional cultural knowledge and Western cultural knowledge.

[0026] Based on the aforementioned English spoken speed quantification parameters, basic vocabulary usage frequency parameters, extended vocabulary usage frequency parameters, and cultural vocabulary usage frequency parameters, English teaching difficulty level parameters are obtained.

[0027] Using the parameters of the proportion of knowledge in traditional Chinese culture and the proportion of knowledge in Western culture as input conditions, we obtain the cultural sensitivity level parameters.

[0028] After standardizing and normalizing the quantitative parameters of spoken English speed, the frequency of use of various words, the difficulty level of English teaching, and the level of cultural sensitivity, a multi-dimensional quantitative parameter system is obtained.

[0029] Optionally, the step of matching multimodal learning state feature data with a multidimensional quantitative parameter system to determine the difference between the student's English proficiency status and the English proficiency status required by the new curriculum standards specifically involves:

[0030] The numerical difference between the actual English speaking speed parameter of students in the multimodal learning state feature data and the English speaking speed quantification parameter in the multidimensional quantification parameter system is calculated to obtain the English speaking speed difference value.

[0031] The difference between the actual vocabulary usage frequency parameter of students in the multimodal learning state feature data and the corresponding basic vocabulary usage frequency parameter, extended vocabulary usage frequency parameter and cultural vocabulary usage frequency parameter in the multidimensional quantitative parameter system is calculated to obtain the difference value of basic vocabulary usage frequency, the difference value of extended vocabulary usage frequency and the difference value of cultural vocabulary usage frequency.

[0032] The difference between the actual proportion of Chinese traditional culture knowledge and the proportion of Western culture knowledge in the multimodal learning state feature data and the parameters of the proportion of Chinese traditional culture knowledge and the proportion of Western culture knowledge in the multidimensional quantitative parameter system are calculated to obtain the difference value of the proportion of Chinese traditional culture knowledge and the difference value of the proportion of Western culture knowledge.

[0033] Using the differences in spoken English speed, frequency of use of basic vocabulary, frequency of use of extended vocabulary, frequency of use of cultural vocabulary, proportion of knowledge of traditional Chinese culture, and proportion of knowledge of Western culture as input conditions, a weighted summation method was used to calculate the comprehensive difference in English proficiency.

[0034] Based on the magnitude of the comprehensive difference value of English proficiency and the preset difference threshold, the degree of gap between students' actual English proficiency and the ability requirements stipulated in the new curriculum standards is determined, and the difference value of English proficiency is obtained.

[0035] Optionally, the step of generating a quantitative index of student anxiety by fusing facial micro-expression recognition algorithm, voice tremor frequency detection algorithm, and writing pressure sensor is as follows:

[0036] A facial micro-expression motion recognition algorithm is used to extract parameters such as the distance of eyebrow position change, the degree of eye muscle contraction, and the amplitude of mouth corner displacement from the student facial image sequence captured by the camera frame by frame. By calculating the numerical difference between the distance of eyebrow position change, the degree of eye muscle contraction, and the amplitude of mouth corner displacement parameters and the preset anxious facial micro-expression template, the facial micro-expression anxiety quantification value is obtained.

[0037] A speech signal frequency domain analysis algorithm is used to perform a fast Fourier transform on the student speech data collected in real time by the microphone, and the harmonic frequency shift and the speech fundamental frequency flutter period parameter are extracted from the speech signal. By calculating the difference between the harmonic frequency shift and the speech fundamental frequency flutter period parameter and the preset speech tremor frequency threshold, the speech tremor frequency anxiety quantification value is obtained.

[0038] The writing pen pressure sensor is used to record the changes in students' writing pressure in real time. The pressure change rate and pressure peak parameters per unit time are extracted. By calculating the difference between the parameters and the preset normal writing pressure range, the writing pressure change anxiety quantification value is obtained.

[0039] The anxiety quantification values ​​of facial micro-expression anxiety, voice tremor frequency anxiety, and writing pressure change anxiety are assigned corresponding weight coefficients, and the weighted sum of the three anxiety quantification values ​​is calculated by a linear weighted fusion method to obtain the initial comprehensive quantification value of anxiety state.

[0040] The initial comprehensive quantitative value of the anxiety state is subjected to min-max normalization to obtain a student anxiety state quantitative index with a uniform numerical range.

[0041] Optionally, based on the differences in English proficiency and the quantitative index of student anxiety, a personalized learning path is generated. When the anxiety index exceeds a threshold, a psychological guidance dialogue is triggered through a virtual interactive tutor, and mindfulness audio is pushed for psychological intervention. Specifically:

[0042] The weak points in students' English learning and the adjustment range of learning intensity are determined based on the magnitude and positive / negative direction of the differences in English proficiency status.

[0043] The student's current psychological state level is determined based on the numerical range of the student anxiety state quantitative index;

[0044] Using the weak points in English learning, the adjustment range of learning intensity, and the student's current psychological state level as input conditions, the system calls the pre-constructed bidirectional adaptive association rule between psychological and learning states to obtain an initial personalized learning path scheme that matches the current learning and psychological state.

[0045] The system monitors the quantitative index of students' anxiety status in real time and compares it with the preset anxiety threshold. When the quantitative index of students' anxiety status exceeds the preset anxiety threshold, the virtual interactive tutor is triggered to initiate a mental health guidance dialogue.

[0046] Based on the differences in students' current anxiety level and English proficiency, corresponding mindfulness audio content is matched and synchronously pushed from a pre-set mindfulness audio content library to provide psychological intervention for students.

[0047] Through the real-time intervention of the virtual interactive tutor's psychological health guidance dialogue and mindfulness audio content, the quantitative index of students' anxiety state is reassessed in real time, the initial plan of personalized learning path is updated, and a personalized learning path with real-time dynamic adjustment is obtained.

[0048] Optionally, the bidirectional adaptive association rule between psychological and learning states is specifically as follows:

[0049] Define the English proficiency status difference level as ,in:

[0050] ;

[0051] Define the quantitative index level of student anxiety as ,in:

[0052] ;

[0053] The intensity of the learning task in the personalized learning path is defined as follows: Psychological intervention frequency is The strength level satisfies: The intervention frequency level meets the following requirements: ,in This refers to high-intensity learning tasks. This refers to a moderate level of learning task intensity. This refers to low-intensity learning tasks. This refers to a high frequency of psychological intervention. This refers to the frequency of psychological interventions. This refers to a low frequency of psychological intervention;

[0054] The mapping relationship between the bidirectional adaptive association rule of psychological and learning states is as follows:

[0055] .

[0056] Optionally, the method based on a bidirectional cultural knowledge graph optimizes the ratio of Western cultural knowledge input to Chinese cultural knowledge output using a Monte Carlo tree search algorithm to dynamically generate a cross-cultural learning scheme, specifically as follows:

[0057] Based on the specific requirements of the new national compulsory education English curriculum standards, a two-way integrated cultural knowledge graph is constructed, which includes knowledge nodes of Chinese culture, knowledge nodes of Western culture, and the mapping relationship between the two.

[0058] Real-time data collection of students' interest in Chinese cultural knowledge and cross-cultural adaptability evaluation data generated during the learning process of cultural knowledge, to determine quantitative parameters of students' interest in Chinese cultural knowledge and cross-cultural adaptability.

[0059] Using the quantitative parameters of Chinese cultural knowledge interest and the quantitative parameters of cross-cultural adaptability as the core input conditions, the Monte Carlo tree search algorithm is called to traverse the cultural knowledge nodes in the bidirectional cultural fusion knowledge graph.

[0060] Based on the results of the Monte Carlo tree search algorithm traversal, the expected learning benefits of different cultural knowledge nodes are calculated and determined.

[0061] Based on the magnitude of the expected learning benefits, the ratio of Western cultural knowledge input nodes to Chinese cultural knowledge output nodes in the student's learning path is dynamically adjusted in real time.

[0062] Based on the adjusted node ratio, personalized cross-cultural learning plans for students are updated and dynamically generated.

[0063] Optionally, the method of real-time feedback of English proficiency status differences, anxiety status quantification index, and cross-cultural learning program execution effect data through cross-terminal data transmission, and real-time optimization and updating of learning paths and psychological intervention programs based on federated learning algorithms, serves as the basis for the next multimodal learning status characteristic data analysis and path generation, specifically as follows:

[0064] The system collects English proficiency assessment data, anxiety state quantification index, and cross-cultural integration learning program implementation effect data in real time through student learning terminals, and uses a differential privacy mechanism to process local differential privacy perturbations.

[0065] The data, after being processed by local differential privacy perturbation, is transmitted to the federated learning server in real time with encryption.

[0066] The federated learning server performs secure aggregation calculations on data from multiple learning terminals without decrypting the original data, thereby obtaining a globally unified encrypted training data set.

[0067] Using the globally unified encrypted data training set as input, the parameters of the multimodal education big model used to analyze the multimodal learning state feature data of students are updated in real time based on the federated average aggregation algorithm.

[0068] The updated parameters of the multimodal education model are securely distributed to each student's learning terminal in real time through a cross-terminal data transmission mechanism, and the real-time update of the terminal model parameters is automatically completed.

[0069] Each student's learning terminal uses the updated multimodal education big data model parameters to analyze the currently collected multimodal learning status characteristics data in real time, and automatically and dynamically generates new personalized learning paths and psychological intervention plans based on the analysis results.

[0070] Optional, a primary and secondary school English AI learning companion system based on a multimodal education model includes:

[0071] The learning terminal is a smart stylus that integrates a camera, microphone, and built-in pressure sensor to collect and fuse multimodal learning state feature data in real time.

[0072] The parameter generation module generates a multi-dimensional quantitative parameter system for English speaking speed, vocabulary frequency, and the proportion of cultural knowledge based on the national English curriculum standards.

[0073] The competency assessment module matches multimodal learning state feature data with multidimensional quantitative parameters item by item to determine the difference value of English proficiency status;

[0074] The anxiety state module uses a combination of facial micro-expression recognition, voice tremor frequency detection, and writing stress detection to generate a quantitative index of anxiety state.

[0075] The path generation and psychological intervention module dynamically generates personalized learning paths and psychological intervention plans based on the differences in English proficiency and the quantitative index of anxiety, and invokes the rules for the association between psychological and learning states.

[0076] The federated learning optimization module optimizes learning paths and psychological intervention programs by updating the parameters of the multimodal education big model in real time through cross-terminal differential privacy data transmission and federated learning algorithms.

[0077] The beneficial effects of this invention are:

[0078] This invention collects and integrates multimodal data such as students' facial expressions, voice, and writing stress in real time to form characteristic data of students' learning status. It constructs a multi-dimensional quantitative parameter system that is precisely adapted to the national English curriculum standards, enabling real-time and accurate assessment and feedback of English proficiency. This effectively solves the problem of disconnect between grading standards and existing technologies, and significantly improves the fit between learning content and domestic curriculum goals. Simultaneously, it generates a quantitative index of anxiety state through facial micro-expression recognition, voice tremor frequency detection, and writing stress change analysis, and dynamically generates personalized psychological intervention plans in real time. This effectively solves the problem of existing technologies lacking real-time attention and intervention for mental health, and significantly alleviates students' anxiety during the learning process. Furthermore, by constructing a bidirectional cultural knowledge graph and using the Monte Carlo tree search algorithm to optimize the ratio of Chinese and Western cultural knowledge content, it achieves effective integration and optimization of cross-cultural learning, breaking through the limitations of weak cultural content in existing technologies, and significantly improving students' cross-cultural expression ability and comprehensive literacy. Attached Figure Description

[0079] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0080] Figure 1This is a flowchart of the AI-powered English learning companion method for primary and secondary schools based on a multimodal education model proposed in this invention. Detailed Implementation

[0081] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0082] refer to Figure 1 A method for creating an AI-powered English learning companion for primary and secondary school students based on a multimodal education model includes:

[0083] The system collects students' facial expression data through cameras, voice data through microphones, and writing trajectory and writing pressure data through smart styluses, and integrates them into multimodal learning state feature data.

[0084] Based on the new national compulsory education English curriculum standards, a multi-dimensional quantitative parameter system was constructed for English speaking speed, vocabulary usage frequency, and the proportion of cultural knowledge.

[0085] By matching multimodal learning state feature data with a multidimensional quantitative parameter system, the difference between students' English proficiency status and the English proficiency status required by the new curriculum standards can be determined.

[0086] By integrating facial micro-expression recognition algorithm, voice tremor frequency detection algorithm, and writing pressure sensor, a quantitative index of student anxiety state is generated.

[0087] Based on the differences in English proficiency and the quantitative index of students' anxiety, a personalized learning path is generated. When the anxiety index exceeds the threshold, a psychological guidance dialogue is triggered by a virtual interactive tutor and mindfulness audio is pushed for psychological intervention.

[0088] Using data on students’ interest in Chinese culture and their cross-cultural adaptability as input, and based on a bidirectional cultural knowledge graph, the Monte Carlo tree search algorithm is used to optimize the ratio of Western cultural knowledge input to Chinese cultural knowledge output, and dynamically generate cross-cultural learning programs.

[0089] By transmitting data across terminals in real time, feedback is provided on differences in English proficiency, a quantitative index of anxiety, and data on the effectiveness of cross-cultural learning programs. Based on federated learning algorithms, learning paths and psychological intervention programs are optimized and updated in real time, serving as the basis for the next analysis of multimodal learning status characteristics and path generation.

[0090] By collecting and integrating multimodal data such as students' facial expressions, voice, and writing stress in real time, a multi-dimensional quantitative parameter system conforming to the new national compulsory education English curriculum standards was constructed, achieving precise matching of English proficiency status. By generating a quantitative index of anxiety status in real time and dynamically triggering virtual tutor psychological intervention, students' anxiety was effectively alleviated. The cross-cultural integration learning program was optimized using the Monte Carlo tree search algorithm, significantly improving students' cross-cultural communication skills. The federated learning algorithm was used to optimize model parameters in real time, improving the personalization and adaptation accuracy of the learning path.

[0091] In this embodiment, the process of collecting students' facial expression data via camera, voice data via microphone, and writing trajectory and pressure data via smart stylus, and then fusing them into multimodal learning state feature data, specifically involves:

[0092] The student's facial expression data collected in real time by the camera is processed frame by frame for noise reduction, and the positions of key facial points and the amplitude of movements are extracted.

[0093] Time-domain framing and spectral feature analysis were performed on the student speech data collected in real time by the microphone to extract the fundamental frequency, harmonic frequency and formant position parameters of the speech.

[0094] The writing trajectory data of students collected in real time by the smart stylus is processed by two-dimensional coordinate mapping to obtain the writing trajectory coordinate parameters, and the pressure value parameters of the writing pressure data are recorded simultaneously.

[0095] Using facial key point location and movement amplitude parameters, fundamental frequency, harmonic frequency and formant location parameters of speech, writing trajectory coordinate parameters and pressure numerical parameters as initial inputs, principal component analysis algorithm is used to determine the weighting factors of each parameter respectively;

[0096] The weight factors of each parameter are multiplied by the quantized values ​​of the corresponding parameters and then summed to obtain the fused multimodal learning state feature data.

[0097] By using cameras, microphones, and smart styluses to collect and integrate students' facial key point positions and movement amplitude parameters, speech spectrum feature parameters, writing trajectory coordinates, and pressure value parameters in real time, and using principal component analysis algorithm to determine the weight of each parameter, the effective integration of multimodal learning state feature data is achieved, improving the accuracy and reliability of learning state assessment.

[0098] In this embodiment, the construction of a multi-dimensional quantitative parameter system for English speaking speed, vocabulary usage frequency, and cultural knowledge proportion specifically includes:

[0099] Based on the English oral proficiency requirements for grades 3 to 9 as stipulated in the new national compulsory education English curriculum standards, the upper and lower limits of the number of words pronounced per minute in English oral expression for each grade are determined and converted into quantitative parameters of English oral speaking speed.

[0100] Based on the specific requirements for the amount of English vocabulary mastery in each grade in the new curriculum standards, the total amount of vocabulary is clearly divided into three independent categories: basic vocabulary, extended vocabulary, and cultural vocabulary, and the frequency parameters of each type of vocabulary in the teaching content are determined.

[0101] Based on the requirement of integrating Chinese and Western cultural knowledge teaching proposed in the new curriculum standards, the teaching time of Chinese traditional cultural knowledge and Western cultural knowledge involved in classroom teaching is quantified to form parameters for the proportion of Chinese traditional cultural knowledge and Western cultural knowledge.

[0102] Based on the aforementioned English spoken speed quantification parameters, basic vocabulary usage frequency parameters, extended vocabulary usage frequency parameters, and cultural vocabulary usage frequency parameters, English teaching difficulty level parameters are obtained.

[0103] Using the parameters of the proportion of knowledge in traditional Chinese culture and the proportion of knowledge in Western culture as input conditions, we obtain the cultural sensitivity level parameters.

[0104] After standardizing and normalizing the quantitative parameters of spoken English speed, the frequency of use of various words, the difficulty level of English teaching, and the level of cultural sensitivity, a multi-dimensional quantitative parameter system is obtained.

[0105] By constructing a multi-dimensional parameter system that includes English speaking speed, frequency of basic vocabulary use, frequency of extended vocabulary and cultural vocabulary use, and the proportion of knowledge of Chinese and Western cultures, and by quantitatively integrating it with the graded requirements of the new curriculum standards and classroom teaching content, a unified expression of language ability assessment and cultural sensitivity evaluation has been achieved. This has improved the accuracy of the multimodal education model in modeling English proficiency and its ability to adapt to cross-cultural teaching.

[0106] In this embodiment, the step of matching multimodal learning state feature data with a multidimensional quantitative parameter system to determine the difference between the student's English proficiency status and the English proficiency status required by the new curriculum standards specifically involves:

[0107] The numerical difference between the actual English speaking speed parameter of students in the multimodal learning state feature data and the English speaking speed quantification parameter in the multidimensional quantification parameter system is calculated to obtain the English speaking speed difference value.

[0108] The difference between the actual vocabulary usage frequency parameter of students in the multimodal learning state feature data and the corresponding basic vocabulary usage frequency parameter, extended vocabulary usage frequency parameter and cultural vocabulary usage frequency parameter in the multidimensional quantitative parameter system is calculated to obtain the difference value of basic vocabulary usage frequency, the difference value of extended vocabulary usage frequency and the difference value of cultural vocabulary usage frequency.

[0109] The difference between the actual proportion of Chinese traditional culture knowledge and the proportion of Western culture knowledge in the multimodal learning state feature data and the parameters of the proportion of Chinese traditional culture knowledge and the proportion of Western culture knowledge in the multidimensional quantitative parameter system are calculated to obtain the difference value of the proportion of Chinese traditional culture knowledge and the difference value of the proportion of Western culture knowledge.

[0110] Using the differences in spoken English speed, frequency of use of basic vocabulary, frequency of use of extended vocabulary, frequency of use of cultural vocabulary, proportion of knowledge of traditional Chinese culture, and proportion of knowledge of Western culture as input conditions, a weighted summation method was used to calculate the comprehensive difference in English proficiency.

[0111] Based on the magnitude of the comprehensive difference value of English proficiency and the preset difference threshold, the degree of gap between students' actual English proficiency and the ability requirements stipulated in the new curriculum standards is determined, and the difference value of English proficiency is obtained.

[0112] By matching students' multimodal learning status characteristic data with a multi-dimensional quantitative parameter system item by item, and quantitatively analyzing the differences in key parameters such as English speaking speed, frequency of basic vocabulary use, frequency of extended vocabulary use, and frequency of cultural vocabulary use, the differences between students' English proficiency and curriculum standards can be accurately calculated. Simultaneously, a two-way comparative analysis of the proportion of traditional Chinese cultural knowledge and the proportion of Western cultural knowledge is introduced to further identify differences in students' cultural understanding and expression abilities. Based on this, a weighted summation of the comprehensive difference value is calculated, combined with a tiered judgment model, to effectively quantify the degree to which students' English proficiency deviates from the curriculum standards, enabling personalized assessment of ability differences.

[0113] In this embodiment, the step of generating a quantitative index of student anxiety by fusing facial micro-expression recognition algorithm, voice tremor frequency detection algorithm, and writing pressure sensor is as follows:

[0114] A facial micro-expression motion recognition algorithm is used to extract parameters such as the distance of eyebrow position change, the degree of eye muscle contraction, and the amplitude of mouth corner displacement from the student facial image sequence captured by the camera frame by frame. By calculating the numerical difference between the distance of eyebrow position change, the degree of eye muscle contraction, and the amplitude of mouth corner displacement parameters and the preset anxious facial micro-expression template, the facial micro-expression anxiety quantification value is obtained.

[0115] A speech signal frequency domain analysis algorithm is used to perform a fast Fourier transform on the student speech data collected in real time by the microphone, and the harmonic frequency shift and the speech fundamental frequency flutter period parameter are extracted from the speech signal. By calculating the difference between the harmonic frequency shift and the speech fundamental frequency flutter period parameter and the preset speech tremor frequency threshold, the speech tremor frequency anxiety quantification value is obtained.

[0116] The writing pen pressure sensor is used to record the changes in students' writing pressure in real time. The pressure change rate and pressure peak parameters per unit time are extracted. By calculating the difference between the parameters and the preset normal writing pressure range, the writing pressure change anxiety quantification value is obtained.

[0117] The anxiety quantification values ​​of facial micro-expression anxiety, voice tremor frequency anxiety, and writing pressure change anxiety are assigned corresponding weight coefficients, and the weighted sum of the three anxiety quantification values ​​is calculated by a linear weighted fusion method to obtain the initial comprehensive quantification value of anxiety state.

[0118] The initial comprehensive quantitative value of the anxiety state is subjected to min-max normalization to obtain a student anxiety state quantitative index with a uniform numerical range.

[0119] By integrating facial micro-expression recognition algorithms, voice tremor frequency detection algorithms, and writing pressure sensors, subtle anxiety characteristics in students' facial expressions, voice signals, and writing actions can be quantified, extracting key parameters reflecting changes in students' emotional states. After weighted fusion of the anxiety quantification values ​​from each sub-dimension, an initial comprehensive quantification value of anxiety state for the current learning situation can be formed, which is then converted into a unified-dimensional anxiety state quantification index through normalization. This index not only comprehensively reflects the degree of psychological stress experienced by students during the learning process but also possesses advantages such as strong real-time performance, high comparability, and applicability to dynamic path control. It provides a reliable quantitative basis for subsequent learning strategy generation and psychological intervention mechanisms, effectively improving the system's accuracy in recognizing students' psychological states and its intervention response capability.

[0120] In this embodiment, the generation of personalized learning paths based on differences in English proficiency and a quantitative index of student anxiety, and the triggering of a psychologically guided dialogue and the delivery of mindfulness audio for psychological intervention when the anxiety index exceeds a threshold, specifically involves:

[0121] The weak points in students' English learning and the adjustment range of learning intensity are determined based on the magnitude and positive / negative direction of the differences in English proficiency status.

[0122] The student's current psychological state level is determined based on the numerical range of the student anxiety state quantitative index;

[0123] Using the weak points in English learning, the adjustment range of learning intensity, and the student's current psychological state level as input conditions, the system calls the pre-constructed bidirectional adaptive association rule between psychological and learning states to obtain an initial personalized learning path scheme that matches the current learning and psychological state.

[0124] The system monitors the quantitative index of students' anxiety status in real time and compares it with the preset anxiety threshold. When the quantitative index of students' anxiety status exceeds the preset anxiety threshold, the virtual interactive tutor is triggered to initiate a mental health guidance dialogue.

[0125] Based on the differences in students' current anxiety level and English proficiency, corresponding mindfulness audio content is matched and synchronously pushed from a pre-set mindfulness audio content library to provide psychological intervention for students.

[0126] Through the real-time intervention of the virtual interactive tutor's psychological health guidance dialogue and mindfulness audio content, the quantitative index of students' anxiety state is reassessed in real time, the initial plan of personalized learning path is updated, and a personalized learning path with real-time dynamic adjustment is obtained.

[0127] By using English proficiency differences and anxiety levels as core inputs, the system dynamically generates personalized learning paths and psychological intervention plans for students. This enables bidirectional adaptive regulation of learning weaknesses and current psychological states, enhancing the relevance and adaptability of path planning. By constructing rules linking psychological and learning states, the system can match learning task intensity and intervention frequency based on ability differences and anxiety levels, achieving a dynamic balance between task intensity and psychological burden. When anxiety exceeds a set threshold, the system automatically triggers a virtual interactive tutor to push personalized mindfulness audio content, providing timely intervention and guiding students to stabilize their emotions and optimize their learning pace. This mechanism achieves a deep integration of dynamic psychological state perception and path generation, enhancing the system's responsiveness to complex learning situations.

[0128] In this embodiment, the bidirectional adaptive association rule between psychological and learning states is specifically as follows:

[0129] Define the English proficiency status difference level as ,in:

[0130] ;

[0131] Define the quantitative index level of student anxiety as ,in:

[0132] ;

[0133] The intensity of the learning task in the personalized learning path is defined as follows: Psychological intervention frequency is The strength level satisfies: The intervention frequency level meets the following requirements: ,in This refers to high-intensity learning tasks. This refers to a moderate level of learning task intensity. This refers to low-intensity learning tasks. This refers to a high frequency of psychological intervention. This refers to the frequency of psychological interventions. This refers to a low frequency of psychological intervention;

[0134] The mapping relationship between the bidirectional adaptive association rule of psychological and learning states is as follows:

[0135] .

[0136] By defining the English proficiency status difference value level E and the anxiety status quantification index level A, and establishing a bidirectional mapping association rule (E,A)→(S,P) according to the intensity and frequency levels of SH>SM>SL and PH>PM>PL, a refined matching of learning task intensity and psychological intervention frequency can be achieved for different combinations of proficiency and anxiety. For the (EH,AH) state, it is mapped to (SL,PH), prioritizing high-frequency psychological intervention; for the (EL,AL) state, it is mapped to (SH,PL), prioritizing the reinforcement of learning tasks; for other state combinations, it is mapped to different schemes such as (SM,PH), (SM,PM), or (SH,PM), achieving a dynamic balance between learning and intervention. After acquiring the E and A values ​​in real time, this mapping mechanism can quickly determine the corresponding (S,P) combination and dynamically apply it to the generation of personalized learning paths and the scheduling of psychological interventions, effectively avoiding the coarse adaptation of a single scheme, improving the pertinence and response speed of learning paths and psychological intervention schemes, enhancing the system's adaptability and stability to diverse learning states, and effectively promoting the synchronous optimization of students' English proficiency and psychological state.

[0137] In this embodiment, the method of optimizing the ratio of Western cultural knowledge input to Chinese cultural knowledge output based on a bidirectional cultural knowledge graph and dynamically generating a cross-cultural learning scheme through a Monte Carlo tree search algorithm is as follows:

[0138] Based on the specific requirements of the new national compulsory education English curriculum standards, a two-way integrated cultural knowledge graph is constructed, which includes knowledge nodes of Chinese culture, knowledge nodes of Western culture, and the mapping relationship between the two.

[0139] Real-time data collection of students' interest in Chinese cultural knowledge and cross-cultural adaptability evaluation data generated during the learning process of cultural knowledge, to determine quantitative parameters of students' interest in Chinese cultural knowledge and cross-cultural adaptability.

[0140] Using the quantitative parameters of Chinese cultural knowledge interest and the quantitative parameters of cross-cultural adaptability as the core input conditions, the Monte Carlo tree search algorithm is called to traverse the cultural knowledge nodes in the bidirectional cultural fusion knowledge graph.

[0141] Based on the results of the Monte Carlo tree search algorithm traversal, the expected learning benefits of different cultural knowledge nodes are calculated and determined.

[0142] Based on the magnitude of the expected learning benefits, the ratio of Western cultural knowledge input nodes to Chinese cultural knowledge output nodes in the student's learning path is dynamically adjusted in real time.

[0143] Based on the adjusted node ratio, personalized cross-cultural learning plans for students are updated and dynamically generated.

[0144] By constructing a bidirectional cultural knowledge graph that integrates Chinese and Western cultural knowledge nodes, and using real-time quantitative parameters of students' interest in Chinese cultural knowledge and their cross-cultural adaptability as core input conditions, the method employs a Monte Carlo tree search algorithm to traverse the knowledge graph nodes and calculate the expected learning benefit value of each node. This allows for the dynamic adjustment of the ratio between Western cultural knowledge input nodes and Chinese cultural knowledge output nodes in the learning path, thereby generating personalized cross-cultural learning plans for students in real time. This method overcomes the limitations of static cultural training in existing technologies, achieving bidirectional interactive optimization and adjustment of cross-cultural knowledge content, ensuring a high degree of match between learning plans and students' interests and adaptability. Simultaneously, the dynamic feedback mechanism based on expected learning benefit values ​​effectively enhances the cultural relevance and learning efficiency of the teaching content, strengthening students' cross-cultural expression abilities and comprehensive literacy.

[0145] In this embodiment, the real-time feedback of English proficiency status differences, anxiety state quantification index, and cross-cultural learning program execution effect data through cross-terminal data transmission, and the real-time optimization and updating of learning paths and psychological intervention programs based on federated learning algorithms, serve as the basis for the next multimodal learning status characteristic data analysis and path generation. Specifically:

[0146] The system collects English proficiency assessment data, anxiety state quantification index, and cross-cultural integration learning program implementation effect data in real time through student learning terminals, and uses a differential privacy mechanism to process local differential privacy perturbations.

[0147] The data, after being processed by local differential privacy perturbation, is transmitted to the federated learning server in real time with encryption.

[0148] The federated learning server performs secure aggregation calculations on data from multiple learning terminals without decrypting the original data, thereby obtaining a globally unified encrypted training data set.

[0149] Using the globally unified encrypted data training set as input, the parameters of the multimodal education big model used to analyze the multimodal learning state feature data of students are updated in real time based on the federated average aggregation algorithm.

[0150] The updated parameters of the multimodal education model are securely distributed to each student's learning terminal in real time through a cross-terminal data transmission mechanism, and the real-time update of the terminal model parameters is automatically completed.

[0151] Each student's learning terminal uses the updated multimodal education big data model parameters to analyze the currently collected multimodal learning status characteristics data in real time, and automatically and dynamically generates new personalized learning paths and psychological intervention plans based on the analysis results.

[0152] By synchronously collecting data on differences in English proficiency, a quantitative index of anxiety, and the effectiveness of cross-cultural integrated learning programs across different terminals, and combining this with a local differential privacy mechanism to perturb the data, student privacy is effectively protected. After the federated learning server performs secure aggregation calculations on the encrypted training set and updates the parameters of the multimodal education model in real time based on the federated averaging algorithm, the updated model parameters can be securely distributed to each student's learning terminal, achieving automatic real-time updates of the terminal model. Each learning terminal uses the updated model parameters to reanalyze the multimodal learning status characteristic data and dynamically generates new personalized learning paths and psychological intervention plans, ensuring that the paths and intervention plans always match the latest learning situation. Under the premise of ensuring data security and privacy, this mechanism achieves a dual improvement in model accuracy and path adaptability through continuous optimization of federated learning, significantly enhancing the system's real-time response capability and personalized service effectiveness.

[0153] In this embodiment, the primary and secondary school English AI learning companion system based on a multimodal education model includes:

[0154] The learning terminal is a smart stylus that integrates a camera, microphone, and built-in pressure sensor to collect and fuse multimodal learning state feature data in real time.

[0155] The parameter generation module generates a multi-dimensional quantitative parameter system for English speaking speed, vocabulary frequency, and the proportion of cultural knowledge based on the national English curriculum standards.

[0156] The competency assessment module matches multimodal learning state feature data with multidimensional quantitative parameters item by item to determine the difference value of English proficiency status;

[0157] The anxiety state module uses a combination of facial micro-expression recognition, voice tremor frequency detection, and writing stress detection to generate a quantitative index of anxiety state.

[0158] The path generation and psychological intervention module dynamically generates personalized learning paths and psychological intervention plans based on the differences in English proficiency and the quantitative index of anxiety, and invokes the rules for the association between psychological and learning states.

[0159] The federated learning optimization module optimizes learning paths and psychological intervention programs by updating the parameters of the multimodal education big model in real time through cross-terminal differential privacy data transmission and federated learning algorithms.

[0160] The learning terminal uses a camera, microphone, and a smart stylus with a built-in pressure sensor to collect and integrate facial expressions, voice, and writing pressure data in real time, forming multimodal learning state feature data to achieve high-precision capture of students' learning behaviors. The parameter generation module generates a multi-dimensional quantitative parameter system based on the national English curriculum standards, including English speaking speed, vocabulary usage frequency, and the proportion of cultural knowledge, effectively improving the matching accuracy between assessment and textbook content. The ability assessment module matches the multimodal learning state feature data with the quantitative parameters item by item to calculate the difference value of English ability status, and the degree of quantification refines the judgment of ability deviations. The anxiety state module integrates facial micro-expression recognition, The voice tremor frequency detection and writing stress detection modules generate an anxiety state quantification index, enabling real-time quantification of learning psychological state. The path generation and psychological intervention module dynamically generates personalized learning paths and psychological intervention plans based on the difference value of English proficiency and the anxiety state quantification index, calling the association rules between psychological and learning states, ensuring synchronous adaptation between learning tasks and psychological interventions. The federated learning optimization module updates the parameters of the large model in real time through cross-terminal differential privacy data transmission and federated averaging algorithm, and securely distributes the updated parameters to the learning terminal, realizing continuous optimization of the model and path, significantly enhancing the system's real-time response capability and personalized service effect.

Claims

1. A method for providing AI-powered English learning companions for primary and secondary schools based on a multimodal education model, characterized in that: include: The system collects students' facial expression data through cameras, voice data through microphones, and writing trajectory and writing pressure data through smart styluses, and integrates them into multimodal learning state feature data. Based on the new national compulsory education English curriculum standards, a multi-dimensional quantitative parameter system was constructed for English speaking speed, vocabulary usage frequency, and the proportion of cultural knowledge. By matching multimodal learning state feature data with a multidimensional quantitative parameter system, the difference between students' English proficiency status and the English proficiency status required by the new curriculum standards can be determined. By integrating facial micro-expression recognition algorithm, voice tremor frequency detection algorithm, and writing pressure sensor, a quantitative index of student anxiety state is generated. Based on the differences in English proficiency and the quantitative index of students' anxiety, a personalized learning path is generated. When the anxiety index exceeds the threshold, a psychological guidance dialogue is triggered by a virtual interactive tutor and mindfulness audio is pushed for psychological intervention. Using data on students’ interest in Chinese culture and their cross-cultural adaptability as input, and based on a bidirectional cultural knowledge graph, the Monte Carlo tree search algorithm is used to optimize the ratio of Western cultural knowledge input to Chinese cultural knowledge output, and dynamically generate cross-cultural learning programs. By transmitting data across terminals in real time, feedback is provided on differences in English proficiency, a quantitative index of anxiety, and data on the effectiveness of cross-cultural learning programs. Based on federated learning algorithms, learning paths and psychological intervention programs are optimized and updated in real time, serving as the basis for the next analysis of multimodal learning status characteristics and path generation.

2. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 1, characterized in that, The process involves collecting students' facial expression data via camera, voice data via microphone, and writing trajectory and pressure data via smart stylus, and then fusing these data into multimodal learning state feature data. Specifically: The student's facial expression data collected in real time by the camera is processed frame by frame for noise reduction, and the positions of key facial points and the amplitude of movements are extracted. Time-domain framing and spectral feature analysis were performed on the student speech data collected in real time by the microphone to extract the fundamental frequency, harmonic frequency and formant position parameters of the speech. The writing trajectory data of students collected in real time by the smart stylus is processed by two-dimensional coordinate mapping to obtain the writing trajectory coordinate parameters, and the pressure value parameters of the writing pressure data are recorded simultaneously. Using facial key point location and movement amplitude parameters, fundamental frequency, harmonic frequency and formant location parameters of speech, writing trajectory coordinate parameters and pressure numerical parameters as initial inputs, principal component analysis algorithm is used to determine the weighting factors of each parameter respectively; The weight factors of each parameter are multiplied by the quantized values ​​of the corresponding parameters and then summed to obtain the fused multimodal learning state feature data.

3. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 1, characterized in that, The proposed multi-dimensional quantitative parameter system for English speaking speed, vocabulary frequency, and the proportion of cultural knowledge is as follows: Based on the English oral proficiency requirements for grades 3 to 9 as stipulated in the new national compulsory education English curriculum standards, the upper and lower limits of the number of words pronounced per minute in English oral expression for each grade are determined and converted into quantitative parameters of English oral speaking speed. Based on the specific requirements for the amount of English vocabulary mastery in each grade in the new curriculum standards, the total amount of vocabulary is clearly divided into three independent categories: basic vocabulary, extended vocabulary, and cultural vocabulary, and the frequency parameters of each type of vocabulary in the teaching content are determined. Based on the requirement of integrating Chinese and Western cultural knowledge teaching proposed in the new curriculum standards, the teaching time of Chinese traditional cultural knowledge and Western cultural knowledge involved in classroom teaching is quantified to form parameters for the proportion of Chinese traditional cultural knowledge and Western cultural knowledge. Based on the aforementioned English spoken speed quantification parameters, basic vocabulary usage frequency parameters, extended vocabulary usage frequency parameters, and cultural vocabulary usage frequency parameters, English teaching difficulty level parameters are obtained. Using the parameters of the proportion of knowledge in traditional Chinese culture and the proportion of knowledge in Western culture as input conditions, we obtain the cultural sensitivity level parameters. After standardizing and normalizing the quantitative parameters of spoken English speed, the frequency of use of various words, the difficulty level of English teaching, and the level of cultural sensitivity, a multi-dimensional quantitative parameter system is obtained.

4. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 1, characterized in that, The process of matching multimodal learning state feature data with a multidimensional quantitative parameter system to determine the difference between students' English proficiency status and the English proficiency status required by the new curriculum standards specifically involves: The numerical difference between the actual English speaking speed parameter of students in the multimodal learning state feature data and the English speaking speed quantification parameter in the multidimensional quantification parameter system is calculated to obtain the English speaking speed difference value. The difference between the actual vocabulary usage frequency parameter of students in the multimodal learning state feature data and the corresponding basic vocabulary usage frequency parameter, extended vocabulary usage frequency parameter and cultural vocabulary usage frequency parameter in the multidimensional quantitative parameter system is calculated to obtain the difference value of basic vocabulary usage frequency, the difference value of extended vocabulary usage frequency and the difference value of cultural vocabulary usage frequency. The difference between the actual proportion of Chinese traditional culture knowledge and the proportion of Western culture knowledge in the multimodal learning state feature data and the parameters of the proportion of Chinese traditional culture knowledge and the proportion of Western culture knowledge in the multidimensional quantitative parameter system are calculated to obtain the difference value of the proportion of Chinese traditional culture knowledge and the difference value of the proportion of Western culture knowledge. Using the differences in spoken English speed, frequency of use of basic vocabulary, frequency of use of extended vocabulary, frequency of use of cultural vocabulary, proportion of knowledge of traditional Chinese culture, and proportion of knowledge of Western culture as input conditions, a weighted summation method was used to calculate the comprehensive difference in English proficiency. Based on the magnitude of the comprehensive difference value of English proficiency and the preset difference threshold, the degree of gap between students' actual English proficiency and the ability requirements stipulated in the new curriculum standards is determined, and the difference value of English proficiency is obtained.

5. The method for AI-powered English learning companions in primary and secondary schools based on a multimodal education model according to claim 1, characterized in that, The process involves fusing facial micro-expression recognition algorithms, voice tremor frequency detection algorithms, and writing pressure sensors to generate a quantitative index of student anxiety. Specifically: A facial micro-expression motion recognition algorithm is used to extract parameters such as the distance of eyebrow position change, the degree of eye muscle contraction, and the amplitude of mouth corner displacement from the student facial image sequence captured by the camera frame by frame. By calculating the numerical difference between the distance of eyebrow position change, the degree of eye muscle contraction, and the amplitude of mouth corner displacement parameters and the preset anxious facial micro-expression template, the facial micro-expression anxiety quantification value is obtained. A speech signal frequency domain analysis algorithm is used to perform a fast Fourier transform on the student speech data collected in real time by the microphone, and the harmonic frequency shift and the speech fundamental frequency flutter period parameter are extracted from the speech signal. By calculating the difference between the harmonic frequency shift and the speech fundamental frequency flutter period parameter and the preset speech tremor frequency threshold, the speech tremor frequency anxiety quantification value is obtained. The writing pen pressure sensor is used to record the changes in students' writing pressure in real time. The pressure change rate and pressure peak parameters are extracted per unit time. By calculating the difference between the pressure change rate and pressure peak parameters and the preset normal writing pressure range, the writing pressure change anxiety value is obtained. The anxiety quantification values ​​of facial micro-expression anxiety, voice tremor frequency anxiety, and writing pressure change anxiety are assigned corresponding weight coefficients, and the weighted sum of the three anxiety quantification values ​​is calculated by a linear weighted fusion method to obtain the initial comprehensive quantification value of anxiety state. The initial comprehensive quantitative value of the anxiety state is subjected to min-max normalization to obtain a student anxiety state quantitative index with a uniform numerical range.

6. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 1, characterized in that, The method generates personalized learning paths based on differences in English proficiency and a quantitative index of student anxiety. When the anxiety index exceeds a threshold, a virtual interactive tutor triggers a psychologically guided dialogue and pushes mindfulness audio for psychological intervention. Specifically: The weak points in students' English learning and the adjustment range of learning intensity are determined based on the magnitude and positive / negative direction of the differences in English proficiency status. The student's current psychological state level is determined based on the numerical range of the student anxiety state quantitative index; Using the weak points in English learning, the adjustment range of learning intensity, and the student's current psychological state level as input conditions, the system calls the pre-constructed bidirectional adaptive association rule between psychological and learning states to obtain an initial personalized learning path scheme that matches the current learning and psychological state. The system monitors the quantitative index of students' anxiety status in real time and compares it with the preset anxiety threshold. When the quantitative index of students' anxiety status exceeds the preset anxiety threshold, the virtual interactive tutor is triggered to initiate a mental health guidance dialogue. Based on the differences in students' current anxiety level and English proficiency, corresponding mindfulness audio content is matched and synchronously pushed from a pre-set mindfulness audio content library to provide psychological intervention for students. Through the real-time intervention of the virtual interactive tutor's psychological health guidance dialogue and mindfulness audio content, the quantitative index of students' anxiety state is reassessed in real time, the initial plan of personalized learning path is updated, and a personalized learning path with real-time dynamic adjustment is obtained.

7. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 6, characterized in that, The bidirectional adaptive association rule between psychological and learning states is as follows: The English proficiency status difference value level is defined as wherein: ; The student anxiety state quantification index grade is defined as wherein: ; The intensity of the learning task in the personalized learning path is defined as follows: The frequency of psychological intervention is The strength level satisfies: The intervention frequency level meets the following requirements: ,in This refers to high-intensity learning tasks. This refers to a moderate level of learning task intensity. This refers to low-intensity learning tasks. This refers to a high frequency of psychological intervention. This refers to the frequency of psychological interventions. This refers to the low frequency of psychological intervention; The mapping relationship between the bidirectional adaptive association rule of psychological and learning states is as follows: 。 8. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 1, characterized in that, The aforementioned knowledge graph based on bidirectional cultural fusion optimizes the ratio of Western cultural knowledge input to Chinese cultural knowledge output using the Monte Carlo tree search algorithm, dynamically generating cross-cultural learning solutions, specifically: Based on the specific requirements of the new national compulsory education English curriculum standards, a two-way integrated cultural knowledge graph is constructed, which includes knowledge nodes of Chinese culture, knowledge nodes of Western culture, and the mapping relationship between the two. Real-time data collection of students' interest in Chinese cultural knowledge and cross-cultural adaptability evaluation data generated during the learning process of cultural knowledge, to determine quantitative parameters of students' interest in Chinese cultural knowledge and cross-cultural adaptability. Using the quantitative parameters of Chinese cultural knowledge interest and the quantitative parameters of cross-cultural adaptability as the core input conditions, the Monte Carlo tree search algorithm is called to traverse the cultural knowledge nodes in the bidirectional cultural fusion knowledge graph. Based on the results of the Monte Carlo tree search algorithm traversal, the expected learning benefits of different cultural knowledge nodes are calculated and determined. Based on the magnitude of the expected learning benefits, the ratio of Western cultural knowledge input nodes to Chinese cultural knowledge output nodes in the student's learning path is dynamically adjusted in real time. Based on the adjusted node ratio, personalized cross-cultural learning plans for students are updated and dynamically generated.

9. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 1, characterized in that, The method involves real-time feedback of English proficiency status differences, anxiety state quantification index, and cross-cultural learning program implementation effect data through cross-terminal data transmission. Based on a federated learning algorithm, the learning path and psychological intervention program are optimized and updated in real time, serving as the basis for the next multimodal learning status characteristic data analysis and path generation. Specifically: The system collects English proficiency assessment data, anxiety state quantification index, and cross-cultural integration learning program implementation effect data in real time through student learning terminals, and uses a differential privacy mechanism to process local differential privacy perturbations. The data, after being processed by local differential privacy perturbation, is transmitted to the federated learning server in real time with encryption. The federated learning server performs secure aggregation calculations on data from multiple learning terminals without decrypting the original data, thereby obtaining a globally unified encrypted training data set. Using the globally unified encrypted data training set as input, the parameters of the multimodal education big model used to analyze the multimodal learning state feature data of students are updated in real time based on the federated average aggregation algorithm. The updated parameters of the multimodal education model are securely distributed to each student's learning terminal in real time through a cross-terminal data transmission mechanism, and the real-time update of the terminal model parameters is automatically completed. Each student's learning terminal uses the updated multimodal education big data model parameters to analyze the currently collected multimodal learning status characteristics data in real time, and automatically and dynamically generates new personalized learning paths and psychological intervention plans based on the analysis results.

10. A primary and secondary school English AI study companion system based on a multi-modal education large model, which executes the primary and secondary school English AI study companion method based on a multi-modal education large model according to any one of claims 1 to 9, characterized in that, include: The learning terminal is an intelligent stylus that integrates a camera, microphone, and built-in pressure sensor to collect and fuse multimodal learning state feature data in real time. The parameter generation module generates a multi-dimensional quantitative parameter system for English speaking speed, vocabulary frequency, and the proportion of cultural knowledge based on the national English curriculum standards. The competency assessment module matches multimodal learning state feature data with multidimensional quantitative parameters item by item to determine the difference value of English proficiency status; The anxiety state module uses a combination of facial micro-expression recognition, voice tremor frequency detection, and writing stress detection to generate a quantitative index of anxiety state. The path generation and psychological intervention module dynamically generates personalized learning paths and psychological intervention plans based on the differences in English proficiency and the quantitative index of anxiety, and invokes the rules for the association between psychological and learning states. The federated learning optimization module optimizes learning paths and psychological intervention programs by updating the parameters of the multimodal education big model in real time through cross-terminal differential privacy data transmission and federated learning algorithms.

Citation Information

Patent Citations

  • Processing method and device for classroom teaching behavior evaluation and storage medium

    CN115358615A

  • Deep academic learning intelligence and deep neural language network system and interfaces

    US20180247549A1