Middle and primary school English AI learning companion system and method based on multi-modal education large model

By collecting and quantitatively evaluating multimodal data, and combining cultural bidirectional fusion knowledge graphs and federated learning algorithms, the shortcomings of primary and secondary school English learning systems in adapting to new curriculum standards and mental health interventions have been addressed, thereby improving English proficiency and intercultural literacy.

CN120997002AActive Publication Date: 2025-11-21SHENZHEN TONGMOJI ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511088044.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-21
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

Existing English learning systems for primary and secondary schools fail to accurately adapt to the new national curriculum standards, lack real-time attention and intervention to students' psychological state, and have a one-dimensional design of cultural education content, failing to effectively improve cross-cultural literacy.

Method used

By collecting multimodal data (facial expressions, voice, writing stress), a multi-dimensional quantitative parameter system is constructed to assess English proficiency and anxiety levels in real time, dynamically generate personalized learning paths, optimize cross-cultural learning programs through a bidirectional cultural knowledge graph, and combine federated learning algorithms to optimize learning paths and psychological interventions in real time.

Benefits of technology

It enables accurate assessment and feedback of English proficiency, significantly alleviates student anxiety, enhances cross-cultural literacy, and improves the fit between learning content and curriculum objectives, as well as the real-time nature of mental health intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997002A_ABST
    Figure CN120997002A_ABST
Patent Text Reader

Abstract

The invention discloses a middle and primary school English AI learning companion system and method based on a multi-modal education large model. The method comprises the following steps of collecting multi-modal learning state data such as facial expressions, voices and writing pressure of students in real time; constructing multi-dimensional quantitative parameters of spoken English speed, vocabulary use frequency and cultural knowledge proportion, which meet the new lesson standard requirements; calculating an English ability state difference value and an anxiety state quantitative index of the student; generating a personalized learning path and a psychological intervention scheme; utilizing a Monte Carlo tree search algorithm to dynamically optimize a cross-culture learning scheme; and learning model parameters are optimized and updated in real time through federated learning. According to the method, the national English new lesson standard requirements are accurately met, and the English ability, the mental health level and the cross-culture fusion ability of students are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent education, and particularly relates to a primary and secondary school English AI study companion system and method based on a multi-modal education large model. BACKGROUND

[0002] With the deep integration of artificial intelligence technology and the field of education, artificial intelligence assisted language learning systems have become an important development direction for primary and secondary school English teaching. At present, most English intelligent learning systems are designed and content matched using internationally recognized language learning grading standards. The design of such standards is intended to serve the globalization of English as a second language teaching, and does not fully consider the specific ability requirements and cultural education goals of students at different stages of the new curriculum standard for primary and secondary school English, resulting in the phenomenon of disconnection between these systems and domestic teaching goals in actual application.

[0003] Most existing language learning systems use students' academic performance as a single evaluation indicator, focusing on language ability training and assessment, and generally lack real-time attention and intervention measures for students' learning process psychological state, especially without effectively considering the widespread learning anxiety problem in the learning process, making it difficult to discover and alleviate students' psychological state and emotional changes in a timely manner, affecting students' overall learning experience and learning effect.

[0004] In addition, the design of cultural education content in existing language learning systems is mostly one-way input of Western cultural content, lacking an effective training mechanism for integrating Chinese culture and achieving bidirectional fusion of Chinese and Western cultural content, resulting in students' inability to effectively improve their English expression ability of Chinese culture during English learning, making it difficult for them to reflect their cultural characteristics and advantages in cross-cultural communication, and not conducive to the overall improvement of students' cross-cultural comprehensive quality.

[0005] Therefore, the current primary and secondary school English artificial intelligence learning system still has obvious deficiencies in terms of precise adaptation of teaching content to the new curriculum standard, real-time attention and intervention of learning psychology, and bidirectional cultural fusion training, and there is an urgent need to propose an intelligent learning solution that can fully and accurately adapt to China's new curriculum standard and combine multi-dimensional psychological counseling and bidirectional cultural training.

[0006] Therefore, how to provide a primary and secondary school English AI study companion system and method based on a multi-modal education large model is a problem that needs to be solved by those skilled in the art. SUMMARY

[0007] One purpose of the present application is to propose a primary and secondary school English AI study companion system and method based on a multi-modal education large model, aiming at the problem that the prior art cannot effectively adapt to the requirements of the new national English curriculum standard and lacks psychological health guidance and cultural integration training, a technical solution is proposed to collect student multi-modal data in real time, accurately quantify English ability and anxiety state, dynamically adjust personalized learning path and psychological intervention, and optimize cross-cultural integration learning scheme, the present application has the advantages of accurately adapting to the new curriculum standard, effectively relieving learning anxiety, and improving cross-cultural comprehensive quality.

[0008] According to the primary and secondary school English AI study companion method based on the multi-modal education large model, the method comprises the following steps:

[0009] The student facial expression data is collected by the camera, the student voice data is collected by the microphone, and the student writing trajectory and writing pressure data are collected by the intelligent handwriting pen, and are fused into multi-modal learning state feature data;

[0010] According to the new curriculum standard of national compulsory education English, a multi-dimensional quantitative parameter system of English oral speed, vocabulary usage frequency and cultural knowledge proportion is constructed;

[0011] The multi-modal learning state feature data is matched with the multi-dimensional quantitative parameter system to determine the difference value between the student English ability state and the English ability state required by the new curriculum standard;

[0012] The student anxiety state quantitative index is generated by fusing the facial micro-expression recognition algorithm, the voice tremor frequency detection algorithm and the writing pressure sensor;

[0013] Based on the English ability state difference value and the student anxiety state quantitative index, a personalized learning path is generated, when the anxiety index exceeds the threshold value, a virtual interactive tutor triggers a psychological guidance dialogue and pushes a mindfulness audio for psychological intervention;

[0014] The student Chinese culture interest degree and cross-cultural adaptability data are input, based on the cultural two-way fusion knowledge graph, the ratio of western culture knowledge input to Chinese culture knowledge output is optimized by the Monte Carlo tree search algorithm, and a cross-cultural learning scheme is dynamically generated;

[0015] The English ability state difference value, the anxiety state quantitative index and the cross-cultural learning scheme execution effect data are fed back in real time through cross-terminal data transmission, the learning path and the psychological intervention scheme are updated in real time based on the federated learning algorithm, and are used as the basis for the next multi-modal learning state feature data analysis and path generation.

[0016] Optionally, the multi-modal learning state feature data is constructed, specifically:

[0017] The student facial expression data collected by the camera in real time is subjected to frame-by-frame image denoising processing, and the facial key point position and action amplitude parameters are extracted;

[0018] The student speech data collected by the microphone in real time is subjected to time-domain framing and spectral feature analysis, and the fundamental frequency, harmonic frequency and formant position parameters of the speech are extracted;

[0019] The student writing trajectory data collected by the intelligent handwriting pen in real time is subjected to two-dimensional coordinate mapping processing to obtain writing trajectory coordinate parameters, and the pressure value parameters of writing pressure data are recorded synchronously;

[0020] The facial key point position and action amplitude parameters, the fundamental frequency, harmonic frequency and formant position parameters of the speech, the writing trajectory coordinate parameters and the pressure value parameters are used as initial inputs, and the principal component analysis algorithm is used to determine the weight factors of each parameter;

[0021] The weight factors of each parameter are multiplied by the corresponding quantitative values of the parameters, and then summed to obtain the fused multi-modal learning state feature data.

[0022] Optionally, the multi-dimensional quantitative parameter system of English oral English speed, vocabulary usage frequency and cultural knowledge proportion is constructed, specifically:

[0023] According to the requirements of the national compulsory education English new curriculum standard for English oral English ability of grades 3 to 9, the upper and lower limits of the number of word pronunciation in the student's English oral English expression per minute are determined year by year, and are converted into English oral English speed quantitative parameters;

[0024] According to the specific requirements of the new curriculum standard for vocabulary mastery of each grade of English learning, the total amount of vocabulary is divided into three independent categories: basic vocabulary, expanded vocabulary and cultural vocabulary, and the usage frequency parameters of each category of vocabulary in teaching content are determined;

[0025] According to the requirements of the new curriculum standard for the integration of Chinese and Western cultural knowledge in teaching, the teaching time proportion of Chinese traditional cultural knowledge and Western cultural knowledge involved in classroom teaching is quantified respectively to form Chinese traditional cultural knowledge proportion parameter and Western cultural knowledge proportion parameter;

[0026] The English teaching difficulty level parameter is obtained according to the English oral English speed quantitative parameter, the basic vocabulary usage frequency parameter, the expanded vocabulary usage frequency parameter and the cultural vocabulary usage frequency parameter respectively;

[0027] The cultural sensitivity level parameter is obtained by taking the Chinese traditional cultural knowledge proportion parameter and the Western cultural knowledge proportion parameter as input conditions;

[0028] The multi-dimensional quantification parameter system is obtained by standardizing and normalizing the English oral speech speed quantification parameter, the vocabulary usage frequency parameter, the English teaching difficulty level parameter and the cultural sensitivity level parameter.

[0029] Optionally, the multi-modal learning state feature data is matched with the multi-dimensional quantification parameter system to determine the English ability state difference value of the student, and the specific process is as follows:

[0030] The actual English oral speech speed parameter of the student in the multi-modal learning state feature data is subjected to numerical difference calculation with the English oral speech speed quantification parameter in the multi-dimensional quantification parameter system to obtain an English oral speech speed difference value.

[0031] The actual vocabulary usage frequency parameter of the student in the multi-modal learning state feature data is subjected to difference calculation with the corresponding basic vocabulary usage frequency parameter, the extended vocabulary usage frequency parameter and the cultural vocabulary usage frequency parameter in the multi-dimensional quantification parameter system to obtain a basic vocabulary usage frequency difference value, an extended vocabulary usage frequency difference value and a cultural vocabulary usage frequency difference value.

[0032] The actual Chinese traditional culture knowledge proportion and the western culture knowledge proportion of the student in the multi-modal learning state feature data are subjected to difference calculation with the Chinese traditional culture knowledge proportion parameter and the western culture knowledge proportion parameter in the multi-dimensional quantification parameter system to obtain a Chinese traditional culture knowledge proportion difference value and a western culture knowledge proportion difference value.

[0033] The English oral speech speed difference value, the basic vocabulary usage frequency difference value, the extended vocabulary usage frequency difference value, the cultural vocabulary usage frequency difference value, the Chinese traditional culture knowledge proportion difference value and the western culture knowledge proportion difference value are taken as input conditions, and a weighted summation method is used to obtain an English ability comprehensive difference value.

[0034] According to the size of the English ability comprehensive difference value and the preset ability difference threshold value, the gap degree between the actual English ability state of the student and the ability requirement of the new curriculum standard is determined to obtain an English ability state difference value.

[0035] Optionally, the student anxiety state quantification index is generated by fusing the face micro-expression recognition algorithm, the voice tremor frequency detection algorithm and the writing pressure sensor, and the specific process is as follows:

[0036] The face micro-expression action recognition algorithm is used to extract the eyebrow position change distance, the eye muscle contraction degree and the mouth corner displacement amplitude parameter in the student face image sequence collected by the camera frame by frame, and the numerical difference value between the parameters and the preset anxiety face micro-expression template is calculated to obtain a face micro-expression anxiety quantification value.

[0037] The speech signal frequency domain analysis algorithm is used to perform fast Fourier transform on the student speech data collected by the microphone in real time, the harmonic frequency offset degree and the speech fundamental frequency tremor period parameter in the speech signal are extracted, and the difference between the parameters and the preset speech tremor frequency threshold is calculated to obtain the speech tremor frequency anxiety quantitative value;

[0038] The writing pressure sensor is used to record the student writing pressure change data in real time, the pressure change rate and the pressure peak value parameters in unit time are extracted, and the difference between the parameters and the preset writing pressure normal interval is calculated to obtain the writing pressure change anxiety quantitative value;

[0039] The face micro-expression anxiety quantitative value, the speech tremor frequency anxiety quantitative value and the writing pressure change anxiety quantitative value are respectively given corresponding weight coefficients, and the weighted sum of the three anxiety quantitative values is calculated by a linear weighting fusion method to obtain an initial comprehensive quantitative value of the anxiety state;

[0040] The initial comprehensive quantitative value of the anxiety state is subjected to minimum-maximum normalization processing to obtain a student anxiety state quantitative index with a unified numerical range.

[0041] Optionally, based on the English ability state difference value and the student anxiety state quantitative index, a personalized learning path is generated, when the anxiety index exceeds a threshold value, a virtual interactive tutor triggers a psychological guidance dialogue and pushes a mindfulness audio for psychological intervention, specifically:

[0042] According to the numerical size and the positive and negative directions of the English ability state difference value, the weak link of the student English learning and the learning intensity adjustment range are determined;

[0043] According to the numerical interval of the student anxiety state quantitative index, the current psychological state grade of the student is determined;

[0044] Taking the weak link of the English learning, the learning intensity adjustment range and the current psychological state grade of the student as input conditions, a pre-constructed psychological and learning state bidirectional adaptive association rule is called to obtain an initial scheme of a personalized learning path matched with the current learning and psychological state;

[0045] The numerical size of the student anxiety state quantitative index and a preset anxiety threshold value is monitored in real time, when the student anxiety state quantitative index exceeds the preset anxiety threshold value, a virtual interactive tutor is triggered to start a psychological health guidance dialogue;

[0046] According to the current anxiety state grade of the student and the English ability state difference value, corresponding mindfulness audio content is matched and pushed from a preset mindfulness audio content library to intervene in the psychology of the student;

[0047] After the instant intervention of the mental health guidance dialogue and the mindfulness audio content of the virtual interactive mentor, the student anxiety state quantitative index is reevaluated in real time, the initial scheme of the personalized learning path is updated, and the real-time dynamically adjusted personalized learning path is obtained.

[0048] Optionally, the bidirectional adaptive association rule between the mental state and the learning state is specifically:

[0049] The English ability state difference value level is defined as E, wherein:

[0050]

[0051] The student anxiety state quantitative index level is defined as A, wherein:

[0052]

[0053] The learning task intensity of the personalized learning path is defined as S, and the psychological intervention frequency is defined as P, wherein the intensity level satisfies: S H >S M >S L , and the intervention frequency level satisfies: P H >P M >P L ;

[0054] The mapping relationship of the bidirectional adaptive association rule between the mental state and the learning state is:

[0055]

[0056] Optionally, based on the bidirectional fusion knowledge graph of culture, the proportion of western culture knowledge input and Chinese culture knowledge output is optimized by a Monte Carlo tree search algorithm, and a cross-cultural learning scheme is dynamically generated, which is specifically:

[0057] According to the specific requirements of the national compulsory education English new curriculum standard, a bidirectional fusion knowledge graph of culture is constructed, which includes Chinese culture knowledge nodes, western culture knowledge nodes and the mapping relationship between them;

[0058] Real-time collection of Chinese culture knowledge interest degree data and cross-cultural adaptability evaluation data generated by students in the process of cultural knowledge learning is performed to determine the Chinese culture knowledge interest quantitative parameter and the cross-cultural adaptability quantitative parameter of the students;

[0059] Taking the Chinese culture knowledge interest quantitative parameter and the cross-cultural adaptability quantitative parameter as the core input conditions, the Monte Carlo tree search algorithm is called to traverse the cultural knowledge nodes in the bidirectional fusion knowledge graph of culture;

[0060] According to the traversal result of the Monte Carlo tree search algorithm, the expected learning benefit values of different cultural knowledge nodes are calculated and determined.

[0061] According to the size of the expected learning benefit value, the proportion relationship of the western culture knowledge input node and the Chinese culture knowledge output node in the student learning path is dynamically adjusted in real time;

[0062] According to the adjusted node proportion relationship, the student's personalized cross-cultural learning scheme is updated and dynamically generated.

[0063] Optionally, the real-time feedback of English ability state difference value, anxiety state quantitative index and cross-cultural learning scheme execution effect data through cross-terminal data transmission is based on the federated learning algorithm to update and optimize the learning path and psychological intervention scheme in real time, which serves as the basis for the next multi-modal learning state feature data analysis and path generation, specifically:

[0064] The English ability evaluation data, anxiety state quantitative index and cross-cultural fusion learning scheme execution effect data are synchronously collected in real time by the student learning terminal, and local differential privacy perturbation processing is performed by using the differential privacy mechanism;

[0065] The data processed by local differential privacy perturbation is transmitted to the federated learning server in real time after being encrypted;

[0066] The federated learning server performs secure aggregation calculation on the data of multiple learning terminals without decrypting the original data, and obtains a globally unified encrypted data training set;

[0067] Taking the globally unified encrypted data training set as input, the multi-modal education big model parameters for analyzing student multi-modal learning state feature data are updated in real time based on the federated average aggregation algorithm;

[0068] The updated multi-modal education big model parameters are securely distributed to each student learning terminal in real time through the cross-terminal data transmission mechanism, and the real-time update of the end-side model parameters is automatically completed;

[0069] Each student learning terminal uses the updated multi-modal education big model parameters to analyze the currently collected multi-modal learning state feature data in real time, and automatically generates a new personalized learning path and psychological intervention scheme based on the analysis results.

[0070] Optionally, the primary and secondary school English AI learning companion system based on the multi-modal education big model comprises:

[0071] The learning terminal is a smart stylus integrating a camera, a microphone and a built-in pressure sensor, which is used to collect and fuse multi-modal learning state feature data in real time;

[0072] The parameter generation module generates a multi-dimensional quantitative parameter system of English oral speech speed, vocabulary frequency and cultural knowledge proportion according to the national English new curriculum standard;

[0073] a capability evaluation module, which matches the multi-modal learning state feature data with multi-dimensional quantitative parameters item by item to determine an English capability state difference value;

[0074] an anxiety state module, which adopts a facial micro-expression recognition module, a speech tremor frequency detection module and a writing pressure detection module to generate an anxiety state quantitative index;

[0075] a path generation and psychological intervention module, which dynamically generates a personalized learning path and a psychological intervention scheme according to the English capability state difference value and the anxiety state quantitative index by calling psychological and learning state association rules;

[0076] a federal learning optimization module, which updates multi-modal education big model parameters in real time through cross-terminal differential privacy data transmission and federal learning algorithm to optimize the learning path and the psychological intervention scheme.

[0077] The beneficial effects of the present application are:

[0078] The present application collects multi-modal data such as facial expressions, speech and writing pressure of students in real time and fuses them to form student learning state feature data, constructs a multi-dimensional quantitative parameter system precisely adapted to the national English new curriculum standard, realizes real-time and accurate evaluation and feedback of English capability state, effectively solves the problem of disconnection of grading standards in the prior art, and significantly improves the adaptability of learning content to domestic curriculum objectives; at the same time, the anxiety state quantitative index is formed by facial micro-expression recognition, speech tremor frequency detection and writing pressure change analysis, and a personalized psychological intervention scheme is dynamically generated in real time, effectively solving the problem of lack of real-time attention and intervention of mental health in the prior art, and significantly relieving the anxiety state of students in the learning process; in addition, by constructing a cultural bidirectional fusion knowledge graph and using a Monte Carlo tree search algorithm to optimize the proportion of Chinese and Western cultural knowledge content, effective fusion and optimization of cross-cultural learning are realized, the limitation of weak cultural connotation in the prior art is broken through, and the cross-cultural expression ability and comprehensive quality of students are significantly improved. BRIEF DESCRIPTION OF DRAWINGS

[0079] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, illustrate the present application and are used to explain the present application, and do not constitute a limitation on the present application. In the drawings:

[0080] Figure 1 A flowchart of the primary and secondary school English AI learning companion method based on the multi-modal education big model proposed by the present application. DETAILED DESCRIPTION

[0081] The application will be described in further detail below with reference to the drawings. These drawings are simplified schematic diagrams and only show the basic structure of the application in a schematic manner, and thus only show the components relevant to the application.

[0082] Reference Figure 1 A primary and secondary school English AI study companion method based on a multi-modal education large model, comprising:

[0083] Student facial expression data is collected through a camera, student voice data is collected through a microphone, and student writing trajectory and writing pressure data are collected through a smart handwriting pen, and are fused into multi-modal learning state feature data;

[0084] A multi-dimensional quantitative parameter system of English oral speech speed, vocabulary usage frequency and cultural knowledge proportion is constructed according to the national compulsory education English new curriculum standard;

[0085] The multi-modal learning state feature data is matched with the multi-dimensional quantitative parameter system to determine the difference value between the student English ability state and the English ability state required by the new curriculum standard;

[0086] Anxiety state quantitative index of the student is generated by fusing a facial micro-expression recognition algorithm, a voice tremor frequency detection algorithm and a writing pressure sensor;

[0087] Based on the English ability state difference value and the anxiety state quantitative index of the student, a personalized learning path is generated, when the anxiety index exceeds a threshold value, a virtual interactive tutor triggers a psychological guidance dialogue and pushes a mindfulness audio for psychological intervention;

[0088] With the student's Chinese culture interest degree and cross-cultural adaptability data as input, based on a culture bidirectional fusion knowledge graph, a western culture knowledge input and Chinese culture knowledge output ratio is optimized through a Monte Carlo tree search algorithm to dynamically generate a cross-cultural learning scheme;

[0089] The English ability state difference value, the anxiety state quantitative index and the cross-cultural learning scheme execution effect data are fed back in real time through cross-terminal data transmission, and the learning path and the psychological intervention scheme are updated in real time based on a federated learning algorithm, which serves as the basis for the next multi-modal learning state feature data analysis and path generation.

[0090] By real-time collection and fusion of students' facial expressions, voices, and writing pressures and other multi-modal data, a multi-dimensional quantitative parameter system conforming to the new curriculum standards of national compulsory education English is constructed, and the precise matching of English ability state is realized; by real-time generation of anxiety state quantitative index and dynamic triggering of virtual tutor psychological intervention, the anxiety state of students is effectively relieved; by Monte Carlo tree search algorithm optimization of cross-cultural fusion learning scheme, the cross-cultural communication ability of students is significantly improved; and by using federated learning algorithm to optimize model parameters in real time, the individualization and adaptation accuracy of learning path are improved.

[0091] In the embodiment, the multi-modal learning state feature data is constructed, specifically as follows:

[0092] The student facial expression data collected by the camera in real time is subjected to frame-by-frame image denoising processing, and the facial key point position and action amplitude parameters are extracted;

[0093] The student voice data collected by the microphone in real time is subjected to time-domain framing and spectral feature analysis, and the fundamental frequency, harmonic frequency and formant position parameters of the voice are extracted;

[0094] The student writing trajectory data collected by the intelligent handwriting pen in real time is subjected to two-dimensional coordinate mapping processing, and the writing trajectory coordinate parameters are obtained, and the pressure value parameters of the writing pressure data are synchronously recorded;

[0095] The facial key point position and action amplitude parameters, the fundamental frequency, harmonic frequency and formant position parameters of the voice, the writing trajectory coordinate parameters and the pressure value parameters are taken as initial inputs, and the principal component analysis algorithm is used to determine the weight factors of the respective parameters;

[0096] The weight factors of the respective parameters are multiplied by the quantitative values of the corresponding parameters respectively, and then summed, to obtain the fused multi-modal learning state feature data.

[0097] The facial key point position and action amplitude parameters, the voice spectral feature parameters, the writing trajectory coordinates and the pressure value parameters of the students are synchronously collected and fused by the camera, the microphone and the intelligent handwriting pen in real time, the principal component analysis algorithm is used to determine the weight of each parameter, and the effective fusion of the multi-modal learning state feature data is realized, and the accuracy and reliability of the learning state evaluation are improved.

[0098] In the embodiment, the multi-dimensional quantitative parameter system of English oral speech speed, vocabulary usage frequency and cultural knowledge proportion is constructed, specifically as follows:

[0099] According to the English oral ability requirements of grades 3 to 9 in the new curriculum standards of national compulsory education English, the upper and lower limits of the number of word pronunciations in the English oral expression of students per minute are determined grade by grade, and are converted into English oral speech speed quantitative parameters;

[0100] According to the specific requirements of the new curriculum standard for the amount of vocabulary mastered by each grade of English learning, the total amount of vocabulary is explicitly divided into three independent categories of basic vocabulary, expanded vocabulary and cultural vocabulary, and the usage frequency parameters of each category of vocabulary in the teaching content are determined respectively;

[0101] According to the requirement of the new curriculum standard for the integration of Chinese and Western cultural knowledge teaching, the teaching time proportion of Chinese traditional cultural knowledge and Western cultural knowledge involved in classroom teaching is quantified respectively to form the Chinese traditional cultural knowledge proportion parameter and the Western cultural knowledge proportion parameter;

[0102] According to the English oral speed quantification parameter, the basic vocabulary usage frequency parameter, the expanded vocabulary usage frequency parameter and the cultural vocabulary usage frequency parameter respectively, the English teaching difficulty level parameter is obtained;

[0103] Taking the Chinese traditional cultural knowledge proportion parameter and the Western cultural knowledge proportion parameter as input conditions, the cultural sensitivity level parameter is obtained;

[0104] After standardizing and normalizing the English oral speed quantification parameter, the usage frequency parameters of various types of vocabulary, the English teaching difficulty level parameter and the cultural sensitivity level parameter, the multi-dimensional quantification parameter system is obtained.

[0105] By constructing the multi-dimensional parameter system of English oral speed, basic vocabulary usage frequency, expanded vocabulary and cultural vocabulary usage frequency, and the proportion of Chinese and Western cultural knowledge, and combining the new curriculum standard grading requirements and classroom teaching content for quantitative integration, the unified expression of language ability evaluation and cultural sensitivity evaluation is realized, and the precision modeling effect of the multi-modal education big model on English ability state and the cross-cultural teaching adaptation ability are improved.

[0106] In the embodiment, the multi-modal learning state feature data is matched with the multi-dimensional quantification parameter system to determine the difference value between the student's English ability state and the English ability state required by the new curriculum standard, specifically:

[0107] The numerical difference value between the student's actual English oral speed parameter in the multi-modal learning state feature data and the English oral speed quantification parameter in the multi-dimensional quantification parameter system is calculated to obtain the English oral speed difference value;

[0108] The actual vocabulary usage frequency parameter of the student in the multi-modal learning state feature data is respectively calculated with the corresponding basic vocabulary usage frequency parameter, expanded vocabulary usage frequency parameter and cultural vocabulary usage frequency parameter in the multi-dimensional quantification parameter system to obtain the basic vocabulary usage frequency difference value, expanded vocabulary usage frequency difference value and cultural vocabulary usage frequency difference value;

[0109] The actual proportion of Chinese traditional culture knowledge and the proportion of western culture knowledge in the multi-modal learning state feature data of the student are respectively subtracted from the proportion of Chinese traditional culture knowledge parameter and the proportion of western culture knowledge parameter in the multi-dimensional quantification parameter system to obtain a difference value of the proportion of Chinese traditional culture knowledge and a difference value of the proportion of western culture knowledge;

[0110] The difference value of the proportion of Chinese traditional culture knowledge and the difference value of the proportion of western culture knowledge are respectively taken as input conditions, and a weighted sum method is used to calculate a comprehensive difference value of English ability.

[0111] According to the size of the comprehensive difference value of English ability and the preset ability difference threshold value, the difference between the actual English ability state of the student and the ability requirement of the new curriculum standard is determined, and an English ability state difference value is obtained.

[0112] By matching the multi-modal learning state feature data of the student with the multi-dimensional quantification parameter system, and quantitatively analyzing the difference values of the key parameters such as English oral speed, basic vocabulary usage frequency, extended vocabulary usage frequency and cultural vocabulary usage frequency, the ability difference between the English ability state of the student and the curriculum standard can be accurately calculated. At the same time, the two-way comparison and analysis of the proportion of Chinese traditional culture knowledge and the proportion of western culture knowledge are introduced, and the ability difference of the student in cultural understanding and expression is further identified. On this basis, the comprehensive difference value is calculated by weighted sum, and the degree of deviation of the English ability state of the student from the curriculum standard is quantified by combining the grading judgment model, and the individualized ability difference evaluation is realized.

[0113] In the embodiment, the student anxiety state quantization index is generated by fusing the face micro-expression recognition algorithm, the voice tremor frequency detection algorithm and the writing pressure sensor, specifically:

[0114] The face micro-expression action recognition algorithm is used to extract the eyebrow position change distance, eye muscle contraction degree and mouth corner displacement amplitude parameters in the student face image sequence collected by the camera frame by frame, and the numerical difference between the parameters and the preset face micro-expression template is calculated to obtain a face micro-expression anxiety quantization value;

[0115] The speech signal frequency domain analysis algorithm is used to perform fast Fourier transform on the student speech data collected by the microphone in real time, and the harmonic frequency offset degree and the speech fundamental frequency tremor cycle parameters in the speech signal are extracted. The difference between the parameters and the preset voice tremor frequency threshold value is calculated to obtain a voice tremor frequency anxiety quantization value;

[0116] The writing pressure sensor is used to record the writing pressure change data of the student in real time, the pressure change rate and the pressure peak value parameters in a unit time are extracted, and the difference between the parameters and a preset normal writing pressure interval is calculated to obtain a writing pressure change anxiety quantitative value;

[0117] The face micro-expression anxiety quantitative value, the voice tremor frequency anxiety quantitative value and the writing pressure change anxiety quantitative value are respectively given corresponding weight coefficients, a weighted sum of the three anxiety quantitative values is calculated by a linear weighting fusion method, and an initial comprehensive quantitative value of the anxiety state is obtained;

[0118] The initial comprehensive quantitative value of the anxiety state is subjected to minimum-maximum normalization processing to obtain a student anxiety state quantitative index with a unified numerical range.

[0119] Through the fusion use of the face micro-expression recognition algorithm, the voice tremor frequency detection algorithm and the writing pressure sensor, the subtle anxiety features of the student in the facial expression, the voice signal and the writing action can be quantified respectively, and the key change parameters reflecting the emotional state of the student can be extracted. After the weighted fusion of the sub-dimension anxiety quantitative values, an initial comprehensive quantitative value of the anxiety state for the current learning state can be formed, and through the normalization processing, a unified dimension anxiety state quantitative index can be converted. The index not only comprehensively reflects the psychological stress degree of the student in the learning process, but also has the advantages of strong real-time, high comparability and being suitable for dynamic path regulation, and provides a quantitative and reliable basis for subsequent learning strategy generation and psychological intervention mechanism, and effectively improves the recognition accuracy and intervention response ability of the system for the psychological state of the student.

[0120] In the embodiment, based on the English ability state difference value and the student anxiety state quantitative index, a personalized learning path is generated, when the anxiety index exceeds a threshold value, a psychological guidance dialogue is triggered by a virtual interactive tutor and a mindfulness audio is pushed for psychological intervention, specifically:

[0121] According to the numerical size and the positive and negative directions of the English ability state difference value, the weak link of the student's English learning and the learning intensity adjustment range are determined;

[0122] According to the numerical range of the student anxiety state quantitative index, the current psychological state grade of the student is determined;

[0123] Taking the weak link of the English learning, the learning intensity adjustment range and the current psychological state grade of the student as input conditions, a pre-constructed psychological and learning state bidirectional adaptive association rule is called to obtain an initial scheme of a personalized learning path matched with the current learning and psychological state;

[0124] The student anxiety state quantitative index is monitored in real time, and the value of the preset anxiety threshold is compared. When the student anxiety state quantitative index exceeds the preset anxiety threshold, the virtual interactive tutor is triggered to start a mental health guidance dialogue;

[0125] According to the current anxiety state level of the student and the English ability state difference value, corresponding mindfulness audio content is matched and pushed from a preset mindfulness audio content library, and the student is intervened psychologically;

[0126] After the real-time intervention of the virtual interactive tutor through the mental health guidance dialogue and the mindfulness audio content, the student anxiety state quantitative index is re-evaluated in real time, the initial scheme of the personalized learning path is updated, and a real-time dynamically adjusted personalized learning path is obtained.

[0127] By taking the English ability state difference value and the anxiety state quantitative index as the core input, the student personalized learning path and the psychological intervention scheme are dynamically generated, which can realize the bidirectional adaptive regulation and control of the weak learning link and the current psychological state, and improve the pertinence and adaptability of the path planning. By constructing the association rules of psychological and learning states, the system can match the learning task intensity and the intervention frequency according to the ability difference and the anxiety level, and realize the dynamic balance of the task intensity and the psychological load. When the anxiety state exceeds the set threshold, the system can automatically trigger the virtual interactive tutor to push the personalized mindfulness audio content, timely intervene and guide the student to restore the emotion, and optimize the learning rhythm. This mechanism realizes the deep fusion of psychological state dynamic perception and path generation, and enhances the response ability of the system to complex learning state.

[0128] In the embodiment, the bidirectional adaptive association rules of psychological and learning states are specifically:

[0129] The English ability state difference value level is defined as E, wherein:

[0130]

[0131] The student anxiety state quantitative index level is defined as A, wherein:

[0132]

[0133] The learning task intensity of the personalized learning path is defined as S, and the psychological intervention frequency is defined as P, wherein the intensity level satisfies: S H >S M >S L , and the intervention frequency level satisfies: P H >P M >P L ;

[0134] The mapping relationship of the bidirectional adaptive association rules of psychological and learning states is:

[0135]

[0136] By defining the English ability state difference value level E and the anxiety state quantitative index level A, and according to the intensity level and frequency level of SH>SM>SL and PH>PM>PL, the bidirectional mapping association rule of (E, A)→(S, P) is established, so that the fine matching of learning task intensity and psychological intervention frequency can be realized for different ability and anxiety combinations. For the (EH, AH) state, it is mapped to (SL, PH), and high-frequency psychological intervention is preferentially provided; for the (EL, AL) state, it is mapped to (SH, PL), and the learning task is preferentially strengthened; for other state combinations, different schemes such as (SM, PH), (SM, PM) or (SH, PM) are respectively mapped, so as to realize the dynamic balance of learning and intervention. After the E and A values are obtained in real time, the corresponding (S, P) combination can be quickly determined and dynamically applied to the generation of personalized learning paths and the scheduling of psychological intervention, which effectively avoids the rough adaptation of a single scheme, improves the pertinence and response speed of the learning path and the psychological intervention scheme, enhances the adaptability and stability of the system to various chemical emotional states, and effectively promotes the synchronous optimization of students' English ability and psychological state.

[0137] In the embodiment, the cultural bidirectional fusion knowledge graph is used to optimize the ratio of western culture knowledge input and Chinese culture knowledge output through the Monte Carlo tree search algorithm, and a cross-cultural learning scheme is dynamically generated, specifically as follows:

[0138] According to the specific requirements of the new national compulsory education English curriculum standard, a cultural bidirectional fusion knowledge graph containing Chinese culture knowledge nodes, western culture knowledge nodes and the mapping relationship therebetween is constructed;

[0139] Chinese culture knowledge interest degree data and cross-cultural adaptability evaluation data generated by students in the process of cultural knowledge learning are collected in real time to determine the quantitative parameters of students' Chinese culture knowledge interest and cross-cultural adaptability;

[0140] The Monte Carlo tree search algorithm is called to traverse the cultural knowledge nodes in the cultural bidirectional fusion knowledge graph with the quantitative parameters of Chinese culture knowledge interest and cross-cultural adaptability as core input conditions;

[0141] According to the traversal result of the Monte Carlo tree search algorithm, the expected learning benefit values of different cultural knowledge nodes are calculated and determined;

[0142] According to the size of the expected learning benefit values, the ratio relationship between the western culture knowledge input nodes and the Chinese culture knowledge output nodes in the student learning path is dynamically adjusted in real time;

[0143] According to the adjusted node proportion relationship, the student's personalized cross-cultural learning scheme is updated and dynamically generated.

[0144] By constructing a bidirectional fusion knowledge graph containing Chinese culture knowledge nodes and western culture knowledge nodes, and real-time collecting student's Chinese culture knowledge interest quantitative parameters and cross-cultural adaptation ability quantitative parameters as core input conditions, calling Monte Carlo tree search algorithm to traverse the knowledge graph nodes to calculate the expected learning return value of each node, the proportion relationship of western culture knowledge input nodes and Chinese culture knowledge output nodes in the learning path can be dynamically adjusted, and then the student's personalized cross-cultural learning scheme is generated in real time. This method breaks through the limitation of static transmission of cultural training in the prior art, realizes the bidirectional interactive optimization adjustment of cross-cultural knowledge content, and makes the learning scheme highly matched with the student's interest and adaptation ability; at the same time, based on the dynamic feedback mechanism of the expected learning return value, the cultural relevance and learning efficiency of the teaching content are effectively improved, and the student's cross-cultural expression ability and comprehensive quality are enhanced.

[0145] In the embodiment, the real-time feedback of the English ability state difference value, the anxiety state quantitative index and the cross-cultural learning scheme execution effect data through cross-terminal data transmission, the real-time optimization and update of the learning path and the psychological intervention scheme based on the federated learning algorithm, and the basis for the next multi-modal learning state feature data analysis and path generation are as follows:

[0146] The English ability evaluation data, the anxiety state quantitative index and the execution effect data of the cross-cultural fusion learning scheme are synchronously collected by the student learning terminal in real time, and the local differential privacy perturbation processing is performed by using the differential privacy mechanism;

[0147] The data processed by the local differential privacy perturbation is transmitted to the federated learning server in real time;

[0148] The federated learning server performs secure aggregation calculation on the data of multiple learning terminals without decrypting the original data, and obtains a globally unified encrypted data training set;

[0149] Taking the globally unified encrypted data training set as input, the multi-modal education big model parameters for analyzing the multi-modal learning state feature data of students are updated in real time based on the federated average aggregation algorithm;

[0150] The updated multi-modal education big model parameters are securely distributed to each student learning terminal in real time through the cross-terminal data transmission mechanism, and the real-time update of the end-side model parameters is automatically completed;

[0151] Each student learning terminal uses the updated multi-modal education big model parameters to analyze the currently collected multi-modal learning state feature data in real time, and automatically generates a new personalized learning path and psychological intervention scheme based on the analysis result.

[0152] The English ability state difference value, the anxiety state quantitative index and the cross-cultural integration learning scheme execution effect data are synchronously collected across terminals, and data perturbation is performed in combination with a local differential privacy mechanism to effectively protect student privacy. After the federated learning server performs secure aggregation calculation on the encrypted training set and updates the multi-modal education large model parameters in real time based on a federated averaging algorithm, the updated model parameters can be safely distributed to each student learning terminal to realize automatic real-time updating of the terminal-side model. Each learning terminal reanalyzes the multi-modal learning state feature data using the updated model parameters and dynamically generates a new individualized learning path and psychological intervention scheme to ensure that the path and the intervention scheme always match the latest learning situation. Under the premise of ensuring data security and privacy, the mechanism realizes the dual improvement of model precision and path adaptability through continuous optimization of federated learning, and significantly enhances the real-time response capability and individualized service effect of the system.

[0153] In the embodiment, the primary and secondary school English AI learning companion system based on the multi-modal education large model comprises:

[0154] The learning terminal is an intelligent stylus integrating a camera, a microphone and a built-in pressure sensor, which is used to collect and fuse multi-modal learning state feature data in real time;

[0155] The parameter generation module generates a multi-dimensional quantitative parameter system of English oral speech speed, vocabulary frequency and cultural knowledge proportion according to the national English new curriculum standard;

[0156] The ability evaluation module determines the English ability state difference value by matching the multi-modal learning state feature data with the multi-dimensional quantitative parameters item by item;

[0157] The anxiety state module generates an anxiety state quantitative index by fusing a facial micro-expression recognition module, a voice tremor frequency detection module and a writing pressure detection module;

[0158] The path generation and psychological intervention module dynamically generates an individualized learning path and a psychological intervention scheme according to the English ability state difference value and the anxiety state quantitative index by calling psychological and learning state association rules;

[0159] The federated learning optimization module updates the multi-modal education large model parameters in real time by cross-terminal differential privacy data transmission and federated learning algorithm, and optimizes the learning path and the psychological intervention scheme.

[0160] The learning terminal collects and fuses facial expression, voice and writing pressure data in real time through the camera, microphone and smart stylus with built-in pressure sensor to form multi-modal learning state feature data, and realizes high-precision capture of student learning behavior; the parameter generation module generates a multi-dimensional quantitative parameter system of English oral speech speed, vocabulary usage frequency and cultural knowledge proportion according to the national English new curriculum standard, effectively improving the matching accuracy of evaluation and teaching content; the ability evaluation module matches and calculates the English ability state difference value item by item according to the multi-modal learning state feature data and the quantitative parameters, and the quantitative degree refines the ability deviation judgment; the anxiety state module generates an anxiety state quantitative index by fusing facial micro-expression recognition, voice tremor frequency detection and writing pressure detection modules, and realizes real-time quantification of learning psychological state; the path generation and psychological intervention module calls psychological and learning state association rules based on the English ability state difference value and the anxiety state quantitative index to dynamically generate individualized learning paths and psychological intervention schemes, ensuring the synchronous adaptation of learning tasks and psychological intervention; the federal learning optimization module updates the large model parameters in real time through cross-terminal differential privacy data transmission and federal average algorithm, and safely issues the updated parameters to the learning terminal, realizes the continuous optimization of the model and the path, and significantly enhances the real-time response ability and individualized service effect of the system.

[0161] Embodiment 1

[0162] In order to verify the feasibility of the present application in implementation, the present application is applied to the eighth grade English classroom learning scene of a middle school. The original teaching mode of the classroom adopts fixed paper teaching materials and online oral English evaluation tools, and the teaching content is evenly distributed in each learning stage, and the training of students' listening, speaking, reading and writing ability lacks dynamic adjustment, and the students' learning anxiety and cultural accomplishment are not intervened in real time, resulting in the phenomenon that students have obvious learning maladjustment, psychological tension and cultural understanding deviation in the test.

[0163] In this scene, the traditional technology relies on teachers to manually set learning tasks and cultural input proportion, and cannot be adjusted individually according to the actual learning ability and psychological state of students, the teaching feedback cycle is as long as one week, the teaching resources are wasted and it is difficult to optimize the learning path in time. The average score of students in oral English fluency test is only 65 points, the average correct rate of reading comprehension test is 72%, the average score of cross-cultural test is 60%, the average anxiety index is 68 (full score 100), and the response time of common artificial intervention strategy is more than 30 seconds, which is difficult to relieve students' anxiety in time.

[0164] In this implementation, a high-definition camera and microphone are first deployed on each student's learning terminal, and a smart stylus with a built-in pressure sensor is equipped for synchronous collection of facial expressions, voice signals, and writing pressure data. The collection frequency is 30 frames of image data per second, 16 kilohertz of voice data, and 100 hertz of pressure sensor data. The multi-modal data is preprocessed by a local fusion module, the image denoising uses a bilateral filtering algorithm, the voice signal extracts the fundamental frequency jitter parameter after fast Fourier transform, and the writing pressure data is normalized by minimum-maximum.

[0165] Next, the parameter generation module determines the upper and lower limits of English oral speech speed (120 to 150 words per minute), the core vocabulary usage frequency parameters (high-frequency words 1800 times per hour, extended words 800 times per hour, and cultural vocabulary 300 times per hour), and the proportion of Chinese and Western cultural knowledge (each accounting for 50% of the initial value) according to the listening, speaking, reading, and writing ability requirements of the eighth grade in the new curriculum standards for compulsory education English. After standardization, a six-dimensional quantitative parameter system is generated.

[0166] The ability evaluation module matches the processed multi-modal learning state feature data with the above six-dimensional quantitative parameters item by item, calculates the oral speech speed difference value, vocabulary usage difference value, and cultural proportion difference value by the difference value method, and weights the sum by the weight coefficients 0.25:0.25:0.25:0.25 to obtain the comprehensive difference value. This process is completed after each collection, taking less than 200 milliseconds.

[0167] The anxiety state module uses improved facial micro-expression action recognition algorithm, voice tremor frequency detection algorithm, and writing pressure detection algorithm to quantify the anxiety indicators. The facial module extracts the inter-brow distance, eyelid amplitude, and lip corner jitter parameters by matching between the pre-set anxiety templates through convolutional neural network, calculates the anxiety facial quantization value; the voice module extracts the harmonic frequency offset and fundamental frequency period jitter parameters to obtain the voice anxiety quantization value; the writing module records the pressure peak value and pressure change rate to extract the writing anxiety quantization value. After linear fusion processing of the three according to the weight 0.4:0.3:0.3, minimum-maximum normalization is performed to obtain the unified dimension anxiety state quantization index. The response delay of this module is less than 300 milliseconds.

[0168] The path generation and psychological intervention module takes the comprehensive difference value and anxiety state quantization index as input, automatically determines the learning task intensity level S and psychological intervention frequency P based on the pre-established bidirectional mapping rules. The mapping relationship follows the grade constraints SH>SM>SL and PH>PM>PL, quickly obtains the corresponding (S, P) combination through table lookup, and calls the virtual interactive tutor interface and mindfulness audio library for immediate intervention. The system triggers a mental health conversation when the anxiety index exceeds the threshold value of 50, with a response delay of less than 4 seconds, which is about 87% shorter than traditional teacher manual intervention.

[0169] In the cross-cultural fusion learning scheme generation process, a bidirectional fusion knowledge graph containing Chinese culture nodes and Western culture nodes is constructed, and the cultural interest degree and cross-cultural adaptation quantitative parameters of students are collected in real time as the iterative input conditions of the MCTS algorithm. The Monte Carlo tree search traverses the graph nodes for 200 times with a tree depth of 5, calculates the expected learning return value of each node, and dynamically adjusts the presentation ratio of Chinese and Western culture knowledge. The algorithm generates a new learning scheme after each iteration, which takes about 1.2 seconds.

[0170] The federal learning optimization module updates the large model parameters through secure aggregation calculation based on the data disturbed by the local differential privacy mechanism in 10 rounds of federal average training. The differential privacy ∈ is set to 1.0, the training set size covers 30 learning terminals, and the converged iteration number of the converged model is 8 rounds. The final model has a multi-modal feature analysis accuracy of 96.2% on the local validation set, which is 5.8 percentage points higher than the initial model. The updated model parameters are distributed to each terminal through an encrypted channel, and the terminal updates in real time and is used again for data analysis and path generation.

[0171] The system tested 40 students in the class for four weeks, compared the comprehensive learning effects of traditional teaching and system-assisted teaching, and randomly selected 5 student samples for prediction and actual measurement comparison. The following table shows the predicted and actual values of some samples in reading comprehension, oral fluency, cross-cultural test, and anxiety index.

[0172] Table 1 Comparison of predicted and actual learning effects and psychological states of eighth-grade students

[0173]

[0174] As can be seen from Table 1, the prediction error of the system on each index is controlled within 5%, among which the maximum prediction error of reading comprehension is 2.35%, the maximum prediction error of oral fluency is 3.57%, the maximum prediction error of cross-cultural test is 2.50%, and the maximum prediction error of anxiety state is 5.13%, showing high accuracy of model prediction. After the implementation of the system for four weeks, the average reading comprehension score of students increased from 72 to 84, the oral fluency increased from 90 words per minute to 124 words per minute, the cross-cultural test increased from 60 points to 78 points, the anxiety index decreased from 68 to 42 on average, and the overall learning efficiency was significantly improved.

[0175] Based on the above example data, the method of the present application has significant technical advantages in dynamically meeting the requirements of national curriculum standards, real-time attention to student psychological state and closed-loop intervention, and bidirectional interactive cultural fusion, providing an efficient, precise, and quantifiable intelligent assistance scheme for primary and secondary school English teaching.

[0176] The above merely describes preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art, according to the technical solution and inventive concept of the present application, makes equivalent replacement or change within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.

Claims

1. A primary and secondary school English AI study companion method based on a multi-modal education large model, characterized in that, The application relates to a multi-modal learning state feature data construction method and system. Student facial expression data is collected through a camera, student speech data is collected through a microphone, and student writing trajectory and writing pressure data are collected through an intelligent handwriting pen, and the data are fused into multi-modal learning state feature data; A multi-dimensional quantitative parameter system of English oral speech speed, vocabulary usage frequency and cultural knowledge proportion is constructed according to the national compulsory education English new curriculum standard; The multi-modal learning state feature data is matched with the multi-dimensional quantitative parameter system to determine the difference value between the student English ability state and the English ability state required by the new curriculum standard; Anxiety state quantitative index of the student is fused and generated through a facial micro-expression recognition algorithm, a speech tremor frequency detection algorithm and a writing pressure sensor; Based on the English ability state difference value and the anxiety state quantitative index of the student, a personalized learning path is generated, and when the anxiety index exceeds a threshold value, a virtual interactive tutor triggers a psychological guidance dialogue and pushes a mindfulness audio for psychological intervention; The student's interest in Chinese culture and cross-cultural adaptability data are taken as input, a cultural bidirectional fusion knowledge graph is used, a Monte Carlo tree search algorithm is used to optimize the input of western culture knowledge and the output of Chinese culture knowledge, and a cross-cultural learning scheme is dynamically generated; The English ability state difference value, the anxiety state quantitative index and the cross-cultural learning scheme execution effect data are fed back in real time through cross-terminal data transmission, a federated learning algorithm is used to optimize and update the learning path and the psychological intervention scheme in real time, and the learning path and the psychological intervention scheme are used as the basis for the next multi-modal learning state feature data analysis and path generation.

2. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 1, characterized in that, The multi-modal learning state feature data is constructed, and specifically, The student facial expression data collected by the camera in real time is subjected to frame-by-frame image denoising processing, and facial key point positions and action amplitude parameters are extracted; The student speech data collected by the microphone in real time is subjected to time domain frame division and spectrum feature analysis, and fundamental frequency, harmonic frequency and formant position parameters of the speech are extracted; The student writing trajectory data collected by the intelligent handwriting pen in real time is subjected to two-dimensional coordinate mapping processing, writing trajectory coordinate parameters are obtained, and pressure value parameters of the writing pressure data are synchronously recorded; The facial key point positions and action amplitude parameters, the fundamental frequency, the harmonic frequency and the formant position parameters of the speech, the writing trajectory coordinate parameters and the pressure value parameters are taken as initial inputs, and principal component analysis algorithm is used to determine weight factors of the parameters; The weight factors of the parameters are multiplied by corresponding quantitative values of the parameters respectively, and then summation is performed to obtain the fused multi-modal learning state feature data.

3. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 1, characterized in that, The multi-dimensional quantitative parameter system of English oral speech speed, vocabulary usage frequency and cultural knowledge proportion is constructed, and specifically, According to the English oral ability requirements of grades 3 to 9 in the national compulsory education English new curriculum standard, upper and lower limits of the number of word pronunciations in the English oral expression of the student per minute are determined year by year, and are converted into English oral speech speed quantitative parameters; According to the specific requirements of the new curriculum standard for the vocabulary mastery of each grade, the total amount of vocabulary is explicitly divided into three independent categories of basic vocabulary, expanded vocabulary and cultural vocabulary, and the usage frequency parameters of the vocabularies in the teaching content are determined respectively. According to the requirement of the new curriculum standard for the integration of Chinese and Western cultural knowledge teaching, the proportion of the teaching time of the Chinese traditional cultural knowledge and the Western cultural knowledge involved in the classroom teaching is quantified to form a Chinese traditional cultural knowledge proportion parameter and a Western cultural knowledge proportion parameter; According to the English oral speech speed quantization parameter, the basic vocabulary usage frequency parameter, the extended vocabulary usage frequency parameter, and the cultural vocabulary usage frequency parameter, an English teaching difficulty level parameter is obtained; Taking the Chinese traditional cultural knowledge proportion parameter and the Western cultural knowledge proportion parameter as input conditions, a cultural sensitivity level parameter is obtained; After standardizing and normalizing the English oral speech speed quantization parameter, the various vocabulary usage frequency parameters, the English teaching difficulty level parameter, and the cultural sensitivity level parameter, a multi-dimensional quantization parameter system is obtained.

4. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 1, characterized in that, The multi-modal learning state feature data is matched with the multi-dimensional quantization parameter system to determine the difference value between the student's English ability state and the English ability state required by the new curriculum standard, which is specifically: The difference value between the student's actual English oral speech speed parameter in the multi-modal learning state feature data and the English oral speech speed quantization parameter in the multi-dimensional quantization parameter system is calculated to obtain an English oral speech speed difference value; The difference values between the student's actual vocabulary usage frequency parameters in the multi-modal learning state feature data and the corresponding basic vocabulary usage frequency parameter, the extended vocabulary usage frequency parameter, and the cultural vocabulary usage frequency parameter in the multi-dimensional quantization parameter system are calculated to obtain a basic vocabulary usage frequency difference value, an extended vocabulary usage frequency difference value, and a cultural vocabulary usage frequency difference value; The difference values between the student's actual Chinese traditional cultural knowledge proportion and Western cultural knowledge proportion in the multi-modal learning state feature data and the Chinese traditional cultural knowledge proportion parameter and the Western cultural knowledge proportion parameter in the multi-dimensional quantization parameter system are calculated to obtain a Chinese traditional cultural knowledge proportion difference value and a Western cultural knowledge proportion difference value; Taking the English oral speech speed difference value, the basic vocabulary usage frequency difference value, the extended vocabulary usage frequency difference value, the cultural vocabulary usage frequency difference value, the Chinese traditional cultural knowledge proportion difference value, and the Western cultural knowledge proportion difference value as input conditions, a weighted summation method is used to calculate an English ability comprehensive difference value; According to the size of the English ability comprehensive difference value and the preset ability difference threshold value, the degree of difference between the student's actual English ability state and the ability requirement specified by the new curriculum standard is determined to obtain an English ability state difference value.

5. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 1, characterized in that, The student's anxiety state quantization index is generated by fusing the face micro-expression recognition algorithm, the voice tremor frequency detection algorithm, and the writing pressure sensor, which is specifically: The face micro-expression action recognition algorithm is used to extract the eyebrow position change distance, eye muscle contraction degree, and mouth corner displacement amplitude parameters from the student's face image sequence collected by the camera frame by frame. By calculating the numerical difference between the parameters and the preset anxiety face micro-expression template, a face micro-expression anxiety quantization value is obtained. The speech signal frequency domain analysis algorithm is adopted to perform fast Fourier transform on the student speech data collected by the microphone in real time, the harmonic frequency offset degree and the speech fundamental frequency tremor period parameters in the speech signal are extracted, the difference between the parameters and the preset speech tremor frequency threshold is calculated, and the speech tremor frequency anxiety quantitative value is obtained; The writing pressure sensor is used to record the student writing pressure change data in real time, the pressure change rate and the pressure peak value parameters in unit time are extracted, the difference between the parameters and the preset writing pressure normal interval is calculated, and the writing pressure change anxiety quantitative value is obtained; The face micro-expression anxiety quantitative value, the speech tremor frequency anxiety quantitative value and the writing pressure change anxiety quantitative value are respectively given corresponding weight coefficients, the weighted sum of the three anxiety quantitative values is calculated by linear weighting fusion method, and the initial comprehensive quantitative value of the anxiety state is obtained; The initial comprehensive quantitative value of the anxiety state is subjected to minimum-maximum normalization processing, and the student anxiety state quantitative index with unified numerical range is obtained.

6. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 1, characterized in that, Based on the English ability state difference value and the student anxiety state quantitative index, a personalized learning path is generated, when the anxiety index exceeds the threshold value, a psychological guidance dialogue is triggered through a virtual interactive tutor and a mindfulness audio is pushed for psychological intervention, specifically: According to the numerical size and the positive and negative directions of the English ability state difference value, the weak link of student English learning and the learning intensity adjustment range are determined; According to the numerical interval of the student anxiety state quantitative index, the current psychological state grade of the student is determined; Taking the weak link of English learning, the learning intensity adjustment range and the current psychological state grade of the student as input conditions, the pre-constructed psychological and learning state bidirectional adaptive association rule is called to obtain the initial scheme of the personalized learning path matched with the current learning and psychological state; The numerical size of the student anxiety state quantitative index and the preset anxiety threshold value is monitored in real time, when the student anxiety state quantitative index exceeds the preset anxiety threshold value, the virtual interactive tutor is triggered to start the psychological health guidance dialogue; According to the current anxiety state grade of the student and the English ability state difference value, the corresponding mindfulness audio content is matched and pushed from the preset mindfulness audio content library, and the student is intervened psychologically; After the real-time intervention of the psychological health guidance dialogue and the mindfulness audio content of the virtual interactive tutor, the student anxiety state quantitative index is re-evaluated in real time, the initial scheme of the personalized learning path is updated, and the real-time dynamically adjusted personalized learning path is obtained.

7. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 6, characterized in that, The psychological and learning state bidirectional adaptive association rule is specifically: The English ability state difference value grade is defined as E, wherein: The student anxiety state quantitative index grade is defined as A, wherein: The learning task intensity of defining the personalized learning path is S, and the psychological intervention frequency is P, wherein the intensity level satisfies: S H >S M >S L , and the intervention frequency level satisfies: P H >P M >P L ; The mapping relationship of the psychological and learning state bidirectional adaptive association rule is:

8. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 1, characterized in that, Based on the cultural bidirectional fusion knowledge graph, the proportion of western culture knowledge input and Chinese culture knowledge output is optimized by the Monte Carlo tree search algorithm, and a cross-cultural learning scheme is dynamically generated, specifically: According to the specific requirements of the national compulsory education English new curriculum standard, a cultural bidirectional fusion knowledge graph containing Chinese culture knowledge nodes, western culture knowledge nodes and the mapping relationship therebetween is constructed; Real-time collection of Chinese culture knowledge interest degree data and cross-cultural adaptation ability evaluation data generated by students in the process of learning cultural knowledge, and determination of the quantitative parameters of the students' interest in Chinese culture knowledge and the quantitative parameters of cross-cultural adaptation ability; Taking the quantitative parameters of Chinese culture knowledge interest and the quantitative parameters of cross-cultural adaptation ability as core input conditions, calling the Monte Carlo tree search algorithm to traverse the cultural knowledge nodes in the culture bidirectional fusion knowledge graph; According to the traversal result of the Monte Carlo tree search algorithm, the expected learning benefit values of different cultural knowledge nodes are calculated and determined; According to the size of the expected learning benefit values, the proportion relationship of the western culture knowledge input node and the Chinese culture knowledge output node in the student learning path is dynamically adjusted in real time; According to the adjusted node proportion relationship, the student's personalized cross-cultural learning scheme is updated and dynamically generated.

9. The primary and secondary school English AI study companion method based on a multi-modal education large model according to claim 1, characterized in that, The real-time feedback of English ability state difference value, anxiety state quantitative index and cross-cultural learning scheme execution effect data through cross-terminal data transmission, real-time optimization and update of learning path and psychological intervention scheme based on federated learning algorithm, as the basis for the next multi-modal learning state feature data analysis and path generation, specifically: Real-time synchronous collection of English ability evaluation data, anxiety state quantitative index and cross-cultural fusion learning scheme execution effect data through the student learning terminal, and local differential privacy perturbation processing is performed by using the differential privacy mechanism; The data processed by local differential privacy perturbation is transmitted to the federated learning server in real time after being encrypted; The federated learning server performs secure aggregation calculation on the data of multiple learning terminals without decrypting the original data, and obtains a globally unified encrypted data training set; Taking the globally unified encrypted data training set as input, the multi-modal education big model parameters for analyzing student multi-modal learning state feature data are updated in real time based on the federated average aggregation algorithm; The updated multi-modal education big model parameters are securely distributed to each student learning terminal through cross-terminal data transmission mechanism in real time, and the real-time update of the end-side model parameters is automatically completed; Each student learning terminal uses the updated multi-modal education big model parameters to analyze the currently collected multi-modal learning state feature data in real time, and automatically generates a new individualized learning path and psychological intervention scheme based on the analysis results.

10. A primary and secondary school English AI study companion system based on a multi-modal education large model, which executes the primary and secondary school English AI study companion method based on a multi-modal education large model according to any one of claims 1 to 9, characterized in that, It includes: Learning terminal, intelligent stylus integrated with camera, microphone and built-in pressure sensor, used for real-time collection and fusion to form multi-modal learning state feature data; Parameter generation module, generating a multi-dimensional quantitative parameter system of English oral speech speed, vocabulary frequency and cultural knowledge proportion according to the national English new curriculum standard; Ability evaluation module, determining the English ability state difference value by matching the multi-modal learning state feature data with the multi-dimensional quantitative parameters item by item; Anxiety state module, using face micro-expression recognition module, voice tremor frequency detection module and writing pressure detection module to generate anxiety state quantitative index; Path generation and psychological intervention module, dynamically generating individualized learning path and psychological intervention scheme according to English ability state difference value and anxiety state quantitative index by calling psychological and learning state association rules; The federal learning optimization module updates the multi-modal education big model parameters in real time through cross-terminal differential privacy data transmission and federal learning algorithm, and optimizes the learning path and psychological intervention scheme.

Citation Information

Patent Citations

  • Processing method and device for classroom teaching behavior evaluation and storage medium

    CN115358615A

  • Person guarantee system

    JP2006277498A

  • Device, method and program for estimating vocabulary learning curve parameter

    JP2013167980A

  • Deep academic learning intelligence and deep neural language network system and interfaces

    US20180247549A1