Middle and primary school multi-person foreign language situational teaching method and system based on VR

By designing a VR-based foreign language contextual teaching system for primary and secondary schools, the problem of lack of a multi-person foreign language teaching system for primary and secondary school students in the existing technology is solved, and a personalized and efficient learning experience is achieved, which enhances the application of VR technology in situation creation and social interaction.

CN120031684AInactive Publication Date: 2025-05-23ZHONGKE LEMU (JIANGSU) TECH CO LTD
View PDF 0 Cites 13 Cited by

Patent Information

Application Number
CN202411927284.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing VR teaching system is mainly aimed at adults, lacks a multi-person foreign language contextual teaching system for primary and secondary school students, and the single-person learning system is difficult to fully utilize the advantages of VR technology in situation creation and social interaction.

Method used

Design a VR-based foreign language contextual teaching method and system for primary and secondary schools. By obtaining teaching tasks and determining the theme, scene and interaction methods of virtual reality situations, combining Octrene space division and voxelization modeling technology for data compression and load optimization, constructing an immersive virtual teaching scenario, and integrating real teaching environment elements through intelligent object recognition and real scene modeling algorithms. Students collect multimodal interactive data through a multi-channel sensor array, use video behavior understanding models and agent reinforcement learning algorithms to generate personalized dialogue feedback, and perform multi-dimensional quantitative evaluation and learning path optimization.

Benefits of technology

It realizes multi-person foreign language contextual teaching for primary and secondary school students, improves learning interest and effect, enhances the application of VR technology in situation creation and social interaction, and provides a personalized and efficient learning experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031684A_ABST
    Figure CN120031684A_ABST
Patent Text Reader

Abstract

The invention provides a middle and primary school multi-person foreign language situational teaching method and system based on VR, and relates to the technical field of virtual reality, and the method comprises the steps: obtaining a teaching task, constructing an immersive virtual teaching scene, setting a situational dialogue task and a self-adaptive difficulty adjustment mechanism, carrying out the semantic description and association mapping to form a multi-level semantic network, and carrying out the virtual teaching. Each student enters an immersive virtual teaching scene, participates in a scene dialogue task, collects multi-modal interaction data and transmits the data back to the edge computing server to obtain mother language voice of the student, translates the mother language voice to obtain a high-quality dialogue text and generates personalized dialogue feedback; performing multi-dimensional quantitative evaluation based on personalized dialogue feedback, generating an ability level portrait, determining the knowledge point mastering degree of each student, generating an knowledge mastering graph, constructing a language learning path optimization model, solving to obtain an optimal learning path, and generating an optimized teaching scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of virtual reality technology, and in particular to a VR-based multi-person foreign language situational teaching method and system for primary and secondary schools. Background Art

[0002] With the acceleration of globalization, the importance of foreign language education has become increasingly prominent. The traditional foreign language teaching method is mainly based on teacher lectures. Students passively accept knowledge and lack language application environment and practice opportunities, which leads to low learning interest and poor learning effect. In order to improve the quality of foreign language teaching, educators at home and abroad have made many attempts and explorations. At present, common foreign language teaching methods include task-based teaching method, communicative teaching method and multimedia-assisted teaching.

[0003] In recent years, with the rapid development of modern information technologies such as computer technology, network technology, and multimedia technology, virtual reality technology has become increasingly mature, providing new opportunities and challenges for foreign language teaching. Most of the existing VR teaching systems are aimed at adult learners, and there is a lack of VR teaching systems for primary and secondary school students. Primary and secondary school students are in a critical period of language learning, and their cognitive abilities, self-control abilities, and language foundations are significantly different from those of adults. It is necessary to design and develop VR teaching systems based on their characteristics. In addition, most of the existing VR teaching systems are single-person learning systems, lacking multi-person collaboration and interactive functions, making it difficult to give full play to the advantages of VR technology in situation creation and social interaction.

[0004] Therefore, a solution is urgently needed to solve the problems existing in the prior art. Summary of the invention

[0005] The embodiments of the present invention provide a VR-based multi-person foreign language situational teaching method and system for primary and secondary schools, which can at least solve some of the problems existing in the prior art.

[0006] A first aspect of an embodiment of the present invention provides a VR-based multi-person foreign language situational teaching method for primary and secondary schools, comprising:

[0007] Acquire teaching tasks and determine the theme, scene, character role setting and interaction mode of the virtual reality scenario based on the teaching tasks, compress the data and optimize the loading of the scene through octree space division and voxel modeling technology, build an immersive virtual teaching scene in combination with a physics-based rendering method and a real-time lighting algorithm, integrate real teaching environment elements into the immersive virtual teaching scene through intelligent object recognition and real-scene modeling algorithms, set situational dialogue tasks and an adaptive difficulty adjustment mechanism in the immersive virtual teaching scene according to the teaching tasks, and combine knowledge graphs and ontology reasoning technologies to semantically describe and associate elements in the immersive virtual teaching scene to form a multi-level semantic network;

[0008] Each student logs in to the multi-person teaching platform through a VR all-in-one machine, sets a virtual image and enters the immersive virtual teaching scene, participates in the situational dialogue task in the immersive virtual teaching scene, collects the multimodal interaction data corresponding to the current student through a multi-channel sensor array and transmits it back to the edge computing server, learns the multimodal interaction data through a video behavior understanding model, generates multi-channel spatiotemporal features and obtains quantitative analysis results, performs speech denoising and feature extraction on the video data collected by the multi-channel sensor array, performs contextual language modeling in combination with the attention mechanism, obtains the student's native language speech, translates the student's native language speech through a speech translation model based on deep migration, generates an initial foreign language text, and screens it through a multilingual bidirectional encoder model to obtain a high-quality dialogue text, and generates personalized dialogue feedback based on the high-quality dialogue text, the quantitative analysis results and the multimodal interaction data through an intelligent agent reinforcement learning algorithm;

[0009] Based on the personalized dialogue feedback and the multi-channel spatiotemporal features, combined with the hierarchical transfer learning algorithm, a multi-dimensional quantitative assessment of the language application ability of each student is performed to generate an ability level portrait, and the development level of each student's skill dimension is tracked based on the knowledge component theory and the multidimensional item response theory. The Bayesian knowledge tracking model is used to determine the degree of mastery of each student's knowledge points and generate a knowledge mastery map. The multi-channel spatiotemporal features are integrated and combined with the long short-term memory network for feature learning to obtain learning behavior characteristics. The ability level portrait, the knowledge mastery map and the learning behavior characteristics are summarized. A language learning path optimization model is constructed based on a factorization machine and combinatorial ranking learning, and the optimal learning path is obtained by combining the multi-objective optimization algorithm. An optimized teaching plan is generated based on the optimal learning path.

[0010] In an optional embodiment,

[0011] The teaching task is obtained and the theme, scene, character role setting and interaction mode of the virtual reality scenario are determined based on the teaching task. The scene is compressed and optimized for loading through octree space division and voxel modeling technology. An immersive virtual teaching scene is constructed by combining a physics-based rendering method and a real-time lighting algorithm. The real teaching environment elements are integrated into the immersive virtual teaching scene through intelligent object recognition and real-scene modeling algorithms. According to the teaching task, a scenario dialogue task and an adaptive difficulty adjustment mechanism are set in the immersive virtual teaching scene. The elements in the immersive virtual teaching scene are semantically described and associated mapped in combination with knowledge graphs and ontology reasoning technology to form a multi-level semantic network including:

[0012] The current learning stage and teaching tasks are obtained from the teaching management system, and the teaching tasks are semantically understood and key information is extracted through natural language processing technology to obtain a structured task description. Based on the structured task description, the teaching content is semantically modeled and associated with mining through ontology reasoning technology to identify the core semantic elements. Based on the hierarchical relationship and logical association between different core semantic elements, the themes and scenes with the highest relevance to the teaching tasks are determined. For each core semantic element, the corresponding element attributes are determined through a knowledge-driven method to obtain a character role setting, and the virtual character dialogue and behavior perception are simulated through speech synthesis technology and gaze tracking technology to set the interaction mode.

[0013] The scene corresponding to each virtual reality scenario is represented in multiple resolutions and processed in detail levels by using the octree-based space division and voxel modeling technology, the number of division levels and node size of the octree are adaptively determined, the scene model is converted into a hierarchical network composed of voxel units by a voxelization algorithm, the detail levels of the nodes in the octree are adjusted in combination with the viewpoint position and the viewing direction, the number of grids is clipped and detail level hierarchical model data of different areas in the cone is generated, the scene map and animation sequence are query-oriented data sharding is performed through a distributed data organization and a multi-level cache mechanism, and redundant storage is performed for loading optimization by setting a multi-level cache architecture;

[0014] Based on the scene after data compression and loading optimization, the lighting, material and reflection properties of the virtual reality scenario are modeled by a physically based rendering method, the ambient lighting is rendered in combination with a pre-calculated radiosity transfer technology, the dynamic light source is smoothly interpolated by a spherical harmonic function and a vertex shader, the height field is modeled by a normal map and a displacement map, a bidirectional scattering distribution function is selected in combination with the reflection characteristics of different materials, and a rendering equation is solved to obtain rendering parameters, and the immersive virtual teaching scene is constructed according to the rendering parameters;

[0015] Through intelligent object recognition algorithm and real scene modeling algorithm, combined with the camera array pre-set in the teaching venue, the teaching environment is determined, and the scene semantics and spatial structure are determined through semantic segmentation and 3D reconstruction algorithm based on deep learning. In combination with plane detection, objects in the real teaching environment are determined, and a 3D model corresponding to each object is generated. The 3D model corresponding to each object is superimposed and fused with the virtual elements in the immersive virtual teaching scene;

[0016] According to the language level of the students and the teaching tasks, scenario dialogue tasks including multiple levels are set in the immersive virtual teaching scene. In each level of dialogue tasks, the students' dialogue performance is used as the environmental state of reinforcement learning, the difficulty parameter of the dialogue task is used as the action space, a reward function is set and the difficulty parameter is optimized, the difficulty of the scenario dialogue task is adaptively adjusted, and the multimodal information in the current scenario is obtained in combination with the knowledge graph technology, a multimodal semantic representation is generated and association mining is performed, the scene, the objects and characters in the real teaching environment are taken as core nodes, triples are constructed through relationship extraction, and a multi-relation heterogeneous knowledge graph is obtained, the multi-level semantic associations are determined through a graph neural network and a tight cluster is generated, and semantic mapping is performed in combination with the association mapping technology to obtain the multi-level semantic network.

[0017] In an optional embodiment,

[0018] The smooth interpolation of dynamic light sources through spherical harmonics and vertex shaders is shown in the following formula:

[0019]

[0020] Among them, P(x) represents the illumination value calculated at vertex x, l represents the index of the order of the spherical harmonic function basis, L represents the order of the spherical harmonic function, and m represents the degree of the spherical harmonic function. Indicates t 1 Spherical harmonic coefficients of the time frame, Indicates t 2 Spherical harmonic coefficients of the time frame, Y lm (θ, φ) represents the spherical harmonic function basis, θ represents the elevation angle in the spherical coordinate system, φ represents the azimuth angle in the spherical coordinate system, and t represents the current time.

[0021] In an optional embodiment,

[0022] Each student logs in to the multi-person teaching platform through a VR all-in-one machine, sets a virtual image and enters the immersive virtual teaching scene, participates in the situational dialogue task in the immersive virtual teaching scene, collects the multimodal interaction data corresponding to the current student through a multi-channel sensor array and transmits it back to the edge computing server, learns the multimodal interaction data through a video behavior understanding model, generates multi-channel spatiotemporal features and obtains quantitative analysis results, performs speech denoising and feature extraction on the video data collected by the multi-channel sensor array, combines the attention mechanism to perform contextual language modeling, obtains the student's native language speech, translates the student's native language speech through a speech translation model based on deep migration, generates an initial foreign language text and screens it through a multilingual bidirectional encoder model to obtain a high-quality dialogue text, and generates personalized dialogue feedback based on the high-quality dialogue text, the quantitative analysis results and the multimodal interaction data through an intelligent agent reinforcement learning algorithm, including:

[0023] Each student logs in to the multi-person teaching platform through a VR all-in-one machine, sets a virtual image and enters the immersive virtual teaching scene, and participates in a pre-set scenario dialogue task in the immersive virtual teaching scene. The multi-channel sensor array collects each student's facial expression and body movement through a video sensor, records each student's voice signal through a microphone array and locates the sound source, obtains the student's hand and head movement data through an inertial measurement unit and an optical motion capture system, records the student's gaze point and gaze time through an eye tracker, determines the focus of each student in the current scene, aligns the data obtained by the multi-channel sensor with timestamps, and performs multimodal data fusion in frames to obtain the multimodal interaction data;

[0024] The multimodal interaction data is added to a pre-set video behavior understanding model, and features are extracted through a pre-trained convolutional neural network to generate a frame-level deep representation and send it to a long short-term memory network, determine the temporal dependency of different frame-level deep representations and generate a hidden state vector, perform attention weighting on the hidden state vector through an attention pooling mechanism and output it in a classified manner, generate a probability distribution of each behavior category and corresponding multi-channel spatiotemporal features, construct a quantitative indicator, and quantitatively analyze the multimodal interaction data based on the probability distribution corresponding to each behavior category to obtain the quantitative analysis result;

[0025] Perform speech noise reduction and feature extraction on the video data recorded by the video sensor and the speech signal recorded by the microphone array, generate a feature sequence and extract features through a recurrent convolutional network, obtain high-order local features and perform sequence modeling, obtain a high-order local sequence, initialize the high-order local features corresponding to the decoder and update them at each decoding moment, generate a context vector in combination with the attention weight, process it through a fully connected layer and a soft normalization function, obtain the probability distribution of the native language words at the current moment, combine the cross entropy loss function and use the real native language words as input for reasoning, and obtain the student's native language speech;

[0026] Based on the native language speech of the student, the semantic mapping relationship between the native language and the foreign language is determined by a speech translation model based on deep transfer, the speech translation model is trained in a pre-set bilingual corpus and the semantic mapping relationship is verified, the hyperparameters in the speech translation model are adjusted in combination with the verification result, the student's native language speech is translated by the adjusted speech translation model, a foreign language text sequence is generated and optimized by a beam search algorithm to obtain an initial foreign language text, the initial foreign language text is encoded by a bidirectional encoder in a multilingual bidirectional encoder model, the encoding result is retrieved in a pre-constructed multilingual corpus, the similarity between each element in the multilingual corpus and the encoding result is calculated, similar dialogues are determined based on the similarity and a candidate set is generated, the elements in the candidate set are scored according to language fluency and relevance, a dialogue score corresponding to each element is obtained, and the similar dialogue with the highest dialogue score is selected as the high-quality dialogue text;

[0027] Based on the high-quality dialogue text, the quantitative analysis results and the multimodal interaction data, the effectiveness of the high-quality dialogue text in the current scenario is determined through an intelligent agent reinforcement algorithm, and feedback optimization is performed according to the quantitative analysis results and the multimodal interaction data to generate the personalized dialogue feedback.

[0028] In an optional embodiment,

[0029] The similarity between each element in the multilingual corpus and the encoding result is calculated as shown in the following formula:

[0030]

[0031] Among them, h i represents the candidate translation, s(h i ) represents the candidate translation h i Similarity with the encoding result, λ 1 represents the probability weight of the language model, p(h i ) represents the candidate translation h i The language model probability, |h i| represents the length of the candidate translation, a represents the length penalty strength, λ 2 represents the coverage weight, n-grams (h i ) represents the candidate translation h i n-grams(f) represents the n-grams of the source language sentence, |f| represents the length of the source language sentence, λ 3 Represents the length ratio weight.

[0032] In an optional embodiment,

[0033] Based on the personalized dialogue feedback and the multimodal interaction data, the language application ability of each student is quantitatively evaluated in multiple dimensions in combination with the hierarchical transfer learning algorithm to generate an ability level portrait. The development level of each student's skill dimension is tracked based on the knowledge component theory and the multidimensional item response theory. The degree of mastery of each student's knowledge points is determined in combination with the Bayesian knowledge tracking model to generate a knowledge mastery map. The multi-channel spatiotemporal features are integrated and feature learning is performed in combination with the long short-term memory network to obtain learning behavior features. The ability level portrait, the knowledge mastery map and the learning behavior features are summarized. A language learning path optimization model is constructed based on a factorization machine and combinatorial ranking learning. The optimal learning path is obtained by solving the multi-objective optimization algorithm. The optimized teaching plan based on the optimal learning path includes:

[0034] Obtain language proficiency test questions from a corpus and annotate knowledge points, add the personalized dialogue feedback and the multi-channel spatiotemporal features to a multi-layer perception mechanism, generate a high-dimensional feature representation and add it to the hierarchical transfer learning algorithm, set a feature extractor for each language proficiency dimension, extract the language proficiency representation corresponding to each modality and map it to a shared semantic space, perform a multi-dimensional quantitative assessment of each student's language application ability based on the mapping results, and splice the mapping results to obtain the ability level portrait;

[0035] Based on the ability level portrait, combined with the knowledge component theory and the multidimensional item response theory, the language test questions are divided into multiple knowledge components and the knowledge component examination weights corresponding to each test question are determined, the knowledge component examination weights and answer records are added to the multidimensional item response theory model, the student's ability level in each dimension is estimated, the correct answer rate of the question is calculated based on the student's ability level in each dimension, the multidimensional item response theory model is optimized in combination with the cross entropy loss function and the real answer record, and the optimization is repeated until the preset maximum number of iterations is reached, and a dynamic evolution sequence corresponding to the ability level is generated. Through the Bayesian model and the dynamic evolution sequence, the current student's mastery of each knowledge point is generated, and the relationship between the knowledge points is combined to generate a knowledge mastery map;

[0036] The multi-channel spatiotemporal features are represented as three-dimensional tensors corresponding to time steps, channel features and original features respectively, and local spatiotemporal features are extracted by combining a three-dimensional convolutional neural network pre-set along the channel dimension and flattened in the channel dimension, and long-term dependency modeling is performed by combining a two-layer long short-term memory network to obtain the learning behavior features;

[0037] Summarize the ability level portrait, the knowledge mastery map and the learning behavior characteristics, score the match between students and learning activities based on a factor decomposition machine, optimize the factor decomposition machine in combination with minimizing the sorting loss, repeat the scoring until the preset maximum number of iterations is reached, and take the first 10 learning activities generated by the last iteration as alternatives. Combined with the genetic algorithm, encode each learning activity into an activity sequence, based on the activity sequence, with the goal of maximizing ability gain and minimizing emotional loss, construct an objective function in combination with pre-set knowledge prerequisite constraints, solve the objective function through a multi-objective optimization algorithm based on the genetic algorithm, generate a Pareto solution set and connect the Pareto solution set to generate an optimal learning path, perform attribute mapping on the optimal learning path, and obtain the optimized teaching plan.

[0038] In an optional embodiment,

[0039] The correct answer rate of the questions is calculated based on the students' ability level in each dimension as shown in the following formula:

[0040]

[0041] Among them, H() represents the probability of students answering correctly, X gj represents the answer of student g on question j, V g represents the ability vector of student g, a j represents the discrimination parameter of question j, b j represents the difficulty parameter of question j, c j represents the guess parameter of question j, q j It represents the weight matrix composed of the weights of knowledge components, and T represents transpose.

[0042] A second aspect of the embodiment of the present invention provides a VR-based multi-person foreign language situational teaching system for primary and secondary schools, including:

[0043] The first unit is used to obtain teaching tasks and determine the theme, scene, character role setting and interaction mode of the virtual reality scenario based on the teaching tasks, compress the data and optimize the loading of the scene through octree space division and voxel modeling technology, build an immersive virtual teaching scene in combination with a physics-based rendering method and a real-time lighting algorithm, integrate real teaching environment elements into the immersive virtual teaching scene through intelligent object recognition and real-scene modeling algorithm, set situational dialogue tasks and an adaptive difficulty adjustment mechanism in the immersive virtual teaching scene according to the teaching tasks, and combine knowledge graphs and ontology reasoning technology to semantically describe and associate elements in the immersive virtual teaching scene to form a multi-level semantic network;

[0044] The second unit is used for each student to log in to the multi-person teaching platform through a VR all-in-one machine, set a virtual image and enter the immersive virtual teaching scene, participate in the situational dialogue task in the immersive virtual teaching scene, collect the multimodal interaction data corresponding to the current student through a multi-channel sensor array and transmit it back to the edge computing server, learn the multimodal interaction data through a video behavior understanding model, generate multi-channel spatiotemporal features and obtain quantitative analysis results, perform speech denoising and feature extraction on the video data collected by the multi-channel sensor array, perform contextual language modeling in combination with the attention mechanism, obtain the student's native language speech, translate the student's native language speech through a speech translation model based on deep migration, generate an initial foreign language text and filter it through a multilingual bidirectional encoder model to obtain a high-quality dialogue text, and generate personalized dialogue feedback based on the high-quality dialogue text, the quantitative analysis results and the multimodal interaction data through an intelligent agent reinforcement learning algorithm;

[0045] The third unit is used to conduct a multi-dimensional quantitative assessment of each student's language application ability based on the personalized dialogue feedback and the multi-channel spatiotemporal features in combination with a hierarchical transfer learning algorithm, generate an ability level portrait, track the development level of each student's skill dimension based on the knowledge component theory and the multidimensional item response theory, determine the degree of mastery of each student's knowledge points in combination with a Bayesian knowledge tracking model, generate a knowledge mastery map, integrate the multi-channel spatiotemporal features and combine them with a long short-term memory network for feature learning to obtain learning behavior characteristics, summarize the ability level portrait, the knowledge mastery map and the learning behavior characteristics, construct a language learning path optimization model based on a factorization machine and combinatorial ranking learning, obtain the optimal learning path in combination with a multi-objective optimization algorithm, and generate an optimized teaching plan based on the optimal learning path.

[0046] According to a third aspect of the embodiments of the present invention,

[0047] An electronic device is provided, comprising:

[0048] processor;

[0049] a memory for storing processor-executable instructions;

[0050] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0051] According to a fourth aspect of the embodiments of the present invention,

[0052] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0053] In the present invention, data compression and loading optimization are performed on the virtual scene through octree space division and voxel modeling technology, which improves the efficiency of scene loading and rendering and reduces delay. Through intelligent object recognition and real-scene modeling algorithms, elements of the real teaching environment are seamlessly integrated into the virtual scene, providing a more realistic interactive experience. Knowledge graphs and ontology reasoning technologies are used to semantically describe and associate elements in the virtual teaching scene to form a multi-level semantic network, providing a semantically rich interactive environment, which helps students better understand and master the teaching content. A language learning path optimization model is constructed based on factor decomposition machines and combinatorial sorting learning, and the optimal learning path is solved in combination with a multi-objective optimization algorithm. An optimized teaching plan is generated to ensure that each student learns on the optimal path, thereby improving learning efficiency and effect. In summary, the present invention organically combines virtual reality technology with intelligent algorithms, provides an efficient, personalized and highly interactive solution for modern education, ensures that each student can learn on the path that best suits him or her, and improves learning effect and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is a flow chart of a VR-based multi-person foreign language situational teaching method for primary and secondary schools according to an embodiment of the present invention;

[0055] Figure 2 This is a schematic diagram of the structure of a VR-based multi-person foreign language situational teaching system for primary and secondary schools according to an embodiment of the present invention. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0057] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0058] Figure 1 FIG. 1 is a flow chart of a method for teaching a foreign language in a VR-based context for multiple students in primary and secondary schools according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0059] S1. Obtain teaching tasks and determine the theme, scene, character role setting and interaction mode of the virtual reality scenario based on the teaching tasks, compress the data and optimize the loading of the scene through octree space division and voxel modeling technology, build an immersive virtual teaching scene in combination with a physics-based rendering method and a real-time lighting algorithm, integrate real teaching environment elements into the immersive virtual teaching scene through intelligent object recognition and real-scene modeling algorithms, set scenario dialogue tasks and an adaptive difficulty adjustment mechanism in the immersive virtual teaching scene according to the teaching tasks, and combine knowledge graphs and ontology reasoning technologies to semantically describe and associate elements in the immersive virtual teaching scene to form a multi-level semantic network;

[0060] The virtual reality scenario refers to a three-dimensional environment generated by a computer, in which the user is immersed and able to interact with it. The octree space partitioning is a data structure that recursively partitions the three-dimensional space. The voxel modeling technology uses voxels to represent objects in the three-dimensional space. Each voxel is a small cube, and the entire three-dimensional model is composed of many voxels. The real-time lighting algorithm dynamically calculates the interaction effect between light and objects during the rendering process. The immersive virtual teaching scene uses virtual reality technology to create a highly interactive and immersive learning environment. The adaptive difficulty adjustment mechanism dynamically adjusts the difficulty of the task according to the user's performance to maintain the user's sense of challenge and interest. The ontology reasoning technology uses ontology to represent concepts and relationships in the knowledge field, and obtains new knowledge or verifies the correctness of existing knowledge through logical reasoning. The association mapping is the process of establishing connections and mappings between different data sets or knowledge bases. The multi-level semantic network is a knowledge representation method that describes concepts and their interactions through multiple levels and relationships.

[0061] In an optional embodiment,

[0062] The teaching task is obtained and the theme, scene, character role setting and interaction mode of the virtual reality scenario are determined based on the teaching task. The scene is compressed and optimized for loading through octree space division and voxel modeling technology. An immersive virtual teaching scene is constructed by combining a physics-based rendering method and a real-time lighting algorithm. The real teaching environment elements are integrated into the immersive virtual teaching scene through intelligent object recognition and real-scene modeling algorithms. According to the teaching task, a scenario dialogue task and an adaptive difficulty adjustment mechanism are set in the immersive virtual teaching scene. The elements in the immersive virtual teaching scene are semantically described and associated mapped in combination with knowledge graphs and ontology reasoning technology to form a multi-level semantic network including:

[0063] The current learning stage and teaching tasks are obtained from the teaching management system, and the teaching tasks are semantically understood and key information is extracted through natural language processing technology to obtain a structured task description. Based on the structured task description, the teaching content is semantically modeled and associated with mining through ontology reasoning technology to identify the core semantic elements. Based on the hierarchical relationship and logical association between different core semantic elements, the themes and scenes with the highest relevance to the teaching tasks are determined. For each core semantic element, the corresponding element attributes are determined through a knowledge-driven method to obtain a character role setting, and the virtual character dialogue and behavior perception are simulated through speech synthesis technology and gaze tracking technology to set the interaction mode.

[0064] The scene corresponding to each virtual reality scenario is represented in multiple resolutions and processed in detail levels by using the octree-based space division and voxel modeling technology, the number of division levels and node size of the octree are adaptively determined, the scene model is converted into a hierarchical network composed of voxel units by a voxelization algorithm, the detail levels of the nodes in the octree are adjusted in combination with the viewpoint position and the viewing direction, the number of grids is clipped and detail level hierarchical model data of different areas in the cone is generated, the scene map and animation sequence are query-oriented data sharding is performed through a distributed data organization and a multi-level cache mechanism, and redundant storage is performed for loading optimization by setting a multi-level cache architecture;

[0065] Based on the scene after data compression and loading optimization, the lighting, material and reflection properties of the virtual reality scenario are modeled by a physically based rendering method, the ambient lighting is rendered in combination with a pre-calculated radiosity transfer technology, the dynamic light source is smoothly interpolated by a spherical harmonic function and a vertex shader, the height field is modeled by a normal map and a displacement map, a bidirectional scattering distribution function is selected in combination with the reflection characteristics of different materials, and a rendering equation is solved to obtain rendering parameters, and the immersive virtual teaching scene is constructed according to the rendering parameters;

[0066] Through intelligent object recognition algorithm and real scene modeling algorithm, combined with the camera array pre-set in the teaching venue, the teaching environment is determined, and the scene semantics and spatial structure are determined through semantic segmentation and 3D reconstruction algorithm based on deep learning. In combination with plane detection, objects in the real teaching environment are determined, and a 3D model corresponding to each object is generated. The 3D model corresponding to each object is superimposed and fused with the virtual elements in the immersive virtual teaching scene;

[0067] According to the language level of the students and the teaching tasks, scenario dialogue tasks including multiple levels are set in the immersive virtual teaching scene. In each level of dialogue tasks, the students' dialogue performance is used as the environmental state of reinforcement learning, the difficulty parameter of the dialogue task is used as the action space, a reward function is set and the difficulty parameter is optimized, the difficulty of the scenario dialogue task is adaptively adjusted, and the multimodal information in the current scenario is obtained in combination with the knowledge graph technology, a multimodal semantic representation is generated and association mining is performed, the scene, the objects and characters in the real teaching environment are taken as core nodes, triples are constructed through relationship extraction, and a multi-relation heterogeneous knowledge graph is obtained, the multi-level semantic associations are determined through a graph neural network and a tight cluster is generated, and semantic mapping is performed in combination with the association mapping technology to obtain the multi-level semantic network.

[0068] The key information extraction is the process of automatically extracting useful information from a large amount of text data. The structured task description is to describe the various steps and requirements of the task in a standardized and systematic way to make it clear, easy to understand and execute. The knowledge-driven method is to use existing knowledge and experience to guide problem solving. The gaze tracking technology is used to detect and record the gaze point position of the user's eyes. The detail level layered processing is a technology that displays or processes objects in layers according to different detail levels. The visual cone is a geometric body that represents the field of view, which is used to determine which objects need to be rendered within the field of view. The pre-calculated radiosity transfer technology pre-calculates the radiosity transfer of light during the rendering process. The invention discloses a method for pre-calculating lighting information, which can reduce the cost of real-time calculation, and realizes fast lighting effect in real-time rendering by pre-calculating lighting information and storing it in texture or other data structures. The spherical harmonics are a set of orthogonal functions defined on a sphere, which are used to represent and process spherical data in three-dimensional space. The normal map is a texture mapping technology, which is used to add details to the surface of a three-dimensional model. The bidirectional scattering distribution function describes the scattering behavior of light on the surface and defines the relationship between the incident and emitted light. The triplet is a data representation form composed of three elements, which usually represents the relationship between entities. The tight cluster refers to a group of data points that are close to each other and have a high similarity in cluster analysis.

[0069] Obtain information about the current learning stage and teaching tasks from the teaching management system, use natural language processing techniques, such as named entity recognition and keyword extraction, to understand the teaching tasks semantically, extract key information, and generate structured task descriptions. Based on the structured task descriptions, use ontology reasoning technology to semantically model the teaching content, build a teaching content ontology, define concepts and relationships, identify the core semantic elements in the teaching content through ontology reasoning, analyze the hierarchical relationships and logical connections between different semantic elements, calculate their relevance to the teaching tasks, and select the most relevant topics and scenes as the entry point for virtual teaching;

[0070] For each core semantic element, use knowledge-driven methods to determine its attributes, such as the character's personality traits, behavioral habits, etc., to form a complete character setting. Use speech synthesis technology, such as parametric speech synthesis and splicing speech synthesis, to generate the virtual character's dialogue speech. Control the rhythm and emotion of the speech according to the character setting. Use gaze tracking technology to simulate the virtual character's attention distribution through the movement of its line of sight. Combine facial expressions, gestures and other actions to improve the realism of the character's behavior. Design human-computer interaction methods, such as speech recognition, gesture recognition, somatosensory interaction, etc., so that students can communicate naturally with virtual characters.

[0071] For each virtual teaching scenario, the octree algorithm is used to divide the scene into spaces. The number of octree division layers and the size of each node are adaptively determined according to the scene complexity. The scene represented by the octree is converted into a regular voxel grid using the voxelization algorithm. According to the viewpoint position and observation direction of the character, the detail level of the node is adaptively adjusted. The detail level of different areas within the cone is divided, and invisible grid patches are cropped to generate scene model data with multiple detail levels. The distributed data organization method is used to slice the scene's textures, animations, etc., to improve the data query and loading speed, and a multi-level cache framework is built to perform redundant caching of frequently accessed data.

[0072] Using physically based rendering methods, the scene's lighting, material, and reflection properties are modeled. Precomputed radiosity transfer technology is used to calculate ambient lighting and bake the results into light maps. For dynamic light sources, spherical harmonics are used to perform real-time calculation and smooth interpolation of irradiance in the vertex shader. Normal maps and displacement maps are used to model the details of the object surface to improve the realism of the surface. For different materials, appropriate bidirectional scattering distribution functions (BRDFs) are selected to represent their reflection characteristics, and the rendering equations are solved to calculate the final rendering results.

[0073] Arrange multiple cameras in the teaching venue, use intelligent object recognition algorithm to detect objects in the scene, segment the camera video stream through the semantic segmentation algorithm based on deep learning, identify the semantic information of the scene, use 3D reconstruction algorithms such as multi-view stereo matching to restore the 3D structure of the real scene, detect the main planes in the real environment (such as the ground, walls, etc.), use this as a constraint to reconstruct the 3D object, generate a mesh model of the object, integrate the reconstructed real object model into the virtual teaching scene, and realize the fusion of virtual and real through coordinate transformation and occlusion processing. Set up multiple levels of situational dialogue tasks according to students' language level and teaching tasks, such as vocabulary exercises, sentence pattern exercises, role-playing, etc., and use students' dialogue performance (such as dialogue scores, completion time, etc.) as reinforcement learning. The environment state of the learning is taken into consideration, and the parameters of dialogue tasks of different difficulty levels (such as vocabulary difficulty, grammatical complexity, etc.) are used as the action space. The reward function is designed to balance the fluency of the dialogue and the effect of learning. The difficulty coefficient of the dialogue task is adaptively adjusted through reinforcement learning algorithms such as policy gradient. The multimodal information in the current dialogue scenario is extracted using knowledge graph technology, and integrated into a consistent semantic representation. With scenes, objects, and characters as central nodes, the relationship extraction technology is used to construct the association between multimodal information, forming a multi-relational heterogeneous knowledge graph. The graph neural network is used for reasoning to discover the multi-layer semantic associations between different concepts, generate semantically tight concept clusters, and use the association mapping technology to map the semantic clusters to a hierarchical semantic network as the basis for situational dialogue generation.

[0074] In this embodiment, the intelligent content generation method can reduce the teacher's lesson preparation pressure and improve the pertinence and relevance of teaching content. Data compression and multi-level caching mechanisms can improve the loading and rendering speed of scenes and ensure the smoothness of interaction. The real teaching environment is reconstructed in three dimensions through computer vision technology and synthesized with virtual elements in real time, which can achieve seamless connection between virtual objects and real environment and improve immersion and realism. By constructing a multi-relation heterogeneous knowledge graph and mining the semantic associations between scenes, objects and characters, rich semantic support can be provided for situational dialogues. Organizing semantic associations into a hierarchical semantic network can simulate the memory structure of the human brain and make knowledge storage and retrieval more efficient. In summary, this embodiment realizes the intelligent generation of teaching content, realistic rendering of virtual scenes, natural integration of human-computer interaction and personalized adaptation of teaching activities, which can greatly improve students' learning interest and efficiency and promote teaching students in accordance with their aptitude and autonomous learning.

[0075] In an optional embodiment,

[0076] The smooth interpolation of dynamic light sources through spherical harmonics and vertex shaders is shown in the following formula:

[0077]

[0078] Among them, P(x) represents the illumination value calculated at vertex x, l represents the index of the order of the spherical harmonic function basis, L represents the order of the spherical harmonic function, and m represents the degree of the spherical harmonic function. Indicates t 1 Spherical harmonic coefficients of the time frame, Indicates t 2 Spherical harmonic coefficients of the time frame, Y lm (θ, φ) represents the spherical harmonic function basis, θ represents the elevation angle in the spherical coordinate system, φ represents the azimuth angle in the spherical coordinate system, and t represents the current time.

[0079] In this embodiment, by calculating the spherical harmonics in the vertex shader, real-time rendering of dynamic light sources can be achieved, which can greatly reduce the amount of calculation and improve rendering efficiency. By linearly interpolating the spherical harmonic coefficients between adjacent frames, a smooth transition of the dynamic light source in time can be achieved, which can reduce jumps and flickers when the lighting changes and improve the continuity of the picture. Compared with storing pre-calculated lighting pixel by pixel, the storage amount of spherical harmonic coefficients is much smaller, and the spherical harmonics can use fewer coefficients to represent spherical functions. Therefore, the amount of pre-calculated spherical harmonic coefficients is small, reducing storage overhead. In summary, this embodiment can provide a more realistic and coherent dynamic lighting effect while ensuring rendering performance.

[0080] S2. Each student logs in to the multi-person teaching platform through a VR all-in-one machine, sets a virtual image and enters the immersive virtual teaching scene, participates in the situational dialogue task in the immersive virtual teaching scene, collects the multimodal interaction data corresponding to the current student through a multi-channel sensor array and transmits it back to the edge computing server, learns the multimodal interaction data through a video behavior understanding model, generates multi-channel spatiotemporal features and obtains quantitative analysis results, performs speech denoising and feature extraction on the video data collected by the multi-channel sensor array, performs contextual language modeling in combination with the attention mechanism, obtains the student's native language speech, translates the student's native language speech through a speech translation model based on deep migration, generates an initial foreign language text and screens it through a multilingual bidirectional encoder model to obtain a high-quality dialogue text, and generates personalized dialogue feedback based on the high-quality dialogue text, the quantitative analysis results and the multimodal interaction data through an intelligent agent reinforcement learning algorithm;

[0081] The multi-channel sensor array is a system composed of multiple sensors, each sensor is responsible for collecting signals of different types or at different locations. The multimodal interaction data refers to interaction information obtained through multiple perception modes (such as vision, hearing, touch, etc.). The contextual language modeling is a technology for generating or understanding natural language by considering contextual information. The deep transfer-based speech translation model uses deep learning and transfer learning technology to directly translate the speech signal of one language into text or speech of another language. The multilingual bidirectional encoder model is a deep learning model that can process multiple languages. It captures the bidirectional dependency of languages ​​by encoding and decoding bidirectional information, and supports multilingual understanding and generation tasks.

[0082] In an optional embodiment,

[0083] Each student logs in to the multi-person teaching platform through a VR all-in-one machine, sets a virtual image and enters the immersive virtual teaching scene, participates in the situational dialogue task in the immersive virtual teaching scene, collects the multimodal interaction data corresponding to the current student through a multi-channel sensor array and transmits it back to the edge computing server, learns the multimodal interaction data through a video behavior understanding model, generates multi-channel spatiotemporal features and obtains quantitative analysis results, performs speech denoising and feature extraction on the video data collected by the multi-channel sensor array, combines the attention mechanism to perform contextual language modeling, obtains the student's native language speech, translates the student's native language speech through a speech translation model based on deep migration, generates an initial foreign language text and screens it through a multilingual bidirectional encoder model to obtain a high-quality dialogue text, and generates personalized dialogue feedback based on the high-quality dialogue text, the quantitative analysis results and the multimodal interaction data through an intelligent agent reinforcement learning algorithm, including:

[0084] Each student logs in to the multi-person teaching platform through a VR all-in-one machine, sets a virtual image and enters the immersive virtual teaching scene, and participates in a pre-set scenario dialogue task in the immersive virtual teaching scene. The multi-channel sensor array collects each student's facial expression and body movement through a video sensor, records each student's voice signal through a microphone array and locates the sound source, obtains the student's hand and head movement data through an inertial measurement unit and an optical motion capture system, records the student's gaze point and gaze time through an eye tracker, determines the focus of each student in the current scene, aligns the data obtained by the multi-channel sensor with timestamps, and performs multimodal data fusion in frames to obtain the multimodal interaction data;

[0085] The multimodal interaction data is added to a pre-set video behavior understanding model, and features are extracted through a pre-trained convolutional neural network to generate a frame-level deep representation and send it to a long short-term memory network, determine the temporal dependency of different frame-level deep representations and generate a hidden state vector, perform attention weighting on the hidden state vector through an attention pooling mechanism and output it in a classified manner, generate a probability distribution of each behavior category and corresponding multi-channel spatiotemporal features, construct a quantitative indicator, and quantitatively analyze the multimodal interaction data based on the probability distribution corresponding to each behavior category to obtain the quantitative analysis result;

[0086] Perform speech noise reduction and feature extraction on the video data recorded by the video sensor and the speech signal recorded by the microphone array, generate a feature sequence and extract features through a recurrent convolutional network, obtain high-order local features and perform sequence modeling, obtain a high-order local sequence, initialize the high-order local features corresponding to the decoder and update them at each decoding moment, generate a context vector in combination with the attention weight, process it through a fully connected layer and a soft normalization function, obtain the probability distribution of the native language words at the current moment, combine the cross entropy loss function and use the real native language words as input for reasoning, and obtain the student's native language speech;

[0087] Based on the native language speech of the student, the semantic mapping relationship between the native language and the foreign language is determined by a speech translation model based on deep transfer, the speech translation model is trained in a pre-set bilingual corpus and the semantic mapping relationship is verified, the hyperparameters in the speech translation model are adjusted in combination with the verification result, the student's native language speech is translated by the adjusted speech translation model, a foreign language text sequence is generated and optimized by a beam search algorithm to obtain an initial foreign language text, the initial foreign language text is encoded by a bidirectional encoder in a multilingual bidirectional encoder model, the encoding result is retrieved in a pre-constructed multilingual corpus, the similarity between each element in the multilingual corpus and the encoding result is calculated, similar dialogues are determined based on the similarity and a candidate set is generated, the elements in the candidate set are scored according to language fluency and relevance, a dialogue score corresponding to each element is obtained, and the similar dialogue with the highest dialogue score is selected as the high-quality dialogue text;

[0088] Based on the high-quality dialogue text, the quantitative analysis results and the multimodal interaction data, the effectiveness of the high-quality dialogue text in the current scenario is determined through an intelligent agent reinforcement algorithm, and feedback optimization is performed according to the quantitative analysis results and the multimodal interaction data to generate the personalized dialogue feedback.

[0089] The inertial measurement unit is an electronic device, usually including an accelerometer, a gyroscope and a magnetometer, used to measure the acceleration, angular velocity and magnetic field strength of an object. The frame-level deep representation refers to generating a high-level feature representation of a deep learning model for each frame or time segment when processing sequence data (such as video, voice). The time dependency describes the dependency and correlation of data in the time series. The bilingual corpus is a data set containing corresponding text pairs in two languages, which is used for tasks such as machine translation, cross-language information retrieval and language learning. The beam search algorithm is a heuristic search algorithm used to find an approximate optimal solution in a larger search space. The bidirectional encoder is a neural network structure that can simultaneously process the forward and backward information of the input sequence.

[0090] Each student uses a VR all-in-one machine to log in to the multi-person teaching platform. Students set their own virtual images in the platform, including personalized features such as appearance and clothing. Students enter pre-designed immersive virtual teaching scenes, such as virtual classrooms and laboratories. Students participate in pre-set situational dialogue tasks in virtual teaching scenes, such as role-playing and group discussions. The multi-channel sensor array collects each student's facial expressions and body movements through video sensors, and the microphone array records each student's voice signal and locates the sound source. The inertial measurement unit and optical motion capture system obtain the student's hand and head movement data. The eye tracker records the student's gaze point and gaze time to determine the focus of each student in the current scene. The data obtained by the multi-channel sensors are time-stamped and fused with multimodal data in frames to obtain multimodal interaction data.

[0091] Add multimodal interaction data to a pre-set video behavior understanding model, extract features through a pre-trained convolutional neural network, generate frame-level deep representations, send the frame-level deep representations to a long short-term memory network, determine the temporal dependencies between different frames, generate hidden state vectors, perform attention weighting on the hidden state vectors through an attention pooling mechanism, perform classification output, generate probability distributions for each behavior category and corresponding multi-channel spatiotemporal features, construct quantitative indicators, and quantitatively analyze multimodal interaction data based on the probability distribution corresponding to each behavior category to obtain quantitative analysis results;

[0092] Perform speech noise reduction and feature extraction on the video data recorded by the video sensor and the speech signal recorded by the microphone array to generate a feature sequence, perform feature extraction through a recurrent convolutional network to obtain high-order local features, perform sequence modeling on the high-order local features to obtain a high-order local sequence, initialize the high-order local features corresponding to the decoder, and update them at each decoding moment, generate a context vector in combination with the attention weight, process it through a fully connected layer and a soft normalization function to obtain the probability distribution of the native language words at the current moment, combine the cross entropy loss function and use the real native language words as input for reasoning to obtain the student's native language speech, based on the student's native language speech, determine the semantic mapping relationship between the native language and the foreign language through a speech translation model based on deep transfer, train the speech translation model on a pre-set bilingual corpus, and verify the semantic mapping relationship, adjust the hyperparameters in the speech translation model based on the verification results, use the adjusted speech translation model to translate the student's native language speech, generate a foreign language text sequence, optimize the foreign language text sequence through a beam search algorithm to obtain an initial foreign language text;

[0093] The initial foreign language text is encoded by the bidirectional encoder in the multilingual bidirectional encoder model, and the encoding result is searched in the pre-built multilingual corpus, and the similarity between each element in the corpus and the encoding result is calculated. Based on the similarity, similar dialogues are determined and a candidate set is generated. Similar dialogues with the highest dialogue scores are selected as high-quality dialogue texts. Based on the high-quality dialogue texts, quantitative analysis results and multimodal interaction data, the effectiveness of the high-quality dialogue texts in the current scenario is determined by the intelligent agent reinforcement algorithm. According to the quantitative analysis results and multimodal interaction data, the intelligent agent is optimized for feedback, and personalized dialogue feedback is generated as a reference and guidance for students in the current scenario dialogue task.

[0094] In this embodiment, multi-modal data such as students' facial expressions, body movements, voices, and sight lines are collected through a multi-channel sensor array, and timestamp alignment and data fusion are performed to provide rich data support for subsequent behavior analysis and personalized feedback. Intelligent behavior analysis methods can help teachers understand students' learning status and behavior patterns more comprehensively. Mother tongue-assisted learning methods can help students better understand and express foreign languages. The dialogue generation method based on a large-scale corpus can provide students with rich and diverse dialogue materials and expand students' language expression ability. Through the intelligent agent enhancement algorithm, combined with high-quality dialogue texts, behavior analysis results and multimodal interaction data, personalized dialogue feedback is generated. Targeted guidance and help can be provided according to the characteristics and needs of each student to improve learning efficiency. In summary, this embodiment realizes the immersion of the learning environment, the intelligence of the interaction method, the personalization of the learning content and the visualization of the learning process, providing students with a more efficient, interesting and personalized learning experience, and at the same time providing teachers with more comprehensive, objective and explainable teaching decision support.

[0095] In an optional embodiment,

[0096] The similarity between each element in the multilingual corpus and the encoding result is calculated as shown in the following formula:

[0097]

[0098] Among them, h i represents the candidate translation, s(h i ) represents the candidate translation h i Similarity with the encoding result, λ 1 represents the probability weight of the language model, p(h i ) represents the candidate translation h i The language model probability, |h i | represents the length of the candidate translation, a represents the length penalty strength, λ 2 represents the coverage weight, n-grams (h i ) represents the candidate translation h i n-grams(f) represents the n-grams of the source language sentence, |f| represents the length of the source language sentence, λ 3 Represents the length ratio weight.

[0099] In this embodiment, by comprehensively considering factors such as language model probability, length penalty, coverage, length ratio, etc., fluent, accurate, and concise translations can be selected to improve the generation quality of high-quality dialogue texts. The coverage factor is introduced to encourage the algorithm to select translations with diverse expressions, avoid overly single or repetitive expressions, and increase the richness of high-quality dialogue texts. Through the length penalty and length ratio factors, the scores of overly long or short translations are suppressed, and the algorithm is prompted to select translations with moderate length and sufficient information. In summary, this embodiment can achieve good results in improving translation quality, enhancing translation diversity, balancing translation length, etc., provides strong support for the generation of high-quality dialogue texts, has high flexibility and computational efficiency, and can adapt to different language pairs and application scenarios.

[0100] S3. Based on the personalized dialogue feedback and the multi-channel spatiotemporal features, combined with the hierarchical transfer learning algorithm, a multi-dimensional quantitative assessment of each student's language application ability is performed to generate an ability level portrait. Based on the knowledge component theory and the multidimensional item response theory, the development level of each student's skill dimension is tracked, and the Bayesian knowledge tracking model is combined to determine the degree of mastery of each student's knowledge points, and a knowledge mastery map is generated. The multi-channel spatiotemporal features are integrated and combined with the long short-term memory network for feature learning to obtain learning behavior characteristics, and the ability level portrait, the knowledge mastery map and the learning behavior characteristics are summarized. A language learning path optimization model is constructed based on a factorization machine and combinatorial ranking learning, and the optimal learning path is obtained by combining a multi-objective optimization algorithm to solve, and an optimized teaching plan is generated based on the optimal learning path.

[0101] The hierarchical transfer learning algorithm optimizes the performance of the model on the target task by gradually transferring knowledge in a multi-level structure. The ability level portrait is a systematic and visual representation of individual ability. The knowledge component theory proposes that knowledge can be decomposed into several basic units or components. Components are the basis for learning and applying knowledge. The multidimensional item response theory is a model for evaluating individual abilities in multiple dimensions. By analyzing the answers of individuals on different questions, their ability levels in various knowledge dimensions are estimated. The Bayesian knowledge tracking model uses the Bayesian method to track and predict the learner's knowledge mastery, and estimates their mastery of each knowledge point based on the learner's historical performance and current learning activities. probability, and dynamically adjust the teaching content. The knowledge mastery map is a visualization tool that shows the learner's mastery of each knowledge point. The learning behavior feature is a feature that analyzes and extracts the learner's behavior data during the learning process. The factor decomposition machine is a machine learning model for prediction and recommendation, which can effectively process sparse data and high-dimensional features. The combined ranking learning is a machine learning method for optimizing the results of multiple ranking tasks at the same time. The language learning path optimization model formulates the optimal learning path by analyzing the learner's language ability and learning behavior. The optimal learning path refers to the most effective learning sequence and strategy provided to learners through optimization algorithms and models.

[0102] In an optional embodiment,

[0103] Based on the personalized dialogue feedback and the multimodal interaction data, the language application ability of each student is quantitatively evaluated in multiple dimensions in combination with the hierarchical transfer learning algorithm to generate an ability level portrait. The development level of each student's skill dimension is tracked based on the knowledge component theory and the multidimensional item response theory. The degree of mastery of each student's knowledge points is determined in combination with the Bayesian knowledge tracking model to generate a knowledge mastery map. The multi-channel spatiotemporal features are integrated and feature learning is performed in combination with the long short-term memory network to obtain learning behavior features. The ability level portrait, the knowledge mastery map and the learning behavior features are summarized. A language learning path optimization model is constructed based on a factorization machine and combinatorial ranking learning. The optimal learning path is obtained by solving the multi-objective optimization algorithm. The optimized teaching plan based on the optimal learning path includes:

[0104] Obtain language proficiency test questions from a corpus and annotate knowledge points, add the personalized dialogue feedback and the multi-channel spatiotemporal features to a multi-layer perception mechanism, generate a high-dimensional feature representation and add it to the hierarchical transfer learning algorithm, set a feature extractor for each language proficiency dimension, extract the language proficiency representation corresponding to each modality and map it to a shared semantic space, perform a multi-dimensional quantitative assessment of each student's language application ability based on the mapping results, and splice the mapping results to obtain the ability level portrait;

[0105] Based on the ability level portrait, combined with the knowledge component theory and the multidimensional item response theory, the language test questions are divided into multiple knowledge components and the knowledge component examination weights corresponding to each test question are determined, the knowledge component examination weights and answer records are added to the multidimensional item response theory model, the student's ability level in each dimension is estimated, the correct answer rate of the question is calculated based on the student's ability level in each dimension, the multidimensional item response theory model is optimized in combination with the cross entropy loss function and the real answer record, and the optimization is repeated until the preset maximum number of iterations is reached, and a dynamic evolution sequence corresponding to the ability level is generated. Through the Bayesian model and the dynamic evolution sequence, the current student's mastery of each knowledge point is generated, and the relationship between the knowledge points is combined to generate a knowledge mastery map;

[0106] The multi-channel spatiotemporal features are represented as three-dimensional tensors corresponding to time steps, channel features and original features respectively, and local spatiotemporal features are extracted by combining a three-dimensional convolutional neural network pre-set along the channel dimension and flattened in the channel dimension, and long-term dependency modeling is performed by combining a two-layer long short-term memory network to obtain the learning behavior features;

[0107] Summarize the ability level portrait, the knowledge mastery map and the learning behavior characteristics, score the match between students and learning activities based on a factor decomposition machine, optimize the factor decomposition machine in combination with minimizing the sorting loss, repeat the scoring until the preset maximum number of iterations is reached, and take the first 10 learning activities generated by the last iteration as alternatives. Combined with the genetic algorithm, encode each learning activity into an activity sequence, based on the activity sequence, with the goal of maximizing ability gain and minimizing emotional loss, construct an objective function in combination with pre-set knowledge prerequisite constraints, solve the objective function through a multi-objective optimization algorithm based on the genetic algorithm, generate a Pareto solution set and connect the Pareto solution set to generate an optimal learning path, perform attribute mapping on the optimal learning path, and obtain the optimized teaching plan.

[0108] The multi-layer perception mechanism refers to the gradual extraction and processing of data features through a multi-layer structure in a neural network. The high-dimensional feature representation refers to mapping data into a high-dimensional space in order to better capture the intrinsic structure and mutual relationship of the data. The shared semantic space refers to mapping different types of data (such as text, images, voice, etc.) into a common semantic space so that they are comparable and operable in the space. The knowledge component examination weight represents the degree of examination of each knowledge dimension by different questions in the multidimensional item response theory. The cross entropy loss function is a loss function for classification tasks that measures the difference between the predicted probability distribution and the true label distribution. The dynamically evolving sequence refers to sequence data that changes over time. In data analysis and In modeling, it is necessary to capture its dynamic characteristics and change patterns. The three-dimensional tensor is a mathematical object containing a three-dimensional array, which is used to represent complex data structures. The local spatiotemporal features refer to data features extracted within a specific time and space range. The ability gain refers to the improvement of an individual's ability in a certain knowledge or skill through learning and training. The emotional loss refers to the decline in performance due to emotional fluctuations or instability. The knowledge prerequisite constraint refers to the basic knowledge that must be mastered before learning new knowledge or skills. The Pareto solution set is a solution set in which there is no other solution that can improve all objectives simultaneously in a multi-objective optimization problem. The attribute mapping refers to mapping the attributes of an object or entity to the corresponding attributes of another object or entity.

[0109] Obtain language proficiency test questions from the corpus, annotate the questions with knowledge points, add personalized dialogue feedback and multi-channel spatiotemporal features to the multi-layer perception mechanism, generate high-dimensional feature representations, add the high-dimensional feature representations to the hierarchical transfer learning algorithm, set up feature extractors for each language proficiency dimension, extract the language proficiency representation corresponding to each modality, map the extracted language proficiency representations to the shared semantic space, conduct multi-dimensional quantitative evaluation of each student's language application ability based on the mapping results, and splice the mapping results to obtain a proficiency level portrait;

[0110] Combining the knowledge component theory and the multidimensional item response theory, the language test questions are divided into multiple knowledge components, the knowledge component examination weights corresponding to each test question are determined, the knowledge component examination weights and answer records are added to the multidimensional item response theory model, the students' ability level in each dimension is estimated, the correct answer rate of the questions is calculated based on the students' ability level in each dimension, the multidimensional item response theory model is optimized by combining the cross entropy loss function and the real answer records, and the optimization is repeated until the preset maximum number of iterations is reached to generate a dynamic evolution sequence corresponding to the ability level. Through the Bayesian model and the dynamic evolution sequence, the current student's mastery of each knowledge point is generated, and the relationship between knowledge points is combined to generate a knowledge mastery map;

[0111] The multi-channel spatiotemporal features are represented as three-dimensional tensors, corresponding to time steps, channel features and original features respectively. The local spatiotemporal features are extracted by combining the three-dimensional convolutional neural network pre-set along the channel dimension, flattened in the channel dimension, and combined with the double-layer long short-term memory network for long-term dependency modeling to obtain learning behavior characteristics. The ability level portrait, knowledge mastery map and learning behavior characteristics are summarized. The matching degree between students and learning activities is scored based on the factorization machine. The factorization machine is optimized by minimizing the sorting loss. The scoring is repeated until the preset maximum number of iterations is reached. The first 10 learning activities generated by the last iteration are taken as alternatives. Each learning activity is encoded into an activity sequence by combining the genetic algorithm. Based on the activity sequence, the objective function is constructed with the ability gain and emotion loss as the goals and the pre-set knowledge prerequisite constraints. The objective function is solved by the multi-objective optimization algorithm based on the genetic algorithm to generate the Pareto solution set and connect the Pareto solution set to generate the optimal learning path. The optimal learning path is attribute mapped to obtain the optimized teaching plan.

[0112] In this embodiment, by annotating the knowledge points of the language proficiency test questions in the corpus, a basis is provided for the subsequent generation of the knowledge mastery map. The personalized dialogue feedback and multi-channel spatiotemporal features are fused by using the multi-layer perception mechanism and the hierarchical transfer learning algorithm to generate a high-dimensional feature representation, thereby improving the accuracy of the language proficiency assessment. The multidimensional item response theory model is used to estimate the students' ability level in each dimension, and a dynamic evolution sequence corresponding to the ability level is generated, thereby realizing the dynamic modeling of the students' language proficiency development process. The Bayesian model and the dynamic evolution sequence are used to generate the current student's mastery of each knowledge point, and the knowledge mastery map is generated by combining the relationship between the knowledge points. The spectrum provides an important basis for personalized teaching. The multi-channel spatiotemporal features are represented as three-dimensional tensors, and the local spatiotemporal features are extracted in combination with a three-dimensional convolutional neural network, which realizes the fine-grained representation of learning behavior. The factorization machine is optimized by minimizing the sorting loss, which improves the accuracy of learning activity recommendation. The learning activities are encoded into activity sequences using a genetic algorithm, and the ability gain and emotion loss are taken as goals. The objective function is constructed in combination with knowledge prerequisite constraints. The optimal learning path is generated through a multi-objective optimization algorithm, which realizes the intelligent planning of the learning path. In summary, this embodiment helps to improve the efficiency and quality of language learning and provide students with a more intelligent and personalized language learning experience.

[0113] In an optional embodiment,

[0114] The correct answer rate of the questions is calculated based on the students' ability level in each dimension as shown in the following formula:

[0115]

[0116] Among them, H() represents the probability of students answering correctly, Xgj represents the answer of student g on question j, V g represents the ability vector of student g, a j represents the discrimination parameter of question j, b j represents the difficulty parameter of question j, c j represents the guess parameter of question j, q j It represents the weight matrix composed of the weights of knowledge components, and T represents transpose.

[0117] In this embodiment, the matching degree between the student's ability and the question difficulty is calculated by the difference between the student's ability vector and the question difficulty parameter, so that the student's mastery of questions of different difficulty levels can be more accurately evaluated, providing a basis for personalized teaching. By introducing the question discrimination parameter, the quality of the question can be more accurately evaluated, and different weights can be given in the ability evaluation. The introduction of the guessing parameter can more reasonably evaluate the student's true ability and reduce the interference of guessing factors. In summary, this embodiment constructs a refined question answering accuracy calculation model, which can more accurately and fine-grainedly evaluate the student's mastery of different knowledge points, providing an important basis for personalized teaching, has good mathematical properties and stability, can improve the reliability of ability evaluation, and can provide more reliable data support for the generation of personalized teaching plans.

[0118] Figure 2 FIG. 1 is a schematic diagram of a structure of a VR-based multi-person foreign language situational teaching system for primary and secondary schools according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0119] The first unit is used to obtain teaching tasks and determine the theme, scene, character role setting and interaction mode of the virtual reality scenario based on the teaching tasks, compress the data and optimize the loading of the scene through octree space division and voxel modeling technology, build an immersive virtual teaching scene in combination with a physics-based rendering method and a real-time lighting algorithm, integrate real teaching environment elements into the immersive virtual teaching scene through intelligent object recognition and real-scene modeling algorithm, set situational dialogue tasks and an adaptive difficulty adjustment mechanism in the immersive virtual teaching scene according to the teaching tasks, and combine knowledge graphs and ontology reasoning technology to semantically describe and associate elements in the immersive virtual teaching scene to form a multi-level semantic network;

[0120] The second unit is used for each student to log in to the multi-person teaching platform through a VR all-in-one machine, set a virtual image and enter the immersive virtual teaching scene, participate in the situational dialogue task in the immersive virtual teaching scene, collect the multimodal interaction data corresponding to the current student through a multi-channel sensor array and transmit it back to the edge computing server, learn the multimodal interaction data through a video behavior understanding model, generate multi-channel spatiotemporal features and obtain quantitative analysis results, perform speech denoising and feature extraction on the video data collected by the multi-channel sensor array, perform contextual language modeling in combination with the attention mechanism, obtain the student's native language speech, translate the student's native language speech through a speech translation model based on deep migration, generate an initial foreign language text and filter it through a multilingual bidirectional encoder model to obtain a high-quality dialogue text, and generate personalized dialogue feedback based on the high-quality dialogue text, the quantitative analysis results and the multimodal interaction data through an intelligent agent reinforcement learning algorithm;

[0121] The third unit is used to conduct a multi-dimensional quantitative assessment of each student's language application ability based on the personalized dialogue feedback and the multi-channel spatiotemporal features in combination with a hierarchical transfer learning algorithm, generate an ability level portrait, track the development level of each student's skill dimension based on the knowledge component theory and the multidimensional item response theory, determine the degree of mastery of each student's knowledge points in combination with a Bayesian knowledge tracking model, generate a knowledge mastery map, integrate the multi-channel spatiotemporal features and combine them with a long short-term memory network for feature learning to obtain learning behavior characteristics, summarize the ability level portrait, the knowledge mastery map and the learning behavior characteristics, construct a language learning path optimization model based on a factorization machine and combinatorial ranking learning, obtain the optimal learning path in combination with a multi-objective optimization algorithm, and generate an optimized teaching plan based on the optimal learning path.

[0122] According to a third aspect of the embodiments of the present invention,

[0123] An electronic device is provided, comprising:

[0124] processor;

[0125] a memory for storing processor-executable instructions;

[0126] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0127] A fourth aspect of the embodiments of the present invention is:

[0128] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0129] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A VR-based multi-person foreign language situational teaching method for primary and secondary schools, characterized by: include: Acquire teaching tasks and determine the theme, scene, character role setting and interaction mode of the virtual reality scenario based on the teaching tasks, compress the data and optimize the loading of the scene through octree space division and voxel modeling technology, build an immersive virtual teaching scene in combination with a physics-based rendering method and a real-time lighting algorithm, integrate real teaching environment elements into the immersive virtual teaching scene through intelligent object recognition and real-scene modeling algorithms, set situational dialogue tasks and an adaptive difficulty adjustment mechanism in the immersive virtual teaching scene according to the teaching tasks, and combine knowledge graphs and ontology reasoning technologies to semantically describe and associate elements in the immersive virtual teaching scene to form a multi-level semantic network; Each student logs in to the multi-person teaching platform through a VR all-in-one machine, sets a virtual image and enters the immersive virtual teaching scene, participates in the situational dialogue task in the immersive virtual teaching scene, collects the multimodal interaction data corresponding to the current student through a multi-channel sensor array and transmits it back to the edge computing server, learns the multimodal interaction data through a video behavior understanding model, generates multi-channel spatiotemporal features and obtains quantitative analysis results, performs speech denoising and feature extraction on the video data collected by the multi-channel sensor array, performs contextual language modeling in combination with the attention mechanism, obtains the student's native language speech, translates the student's native language speech through a speech translation model based on deep migration, generates an initial foreign language text, and screens it through a multilingual bidirectional encoder model to obtain a high-quality dialogue text, and generates personalized dialogue feedback based on the high-quality dialogue text, the quantitative analysis results and the multimodal interaction data through an intelligent agent reinforcement learning algorithm; Based on the personalized dialogue feedback and the multi-channel spatiotemporal features, combined with the hierarchical transfer learning algorithm, a multi-dimensional quantitative assessment of the language application ability of each student is performed to generate an ability level portrait, and the development level of each student's skill dimension is tracked based on the knowledge component theory and the multidimensional item response theory. The Bayesian knowledge tracking model is used to determine the degree of mastery of each student's knowledge points and generate a knowledge mastery map. The multi-channel spatiotemporal features are integrated and combined with the long short-term memory network for feature learning to obtain learning behavior characteristics. The ability level portrait, the knowledge mastery map and the learning behavior characteristics are summarized. A language learning path optimization model is constructed based on a factorization machine and combinatorial ranking learning, and the optimal learning path is obtained by combining the multi-objective optimization algorithm. An optimized teaching plan is generated based on the optimal learning path.

2. The method according to claim 1, characterized in that: The teaching task is obtained and the theme, scene, character role setting and interaction mode of the virtual reality scenario are determined based on the teaching task. The scene is compressed and optimized for loading through octree space division and voxel modeling technology. An immersive virtual teaching scene is constructed by combining a physics-based rendering method and a real-time lighting algorithm. The real teaching environment elements are integrated into the immersive virtual teaching scene through intelligent object recognition and real-scene modeling algorithms. According to the teaching task, a scenario dialogue task and an adaptive difficulty adjustment mechanism are set in the immersive virtual teaching scene. The elements in the immersive virtual teaching scene are semantically described and associated mapped in combination with knowledge graphs and ontology reasoning technology to form a multi-level semantic network including: The current learning stage and teaching tasks are obtained from the teaching management system, and the teaching tasks are semantically understood and key information is extracted through natural language processing technology to obtain a structured task description. Based on the structured task description, the teaching content is semantically modeled and associated with mining through ontology reasoning technology to identify the core semantic elements. Based on the hierarchical relationship and logical association between different core semantic elements, the themes and scenes with the highest relevance to the teaching tasks are determined. For each core semantic element, the corresponding element attributes are determined through a knowledge-driven method to obtain a character role setting, and the virtual character dialogue and behavior perception are simulated through speech synthesis technology and gaze tracking technology to set the interaction mode. The scene corresponding to each virtual reality scenario is represented in multiple resolutions and processed in detail levels by using the octree-based space division and voxel modeling technology, the number of division levels and node size of the octree are adaptively determined, the scene model is converted into a hierarchical network composed of voxel units by a voxelization algorithm, the detail levels of the nodes in the octree are adjusted in combination with the viewpoint position and the viewing direction, the number of grids is clipped and detail level hierarchical model data of different areas in the cone is generated, the scene map and animation sequence are query-oriented data sharding is performed through a distributed data organization and a multi-level cache mechanism, and redundant storage is performed for loading optimization by setting a multi-level cache architecture; Based on the scene after data compression and loading optimization, the lighting, material and reflection properties of the virtual reality scenario are modeled by a physically based rendering method, the ambient lighting is rendered in combination with a pre-calculated radiosity transfer technology, the dynamic light source is smoothly interpolated by a spherical harmonic function and a vertex shader, the height field is modeled by a normal map and a displacement map, a bidirectional scattering distribution function is selected in combination with the reflection characteristics of different materials, and a rendering equation is solved to obtain rendering parameters, and the immersive virtual teaching scene is constructed according to the rendering parameters; Through intelligent object recognition algorithm and real scene modeling algorithm, combined with the camera array pre-set in the teaching venue, the teaching environment is determined, and the scene semantics and spatial structure are determined through semantic segmentation and 3D reconstruction algorithm based on deep learning. In combination with plane detection, objects in the real teaching environment are determined, and a 3D model corresponding to each object is generated. The 3D model corresponding to each object is superimposed and fused with the virtual elements in the immersive virtual teaching scene; According to the language level of the students and the teaching tasks, scenario dialogue tasks including multiple levels are set in the immersive virtual teaching scene. In each level of dialogue tasks, the students' dialogue performance is used as the environmental state of reinforcement learning, the difficulty parameter of the dialogue task is used as the action space, a reward function is set and the difficulty parameter is optimized, the difficulty of the scenario dialogue task is adaptively adjusted, and the multimodal information in the current scenario is obtained in combination with the knowledge graph technology, a multimodal semantic representation is generated and association mining is performed, the scene, the objects and characters in the real teaching environment are taken as core nodes, triples are constructed through relationship extraction, and a multi-relation heterogeneous knowledge graph is obtained, the multi-level semantic associations are determined through a graph neural network and a tight cluster is generated, and semantic mapping is performed in combination with the association mapping technology to obtain the multi-level semantic network.

3. The method according to claim 2, characterized in that The smooth interpolation of dynamic light sources through spherical harmonics and vertex shaders is shown in the following formula: Among them, P(x) represents the illumination value calculated at vertex x, l represents the index of the order of the spherical harmonic function basis, L represents the order of the spherical harmonic function, and m represents the degree of the spherical harmonic function. represents the spherical harmonic coefficients of the t1 time frame, represents the spherical harmonic coefficients of the t2 time frame, Y lm (θ, φ) represents the spherical harmonic function basis, θ represents the elevation angle in the spherical coordinate system, φ represents the azimuth angle in the spherical coordinate system, and t represents the current time.

4. The method according to claim 1, characterized in that: Each student logs in to the multi-person teaching platform through a VR all-in-one machine, sets a virtual image and enters the immersive virtual teaching scene, participates in the situational dialogue task in the immersive virtual teaching scene, collects the multimodal interaction data corresponding to the current student through a multi-channel sensor array and transmits it back to the edge computing server, learns the multimodal interaction data through a video behavior understanding model, generates multi-channel spatiotemporal features and obtains quantitative analysis results, performs speech denoising and feature extraction on the video data collected by the multi-channel sensor array, combines the attention mechanism to perform contextual language modeling, obtains the student's native language speech, translates the student's native language speech through a speech translation model based on deep migration, generates an initial foreign language text and screens it through a multilingual bidirectional encoder model to obtain a high-quality dialogue text, and generates personalized dialogue feedback based on the high-quality dialogue text, the quantitative analysis results and the multimodal interaction data through an intelligent agent reinforcement learning algorithm, including: Each student logs in to the multi-person teaching platform through a VR all-in-one machine, sets a virtual image and enters the immersive virtual teaching scene, and participates in a pre-set scenario dialogue task in the immersive virtual teaching scene. The multi-channel sensor array collects each student's facial expression and body movement through a video sensor, records each student's voice signal through a microphone array and locates the sound source, obtains the student's hand and head movement data through an inertial measurement unit and an optical motion capture system, records the student's gaze point and gaze time through an eye tracker, determines the focus of each student in the current scene, aligns the data obtained by the multi-channel sensor with timestamps, and performs multimodal data fusion in frames to obtain the multimodal interaction data; The multimodal interaction data is added to a pre-set video behavior understanding model, and features are extracted through a pre-trained convolutional neural network to generate a frame-level deep representation and send it to a long short-term memory network, determine the temporal dependency of different frame-level deep representations and generate a hidden state vector, perform attention weighting on the hidden state vector through an attention pooling mechanism and output it in a classified manner, generate a probability distribution of each behavior category and corresponding multi-channel spatiotemporal features, construct a quantitative indicator, and quantitatively analyze the multimodal interaction data based on the probability distribution corresponding to each behavior category to obtain the quantitative analysis result; Perform speech noise reduction and feature extraction on the video data recorded by the video sensor and the speech signal recorded by the microphone array, generate a feature sequence and extract features through a recurrent convolutional network, obtain high-order local features and perform sequence modeling, obtain a high-order local sequence, initialize the high-order local features corresponding to the decoder and update them at each decoding moment, generate a context vector in combination with the attention weight, process it through a fully connected layer and a soft normalization function, obtain the probability distribution of the native language words at the current moment, combine the cross entropy loss function and use the real native language words as input for reasoning, and obtain the student's native language speech; Based on the native language speech of the student, the semantic mapping relationship between the native language and the foreign language is determined by a speech translation model based on deep transfer, the speech translation model is trained in a pre-set bilingual corpus and the semantic mapping relationship is verified, the hyperparameters in the speech translation model are adjusted in combination with the verification result, the student's native language speech is translated by the adjusted speech translation model, a foreign language text sequence is generated and optimized by a beam search algorithm to obtain an initial foreign language text, the initial foreign language text is encoded by a bidirectional encoder in a multilingual bidirectional encoder model, the encoding result is retrieved in a pre-constructed multilingual corpus, the similarity between each element in the multilingual corpus and the encoding result is calculated, similar dialogues are determined based on the similarity and a candidate set is generated, the elements in the candidate set are scored according to language fluency and relevance, a dialogue score corresponding to each element is obtained, and the similar dialogue with the highest dialogue score is selected as the high-quality dialogue text; Based on the high-quality dialogue text, the quantitative analysis results and the multimodal interaction data, the effectiveness of the high-quality dialogue text in the current scenario is determined through an intelligent agent reinforcement algorithm, and feedback optimization is performed according to the quantitative analysis results and the multimodal interaction data to generate the personalized dialogue feedback.

5. The method according to claim 4, characterized in that The similarity between each element in the multilingual corpus and the encoding result is calculated as shown in the following formula: Among them, h i represents the candidate translation, s(h i ) represents the candidate translation h i Similarity with the encoding result, λ1 represents the language model probability weight, p(h i ) represents the candidate translation h i The language model probability, |h i | represents the length of the candidate translation, a represents the length penalty strength, λ2 represents the coverage weight, n-grams(h i ) represents the candidate translation h i n-grams(f) represents the n-grams of the source language sentence, |f| represents the length of the source language sentence, and λ3 represents the length ratio weight.

6. The method according to claim 1, characterized in that Based on the personalized dialogue feedback and the multimodal interaction data, the language application ability of each student is quantitatively evaluated in multiple dimensions in combination with the hierarchical transfer learning algorithm to generate an ability level portrait. The development level of each student's skill dimension is tracked based on the knowledge component theory and the multidimensional item response theory. The degree of mastery of each student's knowledge points is determined in combination with the Bayesian knowledge tracking model to generate a knowledge mastery map. The multi-channel spatiotemporal features are integrated and feature learning is performed in combination with the long short-term memory network to obtain learning behavior features. The ability level portrait, the knowledge mastery map and the learning behavior features are summarized. A language learning path optimization model is constructed based on a factorization machine and combinatorial ranking learning. The optimal learning path is obtained by solving the multi-objective optimization algorithm. The optimized teaching plan based on the optimal learning path includes: Obtain language proficiency test questions from a corpus and annotate knowledge points, add the personalized dialogue feedback and the multi-channel spatiotemporal features to a multi-layer perception mechanism, generate a high-dimensional feature representation and add it to the hierarchical transfer learning algorithm, set a feature extractor for each language proficiency dimension, extract the language proficiency representation corresponding to each modality and map it to a shared semantic space, perform a multi-dimensional quantitative assessment of each student's language application ability based on the mapping results, and splice the mapping results to obtain the ability level portrait; Based on the ability level portrait, combined with the knowledge component theory and the multidimensional item response theory, the language test questions are divided into multiple knowledge components and the knowledge component examination weights corresponding to each test question are determined, the knowledge component examination weights and answer records are added to the multidimensional item response theory model, the student's ability level in each dimension is estimated, the correct answer rate of the question is calculated based on the student's ability level in each dimension, the multidimensional item response theory model is optimized in combination with the cross entropy loss function and the real answer record, and the optimization is repeated until the preset maximum number of iterations is reached, and a dynamic evolution sequence corresponding to the ability level is generated. Through the Bayesian model and the dynamic evolution sequence, the current student's mastery of each knowledge point is generated, and the relationship between the knowledge points is combined to generate a knowledge mastery map; The multi-channel spatiotemporal features are represented as three-dimensional tensors corresponding to time steps, channel features and original features respectively, and local spatiotemporal features are extracted by combining a three-dimensional convolutional neural network pre-set along the channel dimension and flattened in the channel dimension, and long-term dependency modeling is performed by combining a two-layer long short-term memory network to obtain the learning behavior features; Summarize the ability level portrait, the knowledge mastery map and the learning behavior characteristics, score the match between students and learning activities based on a factor decomposition machine, optimize the factor decomposition machine in combination with minimizing the sorting loss, repeat the scoring until the preset maximum number of iterations is reached, and take the first 10 learning activities generated by the last iteration as alternatives. Combined with the genetic algorithm, encode each learning activity into an activity sequence, based on the activity sequence, with the goal of maximizing ability gain and minimizing emotional loss, construct an objective function in combination with pre-set knowledge prerequisite constraints, solve the objective function through a multi-objective optimization algorithm based on the genetic algorithm, generate a Pareto solution set and connect the Pareto solution set to generate an optimal learning path, perform attribute mapping on the optimal learning path, and obtain the optimized teaching plan.

7. The method according to claim 6, characterized in that The correct answer rate of the questions is calculated based on the students' ability level in each dimension as shown in the following formula: Among them, H() represents the probability of students answering correctly, X gj represents the answer of student g on question j, V g represents the ability vector of student g, a j represents the discrimination parameter of question j, b j represents the difficulty parameter of question j, c j represents the guess parameter of question j, q j It represents the weight matrix composed of the weights of knowledge components, and T represents transpose.

8. A VR-based multi-person foreign language situational teaching system for primary and secondary schools, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to obtain teaching tasks and determine the theme, scene, character role setting and interaction mode of the virtual reality scenario based on the teaching tasks, compress the data and optimize the loading of the scene through octree space division and voxel modeling technology, build an immersive virtual teaching scene in combination with a physics-based rendering method and a real-time lighting algorithm, integrate real teaching environment elements into the immersive virtual teaching scene through intelligent object recognition and real-scene modeling algorithm, set situational dialogue tasks and an adaptive difficulty adjustment mechanism in the immersive virtual teaching scene according to the teaching tasks, and combine knowledge graphs and ontology reasoning technology to semantically describe and associate elements in the immersive virtual teaching scene to form a multi-level semantic network; The second unit is used for each student to log in to the multi-person teaching platform through a VR all-in-one machine, set a virtual image and enter the immersive virtual teaching scene, participate in the situational dialogue task in the immersive virtual teaching scene, collect the multimodal interaction data corresponding to the current student through a multi-channel sensor array and transmit it back to the edge computing server, learn the multimodal interaction data through a video behavior understanding model, generate multi-channel spatiotemporal features and obtain quantitative analysis results, perform speech denoising and feature extraction on the video data collected by the multi-channel sensor array, perform contextual language modeling in combination with the attention mechanism, obtain the student's native language speech, translate the student's native language speech through a speech translation model based on deep migration, generate an initial foreign language text and filter it through a multilingual bidirectional encoder model to obtain a high-quality dialogue text, and generate personalized dialogue feedback based on the high-quality dialogue text, the quantitative analysis results and the multimodal interaction data through an intelligent agent reinforcement learning algorithm; The third unit is used to conduct a multi-dimensional quantitative assessment of the language application ability of each student based on the personalized dialogue feedback and the multi-channel spatiotemporal features in combination with a hierarchical transfer learning algorithm, generate an ability level portrait, track the development level of each student's skill dimension based on the knowledge component theory and the multidimensional item response theory, determine the degree of mastery of each student's knowledge points in combination with a Bayesian knowledge tracking model, generate a knowledge mastery map, integrate the multi-channel spatiotemporal features and perform feature learning in combination with a long short-term memory network to obtain learning behavior characteristics, summarize the ability level portrait, the knowledge mastery map and the learning behavior characteristics, construct a language learning path optimization model based on a factor decomposition machine and combinatorial ranking learning, obtain the optimal learning path in combination with a multi-objective optimization algorithm, and generate an optimized teaching plan based on the optimal learning path.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Real-time interactive online Chinese language learning platform

    CN120315598A

  • A real-time interactive online Chinese language learning platform

    CN120315598B

  • Virtual classroom interaction method and system based on knowledge representation and reasoning

    CN120495035A

  • Deep learning-based teaching corpus construction method and system, and medium

    CN120508666A

  • Knowledge graph construction method and system based on large model

    CN120561316A