Education scene intelligent identification system based on multi-modal perception

Through multimodal sensing technology, multi-dimensional data can be collected and processed in real time in educational scenarios, solving the problem of lagging monitoring and feedback in traditional education, realizing dynamic perception and personalized feedback throughout the entire process, and improving teaching quality and efficiency.

CN120632379AInactive Publication Date: 2025-09-12深圳码隆智能科技有限公司

Patent Information

Application Number
CN202511133985.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In traditional education scenarios, it is difficult for teachers to grasp the details of students' operations in real time. Classroom monitoring and test evaluation lack full-process traceability and intelligent feedback. Existing technologies lack multi-dimensional dynamic perception and personalized feedback, making it difficult to meet individualized teaching needs.

Method used

Using multimodal perception technology, multi-dimensional data is collected in real time through AI electronic eyepieces, high-definition video terminals, motion capture sensors and voice acquisition units. Combined with cross-modal embedding algorithms and modal attention fusion networks, multimodal data processing and temporal feature extraction are performed to achieve motion recognition and personalized feedback, supporting intelligent teaching plan generation and process evaluation.

Benefits of technology

It realizes dynamic perception of the entire process of students' experimental operations, improves the real-time and accuracy of the teaching process, breaks the separation of teaching, learning and testing, supports personalized teaching and data security management, and improves teaching efficiency and student operation standardization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632379A_ABST
    Figure CN120632379A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of education informatization and artificial intelligence, and discloses an education scene intelligent identification system based on multi-modal perception. The system is composed of a multi-source education scene data acquisition module, a multi-modal data preprocessing and semantic unified coding module, a high-dimensional time sequence feature extraction module, a multi-modal feature fusion module, a time sequence action intelligent identification module, a real-time misoperation detection and personalized feedback module, and an intelligent teaching plan dynamic generation and process evaluation module. The system is composed of a multi-level security data management and domestic reasoning card adaptation module and a teaching data visualization and intelligent decision support module. Through multi-modal data fusion, time sequence action recognition, edge end real-time reasoning and domestic AI accelerator card adaptation, a whole-process dynamic perception and intelligent feedback system is constructed, experiment teaching, process monitoring, intelligent scoring, teaching plan generation and data visualization closed loop are achieved, the cloud dependence and data safety bottleneck is broken through, and three-level platform linkage of cities, districts and schools is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of educational informatization and artificial intelligence technology, and specifically relates to an intelligent recognition system for educational scenes based on multimodal perception. Background Art

[0002] With the rapid application of artificial intelligence technology in the field of education, traditional education scenarios have exposed many pain points that need to be addressed. On the one hand, it is difficult for teachers to fully grasp the students' operation details in a timely manner during experimental teaching, resulting in insufficient real-time classroom monitoring and process guidance, affecting the quality and efficiency of teaching; on the other hand, the existing examination evaluation system lacks full-process traceability and intelligent feedback, and the examination data cannot be effectively integrated and applied, resulting in the separation of teaching, learning, and examination links.

[0003] In addition, most existing technologies rely on single-modal data collection, which cannot accurately capture multi-source information in complex teaching scenarios. They lack multi-dimensional dynamic perception and personalized feedback, making it difficult to meet the needs of individualized teaching and intelligent management.

[0004] To address the above problems, relying on its core advantages such as multimodal perception technology, real-time interaction of large-model experiments, weakly supervised learning, temporal action recognition, and AIGC desktop redrawing, AIGC has introduced domestic AI inference cards and edge-end real-time inference algorithms to break data silos, solve the problems of feedback delay and high computing power costs, and establish a new type of intelligent recognition system for educational scenarios with real-time data processing, accurate recognition, and intelligent feedback throughout the entire process. Summary of the Invention

[0005] The purpose of the present invention is to provide an intelligent recognition system for educational scenarios based on multimodal perception to solve the problems raised in the above background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solutions: an intelligent recognition system for educational scenarios based on multimodal perception, which consists of a multi-source educational scenario data acquisition module, a multimodal data preprocessing and semantic unified coding module, a high-dimensional temporal feature extraction module, a multimodal feature fusion module, a temporal action intelligent recognition module, a real-time misoperation detection and personalized feedback module, an intelligent teaching plan dynamic generation and process evaluation module, a multi-level security data management and domestic reasoning card adaptation module, and a teaching data visualization and intelligent decision support module; The multi-source education scene data acquisition module collects multimodal data such as images, actions, and voice in real time and transmits it to the edge computing node. The multimodal data preprocessing and semantic unified coding module denoises and cleans the collected data, and the unified coding generates standardized high-dimensional semantic features. The high-dimensional temporal feature extraction module extracts multimodal temporal features and constructs a dynamic action trajectory feature sequence. The multimodal feature fusion module fuses multimodal temporal features through the modal attention mechanism to form the scene state. The temporal action intelligent recognition module identifies action categories and steps in real time and establishes a precise mapping between actions and experimental steps. The real-time error detection and personalized feedback module compares the recognition results with the standard process and generates personalized voice or text feedback. The intelligent lesson plan dynamic generation and process evaluation module generates personalized lesson plans and process evaluations, and supports classroom summarization and review. The multi-level security data management and domestic reasoning card adaptation module encrypts and stores data throughout the process, is compatible with domestic reasoning cards, and supports secure local reasoning. The teaching data visualization and intelligent decision support module analyzes the data throughout the process, generates visual charts, and assists in teaching optimization and decision-making.

[0007] Preferably, the multi-source education scene data acquisition module includes: (1) Multi-source terminal collaborative acquisition: Relying on AI electronic eyepieces, high-definition video terminals, motion capture sensors and voice acquisition units, it realizes real-time synchronous acquisition of multi-dimensional data at the education site, covering multi-modal raw data such as images, hand movements, voice and teaching aid status, and efficiently transmits them to edge computing nodes via a dedicated bus, providing comprehensive data support for subsequent AI processing, which is consistent with the multi-modal AI technology application scenario in its smart experiment program; (2) Multimodal data fusion preprocessing: After noise reduction and filtering, the collected multimodal data is uniformly mapped to the semantic space through a cross-modal embedding algorithm to generate standardized feature vectors. This process integrates multi-source information and connects with the real-time interaction technology of the multimodal large model, laying the data foundation for core functions such as action recognition and AI scoring, and ensuring the accuracy of intelligent analysis of educational scenarios; The expression formula for multimodal data collection is:

[0008] Where, Multimodal datasets, Image data (image), Action trajectory data (Action), : Voice data (Voice), :Device status data (Status).

[0009] Preferably, the multimodal data preprocessing and semantic unified encoding module includes: (1) Deep cleaning of multimodal data: For the images, actions, voices and teaching aids data collected by AI electronic eyepieces and high-definition video terminals, we process and eliminate abnormal information through noise reduction, filtering, and distortion removal, and perform standardization adjustments based on data distribution characteristics to keep the multi-source data consistent in scale, laying the foundation for cross-modal fusion and meeting the stringent data quality requirements of multimodal large models; (2) Cross-modal semantic mapping to generate standardized feature vectors: Using a cross-modal embedding algorithm, the cleaned data is uniformly mapped to the semantic space, and high-dimensional vectors are generated through feature distribution calibration. This process integrates multimodal information and connects with AI action recognition technology to provide standardized input for temporal feature extraction, supporting accurate analysis of the integrated teaching, assessment and evaluation of smart experiments; The standard normalized expression formula is:

[0010] Where, :No. modal normalization features, :No. modal raw data, :No. modal means, :No. modal standard deviation.

[0011] Preferably, the high-dimensional time series feature extraction module includes: (1) Multi-scale network extraction of key temporal features: After receiving the pre-processed standardized feature vector, relying on the multi-scale convolutional neural network based on the temporal transformer, combined with the sliding average method, it accurately captures the temporal regularity of actions, language and experimental steps, and extracts key node features by smoothing the dynamic data fluctuations. This is in line with the multi-frame temporal action recognition technology, laying the foundation for generating dynamic trajectories. The sliding average time series feature extraction expression is:

[0012] Where, time The timing characteristics of Sliding window completion, :No. Frame features, Time position coding; (2) Dynamic trajectory sequence supports scene understanding: In the feature extraction process, multimodal temporal information is integrated to generate a coherent dynamic action trajectory feature sequence. This sequence accurately reflects the experimental operation process and is connected with AI object motion recognition technology to provide structured input for subsequent modal fusion and intelligent recognition, facilitating real-time interactive analysis of educational scenarios.

[0013] Preferably, the multimodal feature fusion module includes: (1) Dynamic weighted fusion of modal attention network: After receiving the temporal feature sequence, the modal attention fusion network adjusts the weights based on the real-time modal confidence and the feature importance distribution mechanism, and dynamically weights and fuses multimodal information such as action and language. This process adapts to the interaction requirements of large multimodal models, strengthens the weight of key information, and provides a precise fusion basis for generating scene state vectors. The weighted average fusion expression formula is:

[0014] Where, The fused state vector, : total number of modes, : modal weight coefficient, Normalized features; (2) Real-time interactive scene state vector generation: By dynamically integrating the temporal features of each modality, a state vector reflecting the real-time interaction of the educational scene is formed. This vector integrates multi-dimensional operation information and connects with AI action recognition technology to provide comprehensive feature input for subsequent action classification and misoperation detection, supporting real-time analysis and decision-making of smart experimental teaching and evaluation.

[0015] Preferably, the time sequence action intelligent recognition module includes: (1) Edge AI reasoning enables real-time action classification: Based on the scene state vector output by the multimodal feature fusion module, edge AI reasoning is combined with a probability judgment mechanism to classify and time-series calibrate actions, dynamically identifying action categories, operation steps, and time-series nodes in student experiments. This, in conjunction with multi-frame time-series action recognition technology, provides a basis for establishing an accurate mapping between actions and experimental steps. The expression formula for probabilistic action recognition is:

[0016] Where, :Action Category The predicted probability of Category weights, : Class bias, :Total number of categories; (2) Accurate mapping supports the perception of the experimental process. During the recognition process, the probability judgment results of multimodal features are integrated to optimize the correspondence between actions and steps. This module is connected with the AI ​​scoring system to output structured recognition results in real time, providing a decision-making basis for subsequent misoperation detection and personalized feedback, and facilitating the precise implementation of smart experimental teaching.

[0017] Preferably, the real-time misoperation detection and personalized feedback module includes: (1) Real-time comparison and identification of experimental operation anomalies: Based on the action step recognition results output by the time-series action intelligent recognition module, they are compared with the preset experimental process standard model in real time. Combined with the feature difference analysis method, students' misoperations or step omissions can be quickly captured. This process is adapted to multimodal AI interaction technology. By dynamically detecting operation deviations, it provides accurate basis for immediate feedback and ensures the standardization of experimental teaching. The dynamic detection expression formula is:

[0018] Where, : The Euclidean distance between the current action and the standard action, : Current action Features, , Standard Action No. Features, : Feature dimension; (2) Personalized prompts help students correct operations: After identifying an anomaly, the system automatically generates personalized voice or text prompts to guide students to correct their operations. This function is connected with the real-time interactive system of smart experiments. By dynamically adjusting the feedback content, it strengthens students' understanding of experimental specifications, helps to accurately improve operational skills during the teaching process, and meets the needs of integrated teaching, assessment and evaluation.

[0019] Preferably, the intelligent teaching plan dynamic generation and process evaluation module refers to the real-time collection of feedback results and experimental scores of the real-time error detection and personalized feedback modules, combined with multi-dimensional performance for comprehensive evaluation, and generation of process evaluation data. Based on the AIGC intelligent teaching plan generation model, it automatically produces experimental summaries, personalized suggestions and evaluation reports, supports instant classroom summaries and post-learning review. This process is consistent with the application of AIGC technology in educational scenarios. By dynamically integrating scoring results, it helps improve the teaching closed loop and reflects the advantages of integrated teaching, assessment and evaluation. The weighted average process score expression formula is:

[0020] Where, Generated process ratings, Total duration (total number of frames), moment weight, Scene state vector.

[0021] Preferably, the multi-level security data management and domestic reasoning card adaptation module refers to the information such as the whole teaching process data and scoring results, which are stored synchronously on the city-district-school three-level platform after encryption processing to ensure transmission and storage security, support local deployment of domestic AI reasoning cards, and adapt to flexible switching between offline and online modes. This mechanism is in line with the multi-level data security assurance system, maintains scoring fairness through standardized encryption processes, and provides safe and reliable technical support for the integration of smart experiment teaching, assessment and evaluation.

[0022] Preferably, the teaching data visualization and intelligent decision-making support module refers to calling the multi-level security data management and domestic reasoning card adapter module to securely store the data set, and with the help of multi-dimensional visualization tools, combined with category proportion statistics, intuitively presenting students' experimental trajectories, teaching quality, scoring data and abnormal nodes. This process is consistent with data visualization analysis capabilities, assisting teachers and managers to accurately adjust teaching and optimize examination arrangements, helping to improve the teaching, examination and evaluation closed loop, and highlighting the advantages of data-driven decision-making in the "AI+Education" scenario.

[0023] The beneficial effects of the present invention are as follows: 1. By introducing multimodal data fusion and time-series motion recognition, the present invention addresses the problems of single monitoring and delayed feedback in existing educational scenes. It designs a multi-source real-time data acquisition architecture based on AI electronic eyepieces, high-definition video acquisition terminals, motion capture sensors and voice acquisition units. The system uses a cross-modal embedding algorithm and a modal attention fusion network to dynamically integrate image, motion, voice and teaching aid status information, and combines it with the motion trajectory extraction based on the time-series Transformer to achieve dynamic perception of the entire process of students' experimental operations and classroom behaviors. Compared with existing teaching monitoring technologies that only rely on a single visual or voice modality, the present invention can comprehensively and accurately capture complex information such as experimental steps, hand movements, and language interactions, greatly improving the accuracy and real-time performance of educational scene recognition, ensuring that errors and omissions in the teaching process can be discovered and fed back immediately, and effectively improving experimental teaching management and student operation standards.

[0024] 2. The present invention breaks through the technical bottlenecks of the existing educational information system's excessive reliance on cloud servers, high computing costs, and difficulty in ensuring data security through deep adaptation of real-time reasoning on the edge and domestic AI acceleration cards. The system encrypts and stores data throughout the process locally, supports flexible switching between offline and online, and fully meets the hierarchical management needs of school, regional, and municipal platforms. Through real-time reasoning and low-latency scoring on the edge, the present invention significantly reduces the computing resources and network burden required for large-scale deployment, and ensures the security and stability of data throughout the entire process of collection, transmission, processing, and storage. This design has been widely used in existing smart experimental teaching systems and AI-enabled assessment processes, meeting the teaching and examination scenarios of different regions and schools across the country. It has good scalability and practical value, and is suitable for large-scale, multi-school joint examination environments.

[0025] 3. The present invention breaks the process management mode of teaching, learning and testing separated in traditional education scenarios through the deep integration of dynamic generation of intelligent teaching plans and process evaluation functions, and realizes a fully closed-loop intelligent system from experimental teaching, process monitoring, intelligent scoring, classroom summary to examination evaluation. The system supports automatic generation of personalized teaching plans, experimental summaries, intelligent comments and personalized learning suggestions based on AIGC, and records the highlights and shortcomings of students' operations in real time to help teachers make differentiated teaching adjustments. Through multi-level data synchronization management, the system supports real-time linkage of teaching and examination platforms at the city, district and school levels, promotes centralized analysis and collaborative optimization of experimental teaching data, significantly improves classroom teaching efficiency, student operation standardization and personalized learning experience, and provides strong support for the continuous implementation of smart education. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a flow chart of the intelligent recognition system for educational scenarios based on multimodal perception of the present invention. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0028] like Figure 1As shown, an embodiment of the present invention provides an intelligent recognition system for educational scenes based on multimodal perception, which consists of a multi-source educational scene data acquisition module, a multimodal data preprocessing and semantic unified coding module, a high-dimensional temporal feature extraction module, a multimodal feature fusion module, a temporal action intelligent recognition module, a real-time misoperation detection and personalized feedback module, an intelligent teaching plan dynamic generation and process evaluation module, a multi-level security data management and domestic reasoning card adaptation module, and a teaching data visualization and intelligent decision support module. The multi-source education scene data acquisition module collects multimodal data such as images, actions, and voice in real time and transmits it to the edge computing node. The multimodal data preprocessing and semantic unified coding module denoises and cleans the collected data, and the unified coding generates standardized high-dimensional semantic features. The high-dimensional temporal feature extraction module extracts multimodal temporal features and constructs a dynamic action trajectory feature sequence. The multimodal feature fusion module fuses multimodal temporal features through the modal attention mechanism to form the scene state. The temporal action intelligent recognition module identifies action categories and steps in real time and establishes a precise mapping between actions and experimental steps. The real-time error detection and personalized feedback module compares the recognition results with the standard process and generates personalized voice or text feedback. The intelligent lesson plan dynamic generation and process evaluation module generates personalized lesson plans and process evaluations, and supports classroom summarization and review. The multi-level security data management and domestic reasoning card adaptation module encrypts and stores data throughout the process, is compatible with domestic reasoning cards, and supports secure local reasoning. The teaching data visualization and intelligent decision support module analyzes the data throughout the process, generates visual charts, and assists in teaching optimization and decision-making.

[0029] Among them, the multi-source education scene data acquisition module refers to the real-time and synchronous acquisition of multi-dimensional data at the education site based on AI electronic eyepieces, high-definition video terminals, motion capture sensors and voice acquisition units, covering multi-modal original data of images, hand movements, voice and teaching aids status, which are efficiently transmitted to the edge computing node via a dedicated bus, providing comprehensive data support for subsequent AI processing, which is in line with the multimodal AI technology application scenario in its smart experiment plan; the collected multimodal data is denoised, filtered and cleaned, and then uniformly mapped to the semantic space through a cross-modal embedding algorithm to generate a standardized feature vector. This process integrates multi-source information and connects with the real-time interaction technology of the multimodal large model, laying a data foundation for core functions such as motion recognition and AI scoring, and ensuring the accuracy of intelligent analysis of educational scenes.

[0030] Among them, the multimodal data preprocessing and semantic unified coding module refers to the image, action, voice and teaching aids data collected by AI electronic eyepieces and high-definition video terminals, which are processed and eliminated through noise reduction, filtering, and distortion removal, and standardized adjustment is performed based on the data distribution characteristics to make the multi-source data consistent in scale, laying the foundation for cross-modal fusion, meeting the strict requirements of multimodal large models on data quality, and using cross-modal embedding algorithms to uniformly map the cleaned data to the semantic space, and generate high-dimensional vectors through feature distribution calibration. This process integrates multimodal information, connects with AI motion recognition technology, provides standardized input for time series feature extraction, and supports accurate analysis of the integrated teaching, assessment and evaluation of smart experiments.

[0031] Among them, the high-dimensional time series feature extraction module includes: the high-dimensional time series feature extraction module refers to receiving the pre-processed standardized feature vector, relying on the multi-scale convolutional neural network based on the time series Transformer, combined with the sliding average method, to accurately capture the time series rules of actions, language and experimental steps, and extract key node features by smoothing dynamic data fluctuations, echoing the multi-frame time series action recognition technology, laying the foundation for generating dynamic trajectories. In the feature extraction process, multi-modal time series information is integrated to generate a coherent dynamic action trajectory feature sequence, which accurately reflects the experimental operation process and connects with AI object action recognition technology to provide structured input for subsequent modal fusion and intelligent recognition, and assist in real-time interactive analysis of educational scenarios.

[0032] Among them, the multimodal feature fusion module refers to the modal attention fusion network that adjusts the weights based on the real-time modal confidence and the feature importance allocation mechanism after receiving the temporal feature sequence, and dynamically weighted fuses multimodal information such as action and language. This process adapts to the interaction needs of the multimodal large model, strengthens the weight of key information, and provides an accurate fusion basis for generating scene state vectors. By dynamically integrating the temporal features of each modality, a state vector reflecting the real-time interaction of the educational scene is formed. This vector integrates multi-dimensional operation information and connects with AI action recognition technology to provide comprehensive feature input for subsequent action classification and misoperation detection, supporting real-time analysis and decision-making of smart experimental teaching and evaluation.

[0033] Among them, the time-series action intelligent recognition module refers to the scene state vector output by the multimodal feature fusion module. Through edge AI reasoning and combined with the probability judgment mechanism, the actions are classified and time-series calibrated, and the action categories, operation steps and time-series nodes in the student experiment are dynamically identified. It echoes the multi-frame time-series action recognition technology and provides a basis for establishing an accurate mapping between actions and experimental steps. In the recognition process, the probability judgment results of multimodal features are integrated to optimize the correspondence between actions and steps. This module is connected with the AI ​​scoring system and outputs structured recognition results in real time, providing a decision-making basis for subsequent misoperation detection and personalized feedback, and helping to implement the precise implementation of smart experimental teaching.

[0034] Among them, the real-time error operation detection and personalized feedback module refers to the action step recognition results output by the time-series action intelligent recognition module, which are compared with the preset experimental process standard model in real time, and combined with the feature difference analysis method to quickly capture students' wrong operations or step omissions. This process is adapted to multimodal AI interaction technology, and through dynamic detection of operation deviations, it provides accurate basis for instant feedback to ensure the standardization of experimental teaching. After identifying the anomaly, the system automatically generates personalized voice or text prompts to guide students to correct the operation. This function is connected with the real-time interactive system of smart experiments. By dynamically adjusting the feedback content, it strengthens students' understanding of experimental specifications, helps to accurately improve operational skills in the teaching process, and meets the needs of integrated teaching, assessment and evaluation.

[0035] Among them, the intelligent lesson plan dynamic generation and process evaluation module refers to the real-time collection of feedback results and experimental scores of the real-time error operation detection and personalized feedback module, combined with multi-dimensional performance for comprehensive evaluation, and generation of process evaluation data. Based on the AIGC intelligent lesson plan generation model, it automatically produces experimental summaries, personalized suggestions and evaluation reports, and supports instant classroom summaries and post-learning review. This process is consistent with the application of AIGC technology in educational scenarios. By dynamically integrating scoring results, it helps improve the teaching closed loop and reflects the advantages of integrated teaching, assessment and evaluation.

[0036] Among them, the multi-level security data management and domestic reasoning card adaptation module refers to the information such as the entire teaching process data and scoring results, which are stored synchronously on the city-district-school three-level platform after encryption to ensure transmission and storage security, support local deployment of domestic AI reasoning cards, and adapt to flexible switching between offline and online modes. This mechanism is in line with the multi-level data security assurance system, maintains scoring fairness through standardized encryption processes, and provides safe and reliable technical support for the integration of teaching, assessment and evaluation of smart experiments.

[0037] Among them, the teaching data visualization and intelligent decision-making support module refers to the data set that is securely stored by calling multi-level secure data management and domestic inference card adaptation modules. With the help of multi-dimensional visualization tools and combined with category proportion statistics, it intuitively presents students' experimental trajectories, teaching quality, scoring data and abnormal nodes. This process is consistent with data visualization analysis capabilities, assisting teachers and managers to accurately adjust teaching and optimize examination arrangements, helping to improve the teaching, examination and evaluation closed loop, and highlighting the advantages of data-driven decision-making in the "AI+Education" scenario.

[0038] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0039] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent recognition system for educational scenarios based on multimodal perception, characterized by: The system consists of a multi-source educational scene data acquisition module, a multimodal data preprocessing and semantic unified coding module, a high-dimensional temporal feature extraction module, a multimodal feature fusion module, a temporal action intelligent recognition module, a real-time misoperation detection and personalized feedback module, an intelligent teaching plan dynamic generation and process evaluation module, a multi-level security data management and domestic reasoning card adaptation module, and a teaching data visualization and intelligent decision support module. The multi-source education scene data acquisition module collects multimodal data such as images, actions, and voice in real time and transmits them to the edge computing node. The multimodal data preprocessing and semantic unified coding module denoises and cleans the collected data, and the unified coding generates standardized high-dimensional semantic features. The high-dimensional temporal feature extraction module extracts multimodal temporal features and constructs a dynamic action trajectory feature sequence. The multimodal feature fusion module fuses multimodal temporal features through the modal attention mechanism to form the scene state. The temporal action intelligent recognition module identifies action categories and steps in real time and establishes a precise mapping between actions and experimental steps. The real-time error detection and personalized feedback module compares the recognition results with the standard process and generates personalized voice or text feedback. The intelligent teaching plan dynamic generation and process evaluation module generates personalized teaching plans and process evaluations, and supports classroom summarization and review. The multi-level security data management and domestic reasoning card adaptation module encrypts and stores data throughout the process, is compatible with domestic reasoning cards, and supports secure local reasoning. The teaching data visualization and intelligent decision support module parses the data throughout the process, generates visual charts, and assists in teaching optimization and decision-making.

2. The intelligent recognition system for educational scenarios based on multimodal perception according to claim 1 is characterized by: The multi-source education scene data acquisition module includes: (1) Multi-source terminal collaborative acquisition: Relying on AI electronic eyepieces, high-definition video terminals, motion capture sensors and voice acquisition units, it realizes real-time synchronous acquisition of multi-dimensional data at the education site, covering multi-modal raw data such as images, hand movements, voice and teaching aid status, and efficiently transmits them to edge computing nodes via a dedicated bus, providing comprehensive data support for subsequent AI processing, which is consistent with the multi-modal AI technology application scenario in its smart experiment program; (2) Multimodal data fusion preprocessing: After noise reduction and filtering, the collected multimodal data is uniformly mapped to the semantic space through a cross-modal embedding algorithm to generate standardized feature vectors. This process integrates multi-source information and connects with the real-time interaction technology of the multimodal large model, laying the data foundation for core functions such as action recognition and AI scoring, ensuring the accuracy of intelligent analysis of educational scenarios; The expression formula for multimodal data collection is: Where, Multimodal datasets, Image data, Motion trajectory data, : Voice data, : Device status data.

3. The intelligent recognition system for educational scenarios based on multimodal perception according to claim 1 is characterized by: The multimodal data preprocessing and semantic unified coding module includes: (1) Deep cleaning of multimodal data: For the images, actions, voices and teaching aids data collected by AI electronic eyepieces and high-definition video terminals, we process and eliminate abnormal information through noise reduction, filtering, and distortion removal, and perform standardization adjustments based on data distribution characteristics to keep the multi-source data consistent in scale, laying the foundation for cross-modal fusion and meeting the stringent data quality requirements of multimodal large models; (2) Cross-modal semantic mapping to generate standardized feature vectors: Using a cross-modal embedding algorithm, the cleaned data is uniformly mapped to the semantic space, and high-dimensional vectors are generated through feature distribution calibration. This process integrates multimodal information and connects with AI action recognition technology to provide standardized input for temporal feature extraction, supporting accurate analysis of the integrated teaching, assessment and evaluation of smart experiments; The standard normalized expression formula is: Where, :No. modal normalization features, :No. modal raw data, :No. modal means, :No. modal standard deviation.

4. The intelligent recognition system for educational scenarios based on multimodal perception according to claim 1 is characterized by: The high-dimensional time series feature extraction module includes: (1) Multi-scale network extraction of key temporal features: After receiving the pre-processed standardized feature vector, relying on the multi-scale convolutional neural network based on the temporal Transformer, combined with the sliding average method, it accurately captures the temporal regularity of actions, language and experimental steps, and extracts key node features by smoothing dynamic data fluctuations, which echoes the multi-frame temporal action recognition technology; The sliding average time series feature extraction expression is: Where, time The timing characteristics of Sliding window completion, :No. Frame features, Time position coding; (2) Dynamic trajectory sequence supports scene understanding: During the feature extraction process, multimodal temporal information is integrated to generate a coherent dynamic action trajectory feature sequence, which accurately reflects the experimental operation process and is connected with AI object action recognition technology.

5. The intelligent recognition system for educational scenarios based on multimodal perception according to claim 1 is characterized by: The multimodal feature fusion module includes: (1) Dynamic weighted fusion of modal attention network: After receiving the temporal feature sequence, the modal attention fusion network adjusts the weights based on the real-time modal confidence and the feature importance distribution mechanism, and dynamically weights and fuses multimodal information such as action and language. This process adapts to the interaction requirements of the multimodal large model, strengthens the weight of key information, and provides a precise fusion basis for generating the scene state vector; The weighted average fusion expression formula is: Where, The fused state vector, : total number of modes, : modal weight coefficient, Normalized features; (2) Real-time interactive scene state vector generation: By dynamically integrating the temporal features of each modality, a state vector reflecting the real-time interaction of the educational scene is formed. This vector integrates multi-dimensional operation information and connects with AI action recognition technology to provide comprehensive feature input for subsequent action classification and misoperation detection, supporting real-time analysis and decision-making of smart experimental teaching and evaluation.

6. The educational scene intelligent recognition system based on multimodal perception according to claim 1 is characterized by: The time sequence action intelligent recognition module includes: (1) Real-time action classification using edge AI reasoning: Based on the scene state vector output by the multimodal feature fusion module, actions are classified and time-series calibrated through edge AI reasoning combined with a probability judgment mechanism. This allows for dynamic identification of action categories, operation steps, and time-series nodes in student experiments, which is consistent with multi-frame time-series action recognition technology. The expression formula for probabilistic action recognition is: Where, :Action Category The predicted probability of Category weights, : Class bias, :Total number of categories; (2) Accurate mapping supports the perception of the experimental process. During the recognition process, the probability judgment results of multimodal features are integrated to optimize the correspondence between actions and steps. This module is connected with the AI ​​scoring system to output structured recognition results in real time.

7. The intelligent recognition system for educational scenarios based on multimodal perception according to claim 1 is characterized by: The real-time misoperation detection and personalized feedback module includes: (1) Real-time comparison and identification of experimental operation anomalies: Based on the action step recognition results output by the time-series action intelligent recognition module, they are compared with the preset experimental process standard model in real time. Combined with the feature difference analysis method, students' misoperations or step omissions can be quickly captured. This process is adapted to multimodal AI interaction technology. By dynamically detecting operation deviations, it provides accurate basis for immediate feedback and ensures the standardization of experimental teaching. The dynamic detection expression formula is: Where, : The Euclidean distance between the current action and the standard action, : Current action Features, , Standard Action No. Features, : Feature dimension; (2) Personalized prompts help students correct operations: After identifying an anomaly, the system automatically generates personalized voice or text prompts to guide students to correct their operations. This function is connected to the real-time interactive system of smart experiments.

8. The intelligent recognition system for educational scenarios based on multimodal perception according to claim 1 is characterized by: The intelligent teaching plan dynamic generation and process evaluation module refers to the real-time collection of feedback results and experimental scores from the real-time error detection and personalized feedback modules, combined with multi-dimensional performance for comprehensive evaluation, and the generation of process evaluation data. Based on the AIGC intelligent teaching plan generation model, it automatically produces experimental summaries, personalized suggestions and evaluation reports, supporting instant classroom summaries and post-learning review. This process is consistent with the application of AIGC technology in educational scenarios. The weighted average process score expression formula is: Where, Generated process ratings, Total duration, moment weight, Scene state vector.

9. The intelligent recognition system for educational scenarios based on multimodal perception according to claim 1, characterized in that: The multi-level security data management and domestic inference card adaptation module refers to information such as the entire teaching process data and scoring results, which are encrypted and synchronously stored on the city-district-school three-level platform to ensure transmission and storage security, support local deployment of domestic AI inference cards, and adapt to flexible switching between offline and online modes. This mechanism is in line with the multi-level data security assurance system.

10. The intelligent recognition system for educational scenarios based on multimodal perception according to claim 1, characterized in that: The teaching data visualization and intelligent decision-making support module refers to calling the multi-level security data management and domestic reasoning card adapter module to securely store the data set. With the help of multi-dimensional visualization tools and combined with category proportion statistics, it intuitively presents students' experimental trajectories, teaching quality, scoring data and abnormal nodes. This process is consistent with data visualization analysis capabilities, assisting teachers and managers to accurately adjust teaching, optimize examination arrangements, and help improve the teaching, examination and evaluation closed loop.

Citation Information

Patent Citations

  • Domestic deep learning algorithm and operation equipment thereof

    CN118397432A

  • AIGC-based multidisciplinary interactive teaching system

    CN119168188A

  • Virtual-real fusion chemical experiment platform for real-time feedback based on multi-modal perception

    CN119516852A

  • Intelligent learning behavior monitoring and abnormity early warning method and device

    CN119723639A

  • AI enabling quantitative teaching evaluation method and system based on multi-modal data portrait

    CN120373971A

Cited By

  • Deep learning-based cooperative cultivation quality intelligent evaluation system and method

    CN120851716A

  • Data standardization processing method for AI ultrasonic large model test

    CN121767352A