Multi-modal operation analysis and real-time projection feedback system based on AI
By combining multimodal data acquisition and AI analysis with light field projection and tactile feedback, the problem of synchronous acquisition and real-time feedback of multimodal data in educational informatization systems has been solved, enabling efficient homework analysis and interactive teaching feedback, thereby improving the system's practicality and teaching effectiveness.
Patent Information
- Application Number
- CN202511003898.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-07
AI Technical Summary
Existing educational information systems face technical bottlenecks in multimodal data acquisition, real-time feedback, and modular hardware design. They struggle to achieve simultaneous acquisition and efficient fusion of visual information, stress data, and physiological signals. Projection feedback is limited in form, lacking tactile interaction and dynamic projection effects. Data management also lacks personalized recommendation mechanisms, hindering the formation of a closed-loop teaching system.
It employs a multimodal data acquisition module, an AI analysis module, a projection feedback module, an edge computing module, and a data management module, combined with an RGB-D camera, a pressure sensor array, and an infrared thermal imaging module. It performs data fusion and analysis through a cross-modal Transformer model, utilizes light field projection technology and microcurrent tactile feedback, and combines edge computing and cloud collaboration to achieve real-time operation analysis and feedback.
It enables real-time analysis and processing of multimodal data, improves the visualization and interactive feedback of homework errors, enhances student interactivity, and improves the system's practicality and teaching relevance through personalized teaching suggestions.
Smart Images

Figure CN120910441A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of educational information technology, and in particular to a multi-modal homework analysis and real-time projection feedback system based on AI. BACKGROUND
[0002] With the development of educational information technology, homework analysis systems have gradually developed from early single-image recognition standalone mode to intelligent systems with multi-modal data collection and cloud collaboration. Early solutions only used cameras to collect homework images for OCR recognition, lacking multi-dimensional data perception such as writing force and facial expressions. With the advancement of sensor technology and AI algorithms, multi-modal data fusion analysis has become the mainstream trend. The introduction of edge computing technology has also improved local processing capabilities to some extent. However, existing systems still have technical bottlenecks in terms of multi-modal data spatio-temporal alignment, real-time feedback interaction, and hardware modular design.
[0003] The main shortcomings of existing technologies are: multi-modal data collection modules cannot achieve synchronous collection and efficient fusion of visual information, pressure data, and physiological signals, resulting in insufficient positioning accuracy of homework errors by AI analysis modules; relying on cloud servers for feature extraction and algorithm reasoning causes high data transmission delay, making it impossible to meet real-time feedback requirements; projection feedback forms are single, lacking tactile interaction and dynamic projection effects, making it difficult to achieve immersive learning experiences; hardware architecture is mostly integrated, with low integration of scanning and projection functions, and the data management module lacks a personalized recommendation mechanism based on knowledge graphs, making it impossible to form a teaching closed loop. SUMMARY
[0004] The present application aims to overcome the above problems and proposes a multi-modal homework analysis and real-time projection feedback system based on AI. To achieve the above purpose, the present application adopts the following technical solutions:
[0005] The multi-modal homework analysis and real-time projection feedback system based on AI includes a multi-modal data collection module, an AI analysis module, a projection feedback module, an edge computing module, and a data management module. The multi-modal data collection module and the AI analysis module are connected through a data transmission link, collect homework multi-modal data, and transmit it to the AI analysis module. The AI analysis module and the edge computing module form a collaborative processing architecture to extract features and analyze multi-modal data. The AI analysis module outputs analysis results to the projection feedback module to generate feedback information. The data management module and the AI analysis module form a bidirectional data interaction to encrypt, store, and manage data.
[0006] Further, the multi-modal data acquisition module includes an RGB-D camera, a pressure sensor array, and an infrared thermal imaging module; the RGB-D camera acquires visual information of the work, the pressure sensor array acquires writing force data, and the infrared thermal imaging module captures facial micro-expression data.
[0007] Further, the AI analysis module includes a multi-modal data fusion unit and an error identification unit; the multi-modal data fusion unit performs spatio-temporal alignment processing on handwriting trajectories, speech explanations, and problem solving steps through a cross-modal Transformer model; the error identification unit performs high-precision identification of handwritten characters, formulas, and symbols based on a convolutional recurrent neural network and a residual network OCR module, and analyzes mathematical problem errors in combination with geometric feature extraction and symbol matching algorithms, and the error identification unit uses a ResNet+Transformer hybrid architecture optimization model.
[0008] Further, the projection feedback module includes a dynamic projection unit and a tactile feedback unit; the dynamic projection unit uses light field projection technology to project the analysis results in 3D highlight form to the work surface, and supports touch interaction operation; the tactile feedback unit is a micro-current tactile unit, and generates pulse feedback signals based on preset error type trigger rules.
[0009] Further, the edge computing module uses an NVIDIA Jetson Nano module to perform multi-modal feature extraction locally, and performs real-time correlation processing on gesture trajectories and speech semantics, reducing dependence on cloud computing.
[0010] Further, the data management module includes a cloud collaboration unit that uploads encrypted work data to the cloud, generates teaching optimization suggestions and teaching analysis data based on collaborative filtering algorithms and knowledge graphs, and performs multi-terminal data synchronization and work data output.
[0011] Further, it also includes a hardware architecture, which includes a detachable scanning projection module, an intelligent adjustment module, and a human-computer interaction module; the scanning projection module includes a flatbed scanner and a short-focus projector; the intelligent adjustment module drives the object table through a motor and performs automatic focusing and position calibration through a photoelectric sensor; the human-computer interaction module includes a touch screen operation panel and a multi-device linkage unit: the touch screen operation panel supports handwriting input and voice instructions for setting analysis parameters; the multi-device linkage unit connects tablets and computers through Wi-Fi or Bluetooth for remote control and data sharing.
[0012] Further, the error identification unit uses a convolutional neural network based on an attention mechanism to locate and classify work error areas, including calculation errors, logical errors, and spatial imagination errors.
[0013] Further, a dynamic desensitization projection algorithm is also included to perform blurring processing on the projection content.
[0014] The present application has the advantages of:
[0015] 1、The present application realizes real-time analysis and processing of task data by using RGB-D camera, pressure sensor array and other means to collect multi-modal data, through cross-modal Transformer model and edge computing local feature extraction, reduces the dependence on cloud computing, and improves the system response speed.
[0016] 2、The present application realizes visual presentation and interactive feedback of task errors by using light field projection technology to project the analysis results in 3D highlight form and support touch interaction through the dynamic projection unit and tactile feedback unit of the projection feedback module, and combining the pulse feedback of the micro-current tactile unit, enhances the interactivity between students and the system, and helps students to understand and correct errors in time.
[0017] 3、The present application realizes multi-terminal data synchronization, personalized teaching suggestion generation and flexible configuration of hardware by using the cloud collaborative unit of the data management module, combining collaborative filtering algorithm and knowledge graph, uploading encrypted task data to the cloud to generate teaching analysis data, and cooperating with the detachable scanning projection module and multi-device linkage unit of the hardware architecture. BRIEF DESCRIPTION OF DRAWINGS
[0018] The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The drawings illustrate embodiments of the present application and, together with the description, serve to explain the principles of the present application.
[0019] In the drawings:
[0020] Figure 1 The system architecture diagram of the AI-based multi-modal task analysis and real-time projection feedback system in embodiment 1. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme of the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.
[0022] The application will be described in detail and specifically below through specific embodiments, so that the application can be better understood, but the following embodiments do not limit the protection scope of the application.
[0023] Embodiment 1
[0024] As Figure 1 shown, the AI-based multi-modal operation analysis and real-time projection feedback system comprises a multi-modal data acquisition module, an AI analysis module, a projection feedback module, an edge computing module and a data management module; the multi-modal data acquisition module and the AI analysis module are connected through a data transmission link, acquire multi-modal operation data and transmit the multi-modal operation data to the AI analysis module; the AI analysis module and the edge computing module constitute a cooperative processing architecture, and perform feature extraction and analysis on the multi-modal data; the AI analysis module outputs analysis results to the projection feedback module to generate feedback information; the data management module and the AI analysis module form bidirectional data interaction, and perform encrypted storage and management of data.
[0025] In specific embodiments, the multi-modal data acquisition module is connected to the AI analysis module through a USB 3.2 interface, the AI analysis module is equipped with an Intel i7-11800H processor, and runs a multi-modal data processing algorithm based on a PyTorch framework; the projection feedback module selects a Mingji W1700 short-focus projector, which supports 4K light field projection and touch interaction; the edge computing module deploys a NVIDIA Jetson Nano module, and communicates with the AI analysis module through an HDMI interface; the data management module realizes AES-256 encrypted storage based on an Ali Cloud OSS, and performs bidirectional interaction with the AI analysis module using a TCP / IP protocol.
[0026] Further, the multi-modal data acquisition module comprises an RGB-D camera, a pressure sensing array and an infrared thermal imaging module; the RGB-D camera acquires visual information of operation, the pressure sensing array acquires writing force data, and the infrared thermal imaging module captures facial micro-expression data.
[0027] In specific embodiments, the multi-modal data acquisition module is specifically configured as follows: the RGB-D camera uses an Intel RealSense D435, has a resolution of 1280×720, and acquires visual information such as operation handwriting track in real time; the pressure sensing array selects a FlexiForce A201 film sensor, has a sampling frequency of 100 Hz, and is integrated on the surface of an operation pad to acquire writing force data; the infrared thermal imaging module uses a FLIR Lepton 3.5, has a resolution of 640×512, is connected to a main control board through a USB interface to capture facial micro-expression changes, and synchronously generates a data stream marked with a time stamp.
[0028] Further, the AI analysis module comprises a multi-modal data fusion unit and an error identification unit; the multi-modal data fusion unit performs spatio-temporal alignment processing on handwriting trajectories, speech explanations and problem solving steps through a cross-modal Transformer model; the error identification unit performs high-precision identification of handwritten characters, formulas and symbols based on a convolutional recurrent neural network and a residual network OCR module, and analyzes mathematical problem errors in combination with geometric feature extraction and symbol matching algorithms, and the error identification unit adopts a ResNet+Transformer hybrid architecture optimization model.
[0029] In specific embodiments, in the AI analysis module, the multi-modal data fusion unit constructs a cross-modal model based on the Hugging Face Transformer library, inputs handwriting trajectory XY coordinate sequences, speech mel spectrum graphs and problem solving step texts, and realizes spatio-temporal alignment through CLS tokens; the error identification unit adopts a CRNN+ResNet-50 architecture, trains an OCR module based on PaddleOCR, and the handwriting character recognition accuracy reaches 98.5%, the mathematical formula analysis is assisted by the SymPy library for symbol matching, and the hybrid architecture realizes model optimization through PyTorch Lightning, and the analysis efficiency of geometric problem errors is improved.
[0030] Further, the projection feedback module comprises a dynamic projection unit and a tactile feedback unit; the dynamic projection unit adopts light field projection technology to project the analysis results in the form of 3D highlights to the work surface and supports touch interaction operation; the tactile feedback unit is a micro-current tactile unit, which generates pulse feedback signals based on preset error type trigger rules.
[0031] In specific embodiments, in the projection feedback module, the dynamic projection unit adopts a Mingji W1700 projector, realizes 10-point touch interaction on the work surface through an infrared touch frame, and projects the analysis results in the form of 3D highlights; the tactile feedback unit selects a TEConnectivity micro-current tactile module, which is built-in with an STM32F103 single-chip microcomputer, and outputs a 3.3V, 0.5ms pulse signal to drive a vibration motor when a logical error is detected, and triggers a 1ms pulse feedback for a calculation error.
[0032] Further, the edge computing module adopts an NVIDIA Jetson Nano module to perform multi-modal feature extraction locally, and performs real-time correlation processing on gesture trajectories and speech semantics, thereby reducing dependence on cloud computing.
[0033] In specific embodiments, the edge computing module adopts NVIDIA Jetson Nano 2GB version, and a lightweight model is deployed locally: the RGB-D data of Intel RealSense D435 is processed by OpenCV to extract HOG features to recognize gesture trajectories; the VGGish model is used to convert speech into a 128-dimensional feature vector, both of which are accelerated by TensorRT to realize real-time correlation processing, and are transmitted to the cloud through Wi-Fi 6, thereby reducing the cloud computing load and controlling the response delay within 200 ms.
[0034] Further, the data management module includes a cloud collaboration unit, which uploads encrypted homework data to the cloud, generates teaching optimization suggestions and teaching analysis data based on collaborative filtering algorithms and knowledge graphs, and performs multi-terminal data synchronization and homework data output.
[0035] In specific embodiments, the cloud collaboration unit of the data management module is built based on the Ali Cloud ECS server and uses the AES-256 encryption protocol; the collaborative filtering algorithm is implemented through the Surprise library, and the subject knowledge graph constructed by Neo4j generates teaching suggestions; multi-terminal synchronization is based on the WebSocket protocol, homework data supports PDF format export, and the parent terminal APP is developed based on Android Studio, integrating the FCM push function to synchronize AI comments and mistake books in real time.
[0036] Further, it also includes a hardware architecture, which includes a detachable scanning and projecting module, an intelligent adjusting module, and a human-computer interaction module; the scanning and projecting module includes a flatbed scanner and a short-focus projector; the intelligent adjusting module drives the stage through a motor and performs automatic focusing and position calibration through a photoelectric sensor; the human-computer interaction module includes a touch screen operation panel and a multi-device linkage unit: the touch screen operation panel supports handwriting input and voice instructions for setting analysis parameters; the multi-device linkage unit connects tablets and computers through Wi-Fi or Bluetooth for remote control and data sharing.
[0037] In specific embodiments, the scanning and projecting module is detachable: the flatbed scanner is Epson Perfection V850, and the short-focus projector is Ming Base W1700; the intelligent adjusting module drives the stage with a NEMA 17 stepper motor, and realizes ±0.5mm automatic focusing with an Omron E3Z-D62 photoelectric sensor; the touch screen of the human-computer interaction module is a 10.1-inch capacitive screen, and multi-device linkage realizes remote control and data sharing through the Broadcom BCM4356 Bluetooth 5.2+Wi-Fi 6 module.
[0038] Further, the error identification unit locates and classifies the work error area based on the attention mechanism of the convolutional neural network, including calculation errors, logical errors, and spatial imagination errors.
[0039] In specific embodiments, the error identification unit is constructed based on TensorFlow 2.5, adopts a ResNet50 backbone network embedded with an SENet attention module, inputs a 512x512 work image, and classifies it into three categories of calculation errors, logical errors, and spatial imagination errors through Softmax; the training data set contains 100,000 annotated works, and the Adam optimizer is used, wherein the recognition accuracy of spatial imagination errors is relatively high.
[0040] Further, it also includes a dynamic desensitization projection algorithm to perform blurring processing on the projection content.
[0041] In specific embodiments, the dynamic desensitization projection algorithm is realized through OpenCV: the name position in the projection content is recognized by PaddleOCR, 11x11 Gaussian blur processing is applied to the name area, only the first letter is kept as a clear character, and the rest of the characters are pixelated; the desensitized image is pushed in real time through the Mingji W1700 projector API, the privacy processing delay is <50ms, and the student's name, student ID, and other sensitive information are ensured to display only the first letter desensitization content when projected.
[0042] The specific embodiments of the present application are described in detail above, but it is only as an example, the present application is not equivalent to the above described specific embodiments. For those skilled in the art, any equivalent modification and replacement of the present application are also within the scope of the present application. Therefore, any equivalent transformation and modification made without departing from the spirit and scope of the present application should be covered within the scope of the present application.
Claims
1. An AI-based multi-modal job analysis and real-time projection feedback system, characterized in that: The system comprises a multi-modal data acquisition module, an AI analysis module, a projection feedback module, an edge computing module, and a data management module. The multi-modal data acquisition module is connected to the AI analysis module through a data transmission link, acquires multi-modal data of the task, and transmits the data to the AI analysis module. The AI analysis module and the edge computing module form a collaborative processing architecture, and perform feature extraction and analysis on the multi-modal data. The AI analysis module outputs the analysis results to the projection feedback module to generate feedback information. The data management module and the AI analysis module form a bidirectional data interaction, and perform encrypted storage and management of data.
2. The AI-based multi-modal job analysis and real-time projection feedback system of claim 1, wherein: The multi-modal data acquisition module comprises an RGB-D camera, a pressure sensor array, and an infrared thermal imaging module. The RGB-D camera acquires visual information of the task, the pressure sensor array acquires writing force data, and the infrared thermal imaging module captures facial micro-expression data. 3.The AI-based multi-modal job analysis and real-time projection feedback system of claim 2, wherein: The AI analysis module comprises a multi-modal data fusion unit and an error identification unit. The multi-modal data fusion unit performs spatio-temporal alignment processing on handwriting trajectories, voice explanations, and problem-solving steps through a cross-modal Transformer model. The error identification unit performs high-precision identification of handwritten characters, formulas, and symbols based on a convolutional recurrent neural network and a residual network OCR module, and analyzes mathematical problem errors by combining geometric feature extraction and symbol matching algorithms. The error identification unit uses a ResNet+Transformer hybrid architecture optimization model.
4. The AI-based multi-modal job analysis and real-time projection feedback system of claim 3, wherein: The projection feedback module comprises a dynamic projection unit and a tactile feedback unit. The dynamic projection unit uses light field projection technology to project the analysis results in 3D highlight form onto the task surface and supports touch interaction operations. The tactile feedback unit is a micro-current tactile unit that generates pulse feedback signals based on preset error type trigger rules. 5.The AI-based multi-modal job analysis and real-time projection feedback system of claim 4, wherein: The edge computing module uses an NVIDIA Jetson Nano module to perform multi-modal feature extraction locally, associate gesture trajectories and voice semantics in real time, and reduce dependence on cloud computing. 6.The AI-based multi-modal job analysis and real-time projection feedback system of claim 5, wherein: The data management module comprises a cloud collaboration unit that uploads encrypted task data to the cloud, generates teaching optimization suggestions and teaching analysis data based on collaborative filtering algorithms and knowledge graphs, and performs multi-terminal data synchronization and task data output. 7.The AI-based multi-modal job analysis and real-time projection feedback system of claim 6, wherein, The system also comprises a hardware architecture that includes a detachable scanning projection module, an intelligent adjustment module, and a human-computer interaction module. The scanning projection module comprises a flatbed scanner and a short-focus projector. The intelligent adjustment module drives the stage through a motor and performs automatic focusing and position calibration through a photoelectric sensor. The human-computer interaction module comprises a touch screen operation panel and a multi-device linkage unit. The touch screen operation panel supports handwriting input and voice commands for setting analysis parameters. The multi-device linkage unit connects tablets and computers through Wi-Fi or Bluetooth for remote control and data sharing. 8.The AI-based multi-modal job analysis and real-time projection feedback system of claim 7, wherein, The error identification unit uses a convolutional neural network based on an attention mechanism to locate and classify error areas, including calculation errors, logical errors, and spatial imagination errors. 9.The AI-based multi-modal job analysis and real-time projection feedback system of claim 8, wherein, Also included is a dynamic desensitization projection algorithm that performs blurring on the projection content.