A study partner system
The learning companion system, with its multi-layered communication architecture, enables efficient transmission of information across multiple terminals and comprehensive analysis of learning data. This solves the problems of poor communication connectivity, limited learning analysis, and fragmented teaching processes in existing teaching systems, thereby improving teaching efficiency and feedback adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI NAN YANG MODEL HIGH SCHOOL
- Filing Date
- 2026-03-31
- Publication Date
- 2026-06-26
Smart Images

Figure CN122285996A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart education technology, and in particular to a learning companion system. Background Technology
[0002] With the deepening of educational informatization, the online-offline integrated teaching model has become an important development direction for basic education. In actual teaching, especially in subjects such as geography that emphasize the combination of text and graphics and real-world exploration, students need to complete learning feedback through multimodal methods such as image taking and voice expression. Teachers, on the other hand, need to develop targeted teaching strategies based on the learning data of all students. This places higher demands on the interactive capabilities, data processing capabilities, and process linkage of the teaching system.
[0003] Currently, while existing interactive teaching systems can achieve basic information transmission and feedback collection, they still have many shortcomings in practical applications and are difficult to adapt to the full-process requirements of large-scale teaching: First, the multi-terminal communication and data transmission logic of the system is not well connected. The information flow between student terminals, teacher terminals, and data processing terminals lacks a clear link design, which easily leads to problems such as information transmission delays and data matching chaos, affecting the overall efficiency of interactive teaching. Second, the processing of student learning feedback is one-dimensional, only able to complete the basic feedback information reception, and unable to form systematic learning-related data and visual analysis reports based on learning feedback at different stages before and during class, making it difficult for teachers to fully grasp the overall learning situation. Third, the generation of teaching instructions has a low correlation with learning data. The formulation of post-class teaching instructions lacks the support of learning data, and the teaching links at the pre-class, in-class, and post-class stages are isolated from each other. The information and data of each link cannot form an effective linkage, making it difficult to build a complete closed loop of interactive teaching information. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this application provides a learning companion system that aims to solve problems such as the narrow coverage of traditional classroom interaction, fragmented processes, insufficient adaptability to subjects that emphasize text and graphics, such as geography, low diagnostic accuracy of learning situations, and low correlation between the generation of teaching instructions and learning data.
[0005] This application provides a learning companion system, the system comprising multiple first terminals, at least one second terminal, and a processing layer; Each of the first terminals and at least one of the second terminals are communicatively connected to the processing layer, and each of the first terminals is communicatively connected to at least one of the second terminals. The first terminal is used to receive pre-class teaching information and / or in-class teaching information, and transmit the pre-class learning feedback information uploaded by the target student in response to the pre-class teaching information and / or the in-class learning feedback information uploaded by the target student in response to the in-class teaching information to the processing layer; The processing layer is used to generate target learning data and target visualized classroom learning report based on the pre-class learning feedback information and / or the in-class learning feedback information, and to send the target learning data and target visualized classroom learning report to the second terminal; The second terminal is used to generate after-class teaching instructions and transmit them to the first terminal based on the learning data and visualized classroom learning reports pushed by the processing layer and in response to the target teacher's operation. The first terminal is also used to receive the after-class teaching instructions and send the collected after-class learning feedback information of the target students in response to the after-class teaching instructions to the processing layer; The processing is used to generate a target feedback result based on the pre-class learning feedback information, and / or the in-class learning feedback information, and the post-class learning feedback information, and send it to the first terminal; The pre-class learning feedback information includes pre-class learning feedback images and / or pre-class learning feedback audio information; the in-class learning feedback information includes in-class learning feedback images and / or in-class learning feedback audio information; and the post-class learning feedback information includes post-class learning feedback images and / or post-class learning feedback audio information.
[0006] In one optional implementation, the first terminal includes an image acquisition module, a sound acquisition module, a controller, an interaction module, a signal transceiver module, and a button module. The image acquisition module, the sound acquisition module, the interaction module, the signal transceiver module, and the button module are all electrically connected to the controller; The controller is used to control the image acquisition module to acquire the pre-class learning feedback image, and / or the in-class learning feedback image, and / or the post-class learning feedback image based on the target student's image acquisition operation on the button module, and to transmit the pre-class learning feedback image, and / or the in-class learning feedback image, and / or the post-class learning feedback image; The controller is used to control the sound acquisition module to acquire the pre-class learning feedback sound information, and / or the in-class learning feedback sound information, and / or the post-class learning feedback sound information based on the target student's sound acquisition operation on the button module, and to transmit the pre-class learning feedback sound information, and / or the in-class learning feedback sound information, and / or the post-class learning feedback sound information. The controller is used to send the pre-class learning feedback image, and / or the in-class learning feedback image, and / or the post-class learning feedback image, and / or the pre-class learning feedback sound information, and / or the in-class learning feedback sound information, and / or the post-class learning feedback sound information to the processing layer through the signal transceiver module; The controller is also configured to receive the target feedback result through the signal transceiver module and display the target feedback result to the target student through the interaction module.
[0007] In one optional implementation, the second terminal is used to generate the pre-class teaching information in response to a second operation by the target teacher, and to generate the in-class teaching information in response to a third operation by the target teacher, and to send the pre-class teaching information and the in-class teaching information to the controller through the signal transceiver module; The controller is used to control the interactive module to display the pre-class teaching information, the in-class teaching information, and the post-class learning information.
[0008] In an optional implementation, the controller is further configured to perform image preprocessing on the pre-class learning feedback image, and / or the in-class learning feedback image, and / or the post-class learning feedback image to obtain the processed pre-class learning feedback image, and / or the processed in-class learning feedback image, and / or the processed post-class learning feedback image, and send the processed pre-class learning feedback image, and / or the processed in-class learning feedback image, and / or the processed post-class learning feedback image to the processing layer through the signal transceiver module; The controller is also configured to perform sound preprocessing on the pre-class learning feedback sound information, and / or the in-class learning feedback sound information, and / or the post-class learning feedback sound information, to obtain processed pre-class learning feedback sound information, and / or processed in-class learning feedback sound information, and / or processed post-class learning feedback sound information, and send the processed pre-class learning feedback sound information, and / or processed in-class learning feedback sound information, and / or processed post-class learning feedback sound information to the processing layer through the signal transceiver module; The processing layer is used to perform a first image recognition and analysis on the processed pre-class learning feedback image, and / or the processed in-class learning feedback image, and / or the processed post-class learning feedback image, to obtain a first image recognition and analysis result; And / or, the processing layer is further configured to perform first speech-to-text processing on the processed pre-class learning feedback audio information, and / or the processed in-class learning feedback audio information, and / or the processed post-class learning feedback audio information, to obtain a first speech-to-text processing result; The processing layer is further configured to perform a first AI processing on the first image recognition and parsing result and / or the first speech-to-text processing result to obtain a first feedback result for the target student, and to transmit the first feedback result back to the controller through the signal transceiver module; wherein, the target feedback result includes the first feedback result; The controller is further configured to sequentially perform first data parsing processing, first format adaptation processing, first classification adaptation processing, and first signal adaptation processing on the first feedback result, and then feed back the processed first feedback result to the corresponding target student through the interaction module.
[0009] In an optional implementation, the image acquisition module is further configured to acquire a response feedback image obtained by the target student in response to the processed feedback result, and send the response feedback image to the controller; And / or, the sound acquisition module is further configured to acquire the response feedback sound information obtained by the target student in response to the processed feedback result, and send the response feedback sound information to the controller; The controller is also configured to perform image preprocessing on the response feedback image to obtain a preprocessed response feedback image, and / or perform sound preprocessing on the response feedback sound information to obtain preprocessed response feedback sound information, and transmit the preprocessed response feedback image, and / or the preprocessed response feedback sound information to the processing layer through the signal transceiver module; The processing layer is also used to perform image recognition and analysis on the preprocessed response feedback image to obtain a second image recognition and analysis result, and to perform second speech-to-text processing on the preprocessed response feedback sound information to obtain a second speech-to-text processing result. The processing layer is further configured to perform a second AI processing on the second image recognition and parsing result and / or the second speech-to-text processing result to obtain a second feedback result for the target student, and to transmit the second feedback result back to the controller through the signal transceiver module; wherein, the target feedback result includes the second feedback result; The controller is also used to sequentially perform second data parsing processing, second format adaptation processing, second classification adaptation processing, and second signal adaptation processing on the second feedback result, and to feed back the processed second feedback result to the corresponding target student through the interaction module until the second feedback result meets the preset feedback result.
[0010] In one optional implementation, the processing layer further includes a feedback unit and a learning analysis unit; The feedback unit is used to output a feedback score and generate feedback information, which includes the reasons for the loss of points and suggestions for improvement; wherein, both the first feedback result and the second feedback result include the corresponding feedback score and the corresponding feedback information. The learning analysis unit is used to summarize the first feedback result and the second feedback result to generate a visualized classroom learning report, and push the visualized classroom learning report to the corresponding second terminal.
[0011] In one optional implementation, the processing layer further includes a processing unit and a storage unit; The storage unit is used to store the knowledge point logic of the target subject, the image features of the target subject, and the student classroom interaction data of the target subject. The processing unit is used to construct a knowledge graph and an image recognition model of the target subject by using machine learning algorithms and combining the knowledge point logic of the target subject, the image features of the target subject, and the classroom interaction data of students in the target subject. Based on the knowledge graph of the target subject, the unit uses the image recognition model of the target subject to perform a first AI processing on the first image recognition analysis result to obtain a first feedback result for the target student, and performs a second AI processing on the second image recognition analysis result to obtain a second feedback result for the target student. The knowledge point logic of the target discipline is determined based on the knowledge point name, knowledge point level, inclusion relationship, prerequisite relationship, association relationship and difficulty level of the knowledge points in the target discipline. The image features of the target subject are determined based on the pre-class learning feedback image corresponding to the target subject, and / or the in-class learning feedback image, and / or the post-class learning feedback image; The student classroom interaction data for the target subject includes pre-class learning feedback images, in-class learning feedback images, and post-class learning feedback images for the target subject, as well as pre-class learning feedback audio information, in-class learning feedback audio information, and post-class learning feedback audio information for the target subject, along with response feedback images and response audio information for the target subject.
[0012] In an optional implementation, the processing unit is further configured to push suitable learning resources to the corresponding target student based on the feedback score after the second feedback result satisfies the preset feedback result or after the target student stops answering.
[0013] In one alternative implementation, the first terminal further includes a processor; The image acquisition module is used to acquire free-response images in response to the free-response operation of the target student; The controller is used to respond to the free-response question and answer operation of the target student, control the image acquisition module to acquire free-response question and answer images, and control the sound acquisition module to acquire free-response question and answer sound information; The controller is also used to transmit the free-response image and the free-response audio information to the processor; The processor is used to process the free-response image and / or the free-response sound information according to the preset computing power model and preset image recognition model built into the first terminal, determine the knowledge points of the corresponding target subject, generate progressive questions according to the knowledge points of the target subject and according to the preset questioning logic, and output them to the target student through the interaction module.
[0014] In an optional implementation, the processor is further configured to obtain the language behavior characteristics and emotional expression characteristics of the target student based on the free-response image and the free-response sound information, and determine the stress information of the target student based on the language behavior characteristics and emotional expression characteristics. When the stress information of the target student meets the preset warning conditions, the processing unit sends the warning information to the corresponding second terminal through the signal transceiver module.
[0015] In one optional implementation, the signal transceiver module supports WS protocol, HTTP protocol, Socket protocol, MQTT protocol, TCP protocol, UDP protocol and HTTPS protocol.
[0016] This application provides a learning companion system that addresses the technical shortcomings of existing interactive teaching systems. It establishes a three-layer communication architecture consisting of multiple first terminals, at least one second terminal, and a processing layer. The system clearly defines the functions of each terminal and the processing layer, as well as the information flow links. This adapts to the teaching needs of subjects like geography that emphasize multimodal feedback, effectively solving problems such as poor multi-terminal communication connectivity, a single dimension of learning information processing, fragmented teaching processes, and a disconnect between teaching instructions and learning data in existing systems. It achieves a closed-loop interaction of teaching information and learning feedback throughout the entire process—before, during, and after class. In this learning companion system, each first terminal and at least one second terminal establish a communication connection with the processing layer, and each first terminal establishes a communication connection with at least one second terminal. This replaces the ambiguous information flow of the existing system, avoids the problems of information transmission lag and data matching chaos from the architecture, and ensures the efficient and stable transmission of pre-class teaching information, in-class teaching information, post-class teaching instructions and various learning feedback information among multiple terminals, adapting to the multi-terminal concurrent interaction needs of large-scale teaching.
[0017] The processing layer can generate target learning data and a target visualized classroom learning report based on pre-class and in-class learning feedback information and send them to the second terminal. This breaks through the limitation of the existing system, which can only receive feedback information and cannot form systematic learning analysis results. It allows teachers to fully grasp the overall learning situation of all students' pre-class preparation and in-class learning through the second terminal, solving the problem of incomplete and unsystematic acquisition of teachers' learning information in the current teaching, and laying a data foundation for teachers to formulate targeted teaching strategies.
[0018] Subsequently, the second terminal can generate after-class teaching instructions based on the learning data pushed by the processing layer and the visualized classroom learning report. This changes the current situation where after-class teaching instructions lack learning support and rely solely on the teacher's subjective judgment. It allows the formulation of after-class teaching instructions to be deeply linked to the students' actual learning situation before and during class, ensuring the pertinence and effectiveness of after-class teaching instructions and effectively solving the technical problem of the disconnect between teaching instructions and learning data.
[0019] Furthermore, this learning companion system collects and transmits pre-class / in-class learning feedback through a first terminal, generates a learning report at the processing layer and sends it to a second terminal, which then generates post-class teaching instructions based on the learning situation. The first terminal collects and transmits post-class learning feedback, and the processing layer integrates feedback from all stages to generate a complete chain of target feedback results. This solves the technical problem of the existing system where the teaching links before, during, and after class are isolated from each other. It realizes the orderly flow and effective linkage of teaching information and learning feedback data in each link, forming a closed loop of teaching interaction from the distribution of teaching information and the collection of learning feedback to the analysis of learning situation, the generation of teaching instructions, and personalized feedback.
[0020] Secondly, the learning feedback information in this learning companion system includes images and / or sound in the form of images, during class, and after class. It is specifically adapted to the actual teaching needs of students in subjects such as geography that emphasize the combination of text and images and real-world exploration, where students complete learning feedback through image shooting and voice expression. This solves the problem of insufficient adaptability of existing systems to multimodal feedback, making the form of learning feedback collection more in line with the characteristics of subject teaching and enhancing the practical application value of the system.
[0021] Finally, the processing layer can integrate learning feedback information from all stages—before, during, and after class—to generate target feedback results and send them to the first terminal. This solves the limitation of the existing system's single-dimensional feedback processing, allowing the feedback results to comprehensively reflect the students' problems and shortcomings throughout the learning process, ensuring the comprehensiveness and accuracy of the feedback results. At the same time, the feedback results can be directly pushed to the corresponding target students through the first terminal, realizing personalized delivery of learning feedback and meeting the needs of personalized feedback in large-scale teaching. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of the structure of a learning companion system provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a first terminal provided in an embodiment of this application. Detailed Implementation
[0023] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. It is understood that the terms “first,” “second,” etc., as used herein may be used to describe various information or data, but these elements are not limited by these terms. These terms are only used to distinguish first information from another type of information. For example, without departing from the scope of this application, first action information may be referred to as second action information, and similarly, second action information may be referred to as first action information. Both first action information and second action information are action information, but they are not the same action information.
[0025] With the deepening of educational informatization, the online-offline integrated teaching model has become an important development direction for basic education. In actual teaching, especially in subjects such as geography that emphasize the combination of text and graphics and real-world exploration, students need to complete learning feedback through multimodal methods such as image taking and voice expression. Teachers, on the other hand, need to develop targeted teaching strategies based on the learning data of all students. This places higher demands on the interactive capabilities, data processing capabilities, and process linkage of the teaching system.
[0026] Currently, while existing interactive teaching systems can achieve basic information transmission and feedback collection, they still have many shortcomings in practical applications and are difficult to adapt to the full-process requirements of large-scale teaching: First, the multi-terminal communication and data transmission logic of the system is not well connected. The information flow between student terminals, teacher terminals, and data processing terminals lacks a clear link design, which easily leads to problems such as information transmission delays and data matching chaos, affecting the overall efficiency of interactive teaching. Second, the processing of student learning feedback is one-dimensional, only able to complete the basic feedback information reception, and unable to form systematic learning-related data and visual analysis reports based on learning feedback at different stages before and during class, making it difficult for teachers to fully grasp the overall learning situation. Third, the generation of teaching instructions has a low correlation with learning data. The formulation of post-class teaching instructions lacks the support of learning data, and the teaching links at the pre-class, in-class, and post-class stages are isolated from each other. The information and data of each link cannot form an effective linkage, making it difficult to build a complete closed loop of interactive teaching information.
[0027] The technical solutions shown in this application will now be described in detail through specific embodiments. It should be noted that the following embodiments may exist independently or in combination with each other; for identical or similar content, the description will not be repeated in different embodiments.
[0028] Reference Figure 1 This application provides a learning companion system, which includes multiple first terminals 100, at least one second terminal 200, and a processing layer 300.
[0029] Each of the first terminals 100 and at least one of the second terminals 200 are communicatively connected to the processing layer 300, and each of the first terminals 100 is communicatively connected to at least one of the second terminals 200.
[0030] The first terminal 100 is used to receive pre-class teaching information and / or in-class teaching information, and transmit the pre-class learning feedback information uploaded by the target student in response to the pre-class teaching information and / or the in-class learning feedback information uploaded by the target student in response to the in-class teaching information to the processing layer 300.
[0031] The processing layer 300 is used to generate target learning data and target visualized classroom learning report based on the pre-class learning feedback information and / or the in-class learning feedback information, and to send the target learning data and the target visualized classroom learning report to the second terminal 200.
[0032] The second terminal 200 is used to generate after-class teaching instructions and transmit them to the first terminal 100 in response to the operation of the target teacher, based on the learning data and visualized classroom learning report pushed by the processing layer 300.
[0033] The first terminal 100 is also used to receive the after-class teaching instructions and send the collected after-class learning feedback information of the target students in response to the after-class teaching instructions to the processing layer 300.
[0034] The processing is used to generate a target feedback result based on the pre-class learning feedback information, and / or the in-class learning feedback information, and the post-class learning feedback information, and send it to the first terminal 100.
[0035] The pre-class learning feedback information includes pre-class learning feedback images and / or pre-class learning feedback audio information; the in-class learning feedback information includes in-class learning feedback images and / or in-class learning feedback audio information; and the post-class learning feedback information includes post-class learning feedback images and / or post-class learning feedback audio information.
[0036] In this embodiment, the learning companion system is a smart teaching interaction system adapted to large-scale teaching scenarios in basic education (especially suitable for subjects such as geography that emphasize the combination of text and graphics and real-world exploration). Through a three-layer architecture of multiple student terminals (first terminal 100), at least one teacher terminal (second terminal 200), and a processing layer 300, it establishes a full-process teaching information interaction and learning feedback system for pre-class preparation, in-class interaction, and post-class consolidation. It solves the problems of unclear communication links, unsystematic learning analysis, fragmented teaching links, and insufficient multimodal feedback adaptation in traditional teaching interaction systems, and realizes intelligent linkage of the entire process of teaching information distribution, efficient collection of learning feedback, scientific analysis of learning data, and targeted generation of teaching instructions.
[0037] Among them, the first terminal 100 is a terminal device for students (which can be a customized interactive terminal, tablet, computer, etc.). It is the operation end and collection end for students to receive teaching information and complete learning feedback. The system is configured with multiple first terminals 100, each corresponding to a different target student, to meet the needs of full participation in large-scale teaching.
[0038] The first terminal 100 is used to receive pre-class teaching information (such as pre-class tasks, pre-class tests, subject pre-class materials, etc.), in-class teaching information (such as classroom questions, classroom exercises, landscape exploration tasks, etc.), and post-class teaching instructions (such as targeted homework, remedial tasks, extension exercises, etc.) issued by the second terminal 200. It is also used to collect multimodal learning feedback information (image and sound dual forms, adapted to the needs of real-world exploration and text-based answers in subjects such as geography) from the target students for teaching information / instructions at different stages, so that students do not have to answer only with text, which is more in line with the actual teaching scenario. It also transmits the collected pre-class, in-class, and post-class learning feedback information to the processing layer 300 in real time for subsequent data processing and analysis.
[0039] The second terminal 200 is a terminal device for teachers (which can be a computer, tablet, interactive whiteboard, etc.). The system is equipped with at least one terminal (multiple terminals can be configured according to teaching needs, such as the main lecturer terminal and the teaching assistant terminal). It is the operation terminal and decision-making terminal for teachers to issue teaching information, view learning data, and generate teaching instructions.
[0040] The second terminal 200 is used to receive target learning data (such as the whole class's pre-class preparation accuracy rate, in-class answer error rate, individual knowledge weaknesses, and other quantitative data) and target visualized classroom learning reports (such as data charts, high-frequency error analysis, learning distribution maps, and other visualized analysis results) generated and pushed by the processing layer 300, allowing teachers to intuitively and comprehensively grasp the learning situation of all students; it is also used to respond to teachers' manual operations (such as teachers adjusting the difficulty of homework and designing remedial tasks according to the learning situation) based on the learning data and reports pushed by the processing layer 300, generate targeted after-class teaching instructions, and ensure that the teaching instructions are deeply linked to the students' actual learning situation; and it can also distribute pre-class teaching information and in-class teaching information to all first terminals 100, and distribute after-class teaching instructions to the corresponding first terminals 100, thereby realizing the distribution of teaching information.
[0041] The processing layer 300 is the core data processing center and information transfer hub of the learning companion system (which can be deployed on a local server or a cloud server). It is a key node connecting the first terminal 100 and the second terminal 200, and is also the core technology for realizing learning analysis and feedback generation. It is responsible for background data processing and information distribution.
[0042] The processing layer 300 is used to extract, statistically analyze, and generate target learning data (quantified raw analysis data) and target visualized classroom learning reports (visualized analysis results for easy viewing by teachers) from the pre-class / in-class learning feedback information uploaded by the first terminal 100, and send both to the second terminal 200. It is also used to integrate, analyze, and personalize the pre-class, in-class, and post-class learning feedback information uploaded by the first terminal 100, generating target feedback results (such as students' knowledge weaknesses, reasons for answering questions incorrectly, and suggestions for learning improvement), and sending the results to the first terminal 100 of the corresponding target student to achieve personalized feedback. Furthermore, as an information relay hub, it processes the learning feedback information from the first terminal 100 and distributes it to the second terminal 200; it also distributes teaching instructions from the second terminal 200 to the corresponding first terminal 100, ensuring that the information flow between the student and teacher ends is unbiased and unconfused.
[0043] In this embodiment of the application, the processing procedure of the learning companion system is as follows: The second terminal 200 sends pre-class teaching information (such as geography pre-class assignments, pre-class landscape observation assignments, pre-class tests, etc.) to all first terminals 100. The first terminal 100 receives pre-class teaching information. After the target students complete the pre-class preparation, the first terminal 100 collects the students' pre-class learning feedback information (such as photos of geographical landscapes, photos of pre-class assignments, and pre-class answers given by voice). The first terminal 100 uploads the collected pre-class learning feedback information to the processing layer 300 for subsequent analysis.
[0044] The second terminal 200 sends in-class teaching information (such as classroom questions, classroom landscape exploration tasks, classroom exercises, etc.) to all first terminals 100. The first terminal 100 receives teaching information during class. After the target students complete the classroom tasks, the first terminal 100 collects the students' learning feedback information during class (such as voice answers to classroom questions, photos of classroom exploration landscapes, photos of homework segments, etc.). The first terminal 100 uploads the collected in-class learning feedback information to the processing layer 300; The processing layer 300 performs data statistics and analysis based on pre-class and in-class learning feedback information, and generates target learning data and target visualized classroom learning reports. The processing layer 300 sends the target learning data and the target visualized classroom learning report to the second terminal 200, through which the teacher can view the learning situation of the whole class.
[0045] Teachers manually operate based on the learning data and reports pushed by the processing layer 300 (such as designing targeted remedial tasks and adjusting the difficulty of homework) to generate after-class teaching instructions, and then send the instructions to all first terminals 100 through the second terminal 200. The first terminal 100 receives after-class teaching instructions. After the target students complete the after-class tasks, the first terminal 100 collects the students' after-class learning feedback information (such as photos of after-class assignments, landscape pictures of after-class explorations, and after-class answers given by voice). The first terminal 100 uploads the collected after-school learning feedback information to the processing layer 300; The processing layer 300 integrates learning feedback information from all stages—before, during, and after class—to conduct personalized analysis for each target student and generate target feedback results (such as individual knowledge weaknesses, reasons for errors, and suggestions for improvement). The processing layer 300 sends the target feedback results to the first terminal 100 of the corresponding target student, and the student views the personalized feedback, completing the entire learning loop.
[0046] Based on the above description, the learning feedback information in the learning companion system provided in this application embodiment includes images and / or sound in the form of images, in-class and post-class, rather than the traditional single form of text. It is specifically adapted to the needs of subjects such as geography that emphasize the combination of text and images and real-world exploration: students can complete image feedback by taking pictures of landscapes in subjects such as geography and handwritten images of assignments, and complete answer feedback by expressing themselves through voice. They are not limited to text input, which is more in line with the actual teaching scenarios of the subject and improves the convenience and effectiveness of students' learning feedback.
[0047] Furthermore, the processing layer 300 does not only analyze feedback during class, but integrates feedback information from before and during class to generate learning data and reports. This allows teachers to grasp the complete learning status of students from pre-class preparation to class, avoiding one-sided judgments based solely on classroom performance and ensuring the targeted nature of post-class teaching instructions. In addition, the processing layer 300 integrates feedback information from all stages of class—before, during, and after class—to generate personalized target feedback results. This ensures that the feedback results comprehensively reflect students' learning problems, rather than just error prompts at a single stage, achieving personalized guidance with one feedback per student and adapting to the needs of differentiated instruction in large-scale teaching.
[0048] Furthermore, the embodiments of this application link the three teaching stages of pre-class, in-class, and post-class through the information flow of teaching information, learning feedback, learning data, teaching instructions, and personalized feedback. This breaks the problem of the traditional teaching system where each stage is independent and data is not shared. It transforms teaching from separate pre-class preparation, independent in-class interaction, and blind post-class consolidation into a systematic teaching process with full linkage, achieving a dual improvement in teaching efficiency and relevance.
[0049] The following explanations are provided for some of the terms mentioned above: Target students: Students who correspond one-to-one with a single first terminal 100, are the main operators of the first terminal 100, and are also the main generators of learning feedback; Target learning data: The quantitative raw data (such as accuracy, error rate, and mastery of knowledge points) obtained by the processing layer 300 based on the analysis of students' pre-class / in-class learning feedback is the basic data for teachers to formulate teaching instructions; Visualized Classroom Learning Report: The processing layer 300 visualizes the learning data and analyzes the results (such as bar charts, pie charts, high-frequency error lists, learning distribution tables, etc.) to help teachers grasp the overall learning situation of the class more intuitively and efficiently. Target feedback results: The processing layer 300 provides personalized assessment results based on the analysis of students' learning feedback throughout the entire process (such as individual knowledge weaknesses, reasons for answering questions incorrectly, and suggestions for learning improvement). This is exclusive feedback for a single target student. "and / or" indicates any combination of multiple forms / information, such as pre-class learning feedback images and / or pre-class learning feedback audio information. This means that students can upload only images, only audio, or both images and audio. All three forms are supported by the system, improving the flexibility of feedback.
[0050] Reference Figure 2 In one optional implementation, the first terminal 100 includes an image acquisition module 101, a sound acquisition module 102, a controller 103, an interaction module 104, a signal transceiver module 105, and a button module 106. The image acquisition module 101, the sound acquisition module 102, the interaction module 104, the signal transceiver module 105, and the button module 106 are all electrically connected to the controller 103. The controller is used to control the image acquisition module 101 to acquire the pre-class learning feedback image, and / or the in-class learning feedback image, and / or the post-class learning feedback image based on the target student's image acquisition operation on the button module, and to obtain the acquired pre-class learning feedback image, and / or the in-class learning feedback image, and / or the post-class learning feedback image; The controller is used to control the sound acquisition module 102 to acquire the pre-class learning feedback sound information, and / or the in-class learning feedback sound information, and / or the post-class learning feedback sound information based on the target student's image acquisition operation on the button module, and to obtain the acquired pre-class learning feedback sound information, and / or the in-class learning feedback sound information, and / or the post-class learning feedback sound information; The controller 103 is used to send the pre-class learning feedback image, and / or the in-class learning feedback image, and / or the post-class learning feedback image, and / or the pre-class learning feedback sound information, and / or the in-class learning feedback sound information, and / or the post-class learning feedback sound information to the processing layer 300 through the signal transceiver module 105. The controller 103 is also used to receive the target feedback result through the signal transceiver module 105 and display the target feedback result to the target student through the interaction module 104.
[0051] In this embodiment, the image acquisition module 101 is a hardware unit with image shooting / acquisition function (such as a high-definition camera, barcode scanning module, image scanner, etc.), which is the core hardware for the first terminal 100 to realize image-based learning feedback acquisition, and is specifically used for geographical and other applications that require real-scene shooting and operation photography.
[0052] For example, the image acquisition module 101 is used to capture images such as landscape maps, handwritten pre-class assignment photos, and pre-class test answer sheets (pre-class learning feedback images) for pre-class geography pre-study tasks; during class, it captures images such as real-world geographical scenes for classroom exploration, handwritten answers for classroom exercises, and on-site photos for landscape recognition tasks (in-class learning feedback images); and after class, it captures images such as handwritten homework, geographical landscape maps for post-class extended exploration, and photos of the completion status of supplementary tasks (post-class learning feedback images). After acquiring the images, the image acquisition module 101 sends the image data (raw shooting data, such as JPG / PNG format) to the controller 103 without requiring manual intervention from students, ensuring the timeliness of data acquisition. This image acquisition module 101 can support common image acquisition needs in subjects such as geography (such as wide-angle shooting and high-definition macro shooting), and can clearly capture landscape details and handwritten text, avoiding recognition failures in subsequent processing layer 300 due to image blurring.
[0053] The sound acquisition module 102 is a hardware unit with sound pickup / recording functions (such as a noise-canceling microphone, array microphone, voice recording module, etc.). It is the hardware for the first terminal 100 to realize voice-based learning feedback acquisition and is adapted to the voice acquisition needs in noisy classroom environments.
[0054] For example, the sound acquisition module 102 is used to collect students' oral answers to pre-class assignments, voice questions about pre-class content, and voice descriptions of landscape observations before class (pre-class learning feedback sound information); to collect voice answers to classroom questions, voice analysis of geographical landscape features, and voice interactions with teachers during class (in-class learning feedback sound information); and to collect voice answers to homework assignments, voice reports on remedial tasks, and voice questions about feedback results after class (after-class learning feedback sound information). Furthermore, it incorporates a basic noise reduction algorithm to filter out invalid background noise and retain only the clear voice of the target students, improving the accuracy of the subsequent speech-to-text conversion in the 300-level processing layer. After acquiring the voice data, the audio data (such as MP3 / WAV format) is sent to the controller 103 to ensure real-time voice feedback.
[0055] Button module 106 is the physical interaction entry point for students in the first terminal. It is a physical switch that triggers core operations such as image acquisition and sound acquisition. Its function is to allow students to quickly initiate the learning feedback acquisition process through simple button operations, reduce the operation threshold of classroom interaction, and adapt to the high-efficiency use needs in large class scenarios.
[0056] Furthermore, in this embodiment of the application, the only triggering condition for starting the image acquisition module and the sound acquisition module is when the student performs an image acquisition operation on the button module (such as pressing the photo / record button).
[0057] The controller will not automatically activate the camera or microphone without a button press, preventing silent background image or audio capture of students. In other words, the button module and controller operate on a one-way communication channel: an electrical signal is only generated when a student presses a button, at which point the controller executes the capture logic. The controller cannot bypass the button module to directly send capture commands to the image / audio capture module, thus eliminating the possibility of automatic background capture at the hardware level.
[0058] The controller 103 is the central processing unit of the first terminal 100 (such as MCU microcontroller 103, single-chip microcomputer, embedded processor such as ESP32 / STM32, Android motherboard, etc.), and is the scheduling core of all modules, responsible for data transfer, instruction parsing, and flow control.
[0059] For example, the controller 103 is used to: The system receives image data from each stage uploaded by the image acquisition module 101 and voice data from each stage uploaded by the sound acquisition module 102, and temporarily caches them (to avoid data loss). According to the preset rules, the cached image / audio data is sent to the processing layer 300 through the signal transceiver module 105, supporting single-type data transmission (such as transmitting only pre-class images) and multi-type data combination transmission (such as transmitting in-class images and post-class audio at the same time). The signal transceiver module 105 receives the target feedback results (such as personalized feedback in the form of text, images, and voice) sent by the processing layer 300, and parses the data format (such as converting the encrypted data of the processing layer 300 into a format that the terminal can recognize). The parsed target feedback result is sent to the interaction module 104, and a display command is issued to control the display content and timing of the interaction module 104 (such as pop-up display or full-screen display). The system centrally schedules the start and stop of each module (e.g., when a student triggers a data collection operation, the system controls the image / sound module to start; after data transmission is completed, the system controls the module to go into sleep mode), avoiding unnecessary power consumption and improving the terminal's battery life.
[0060] The interactive module 104 is a human-computer interaction interface hardware unit (such as a touch screen, LED screen, voice broadcaster, etc., with a preference for touch screens) for target students. It is a visual / interactive carrier for students to receive teaching information and view feedback results.
[0061] For example, the interaction module 104 is used to receive instructions from the controller 103, display pre-class teaching information, in-class teaching information, and post-class teaching instructions (such as pre-class task text, classroom questions, and homework requirements) issued by the second terminal 200 / processing layer 300; and receive target feedback results issued by the controller 103, displaying them to students in an intuitive form, including: Text-based: Reasons for lost points, suggestions for improvement, key knowledge points, etc.; Image-based: Images of assignments with incorrect answers marked, landscape feature analysis maps, etc.; Voice-related: Feedback and suggestions for voice broadcast (optional, adapted to different students' usage habits); The interactive module 104 supports students' touch operations (such as clicking the image acquisition button, sending feedback button, and viewing details button). The operation commands are transmitted back to the controller 103 via electrical connection to trigger the corresponding module's action (such as starting the image acquisition module 101). The interactive module 104 can display the terminal's working status in real time (such as data uploading, successful reception, network error), allowing students to clearly understand the operation results.
[0062] The signal transceiver module 105 is a hardware unit (such as a Wi-Fi module, 4G / 5G module, Bluetooth module, Ethernet interface, etc.) used to realize data communication between the first terminal 100 and the outside (processing layer 300, second terminal 200), and is the data transmission channel of the first terminal 100.
[0063] For example, the signal transceiver module 105 is used to receive image / voice data sent by the controller 103 and send it to the processing layer 300 through a preset communication protocol (such as HTTP, MQTT, TCP / IP). It supports the parallel transmission of multiple types of data (such as simultaneous uploading of images and voice) and has a breakpoint resume function (automatic resume transmission when the network is restored after an interruption). It receives the target feedback results sent by the processing layer 300 and the teaching information / instructions sent by the second terminal 200 and transmits the data to the controller 103 in real time. It is compatible with the communication protocols of the processing layer 300 and the second terminal 200 to ensure the compatibility of cross-device data transmission (such as the terminal can automatically adapt if the processing layer 300 uses the MQTT protocol). The signal transceiver module 105 can have a built-in signal enhancement / anti-interference design to adapt to the scenario of concurrent communication of multiple terminals in the classroom and avoid data transmission lag and loss.
[0064] In one optional example, the signal transceiver module 105 supports WS protocol, HTTP protocol, Socket protocol, MQTT protocol, TCP protocol, UDP protocol and HTTPS protocol.
[0065] Optionally, the second terminal 200 is used to generate the pre-class teaching information in response to the second operation of the target teacher, and to generate the in-class teaching information in response to the third operation of the target teacher, and to send the pre-class teaching information and the in-class teaching information to the controller 103 through the signal transceiver module 105; The controller 103 is used to control the interactive module 104 to display the pre-class teaching information, the in-class teaching information, and the post-class learning information.
[0066] In this embodiment of the application, the specific operation of the second operation may include: the teacher inputting a pre-class task (such as "taking a picture of the river landscape in the hometown and describing its features") in the editing interface of the second terminal 200; uploading pre-class materials (such as micro-lessons on geographical knowledge points and a pre-class test bank); selecting a preset pre-class task template and confirming its distribution; and adjusting parameters such as the deadline and display format of the pre-class task.
[0067] The specific operation of the third operation may include: teachers inputting classroom questions (such as analyzing the distribution pattern of this climate type) on the classroom interaction interface of the second terminal 200; uploading in-class exploration materials (such as geographical landscape maps and classroom exercise documents); selecting classroom interaction methods (such as quiz-and-answer, all students answering) and generating corresponding instructions; and adjusting the difficulty of in-class tasks for different student groups (such as pushing basic questions to struggling students and extension questions to high-achieving students).
[0068] For example, the second terminal 200 may have a built-in teaching information editing / generation module (software level, such as a web page editor or client operation interface) to respond to the teacher's second / third operation, converting the teacher's manual input / selection into standardized teaching information data. The generated pre-class / in-class teaching information is converted into a format that the first terminal 100 can recognize (such as text, images, audio, structured instructions, etc.) to avoid the inability to display the information on the student's end due to format incompatibility.
[0069] For example, the Word version of the pre-class assignment uploaded by the teacher will be automatically converted into a graphic and text format that can be displayed by the interactive module 104 of the first terminal 100; the voice classroom instructions recorded by the teacher will be converted into a standard audio format (MP3).
[0070] Teachers can specify the distribution scope of teaching information through the second / third operation. They can choose to distribute it to all students (all first terminals 100) or to distribute it personalized (only to the first terminals 100 of specific students), adapting to the needs of differentiated instruction. The generated teaching information supports subject-specific formats such as landscape icon annotations, topographic profiles, and climate data charts, without requiring teachers to convert the formats, directly adapting to the display needs of the first terminal 100.
[0071] In this embodiment of the application, the presentation of pre-class teaching information may include: text: description of pre-class tasks (such as "take pictures and describe the terrain features of your hometown"); pictures: schematic diagrams of geographical knowledge points and pre-class reference landscape pictures; audio: pre-class guidance voice recorded by the teacher; attachments: pre-class micro-lessons and pre-class test question bank that can be clicked to view.
[0072] The presentation of teaching information during class may take the following forms: full-screen text: core classroom questions (such as "analyze the causes of tropical rainforest climate"); combination of text and images: landscape images and questions (such as showing an image of the Amazon rainforest and asking questions about the ecological role of vegetation in the region); countdown: deadline for answering classroom questions; interactive buttons: start answering, submit answers, and view examples.
[0073] The presentation of after-class teaching information may include: list format: personalized homework list (e.g., 1. Review the knowledge points of terrain types; 2. Take 3 landscape pictures of different terrains); annotation format: screenshots of wrong questions with correction requirements; progress display: progress bar of after-class task completion; resource links: links to micro-lessons and exercises for targeted supplementation.
[0074] In an optional implementation, the controller 103 is further configured to perform image preprocessing on the pre-class learning feedback image, and / or the in-class learning feedback image, and / or the post-class learning feedback image to obtain the processed pre-class learning feedback image, and / or the processed in-class learning feedback image, and / or the processed post-class learning feedback image, and send the processed pre-class learning feedback image, and / or the processed in-class learning feedback image, and / or the processed post-class learning feedback image to the processing layer 300 through the signal transceiver module 105; The controller 103 is further configured to perform sound preprocessing on the pre-class learning feedback sound information, and / or the in-class learning feedback sound information, and / or the post-class learning feedback sound information, to obtain processed pre-class learning feedback sound information, and / or processed in-class learning feedback sound information, and / or processed post-class learning feedback sound information, and send the processed pre-class learning feedback sound information, and / or processed in-class learning feedback sound information, and / or processed post-class learning feedback sound information to the processing layer 300 through the signal transceiver module 105; The processing layer 300 is used to perform a first image recognition and analysis on the processed pre-class learning feedback image, and / or the processed in-class learning feedback image, and / or the processed post-class learning feedback image, to obtain a first image recognition and analysis result; And / or, the processing layer 300 is further configured to perform first speech-to-text processing on the processed pre-class learning feedback sound information, and / or the processed in-class learning feedback sound information, and / or the processed post-class learning feedback sound information, to obtain a first speech-to-text processing result; The processing layer 300 is further configured to perform a first AI processing on the first image recognition and parsing result and / or the first speech-to-text processing result to obtain a first feedback result for the target student, and to transmit the first feedback result back to the controller 103 through the signal transceiver module 105; wherein, the target feedback result includes the first feedback result; The controller 103 is further configured to sequentially perform first data parsing processing, first format adaptation processing, first classification adaptation processing, and first signal adaptation processing on the first feedback result, and then feed back the processed first feedback result to the corresponding target student through the interaction module 104.
[0075] In this embodiment, the controller 103 is used to optimize the original feedback data, reduce the computational pressure on the processing layer 300, and improve the accuracy of subsequent intelligent analysis. The preprocessing is only a basic standardization operation and does not involve complex recognition and analysis. For learning feedback images (such as geographical landscape images, handwritten homework images, etc.) before, during, and after class, the controller 103 will perform format standardization conversion (e.g., converting to JPG / PNG and other preset formats of the processing layer 300), basic noise reduction and deblurring (filtering paper spots, sharpening handwritten words and landscape details), angle correction and cropping (correcting oblique images, removing invalid areas such as desktops / backgrounds), lossless compression (reducing file size), etc. The processing combination is automatically matched according to the image type. For example, landscape images focus on noise reduction and compression, while homework images focus on correction and sharpening. For audio feedback information (students' oral answers, landscape analysis, etc.) received before, during, and after class, the controller 103 performs format standardization (converting to MP3 / WAV format and unifying the sampling rate), classroom environment noise reduction (filtering background noise and retaining core audio segments), silent segment clipping (removing blank segments due to pauses in answers), and volume normalization (unifying the speech standard at different volumes), with a focus on ensuring the clear extraction of professional terms in subjects such as geography (e.g., karst landforms, contour lines). After preprocessing, the controller 103 can also add terminal and stage identifiers to the processed images / audio, and send them to the processing layer 300 through the signal transceiver module 105, supporting single-type data transmission or batch transmission of multiple types of data.
[0076] The processing layer 300 receives preprocessed data transmitted from the first terminal 100. It generates the first feedback result through image recognition and analysis, speech-to-text processing, and AI processing. This first feedback result is also a component of the target feedback result. All analyses are based on the target subject's proprietary model and lexicon, ensuring professionalism and accuracy. First, it performs image recognition and analysis. The processing layer 300, relying on computer vision algorithms and the target subject's proprietary recognition model, performs structured analysis on the preprocessed image. For example, for a geographical landscape map, it identifies the landscape type, core features, and student annotation information; for a handwritten homework map, it identifies the handwritten answer content, answer completeness, and geographical symbol annotations. Finally, it outputs the first image recognition and analysis result that can be used for AI analysis. Simultaneously, the processing layer 300 performs first speech-to-text processing on the preprocessed audio information. Through speech recognition algorithms and the target subject's proprietary terminology lexicon, it converts the student's speech feedback into text content. It prioritizes the recognition of subject-specific terminology to avoid transcription errors, outputting the first speech-to-text processing result. It also supports flexible processing modes such as image-only analysis, speech-only transcription, or a combination of both.
[0077] Subsequently, the processing layer 300 performs the first AI processing based on the above two results. Using a dedicated AI model for the target subject (learning analysis, error assessment, and knowledge point matching model), it completes knowledge point matching (determining the student's knowledge mastery and weaknesses), error assessment (analyzing error types and core causes), and personalized suggestion generation (providing targeted improvement plans based on weaknesses). All information is integrated into a structured first feedback result, including student identification, feedback stage, knowledge point mastery status, error causes, improvement suggestions, and related learning resources, ensuring the feedback's relevance and practicality. Finally, based on the terminal identification, the processing layer 300 transmits the first feedback result back to the corresponding target student's first terminal 100 controller 103 via the signal transceiver module 105.
[0078] After receiving the first feedback result from the processing layer 300, the controller 103 of the first terminal 100 transforms the structured raw feedback data into a form that can be displayed on the terminal and is easy for students to understand, without changing the core content of the feedback, only optimizing the format and form. The first step is data parsing and processing. This involves parsing the structured data (such as JSON format) returned by the processing layer 300, extracting core information such as knowledge points, error reasons, and improvement suggestions, and filtering redundant fields such as internal identifiers of the processing layer 300 to simplify the data structure. The second step is format adaptation processing. This involves converting the parsed core information into a display format supported by the interaction module 104, such as converting text suggestions into rich text, knowledge point links into clickable buttons, and matching error points with corresponding geographical maps to adapt to the terminal's display capabilities. The third step is category adaptation processing. Based on the stage attributes of the feedback results (pre-class / in-class / post-class) and the knowledge point type (terrain / climate / river), corresponding display templates are matched. For example, in-class errors are marked in red, and post-class supplementation suggestions are displayed in a list, making the display format fit the learning scenario. The fourth step is signal adaptation processing. This involves converting the optimized feedback results into hardware-recognizable signals such as display signals and voice broadcast signals from the interaction module 104, while simultaneously verifying data integrity to ensure no missing data.
[0079] Afterwards, the controller 103 displays the processed first feedback result to the target student through the interaction module 104. The display format is based on a combination of text and images, aligning with the characteristics of the geography subject, and also supports personalized extensions, making the feedback result more intuitive and practical. The interaction module 104 will display the error reasons and improvement suggestions in segments, accompanied by corresponding diagrams (such as climate type comparison diagrams and terrain feature maps), and mark the errors in the student's feedback in red (such as marking "should be hills, not plains" on the landscape map); it also supports optional voice broadcast function to meet the needs of lower grade students or visually impaired students; it also has interactive extension buttons, such as "related knowledge point exercises", "micro-lesson links", and "error notebook collection", to realize extended learning based on the feedback result; the feedback result is automatically stored in the terminal's local cache, and students can view the history at any time for easy review and improvement.
[0080] Overall, the first terminal 100 handles lightweight preprocessing and result adaptation, controlling hardware costs and power consumption, and adapting to the terminal deployment needs of large-scale teaching. The processing layer 300 focuses on high-performance intelligent analysis, relying on subject-specific models to ensure analytical accuracy and solve the problem of insufficient subject adaptation of general models. From the collection of multimodal student feedback to the final personalized feedback display, a seamless closed-loop process is formed, upgrading the system from simple data transmission to intelligent analysis and personalized feedback. This truly realizes intelligent processing of learning feedback at all stages of learning in each subject before, during, and after class, providing technical support for individualized instruction under large-scale teaching.
[0081] Furthermore, the image acquisition module 101 is also used to acquire the response feedback image obtained by the target student in response to the processed feedback result, and send the response feedback image to the controller 103; And / or, the sound acquisition module 102 is further configured to acquire the response feedback sound information obtained by the target student in response to the processed feedback result, and send the response feedback sound information to the controller 103; The controller 103 is also used to perform image preprocessing on the response feedback image to obtain a preprocessed response feedback image, and / or to perform sound preprocessing on the response feedback sound information to obtain preprocessed response feedback sound information, and to transmit the preprocessed response feedback image, and / or the preprocessed response feedback sound information to the processing layer 300 through the signal transceiver module 105; The processing layer 300 is also used to perform image recognition and analysis on the preprocessed response feedback image to obtain a second image recognition and analysis result, and to perform second speech-to-text processing on the preprocessed response feedback sound information to obtain a second speech-to-text processing result. The processing layer 300 is further configured to perform a second AI processing on the second image recognition and parsing result and / or the second speech-to-text processing result to obtain a second feedback result for the target student, and to transmit the second feedback result back to the controller 103 through the signal transceiver module 105; wherein, the target feedback result includes the second feedback result; The controller 103 is also used to perform second data parsing processing, second format adaptation processing, second classification adaptation processing and second signal adaptation processing on the second feedback result in sequence, and to feed back the processed second feedback result to the corresponding target student through the interaction module 104 until the second feedback result meets the preset feedback result or the target student stops answering.
[0082] In this embodiment, after viewing the first feedback result, the target student can upload a reply feedback image through the image acquisition module 101 or upload reply feedback audio information through the audio acquisition module 102. The controller 103 performs the same lightweight preprocessing on the reply image / audio as the initial feedback and then transmits it to the processing layer 300. The processing layer 300 performs secondary image recognition and analysis, secondary speech-to-text processing on the preprocessed reply data, and then generates a second feedback result through second AI processing and sends it back. The controller 103 performs adaptation processing on the second feedback result and displays it to the student. This process continues to iterate until the second feedback result meets the preset feedback result (such as correct knowledge points or answers that meet the requirements), or the target student stops replying.
[0083] Based on the above description, after a student answers incorrectly on the first attempt, they can correct their answer based on the feedback and try again. The system then analyzes the feedback and provides further suggestions for improvement until the student masters the knowledge point. This mechanism solves the problem that traditional teaching methods cannot completely resolve students' cognitive errors with a single feedback, achieving truly personalized tutoring.
[0084] Furthermore, this iterative feedback mechanism focuses on students' core errors. Through a process of error discovery, student correction, effect verification, and further correction, it forces students to repeatedly refine their weak knowledge points, avoiding the problems of forgetting feedback immediately and repeating errors. For example, if a student initially misidentifies a landscape type, they can retake the photo and describe it during the review. The system then performs a second analysis to verify whether the student has mastered the method for determining landscape characteristics, until the feedback result meets the standard, significantly improving the efficiency of mastering core geographical knowledge points.
[0085] Furthermore, the image / sound preprocessing logic for the response feedback, the recognition and analysis / AI processing logic of the processing layer 300, and the adaptation and display logic of the controller 103 all reuse the technical architecture of the initial feedback (only adding a "second" identifier to distinguish the stage), without the need to redesign the core modules. This ensures the scalability of the system functions, avoids the increased costs caused by repeated development, and ensures the consistency of data processing across multiple terminals and stages, adapting to the system stability requirements of large-scale teaching.
[0086] Finally, the setting of "preset feedback results" in this embodiment (such as a knowledge point recognition accuracy rate of ≥90% and no core errors in the answer description) transforms the student's learning effect from subjective judgment to objective quantification. The processing layer 300 can automatically determine whether to terminate the iteration based on the preset standards, which not only reduces the workload of teachers' manual verification, but also allows teachers to grasp the knowledge point attainment status of each student, providing more refined learning data support for teachers to adjust their subsequent teaching strategies.
[0087] In one optional implementation, the processing layer 300 further includes a feedback unit and a learning analysis unit; The feedback unit is used to output a feedback score and generate feedback information, which includes the reasons for the loss of points and suggestions for improvement; wherein, both the first feedback result and the second feedback result include the corresponding feedback score and the corresponding feedback information. The learning analysis unit is used to summarize the first feedback result and the second feedback result to generate a visualized classroom learning report, and push the visualized classroom learning report to the corresponding second terminal 200.
[0088] In this embodiment, the feedback unit is used to output quantifiable feedback scores (such as percentage or grade system) for the student's initial feedback and follow-up feedback, based on the first / second image recognition and parsing results, the speech-to-text results, and the corresponding subject knowledge point weights (e.g., higher scores for core knowledge points). The scoring logic closely follows the preset achievement standards, for example: Initial feedback (first feedback result): The student answered incorrectly about the characteristics of tropical rainforest climate. The core knowledge point accounts for 60% of the score, and the feedback score is 40 points. Second Feedback Result: After correcting their mistakes, the student answers correctly and receives a score of 100. It should be understood that the score directly reflects the student's mastery of the corresponding knowledge point, and both the first and second feedback results carry their own independent scores, clearly demonstrating the improvement before and after the second feedback.
[0089] The feedback unit is also used to generate structured feedback information based on the scoring, including reasons for the points lost and suggestions for improvement, and the feedback information is strongly correlated with the feedback score: Reasons for point deduction: The core logic of the point deduction is based on positioning (e.g., 50 points were lost due to incorrect landscape type identification, failure to grasp the characteristics of hills with an elevation of 200-500 meters), and 20 points were lost due to incorrect terminology, writing "karst landform" as "karst landform"), rather than a general description; Improvement suggestions: Provide actionable guidance for the reasons for lost points (e.g., if the loss is due to confusion about climate characteristics, suggest comparing the precipitation bar charts for tropical rainforest / tropical monsoon climates and completing 3 specific practice questions). Furthermore, the improvement suggestions in the first and second feedback results should be progressive (the suggestions after the response should focus more on the remaining errors). Ultimately, both the first and second feedback results should integrate feedback scores and information to form complete quantitative and qualitative feedback data.
[0090] The learning analysis unit is used to summarize the first and second feedback results of all students in the class. For example, it may include: Scoring dimensions: average score of the whole class in the initial feedback, average score after the re-answer, score improvement, percentage of students in each score range (0-60 / 60-80 / 80-100); Knowledge point dimension: Error rate and correction rate of each subject's knowledge points; Error type dimension: the percentage distribution of types such as terminology errors and feature description errors; Individual dimension: Changes in scores before and after a single student's response, a list of students who did not meet the standard, and key errors.
[0091] The learning analysis unit can also be used to transform the summarized data into a visual form to adapt to different subject teaching scenarios. The learning analysis unit can also automatically push the generated visual classroom learning report to the corresponding teacher's second terminal 200 according to the classroom identifier and teacher terminal identifier. It supports real-time push (generated immediately after class) or timed push (such as 10 minutes after class) to ensure that teachers can grasp the learning situation of the whole class in a timely manner.
[0092] Based on the above description, the feedback score output by the feedback unit in this embodiment transforms the student's learning effect from a vague right-or-wrong judgment into a quantifiable score assessment. Teachers can intuitively judge the student's mastery of knowledge points through the score, and the comparison of scores before and after the response can clearly reflect the student's corrective effect. For example, teachers can quickly identify students whose scores have improved by less than 20 points and provide targeted one-on-one tutoring; by analyzing the score distribution, teachers can pinpoint the knowledge points where the whole class generally scores low and determine the key content to be explained in subsequent classes, thus solving the technical problem that traditional feedback is mainly qualitative and difficult to focus on.
[0093] Furthermore, the feedback unit generates structured information including reasons for point deductions and improvement suggestions, closely adhering to the logic of score reduction and avoiding the vagueness of traditional feedback suggestions. Students can clearly understand where they lost points, why they lost points, and how to improve. For example, in geography, if a student sees that the reason for a point deduction is an error in interpreting contour lines, they can directly refer to the contour line interpretation micro-lesson and complete 5 specific questions to correct their mistakes, significantly improving the effectiveness of their responses and accelerating their mastery of knowledge points.
[0094] Furthermore, the learning analysis unit aggregates and transforms scattered individual feedback data into visual reports. Teachers no longer need to manually calculate class scores and error types; they can quickly grasp core information simply through charts. For example, a heat map clearly shows that tropical rainforest climate characteristics are a high-error-rate knowledge point for the whole class, and a bar chart shows that the average score increased by 15 points after resubmission (verifying the effectiveness of the feedback mechanism). This reduces teachers' workload in data analysis, allowing them to focus their energy on adjusting teaching strategies rather than data processing.
[0095] Visualized reports can be designed with specific dimensions for image-dependent subjects such as geography (e.g., error types in landscape recognition, error rates in topography and climate knowledge points), rather than generalized learning analysis. For example, the report can show a comparison of recognition error rates for different topographic landscape maps, which teachers can use to supplement the explanation of corresponding landscape real-world examples. This makes the learning analysis more in line with the teaching characteristics of geography, which emphasizes real-world observation and the combination of text and images, and improves the pertinence of teaching adjustments.
[0096] The feedback data and visualization reports compiled by the learning analysis unit can be stored for a long time. Teachers can compare the learning reports from different classes and different stages to evaluate the teaching effectiveness. For example, by comparing the pre-class preparation feedback report and the post-class review feedback report, teachers can determine whether the design of the pre-class preparation task is effective. By comparing the visualization reports of different classes, teachers can analyze the suitability of teaching methods and provide data support for long-term teaching optimization.
[0097] Optionally, the processing layer 300 further includes a processing unit and a storage unit; The storage unit is used to store the knowledge point logic of the target subject, the image features of the target subject, and the student classroom interaction data of the target subject. The processing unit is used to construct a knowledge graph and an image recognition model of the target subject by using machine learning algorithms and combining the knowledge point logic of the target subject, the image features of the target subject, and the classroom interaction data of students in the target subject. Based on the knowledge graph of the target subject, the unit uses the image recognition model of the target subject to perform a first AI processing on the first image recognition analysis result to obtain a first feedback result for the target student, and performs a second AI processing on the second image recognition analysis result to obtain a second feedback result for the target student. The knowledge point logic of the target discipline is determined based on the knowledge point name, knowledge point level, inclusion relationship, prerequisite relationship, association relationship and difficulty level of the knowledge points in the target discipline. The image features of the target subject are determined based on the pre-class learning feedback image corresponding to the target subject, and / or the in-class learning feedback image, and / or the post-class learning feedback image; The student classroom interaction data for the target subject includes pre-class learning feedback images, in-class learning feedback images, and post-class learning feedback images for the target subject, as well as pre-class learning feedback audio information, in-class learning feedback audio information, and post-class learning feedback audio information for the target subject, along with response feedback images and response audio information for the target subject.
[0098] In this embodiment, the storage unit is used to realize the centralized and structured storage of target subject-specific data, providing data support for the model building and AI analysis of the processing unit.
[0099] Optionally, the knowledge point logic of the target subject may include: Basic attributes; Hierarchical relationship: hierarchical division of knowledge points; Relationship: inclusion relationship (e.g., contour line interpretation is included in terrain identification), prerequisite relationship (e.g., latitude and longitude positioning is a prerequisite knowledge point for regional climate analysis), association relationship (e.g., river hydrological characteristics are associated with terrain type); Difficulty level: difficulty level of knowledge points (e.g., climate formation analysis is high difficulty, landscape type identification is basic difficulty). All knowledge point logic is stored in a structured format and can be directly called by the processing unit.
[0100] Image features for the target subject may include: subject-specific image feature data, sourced from student-uploaded pre-class / in-class / post-class learning feedback images (such as homework images), for example: Topographical features: Visual features of mountains (elevation > 500 meters, steep slope) and hills (200-500 meters, gentle slope); Climate characteristics: Color and vegetation distribution characteristics of landscapes in different climate zones (tropical rainforest, temperate grassland); Symbolic features: visual features such as contour lines, latitude and longitude labels, and climate charts. It stores structured data (such as feature vectors) after feature extraction, rather than the original image, significantly reducing storage costs and facilitating model retrieval.
[0101] The target subject's student classroom interaction data is a full-scale storage of multimodal interaction data from geography students, which may include: Image-based: pre-class / in-class / post-class learning feedback images, and response feedback images; Audio data includes pre-class, in-class, and post-class learning feedback audio information and response feedback audio information. All data is linked to student identifiers, knowledge point identifiers, and feedback stages to form a traceable interactive dataset.
[0102] Based on the data stored in the storage unit, the processing unit constructs a knowledge graph and an image recognition model for the target discipline. Knowledge graph construction for the target subject: Utilizing machine learning algorithms (such as knowledge graph embedding and association rule mining), a knowledge graph specific to the geography subject is constructed based on the knowledge point logic of storage units. The graph uses knowledge points as nodes and relationships such as inclusion, prerequisite, and association as edges, intuitively presenting the hierarchy and association logic of the target subject's knowledge.
[0103] Image recognition model construction for the target subject: Based on image features stored in the memory and student classroom interaction image data, a subject-specific image recognition model is trained using deep learning algorithms (such as CNN and ResNet). This model is optimized for the target subject; for example, the image recognition model for geography can achieve the following: Prioritize identifying the geographical features of the landscape (such as river width and terrain slope) rather than general visual features; By incorporating error cases from student interaction data (such as images that mistakenly label hills as plains), the model's ability to identify common mistakes can be enhanced.
[0104] Next, the processing unit combines the constructed geographic knowledge graph and image recognition model to perform AI processing on the students' feedback data: The first AI processing step involves analyzing the initial image recognition results and combining them with the knowledge graph to determine the knowledge point level, prerequisite / related knowledge points, and verifying the accuracy of the answer using the image recognition model. Finally, the first feedback result is generated (including score, reasons for errors, and improvement suggestions). For example, if a student misjudges the characteristics of a tropical rainforest climate, the knowledge graph identifies the prerequisite knowledge point as "temperature / precipitation determination." The AI will prioritize checking for issues with the student's understanding of this prerequisite knowledge point. The second AI processing step involves analyzing the second image recognition result of the response feedback, combining it with knowledge graph analysis to determine if the response content covers the knowledge points initially incorrect. The image recognition model is then used to verify the correctness of the response image, generating a second feedback result. This AI analysis of the response feedback focuses more on whether errors have been corrected and whether related knowledge points have been mastered.
[0105] Based on the above description, the knowledge graph in this application embodiment enables AI processing to locate the hierarchy and relationships of knowledge points. For example, if a student answers "river hydrological characteristics" incorrectly, AI processing can trace back to the previous knowledge point "topographic slope analysis" through the knowledge graph, and the improvement suggestions given are more in line with the logic of geographical knowledge, rather than the general "review hydrological characteristics".
[0106] Furthermore, as student interaction data increases, the image recognition model can learn more geographical error-prone cases (such as the differentiated characteristics of river landscapes in different regions), and the recognition accuracy continues to improve; the knowledge graph can also supplement the knowledge point associations of students' high-frequency errors (such as finding that errors in "climate characteristics" are often associated with insufficient knowledge of "precipitation type"), making the improvement suggestions from AI feedback more targeted.
[0107] Furthermore, the storage unit stores lightweight data such as feature vectors and structured knowledge point logic, rather than raw image / speech data, thus reducing storage costs. Structured storage allows the processing unit to access data without repeated parsing, directly connecting to model training and AI analysis processes, reducing computing power consumption and improving real-time feedback efficiency in classroom scenarios.
[0108] Optionally, the processing unit is further configured to push suitable learning resources to the corresponding target student based on the feedback score.
[0109] In this embodiment of the application, the processing unit first performs learning situation stratification and labeling of target students based on the feedback score (feedback score of the first / second feedback result), combined with the preset score threshold and the difficulty of subject knowledge points.
[0110] For example, thresholds are applied based on the difficulty of subject knowledge points (higher thresholds for basic knowledge points and lower thresholds for advanced knowledge points). For instance: Basic knowledge points: ≥90 points = "Mastery", 70-89 points = "Basic Mastery", <70 points = "Not Mastery"; Advanced knowledge points: ≥80 points = "Mastery", 60-79 points = "Basic Mastery", <60 points = "Not Mastery". In addition to hierarchical labels, the processing unit also generates labels based on the reasons for lost marks.
[0111] Next, the processing unit connects to the pre-set geography subject learning resource library (the knowledge point logic and knowledge graph of the resource library and storage unit are linked), first classifies the resources in a structured way, and then establishes matching rules based on students' score stratification and learning status tags.
[0112] The resource classification dimensions include: Knowledge point dimension: categorized by subject knowledge graph nodes; Difficulty dimension: basic, intermediate, advanced; Format dimension: illustrated explanations, micro-lecture videos, specialized practice questions, real-world inquiry tasks, and case studies analyzing incorrect answers; Adaptable scenario dimension: pre-class preparation, in-class consolidation, post-class review, and quiz reinforcement.
[0113] The implementation methods for pushing adapted learning resources may include: Not mastered (low score): Push basic explanation and low-difficulty practice resources, prioritizing the prerequisite basic content of the knowledge points; Basic Mastery (Medium Score): We provide analysis of common mistakes and advanced practice questions, focusing on the specific areas where students lose marks. Master (High Score): Push out extended exploration + cross-knowledge point related resources to meet the improvement needs of high-achieving students; Score improvement after resubmission: If the score in the second feedback is ≥20 points higher than the score in the first feedback, knowledge point reinforcement + similar expansion resources will be provided; if the improvement is <10 points, targeted reinforcement of the points lost will be provided + one-on-one Q&A guidance resources.
[0114] After the processing unit completes resource matching, it integrates the adapted learning resource information (resource link, description, adaptation purpose) into the second feedback result (or supplements and pushes it to the first feedback result), and pushes it through the following methods: Terminal push: Resource information is synchronously transmitted back to the target student's first terminal 100 controller 103 along with the feedback results. After the controller 103 completes the adaptation process, the "Personalized Learning Resources" section is displayed in the interaction module 104, for example: Below the feedback results marked "Not Mastered" in red, it shows "Recommended Learning: Contour Line Interpretation Basic Micro-Lesson (5 minutes) + 3 Basic Interpretation Questions"; Push timing: Initial feedback (first feedback result) pushes supplementary resources, follow-up feedback (second feedback result) pushes reinforcement / expansion resources; if the second feedback result meets the preset criteria (e.g., score ≥ 90 points), stop pushing supplementary resources and only push expansion resources; All pushed resources are linked to student classroom interaction data in the storage unit, recording the student's resource viewing / completion status. The push strategy can be adjusted based on resource usage (e.g., push again if micro-lessons are not viewed, push advanced resources if exercises are completed).
[0115] Based on the above description, this application embodiment uses scores as the quantitative basis and combines subject characteristics to achieve hierarchical matching of resources, solving the technical problem of resource push not matching students' learning situation; it strengthens the closed-loop effect of feedback-learning-reply, accelerating students' mastery of knowledge points; while reducing teachers' labor costs, it realizes personalized resource push under large-scale teaching, further enhancing the practical value and intelligence level of the learning companion system in the teaching scenario.
[0116] In one alternative implementation, the first terminal 100 further includes a processor; The image acquisition module 101 is used to acquire free-response images in response to the free-response operation of the target student; The controller 103 is used to respond to the free-response operation of the target student, control the image acquisition module 101 to acquire free-response images, and control the sound acquisition module 102 to acquire free-response sound information; The controller 103 is also used to transmit the free-response image and the free-response audio information to the processor; The processor is used to process the free-response image and / or the free-response sound information according to the preset computing power model and preset image recognition model built into the first terminal 100, determine the knowledge points of the corresponding target subject, generate progressive questions according to the knowledge points of the target subject and according to the preset questioning logic, and output them to the target student through the interaction module 104.
[0117] In this embodiment, the target student performs a free-response operation through the interaction module 104 of the first terminal 100 (such as clicking the free question button, speaking "I want to ask a question," uploading a question image, and then selecting to initiate an investigation). The controller 103 responds to this operation and immediately triggers a multimodal acquisition command: controlling the image acquisition module 101 to capture / upload the free-response image specified by the student; and controlling the sound acquisition module 102 to record the student's free-response audio information.
[0118] The controller 103 transmits the collected free-response image and sound information to the processor built into the first terminal 100. The processor performs localized analysis based on two types of core models stored locally on the terminal: a preset computing power model: a lightweight AI model with core capabilities of speech-to-text, semantic understanding, and knowledge point matching; and a preset image recognition model: a dedicated lightweight image recognition model for each target subject, with core capabilities of recognizing image features, handwritten text, and symbols.
[0119] The processor's specific processing includes: if only free-response question-and-answer images are received, then image features are extracted using a preset image recognition model, matched with the knowledge point database of the corresponding subject stored locally, and the corresponding target subject knowledge point is determined; if only free-response question-and-answer audio information is received, then speech-to-text conversion and semantic understanding are completed using a preset computing power model, and the target knowledge point is determined by matching the knowledge point database; if images and audio are received, then image features and audio semantics are fused, and the target knowledge point is determined after cross-validation.
[0120] After the processor determines the target knowledge point, it generates progressive questions (instead of directly giving the answer) based on the preset questioning logic stored locally (questioning rules that conform to the inquiry-based learning patterns of the corresponding subject).
[0121] For example, progressive questions can range from basic understanding to in-depth exploration. For instance, based on the target knowledge point, the logic of the questions could be to first identify the phenomenon, then analyze the influencing factors, and finally guide the student to summarize independently.
[0122] Finally, the processor transmits the generated progressive questions to the controller 103, which outputs them to the target students in the form of text / images and voice through the interaction module 104 (such as the interaction module 104 displaying the question chain and broadcasting the questions), guiding students to explore independently.
[0123] Based on the above description, the processor in this application embodiment does not generate a standard answer, but a progressive chain of questions, which can cultivate students' independent inquiry ability; it is suitable for students with different cognitive levels, with progressive questions ranging from basic to advanced, allowing students with low cognitive levels to complete basic questions and students with high cognitive levels to challenge in-depth questions, thus taking into account the needs of tiered learning; it strengthens knowledge internalization: by asking questions, students are guided to actively call upon existing knowledge, which makes it easier to achieve a deep understanding of knowledge points than passively receiving feedback.
[0124] Optionally, the processor is further configured to obtain the language behavior characteristics and emotional expression characteristics of the target student based on the free-response image and the free-response sound information, and determine the stress information of the target student based on the language behavior characteristics and emotional expression characteristics. When the stress information of the target student meets the preset warning conditions, the processing unit sends the warning information to the corresponding second terminal 200 through the signal transceiver module 105.
[0125] In this embodiment, the free-response audio information (after being converted to text) and handwritten text / lip movements in images can be analyzed based on a preset computing power model to extract language behavior features reflecting the learning state: Content characteristics: The complexity of the question (e.g., questions based on the terrain, or questions about how to analyze the differences in river flood seasons by combining terrain and climate are considered advanced questions), the completeness of the expression (whether the question is clearly described, e.g., "I don't know how to answer this" is considered vague), and the degree of repetition (whether the same knowledge point is asked multiple times). Sentence structure characteristics: Whether negative / anxious sentences are used (e.g., I'm sure I can't learn this, this question is too difficult for me to answer), whether interrogative / uncertain sentences are used frequently (e.g., are they all wrong? Did I misunderstand?). Behavioral characteristics: the degree of illegibility of handwritten text in images (e.g., extremely illegible handwriting, ≥5 corrections), speech rate (too fast / too slow), and pause frequency (≥3 pauses per sentence).
[0126] In this embodiment, multimodal data can be parsed based on a lightweight emotion recognition model (integrating voice emotion recognition and facial expression recognition): Voice dimension: Extracting the tone (sharp / low), volume (fluctuating), and emotional feature vectors (such as voice features of anxiety, irritability, and frustration). Image dimension: If the free-response image contains the student's face (such as a student taking a picture of themselves asking a question), identify facial expressions (frowning, pouting, avoiding eye contact, and other negative emotional features); if it is a handwritten / landscape image, combine it with correction marks and negative symbols to help determine the emotion.
[0127] The processor inputs the extracted language behavior features and emotional expression features into a locally preset stress assessment model to quantify stress information and determine early warnings. Stress information quantification: Stress scores are generated from 0 to 100.
[0128] Language behavior (60%): Repeatedly asking the same knowledge point (+20 points), vague and negative expression (+15 points), extremely illegible handwriting (+10 points); Emotional expression (40%): Speech recognition is anxiety / frustration (+25 points), facial expression is irritability (+10 points), and image is labeled with negative symbols (+5 points); For example: if a student asks the same question 3 times, the stress score for speech anxiety and illegible handwriting is 20+25+10=55 points.
[0129] For example, preset warning conditions: three levels of warning can be set according to pressure scores, adapting to classroom teaching scenarios: Level 1 Warning (Low Risk): 50-69 points, characterized by mild anxiety and difficulty in mastering knowledge points; Level 2 Warning (Medium Risk): 70-89 points, characterized by obvious irritability and lack of confidence in learning; Level 3 Warning (High Risk): ≥90 points, characterized by severe depression and resistance to continuing to learn.
[0130] When student stress information meets any of the warning conditions, the warning process is triggered: The controller 103 of the first terminal 100 implicitly marks the student's free question and answer interface in the interaction module 104 (not to be displayed to the student, so as to avoid increasing the pressure). The controller 103 transmits student identification, stress score, warning level, and core triggers (such as repeated questioning contour line interpretation and voice anxiety) to the processing unit of the processing layer 300 through the signal transceiver module 105. Upon receiving the information, the processing unit immediately pushes the standardized early warning information (including student information, stress level, triggers, and suggested intervention methods) to the corresponding teacher's second terminal 200.
[0131] Based on the above description, the embodiments of this application can promptly capture students' hidden learning obstacles: for example, students may not be able to express the knowledge they know clearly due to excessive pressure. The system can distinguish between knowledge gaps and expression errors caused by pressure, thus preventing teachers from misjudging students' knowledge mastery levels.
[0132] In one specific embodiment, a learning companion system is provided, which is adapted to the large-class teaching scenario of middle school geography (with 50 students per class, corresponding to 50 first terminals, 2 second terminals configured on the teacher's end, and the processing layer deployed locally on the school's smart education server). The system realizes the whole process of teaching before, during and after class through a three-layer architecture design of first terminals, second terminals and processing layer. At the same time, it supports students' free question and answer and learning pressure monitoring, which is in line with the teaching characteristics of geography that combines text and graphics and real-world exploration. The specific implementation details are as follows.
[0133] The first terminal is a customized ESP32 motherboard integrated terminal, equipped with an image acquisition module (5-megapixel high-definition wide-angle camera), an audio acquisition module (noise-canceling array microphone), a controller (customized ESP32 motherboard, optimized for geographical data processing computing power and interface configuration), an interaction module (3.5-inch touch screen), a signal transceiver module (multi-protocol Wi-Fi / Bluetooth module, supporting WS / HTTP / Socket / MQTT protocols), a processor (lightweight embedded AI processor), and a 3D-printed shell (with reserved camera shooting window and microphone pickup hole, handheld design adapted for classroom scene shooting). Each first terminal is uniquely bound to one target student, realizing the association and traceability of student learning data.
[0134] The controller uses a customized ESP32 motherboard, which can be modified and optimized in terms of circuitry, interfaces, and computing power to better meet the specific needs of the learning companion system's classroom interaction and image acquisition and processing requirements. This allows for better adaptation to the unique needs of geography subjects, such as voice interaction, image transmission and processing, and low-latency response. Alternatively, it can be replaced with other embedded motherboards of comparable performance, maintaining interface compatibility and control logic without affecting the operation of classroom interaction and image processing workflows.
[0135] The sound acquisition module can use a noise-canceling microphone to improve the accuracy of voice acquisition in noisy classroom environments; it is used to collect students' voice answers in class, capture the content of their answers, provide support for the voice input process in classroom interaction, and adapt to the clear acquisition requirements in noisy classroom environments.
[0136] The camera module can be replaced with a high-definition, wide-angle model to improve the clarity and coverage of landscape recognition and assignment photography, ensuring the accuracy of interactive data. It is used to collect geographical landscape maps and student assignment images, capturing geographical features and assignment responses within the images, and transmitting them to the core processing layer for recognition, analysis, and grading. This enriches the system's perception capabilities and adapts to the characteristics of geography teaching and assignment grading needs.
[0137] The casing can be customized to meet the needs of different schools, and mass-produced using injection molding and other processing techniques to suit large-scale classroom use. Designed according to school culture and usage habits, the casing protects internal components and includes reasonable provisions for camera shooting windows and microphone / screen mounting positions, enhancing portability and aesthetics, and meeting the needs of students holding the device for answering questions and taking photos in classroom settings.
[0138] The second terminal is a teacher-specific all-in-one teaching machine + tablet terminal, equipped with a high-definition display module, a multi-protocol signal transceiver module, a teaching information editing module, and a student learning report display module. It supports teachers in initiating teaching instructions, editing pre-class / in-class teaching information, viewing visualized student learning reports, and receiving student stress warning information, realizing full-process operation and control of classroom teaching.
[0139] The processing layer is deployed on the school's local smart education server and includes a feedback unit, a learning analysis unit, a processing unit, and a storage unit. It is equipped with high-performance computing chips and subject-specific algorithm models. The storage unit pre-stores the logic of knowledge points for each subject, image features, and student classroom interaction data. The processing unit integrates machine learning algorithms and a subject-specific model for geography, supports concurrent data processing and real-time AI analysis across multiple terminals, and ensures smooth system operation in large-class teaching scenarios.
[0140] Pre-class teaching information generation and distribution: Geography teachers perform a second operation through a second terminal, editing pre-class preparation tasks (such as taking pictures of the terrain and landscape around their hometown and describing its characteristics in conjunction with the textbook), generating pre-class teaching information that includes text requirements, landscape picture examples, and links to pre-class knowledge points, and then distributing it to 50 first terminals through the signal transceiver module.
[0141] Pre-class teaching information display: The first terminal controller receives pre-class teaching information through the signal transceiver module and controls the interactive module to display it to the target students in the form of pictures and text. Students can view the pre-class requirements and download reference materials through the interactive module.
[0142] Pre-class learning feedback collection and preprocessing: After completing the pre-class preparation, students take pictures of their hometown's topographic landscape through the first terminal image acquisition module (pre-class learning feedback images), and verbally describe the landscape features through the sound acquisition module (pre-class learning feedback sound information); the controller performs angle correction, noise reduction, and compression preprocessing on the collected images, and performs noise reduction, silent segment cropping, and volume normalization preprocessing on the sound information to generate processed pre-class learning feedback data.
[0143] Pre-class data upload and preliminary learning analysis: The controller uploads the processed pre-class learning feedback data to the processing layer via the signal transceiver module; the processing unit in the processing layer analyzes the landscape map using a geography image recognition model, converts the spoken content into text through speech-to-text processing (integrating a geography terminology database), and completes the first AI processing in conjunction with a geography knowledge graph to generate the first feedback result (including feedback score, reasons for lost points such as terrain type recognition errors, and improvement suggestions); the feedback unit sends the first feedback result back to the first terminal, and after data parsing and format adaptation by the controller, it is displayed to the students through the interactive module, allowing students to independently identify and fill in any gaps in their knowledge.
[0144] Pre-class learning summary: The learning analysis unit at the processing layer summarizes the first feedback results of 0 students and generates a pre-class visual learning report (including the average score of the whole class in pre-class preparation, the error rate of terrain type identification, and frequently missed knowledge points, etc.), which is pushed to the second terminal. Teachers generate in-class teaching content based on the report.
[0145] Teachers perform a third operation through a second terminal, generating in-class teaching information based on pre-class learning information, and sending it to all first terminals through a signal transceiver module. The controller controls the interactive module to display the in-class teaching information in full screen.
[0146] Students use the first terminal image acquisition module to capture images of the tropical rainforest landscape in the textbook, and use the sound acquisition module to verbally analyze the content; the controller performs lightweight preprocessing on the collected data and then quickly uploads it to the processing layer.
[0147] The processing unit performs second image recognition analysis and second speech-to-text processing on the learning feedback data in class. It combines the geographical knowledge graph and image recognition model to complete the second AI processing. The feedback unit generates the second feedback result (such as a feedback score of 70 points, the reason for the loss of points "climate causes were not combined with latitude location analysis", and the improvement suggestion "review the latitudinal characteristics of tropical rainforest climate distribution"), and sends it back to the first terminal for display.
[0148] After reviewing the second feedback results, students correct their mistakes by adding latitude information to the landscape map using the image acquisition module and verbally reciting the analysis content using the sound acquisition module to complete the reply operation. The first terminal collects the reply feedback image / sound information, which is then preprocessed by the controller and uploaded to the processing layer. The processing layer repeats the above analysis process to generate a new second feedback result until the student's feedback result meets the preset standard (e.g., feedback score ≥ 90 points).
[0149] The processing layer's learning analysis unit summarizes in real time the in-class answer data, response status, score distribution, and frequently missed knowledge points of all students, generating a visualized in-class learning report (including a bar chart showing the error rate of each knowledge point, a pie chart showing the score distribution, and screenshots of typical error cases), which is then pushed to the second terminal. Teachers conduct targeted classroom reviews based on the report, focusing on explaining common errors made by the whole class and providing one-on-one guidance to individual students who have not met the standards.
[0150] In one example, the classroom interaction process (including image acquisition scenarios) proceeds in a progressive manner according to the following steps, adapting to the classroom scenario under the personalized guidance mode: 1. Teachers can initiate geography-related questions in class, including those involving images and text, and landscape identification, which will then be simultaneously uploaded to the learning partner system. 2. Students can input voice answers via the terminal microphone module, including answer instructions, name and answer content, or capture and upload geographical landscape maps / assignment images via the camera module. The customized ESP32 motherboard will receive voice and image data in real time and complete preliminary preprocessing, image cropping and format conversion. 3. The signal transceiver module quickly transmits the preprocessed voice and image data to the processing layer via the WS protocol; 4. The processing layer completes speech-to-text conversion, image recognition and parsing, AI scoring, and personalized feedback generation; 5. Feedback results are transmitted back to the terminal via HTTP protocol. After being processed by the customized motherboard, the screen module displays the score, scoring criteria, and improvement suggestions in real time. Feature analysis results are displayed simultaneously for landscape recognition questions, and incorrect question annotations are displayed simultaneously for homework grading. 6. After reviewing the feedback, students can choose to answer the questions again, repeating steps two through five to optimize their answering strategies or add more images. 7. The customized motherboard synchronously records students' answers, replies, score changes, and image acquisition and processing data, and uploads them to the core processing layer for aggregation, providing support for teachers' classroom lectures and reviews.
[0151] The entire process forms a complete chain of interactive classroom actions, ensuring that all participants participate synchronously and the process is coherent. When switching to the free question and answer mode, the terminal can quickly switch to the dialogue interface, and the camera module still supports taking pictures of the landscape to ask questions, maintaining the continuity of operation.
[0152] In this example, the processing layer is the core processing unit for classroom interaction, dual-mode operation, and image processing. It is responsible for data parsing, image recognition, AI scoring, learning analysis, and feedback generation, connecting the terminal interaction with the teacher's feedback. The technical process revolves around the requirements of classroom interaction, image processing, and dual-mode operation, and consists of three steps: 1. Data Preprocessing: Based on the optimized computing power and transmission adaptation capabilities of the ESP32 motherboard, it efficiently receives the voice and image data transmitted from the signal transceiver module of the first terminal. After optimization using speech-to-text technology, the data is integrated into a core terminology database of geography, transforming the voice data into standardized digital text. For geographical landscape maps captured by cameras, core geographical features such as terrain, vegetation, and climate are extracted using image recognition technology and classified and labeled according to a geographical knowledge graph. For homework images captured by cameras, the answer area is cropped using image recognition technology and converted into digital text using OCR technology. The customized motherboard can further improve the speed and accuracy of data preprocessing and image analysis through hardware optimization, ensuring the timeliness of classroom interaction and providing computing power support for real-time image analysis and dialogue analysis in the free question-and-answer mode.
[0153] 2. Semantic Analysis: Import the teacher-prepared detailed scoring rules for geography, and use generative AI as the engine to perform topic identification, key information extraction, and logical judgment of geographical element associations in student answers using natural language processing technology; for image-text and landscape recognition questions, combine the image analysis results with the text answers for analysis to match the key points of classroom questions; for homework grading, score each OCR-converted text answer according to the scoring rules, and identify the types of wrong questions and the reasons for losing marks.
[0154] 3. Feedback Generation: While outputting scores, personalized feedback including reasons for lost marks and suggestions for improvement is generated. For landscape recognition questions, feature analysis and suggestions related to subject knowledge points are output simultaneously. Homework is graded simultaneously, with the location of wrong questions and correction ideas marked to ensure that the feedback content is relevant to classroom questions, image information, and students' actual answers. Simultaneously, all students' answer data, reply status, score distribution, high-frequency errors, image analysis results, and other information are summarized to form a visualized classroom learning report, which is pushed to the teacher's end to provide data support for real-time classroom review.
[0155] In terms of algorithm support, machine learning algorithms are integrated to construct a geographical subject knowledge graph and image recognition model. Personalized recommendation algorithms are designed by combining knowledge point logic, image features and student classroom interaction data. A dual mode of offline caching and real-time uploading is adopted to optimize network adaptation and avoid data transmission delays from affecting classroom interaction and image processing.
[0156] The primary terminal enables software and hardware collaboration and serves as a crucial link between students and the processing layer in the classroom interaction process. It comprises four main functional modules: 1. Voice capture module: Captures students' voice answers through a microphone module, supports integrated input of answer instructions, name and answer, and adapts to the needs of rapid classroom response; 2. Image acquisition module: Captures geographical landscape maps and work images through a camera module, supports real-time capture and upload as well as import of historical images, and is suitable for geographical landscape observation and work submission scenarios; 3. Signal transceiver module: Relying on the conversion between WS protocol, HTTP and Socket protocol, it realizes high-speed transmission of voice data, image data, response information and feedback results, and connects the customized motherboard and the core processing layer to ensure uninterrupted interaction. 4. Interactive Module: The screen module intuitively presents information such as scores, scoring criteria, improvement suggestions, class answer statistics summary, and image analysis results. It supports students to view, answer, and quickly switch between two modes. It also provides quick answer and re-capture buttons to optimize the classroom interaction experience.
[0157] The operation process strictly follows classroom interaction scenarios, while being compatible with dual-mode switching and image acquisition logic: Teachers initiate classroom questions (including landscape recognition and homework grading), students receive the questions and input voice answers through the terminal, or capture images through the camera and upload them, the customized motherboard preprocesses the voice / image data and transmits it to the core processing layer, the system transmits and processes the data in real time (speech-to-text, image recognition and analysis, AI scoring), the terminal displays personalized feedback and image analysis results, students can adjust their thinking based on the feedback and answer again (reply mode) or supplement the captured images, the terminal synchronously uploads the reply data, and supplements the images to the system, forming a secondary optimization link for classroom interaction; if students switch to free question and answer mode, the terminal interface automatically jumps, and the voice acquisition, image acquisition and data processing modules are seamlessly connected to maintain the consistency of operation and data transmission.
[0158] Free Q&A Interaction: During the independent exploration phase in class, if students have questions about geographical knowledge points, they can click the "Ask a Question" button on the first terminal interaction module to perform a free Q&A operation; the controller responds to the operation, controlling the image acquisition module to take pictures of relevant images (such as contour maps in the textbook) and controlling the sound acquisition module to record the voice of the question (such as "How to determine the slope of the terrain based on contour lines?"), and transmits the data to the terminal's built-in processor.
[0159] Localized knowledge point matching and progressive question generation: Based on the built-in preset computing power model and the lightweight image recognition model of geography, the processor analyzes the free question and answer image / sound information, determines the target knowledge points, generates a progressive question chain according to the Socratic preset question logic, and displays it to students through the interactive module to guide students to explore independently. The entire process does not rely on the processing layer and achieves localized real-time response.
[0160] Learning stress monitoring and early warning: While analyzing free-response data, the processor extracts students' language behavior information (such as slow speaking speed, repetitive expressions, and illegible handwriting) and emotional expression information (such as low voice tone, frowning facial expression, and multiple correction marks in the image). This information is quantified into a stress score (e.g., 85 points, Level 2 warning) through a built-in stress assessment model. When the stress score meets the preset warning conditions, the controller uploads the student's identifier, stress score, warning level, and core triggers (e.g., "repeated questioning of contour line interpretation, voice anxiety, and 6 corrections in handwriting") to the processing unit. The processing unit immediately sends the warning information to the second terminal through the signal transceiver module. After receiving the warning, the teacher provides targeted guidance and explanation of knowledge points to the student during class breaks to prevent stress accumulation.
[0161] Post-class teaching instructions generation and distribution: Teachers view the visualized learning progress report during class through the second terminal, and respond to the teacher's operation to generate post-class teaching instructions (differentiated design: struggling students: complete basic exercises on terrain type identification + take a picture of the local terrain landscape and label the type; average students: analyze the relationship between climate type and landscape characteristics, and complete 3 comprehensive questions; high-achieving students: explore the impact of local climate on terrain landscape, write a short research report, and distribute it to the corresponding first terminal through the signal transceiver module).
[0162] After-class learning feedback collection and analysis: After students complete their after-class tasks, they collect image / sound information (such as homework photos and oral content of inquiry reports) through the first terminal. After being pre-processed by the controller, the information is uploaded to the processing layer. The processing unit of the processing layer performs comprehensive analysis on the after-class data and combines the students' learning data throughout the entire process of pre-class, in-class, and post-class learning to create a three-dimensional digital profile for each student (covering knowledge mastery, thinking characteristics, and core competency achievement level).
[0163] Personalized learning resource delivery: The processing unit in the processing layer matches suitable learning resources from the geography subject resource database based on the student's feedback score and 3D digital profile, and pushes them to the first terminal through the signal transceiver module. The controller displays resource links and learning guides through the interactive module, and students can click to learn independently, realizing personalized after-class supplementation and expansion.
[0164] After-class learning progress summary: The learning progress analysis unit at the processing layer summarizes the learning data of all students before, during, and after class, and generates a visualized learning progress report after class (including changes in students' scores throughout the process, trends in knowledge point mastery, and personalized resource usage). This report is then pushed to the second terminal, where teachers can summarize the effectiveness of classroom teaching based on the report and optimize subsequent teaching plans and content design.
[0165] The storage unit structurally stores the logical knowledge points of geography (including knowledge point names, levels, inclusion / precedence / association relationships, such as "climate type", "tropical rainforest climate", "climate formation", with "latitude location" as the preceding knowledge point), geographical image features (extracting core visual features of landscapes such as topography, climate, and rivers, such as the visual features of "high temperature and abundant rainfall, dense vegetation" of tropical rainforest landscapes), and student classroom interaction data (covering all image and sound information before / during / after class / answer / free Q&A, and associating student identifiers with knowledge point identifiers), providing data support for the construction of processing unit models and AI analysis.
[0166] The processing unit utilizes machine learning algorithms, combined with the geographical knowledge point logic, image features, and student classroom interaction data from the storage unit, to construct a geographical knowledge graph and a geographically specific image recognition model. The knowledge graph intuitively presents the logical connections between geographical knowledge points, while the image recognition model is specifically optimized for geographical landscapes, handwritten assignments, and geographical symbols to improve recognition accuracy. During AI processing, the processing unit links the image recognition and analysis results / speech-to-text results with the knowledge graph to pinpoint students' knowledge weaknesses and the reasons for their errors, generating targeted feedback results. Simultaneously, it matches and pushes learning resources based on the feedback scores.
[0167] The feedback unit focuses on individual students, enabling the quantitative output of feedback scores and the structured generation of reasons for point deductions and improvement suggestions, ensuring the authenticity of individual feedback. The learning analysis unit focuses on the class as a whole, summarizing individual feedback results from multiple dimensions and generating a visualized learning report to provide data support for teachers' teaching strategies. The two work together to achieve a closed loop of "individual feedback - class commonality analysis - teaching strategy optimization", adapting to the needs of differentiated instruction in large-class teaching.
[0168] The learning buddy system in this embodiment can achieve the following implementation effects: Full coverage of classroom interaction: 50 primary terminals support the participation of all students in the class in pre-class, in-class, and post-class interactions, with an interaction coverage rate of 100%. This breaks through the limitation of a small number of students participating in traditional classrooms, and significantly increases students' opportunities for classroom expression and feedback.
[0169] Highly adaptable to geography subjects: The image acquisition module and the geography-specific image recognition model are adapted to the needs of real-scene shooting and assignment photo grading. The speech-to-text processing incorporates a geography terminology database, achieving a professional terminology recognition accuracy rate of over 98%, perfectly aligning with the text-and-image teaching characteristics of geography. (It is worth noting that geography is only an example in this embodiment, and this application embodiment is also applicable to other subjects that adapt to text-and-image combinations.)
[0170] The teaching loop is complete and efficient: it realizes the full-process teaching linkage of pre-class diagnosis, in-class interaction and post-class adaptation, and the learning data is transferred and analyzed in real time. Teachers can design teaching content based on the visualized learning reports, which significantly improves the pertinence of teaching.
[0171] Personalized tutoring implemented: Through closed-loop question-and-answer sessions, three-dimensional digital profiling, and personalized learning resource delivery, a customized approach is adopted for each student. This has improved the efficiency of knowledge gap filling for struggling students and significantly increased the depth of exploration for high-achieving students.
[0172] Dual-dimensional management of mind and body: The free question and answer function cultivates students' independent inquiry ability, while the learning pressure monitoring and early warning function realizes imperceptible and holistic care for mind and body, effectively reducing students' geography learning anxiety and significantly improving classroom learning experience and participation.
[0173] Improved teacher productivity: The processing layer automates data parsing, AI analysis, and student learning summary, reducing the workload of teachers in manual grading and data statistics, allowing them to focus more on classroom teaching and personalized guidance.
[0174] The learning companion system in this embodiment can flexibly adjust the number of first terminals deployed, the content of teaching information and model parameters according to the geography teaching needs of different schools and grades. At the same time, it can be extended to other subjects such as Chinese, history, and biology that emphasize the combination of text and graphics and inquiry-based learning, and has good versatility and scalability.
[0175] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A learning companion system, characterized in that, The system includes multiple first terminals, at least one second terminal, and a processing layer; Each of the first terminals and at least one of the second terminals are communicatively connected to the processing layer, and each of the first terminals is communicatively connected to at least one of the second terminals. The first terminal is used to receive pre-class teaching information and / or in-class teaching information, and transmit the pre-class learning feedback information uploaded by the target student in response to the pre-class teaching information and / or the in-class learning feedback information uploaded by the target student in response to the in-class teaching information to the processing layer; The processing layer is used to generate target learning data and target visualized classroom learning report based on the pre-class learning feedback information and / or the in-class learning feedback information, and to send the target learning data and target visualized classroom learning report to the second terminal; The second terminal is used to generate after-class teaching instructions and transmit them to the first terminal based on the learning data and visualized classroom learning reports pushed by the processing layer and in response to the target teacher's operation. The first terminal is also used to receive the after-class teaching instructions and send the collected after-class learning feedback information of the target students in response to the after-class teaching instructions to the processing layer; The processing is used to generate a target feedback result based on the pre-class learning feedback information, and / or the in-class learning feedback information, and the post-class learning feedback information, and send it to the first terminal; The pre-class learning feedback information includes pre-class learning feedback images and / or pre-class learning feedback audio information; the in-class learning feedback information includes in-class learning feedback images and / or in-class learning feedback audio information; and the post-class learning feedback information includes post-class learning feedback images and / or post-class learning feedback audio information.
2. The system according to claim 1, characterized in that, The first terminal includes an image acquisition module, a sound acquisition module, a controller, an interaction module, a signal transceiver module, and a button module; The image acquisition module, the sound acquisition module, the interaction module, the signal transceiver module, and the button module are all electrically connected to the controller; The controller is used to control the image acquisition module to acquire the pre-class learning feedback image, and / or the in-class learning feedback image, and / or the post-class learning feedback image based on the target student's image acquisition operation on the button module; The controller is used to control the sound acquisition module to acquire the pre-class learning feedback sound information, and / or the in-class learning feedback sound information, and / or the post-class learning feedback sound information based on the target student's sound acquisition operation on the button module; The controller is used to send the pre-class learning feedback image, and / or the in-class learning feedback image, and / or the post-class learning feedback image, and / or the pre-class learning feedback sound information, and / or the in-class learning feedback sound information, and / or the post-class learning feedback sound information to the processing layer through the signal transceiver module; The controller is also configured to receive the target feedback result through the signal transceiver module and display the target feedback result to the target student through the interaction module.
3. The system according to claim 2, characterized in that, The second terminal is used to generate the pre-class teaching information in response to the second operation of the target teacher, and to generate the in-class teaching information in response to the third operation of the target teacher, and to send the pre-class teaching information and the in-class teaching information to the controller through the signal transceiver module; The controller is used to control the interactive module to display the pre-class teaching information, the in-class teaching information, and the post-class learning information.
4. The system according to claim 2, characterized in that, The controller is also configured to perform image preprocessing on the pre-class learning feedback image, and / or the in-class learning feedback image, and / or the post-class learning feedback image to obtain the processed pre-class learning feedback image, and / or the processed in-class learning feedback image, and / or the processed post-class learning feedback image, and send the processed pre-class learning feedback image, and / or the processed in-class learning feedback image, and / or the processed post-class learning feedback image to the processing layer through the signal transceiver module; The controller is also configured to perform sound preprocessing on the pre-class learning feedback sound information, and / or the in-class learning feedback sound information, and / or the post-class learning feedback sound information, to obtain processed pre-class learning feedback sound information, and / or processed in-class learning feedback sound information, and / or processed post-class learning feedback sound information, and send the processed pre-class learning feedback sound information, and / or processed in-class learning feedback sound information, and / or processed post-class learning feedback sound information to the processing layer through the signal transceiver module; The processing layer is used to perform a first image recognition and analysis on the processed pre-class learning feedback image, and / or the processed in-class learning feedback image, and / or the processed post-class learning feedback image, to obtain a first image recognition and analysis result; And / or, the processing layer is further configured to perform first speech-to-text processing on the processed pre-class learning feedback audio information, and / or the processed in-class learning feedback audio information, and / or the processed post-class learning feedback audio information, to obtain a first speech-to-text processing result; The processing layer is further configured to perform a first AI processing on the first image recognition and parsing result and / or the first speech-to-text processing result to obtain a first feedback result for the target student, and to transmit the first feedback result back to the controller through the signal transceiver module; wherein, the target feedback result includes the first feedback result; The controller is further configured to sequentially perform first data parsing processing, first format adaptation processing, first classification adaptation processing, and first signal adaptation processing on the first feedback result, and then feed back the processed first feedback result to the corresponding target student through the interaction module.
5. The system according to claim 4, characterized in that, The image acquisition module is also used to acquire the response feedback image obtained by the target student in response to the processed feedback result, and send the response feedback image to the controller; And / or, the sound acquisition module is further configured to acquire the response feedback sound information obtained by the target student in response to the processed feedback result, and send the response feedback sound information to the controller; The controller is also configured to perform image preprocessing on the response feedback image to obtain a preprocessed response feedback image, and / or perform sound preprocessing on the response feedback sound information to obtain preprocessed response feedback sound information, and transmit the preprocessed response feedback image, and / or the preprocessed response feedback sound information to the processing layer through the signal transceiver module; The processing layer is also used to perform image recognition and analysis on the preprocessed response feedback image to obtain a second image recognition and analysis result, and to perform second speech-to-text processing on the preprocessed response feedback sound information to obtain a second speech-to-text processing result. The processing layer is further configured to perform a second AI processing on the second image recognition and parsing result and / or the second speech-to-text processing result to obtain a second feedback result for the target student, and to transmit the second feedback result back to the controller through the signal transceiver module; wherein, the target feedback result includes the second feedback result; The controller is also used to sequentially perform second data parsing processing, second format adaptation processing, second classification adaptation processing, and second signal adaptation processing on the second feedback result, and to feed back the processed second feedback result to the corresponding target student through the interaction module until the second feedback result meets the preset feedback result or the target student stops answering.
6. The system according to claim 5, characterized in that, The processing layer also includes a feedback unit and a learning analysis unit; The feedback unit is used to output a feedback score and generate feedback information, which includes the reasons for the loss of points and suggestions for improvement; wherein, both the first feedback result and the second feedback result include the corresponding feedback score and the corresponding feedback information. The learning analysis unit is used to summarize the first feedback result and the second feedback result to generate a visualized classroom learning report, and push the visualized classroom learning report to the corresponding second terminal.
7. The system according to claim 6, characterized in that, The processing layer further includes a processing unit and a storage unit; The storage unit is used to store the knowledge point logic of the target subject, the image features of the target subject, and the student classroom interaction data of the target subject. The processing unit is used to construct a knowledge graph and an image recognition model of the target subject by using machine learning algorithms and combining the knowledge point logic of the target subject, the image features of the target subject, and the classroom interaction data of students in the target subject. Based on the knowledge graph of the target subject, the unit uses the image recognition model of the target subject to perform a first AI processing on the first image recognition analysis result to obtain a first feedback result for the target student, and performs a second AI processing on the second image recognition analysis result to obtain a second feedback result for the target student. The knowledge point logic of the target discipline is determined based on the knowledge point name, knowledge point level, inclusion relationship, prerequisite relationship, association relationship and difficulty level of the knowledge points in the target discipline. The image features of the target subject are determined based on the pre-class learning feedback image corresponding to the target subject, and / or the in-class learning feedback image, and / or the post-class learning feedback image; The student classroom interaction data for the target subject includes pre-class learning feedback images, in-class learning feedback images, and post-class learning feedback images for the target subject, as well as pre-class learning feedback audio information, in-class learning feedback audio information, and post-class learning feedback audio information for the target subject, along with response feedback images and response audio information for the target subject.
8. The system according to claim 7, characterized in that, The processing unit is further configured to push suitable learning resources to the corresponding target student based on the feedback score after the second feedback result meets the preset feedback result or after the target student stops answering.
9. The system according to claim 2, characterized in that, The first terminal also includes a processor; The image acquisition module is used to acquire free-response images in response to the free-response operation of the target student; The controller is used to respond to the free-response question and answer operation of the target student, control the image acquisition module to acquire free-response question and answer images, and control the sound acquisition module to acquire free-response question and answer sound information; The controller is also used to transmit the free-response image and the free-response audio information to the processor; The processor is used to process the free-response image and / or the free-response sound information according to the preset computing power model and preset image recognition model built into the first terminal, determine the knowledge points of the corresponding target subject, generate progressive questions according to the knowledge points of the target subject and according to the preset questioning logic, and output them to the target student through the interaction module.
10. The system according to claim 9, characterized in that, The processor is further configured to acquire the language behavior characteristics and emotional expression characteristics of the target student based on the free-response image and the free-response sound information, and determine the stress information of the target student based on the language behavior characteristics and emotional expression characteristics. When the stress information of the target student meets the preset warning conditions, the processing unit sends the warning information to the corresponding second terminal through the signal transceiver module.