Low-cost real-time understanding degree identification system based on pupil data
Through a low-cost real-time understanding degree recognition system based on pupil data, the user's pupil data is analyzed using deep learning models, and the problems of high costs, single mode limitations and poor user experience in the existing technology are solved, and automated and real-time recognition in multimodal scenarios are realized, improving the accuracy and user experience of recognition.
Patent Information
- Application Number
- CN202510037722.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-10
AI Technical Summary
The existing level of understanding identification technology has problems such as high cost, single mode limitations, poor user experience, low automation level and lack of real-time feedback.
A low-cost real-time understanding degree recognition system based on pupil data is adopted to collect user's pupil data in real time through a webcam, and a deep learning model (such as CNN, RNN, LSTM, BiLSTM and Transformer) is used for analysis and prediction, so as to realize automation and real-time recognition in multimodal scenarios.
It reduces the cost of using the understanding degree identification system, enhances real-time and multimodal support capabilities, improves user experience and automation, and provides accurate understanding degree assessment in a variety of application scenarios.
Smart Images

Figure CN120011100A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computers, and more particularly, to a low-cost real-time understanding degree recognition system based on pupil data. Background Art
[0002] In the current development of human-computer interaction systems, the user's ability to understand information (i.e., the degree of understanding) has become a key factor affecting the quality and effectiveness of the interactive experience. Accurate identification and analysis of the degree of understanding not only helps to optimize the user experience, but also supports the provision of personalized services. However, existing methods for identifying the degree of understanding have many technical limitations and are difficult to be widely used in the real world.
[0003] First, traditional comprehension level identification methods mainly rely on self-report and performance tests, such as listening tests, reading comprehension tests, and multiple-choice tests (technical routes such as Figure 1 These methods can provide valuable data and insights in a controlled experimental environment, but their results are often not available in real time and lack the ability to reflect the user's immediate status. In addition, these methods rely on the user's active participation and self-reporting. This subjectivity may lead to biased results, such as "initial rise bias" (i.e., users may make inaccurate self-assessments of their understanding in the early stages). These subjective factors not only affect the reliability of the data, but may also interrupt the user's normal interaction process in actual applications, reducing the user experience.
[0004] Secondly, in order to overcome these limitations, some understanding degree recognition technologies based on physiological signals and machine learning have been developed in recent years (technical routes such as Figure 2 These technologies collect users' physiological indicators, such as electroencephalogram (EEG) signals and eye movement data, to assess their understanding of information. Although these methods can achieve real-time recognition to a certain extent and provide a more objective data basis, they still face the challenges of high cost and inconvenience in use. EEG devices require complex hardware support and professional operators. Although eye trackers can capture users' gaze points and line of sight, they are usually expensive and not easy to carry, which limits the application of these technologies in daily life.
[0005] In addition, existing comprehension recognition technologies are mostly limited to single-mode scenarios, such as measuring only the comprehension of text or speech. This single-mode limitation makes it impossible to provide effective comprehension assessment in more complex multimodal environments, such as scenarios that combine multiple input forms such as text, speech, and images. This limitation limits the applicability of the technology in a variety of application scenarios, such as education, medical care, and intelligent assistants, which require multimodal interaction.
[0006] Existing technical methods have not been effectively promoted in a wide range of real-world applications due to their reliance on expensive and complex external equipment, lack of real-time feedback capabilities, limitations of a single modality, and negative impact on user experience. In order to solve these problems, a new technical solution is urgently needed, which should have the characteristics of low cost, real-time, user-friendly, and multi-modal support, so as to provide users with accurate understanding assessment in a wider range of application scenarios. Summary of the invention
[0007] The purpose of the embodiments of the present disclosure is to provide a low-cost real-time understanding degree recognition system based on pupil data to solve the many defects of the prior art in understanding degree recognition, including high cost, single modality limitation, poor user experience, low automation level and lack of real-time feedback. These problems mainly stem from the high dependence of the existing methods on expensive external equipment and excessive focus on single modality data, and the failure to effectively integrate multiple information sources and interaction modes.
[0008] This technical solution provides a low-cost real-time understanding degree recognition system based on pupil data, including two parts: front-end and back-end:
[0009] The front end includes a display and interaction unit of the user interface, which captures the user's facial video stream data by calling the network camera, collects and generates the user's pupil data in real time, and transmits the pupil data to the back end through a Socket connection. Specifically, the front end identifies the user's facial area through a facial detection algorithm (MediaPipe) and then locates the eye area. On this basis, a pupil tracking algorithm is used to identify the position and size of the pupil. First, the system obtains video stream data through a network camera, extracts the eye area in each frame, and accurately calculates the pupil center and diameter through edge detection and image segmentation technology. Secondly, the front end tracks the movement of the user's pupil in real time and captures the dynamic changes of the pupil, including the enlargement and contraction of the pupil (i.e., pupil reaction). This process uses image processing technology, including circle detection and contour extraction, to obtain the exact position of the pupil and update the data. Pupil data (including position, diameter, change speed, etc.) is transmitted to the back end in real time through a Socket connection, and the back end further analyzes these data;
[0010] The back end receives the pupil data transmitted by the front end and performs preprocessing, and performs real-time analysis and prediction through multiple modules, the multiple modules including:
[0011] The face detection module transfers the ratio to the pupil extraction module through the calibration module, and transfers the processed frame data to the pupil extraction module through the face recognition and feature extraction module;
[0012] The pupil extraction module processes the frame data using the calibration ratio, performs pupil capture and pupil collection, obtains pupil sequence data and transmits it to the model recognition module;
[0013] The model recognition module takes pupil sequence data as input and adopts pupil extraction deep learning model. By embedding convolutional neural network (CNN), recurrent neural network (RNN), long short-term memory network (LSTM), bidirectional long short-term memory network (BiLSTM) and Transformer model, a complete analysis method is constructed. In order to improve the prediction accuracy and stability, the soft voting strategy in ensemble learning is adopted to integrate the prediction results of different network models. The adaptability and prediction ability of the model in different situations are further enhanced through ensemble learning. The time series changes of pupil data are analyzed to predict the user's subjective and objective understanding. Specifically, the model recognition module first receives the pupil data sequence after front-end collection and preprocessing. These data include pupil size, position and dynamic change information. Subsequently, it is processed through the following deep learning network:
[0014] Convolutional Neural Network (CNN): This model uses a two-layer convolutional neural network structure, each layer contains 64 convolution kernels, the convolution kernel size is 3x3, and the ReLU activation function is used to enhance nonlinear processing capabilities. The first convolution layer reduces the data dimension through max pooling, with a step size of 2, and the data dimension is halved. After convolution processing, the data is passed to the fully connected layer, and finally the comprehensibility score is generated through the Sigmoid activation function. This structure can quickly extract local features and is suitable for processing high-dimensional time series data.
[0015] Recurrent Neural Network (RNN): This model uses a two-layer RNN structure, with each layer containing 256 hidden units. RNN has a strong ability to process time series data and is suitable for capturing the temporal dynamic characteristics in data. When data is input, the RNN layer will gradually update its internal state and generate the user's understanding score through the fully connected layer at the final time point.
[0016] Long Short-Term Memory (LSTM): LSTM has a two-layer structure with 128 hidden units in each layer, which can effectively handle dependencies in long time series. In pupil data processing, LSTM is particularly suitable for capturing the long-term dependency characteristics of pupil size over time. LSTM uses a dropout rate of 0.5 to prevent overfitting. The output is mapped to between 0 and 1 through the Sigmoid function to generate the user's understanding score.
[0017] Bidirectional Long Short-Term Memory (BiLSTM): BiLSTM uses a bidirectional structure (i.e., forward and backward LSTM layers) to enable the model to consider both past and future time steps, thereby providing a comprehensive understanding of pupil data. Each layer contains 128 hidden units and uses a dropout rate of 0.5. The output of the BiLSTM model is converted into a comprehension score through a sigmoid function.
[0018] Transformer model: The Transformer model uses the self-attention mechanism, through a multi-layer encoder (16 layers), 256 hidden units in each layer, and 8 attention heads. It can process large-scale time series data in parallel and is suitable for processing longer time series dependencies in pupil data. When generating comprehension scores, the model maps the final encoder output to the score space through a fully connected layer. The advantage of the Transformer model lies in its ability to handle long-term dependencies and its ability to better handle large-scale data sets.
[0019] The user's understanding performance predicted by the model includes subjective understanding and objective understanding. Subjective understanding refers to the participant's self-reported level of understanding, which is usually collected through questionnaires, such as the participant's perception of understanding of a task or information. Objective understanding refers to the proportion of information that users can accurately repeat or reproduce, reflecting the user's actual mastery of the content. It is usually measured through tests or quizzes, such as what proportion of knowledge points users can correctly answer or the accuracy of repeating given content. Together, these two indicators provide a comprehensive picture of the user's understanding of the content. The subjective understanding reflects the user's cognitive perception, while the objective understanding focuses on actual accuracy. The calculation formula for subjective understanding can be expressed as:
[0020]
[0021] Among them, the understanding score is the understanding score reported by the participants themselves, and the maximum possible score is the highest value of the score, usually 9 points. This indicator reflects the user's subjective feelings about the task or information.
[0022] The calculation formula of objective understanding degree can be expressed as:
[0023]
[0024] The correct content is the part of the content that the user can accurately understand and repeat, and the total content is the entire content that needs to be repeated. This indicator can objectively reflect the user's actual grasp of the information.
[0025] Finally, through the adaptive system, the degree of understanding is applied to improve the user experience. Specifically, these two metrics are the core of understanding performance evaluation. The subjective degree of understanding focuses on the user's perception, while the objective degree of understanding emphasizes accuracy and repetition ability. In the field of artificial intelligence (AI), the measurement of understanding is increasingly important because users need to understand text, voice, images and other information in order to effectively interact with AI systems. By measuring and analyzing the user's understanding performance, the AI system can adjust the interactive content according to the user's level of understanding, ensuring that the information presentation matches the user's understanding ability, thereby optimizing the interaction effect.
[0026] The preprocessing methods include denoising and normalization.
[0027] The technical effects to be achieved by the embodiments of the present invention are:
[0028] A low-cost, real-time, deep learning and pupil data-based understanding recognition system is proposed. The system uses common webcams to collect user pupil data without any external complex equipment, and can automatically and real-time evaluate the user's understanding in multimodal scenarios. By introducing advanced deep learning algorithms, the system can not only significantly reduce hardware costs and operational complexity, but also significantly improve recognition accuracy and user experience. This innovative technical path overcomes many shortcomings of existing technologies and provides a more efficient and practical solution for the field of intelligent human-computer interaction.
[0029] And can further achieve the following technical effects:
[0030] 1. Reduce the cost of using the understanding degree recognition system. By collecting the user's pupil data through a common webcam, it no longer relies on expensive physiological sensors, significantly reducing the cost of equipment procurement and use. In addition, by using a pre-trained deep learning model, the computing resources and energy consumption required for model training can be reduced, thereby reducing the overall cost of the system.
[0031] 2. Enhance the real-time performance of the understanding recognition system, collect and process the user's pupil data in real time, and use a variety of deep learning models (such as CNN, RNN, LSTM, BiLSTM and Transformer) for predictive analysis to ensure that instant feedback is provided during the user's interaction with the system. This real-time performance allows the system to dynamically adjust the complexity and presentation of information, thereby improving the user's understanding and experience.
[0032] 3. Enhance multimodal scenario support to strengthen the flexibility of understanding recognition, and be able to process multiple modal inputs including text, voice, images and their combinations, so that it has a wider application prospect in various scenarios such as education, medical care, and entertainment.
[0033] 4. Reduce the user's external wearable burden to improve user experience. Based on the data collection method based on the webcam, there is no need to wear additional equipment. Users can interact in a natural state, which greatly improves the convenience and comfort of use.
[0034] 5. The whole process is systematized to improve the level of automation. The deep learning model automatically analyzes the user's pupil data and identifies the degree of understanding, avoiding the tedious experimental operations and human intervention in the existing methods, improving the system's automation level and enabling it to be seamlessly applied in real-world environments.
[0035] 6. Enhance the scalability and adaptability of the system with full-stack web system development. The system is designed as a lightweight network plug-in that can be easily embedded in existing human-computer interaction systems (such as intelligent education systems and intelligent medical systems). In addition, the system does not require external devices, has strong adaptability and scalability, and can be applied to a variety of hardware platforms (such as smartphones, tablets, and personal computers). BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The above and other objects and features of the present disclosure will become more apparent from the following description in conjunction with the accompanying drawings.
[0037] Figure 1 is a schematic diagram showing a technology roadmap identified based on the degree of understanding of conventional methods of the prior art;
[0038] Figure 2 is a schematic diagram showing a technology roadmap based on physiological instruments and machine learning according to the prior art;
[0039] Figure 3 is a schematic diagram showing a low-cost real-time understanding degree recognition system based on pupil data according to an embodiment of the present disclosure;
[0040] Figure 4 It is a schematic diagram showing a system architecture diagram of a low-cost real-time understanding degree recognition system technology based on pupil data according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0041] The following specific embodiments are provided to help the reader obtain a comprehensive understanding of the methods, devices and / or systems described herein. However, after understanding the disclosure of the present application, various changes, modifications and equivalents of the methods, devices and / or systems described herein will be clear. For example, the order of operations described herein is only an example and is not limited to those orders set forth herein, but can be changed as will be clear after understanding the disclosure of the present application, except for operations that must occur in a specific order. In addition, for greater clarity and simplicity, the description of features known in the art may be omitted.
[0042] The features described herein can be implemented in different forms and should not be construed as being limited to the examples described herein. Rather, the examples described herein have been provided to illustrate only some of the many possible ways to implement the methods, devices, and / or systems described herein, which will be clear after understanding the disclosure of the present application.
[0043] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more.
[0044] Although terms such as "first", "second", and "third" may be used herein to describe various members, components, regions, layers, or portions, these members, components, regions, layers, or portions should not be limited by these terms. Instead, these terms are only used to distinguish one member, component, region, layer, or portion from another member, component, region, layer, or portion. Therefore, without departing from the teachings of the examples described herein, the first member, first component, first region, first layer, or first portion referred to in the examples may also be referred to as the second member, second component, second region, second layer, or second portion.
[0045] In the specification, when an element (such as a layer, a region or a substrate) is described as being “on”, “connected to” or “coupled to” another element, the element may be directly “on”, “connected to” or “coupled to” another element, or one or more other elements may be present therebetween. Conversely, when an element is described as being “directly on”, “directly connected to” or “directly coupled to” another element, there may be no other elements present therebetween.
[0046] The terms used herein are only used to describe various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. The terms "comprise", "include" and "have" indicate the presence of the described features, quantities, operations, components, elements and / or combinations thereof, but do not exclude the presence or addition of one or more other features, quantities, operations, components, elements and / or combinations thereof.
[0047] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by a person of ordinary skill in the art to which the present disclosure belongs after understanding the present disclosure. Unless explicitly defined as such herein, terms (such as those defined in a general dictionary) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and should not be interpreted in an idealized or overly formal manner.
[0048] Furthermore, in the description of examples, when it is considered that a detailed description of a well-known related structure or function would cause vague interpretation of the present disclosure, such a detailed description will be omitted.
[0049] Figure 3 is a schematic diagram showing a low-cost real-time understanding degree recognition system based on pupil data according to an embodiment of the present disclosure.
[0050] In order to achieve the above-mentioned invention object, the technical framework adopted by the present invention is as follows Figure 3 A real-time understanding level recognition system (AI-Mind-Reader) based on deep learning and pupil data is shown.
[0051] The solution aims to achieve efficient user understanding measurement in multimodal scenarios without the need for external equipment.
[0052] The overall design of the system includes the front-end (Vue), the back-end (Flask), and the Socket connection for data transmission. The front-end is developed using the Vue framework and is mainly responsible for the display and interaction of the user interface, including the presentation of text, images, and audio materials, as well as capturing the user's facial video stream data. By calling the webcam, the front-end collects the user's pupil data in real time and transmits this data to the back-end through the Socket connection. The application of the Socket connection ensures the real-time and efficient data transmission between the front-end and the back-end, enabling the system to provide instant feedback during the user's interaction with the system.
[0053] At the back end, the system is built with the Flask framework, which is responsible for receiving pupil data transmitted by the front end and performing real-time analysis and prediction. The back end system contains multiple modules, including face detection module, pupil extraction module and model recognition module. Each module ensures the real-time performance of data processing through a multi-threaded processing mechanism. The system uses tools such as MediaPipe to detect facial features and capture and track pupil data, and generates pupil sequence data in real time. First, the system captures the user's facial video stream data by calling the network camera, collects and generates the user's pupil data in real time, and transmits the pupil data to the back end through a Socket connection. Specifically, the front end identifies the user's facial area through a facial detection algorithm (MediaPipe) and then locates the eye area. On this basis, a pupil tracking algorithm is used to identify the position and size of the pupil. First, the system obtains video stream data through a network camera, extracts the eye area in each frame, and accurately calculates the pupil center and diameter through edge detection and image segmentation technology. Secondly, the front end tracks the movement of the user's pupil in real time and captures the dynamic changes of the pupil, including the enlargement and contraction of the pupil (i.e., pupil reaction). This process uses image processing technology, including circle detection and contour extraction, to obtain the exact position of the pupil and update the data. Pupil data (including position, diameter, change speed, etc.) is transmitted to the backend in real time via a socket connection. In order to ensure the accuracy and consistency of the data, after being transmitted to the backend, the data undergoes preprocessing steps such as denoising and standardization to provide high-quality input for subsequent deep learning model analysis.
[0054] The deep learning model is one of the core technical means of the present invention. The system embeds five different types of deep learning models, including convolutional neural network (CNN), recurrent neural network (RNN), long short-term memory network (LSTM), bidirectional long short-term memory network (BiLSTM) and Transformer model. At the same time, the soft voting strategy in ensemble learning is adopted to integrate the prediction results of different network models, enhancing the adaptability and prediction ability of the model in different situations. These models predict the user's understanding by analyzing the time series changes of pupil data. In order to reduce computing costs and time consumption, these models have been pre-trained and optimized for specific tasks. Through the collaborative work of multiple models, the system can provide more accurate understanding degree recognition results in various complex multimodal scenarios, thereby improving the user's experience and the breadth of application. Specifically, the model recognition module first receives the pupil data sequence after front-end collection and preprocessing, which includes the size, position and dynamic change information of the pupil. Subsequently, it is processed by the following deep learning network:
[0055] Convolutional Neural Network (CNN): This model uses a two-layer convolutional neural network structure, each layer contains 64 convolution kernels, the convolution kernel size is 3x3, and the ReLU activation function is used to enhance nonlinear processing capabilities. The first convolution layer reduces the data dimension through max pooling, with a step size of 2, and the data dimension is halved. After convolution processing, the data is passed to the fully connected layer, and finally the comprehensibility score is generated through the Sigmoid activation function. This structure can quickly extract local features and is suitable for processing high-dimensional time series data.
[0056] Recurrent Neural Network (RNN): This model uses a two-layer RNN structure, with each layer containing 256 hidden units. RNN has a strong ability to process time series data and is suitable for capturing the temporal dynamic characteristics in data. When data is input, the RNN layer will gradually update its internal state and generate the user's understanding score through the fully connected layer at the final time point.
[0057] Long Short-Term Memory (LSTM): LSTM has a two-layer structure with 128 hidden units in each layer, which can effectively handle dependencies in long time series. In pupil data processing, LSTM is particularly suitable for capturing the long-term dependency characteristics of pupil size over time. LSTM uses a dropout rate of 0.5 to prevent overfitting. The output is mapped to between 0 and 1 through the Sigmoid function to generate the user's understanding score.
[0058] Bidirectional Long Short-Term Memory (BiLSTM): BiLSTM uses a bidirectional structure (i.e., forward and backward LSTM layers) to enable the model to consider both past and future time steps, thereby providing a comprehensive understanding of pupil data. Each layer contains 128 hidden units and uses a dropout rate of 0.5. The output of the BiLSTM model is converted into a comprehension score through a sigmoid function.
[0059] Transformer model: The Transformer model uses the self-attention mechanism, through a multi-layer encoder (16 layers), 256 hidden units in each layer, and 8 attention heads. It can process large-scale time series data in parallel and is suitable for processing longer time series dependencies in pupil data. When generating comprehension scores, the model maps the final encoder output to the score space through a fully connected layer. The advantage of the Transformer model lies in its ability to handle long-term dependencies and its ability to better handle large-scale data sets.
[0060] The user's understanding performance predicted by the model includes subjective understanding and objective understanding. Subjective understanding refers to the participant's self-reported level of understanding, which is usually collected through questionnaires, such as the participant's perception of understanding of a task or information. Objective understanding refers to the proportion of information that users can accurately repeat or reproduce, reflecting the user's actual mastery of the content. It is usually measured through tests or quizzes, such as what proportion of knowledge points users can correctly answer or the accuracy of repeating given content. Together, these two indicators provide a comprehensive picture of the user's understanding of the content. The subjective understanding reflects the user's cognitive perception, while the objective understanding focuses on actual accuracy. The calculation formula for subjective understanding can be expressed as:
[0061]
[0062] Among them, the understanding score is the understanding score reported by the participants themselves, and the maximum possible score is the highest value of the score, usually 9 points. This indicator reflects the user's subjective feelings about the task or information.
[0063] The calculation formula of objective understanding degree can be expressed as:
[0064]
[0065] The correct content is the part of the content that the user can accurately understand and repeat, and the total content is the entire content that needs to be repeated. This indicator can objectively reflect the user's actual grasp of the information.
[0066] Experimental verification:
[0067] The system generates the user's understanding level recognition results through real-time analysis of pupil data, and transmits the results back to the front end through a Socket connection for dynamic display. The user can obtain real-time feedback during the interaction with the system, and the system can dynamically adjust the complexity and presentation of information according to the user's understanding level to further optimize the user experience. In this way, the present invention realizes the automation, real-time monitoring and feedback functions of the user's understanding level in a multimodal interaction scenario, providing innovative solutions for multiple application fields such as smart education and smart medical care.
[0068] Figure 3The user interaction process of the AI-Mind-Reader system and the working mechanism of its adaptive system are demonstrated in detail. The system adjusts the interaction strategy according to the user groups (such as regular users, elderly groups, non-native users, etc.), and uses pupil data and facial feature capture to continuously monitor and collect user understanding data. The adaptive system dynamically adjusts the complexity and presentation of the content to improve the user experience. The system is not only suitable for a variety of application fields such as conversational AI (such as ChatGPT) and smart education (such as Khan Academy), but can also adapt to autonomous driving and other complex scenarios.
[0069] Figure 4 The overall structure of the AI-Mind-Reader comprehension level recognition system is shown in Figure 1. The front-end (Vue) is responsible for stimulus display (such as the presentation of text, images, and audio materials) and the selection of the camera area. The captured pupil sequence data is transmitted to the back-end (Flask) via a Socket connection. Multiple modules in the back-end (such as face detection, pupil extraction, and model recognition) work in parallel to process input data in real time and generate comprehension level results. Finally, the results are returned to the front-end for display via a Socket connection. The entire system realizes real-time data interaction between the front-end and the back-end through a Socket connection, effectively supporting multi-modal comprehension level recognition.
[0070] The present invention realizes a real-time, low-cost, multi-modal, pupil data-based understanding degree recognition system. Compared with the prior art, the present invention has the following five advantages:
[0071] The present invention provides real-time user understanding degree recognition and measurement, which has higher real-time use value. The system captures the user's facial activities through a network camera, captures pupil sequence data based on MediaPipe, and passes it into an embedded deep learning model. In the present invention, five deep neural network models such as CNN, RNN, LSTM, BiLSTM, and Transformer are selected, and the model training is completed through laboratory pupil sequence data (720 pupil sequences, each sequence has 16381 pupil time point data), which can capture complex features from the pupil sequence data and accurately predict the corresponding understanding degree. Based on the incoming pupil data, the embedded model can calculate the understanding degree in real time, and pass the calculation results to the web front end for real-time update and display. Traditional understanding degree measurement methods, such as self-report, understanding performance test, etc., adopt the form of regular evaluation of understanding degree, and the user's understanding process will be interrupted by the evaluation content, so it does not have good continuity and real-time performance. Compared with the existing understanding degree measurement method, the present invention can ensure real-time feedback of understanding degree through real-time capture of network cameras, continuous transmission of sockets, and real-time calculation of deep learning models. The real-time performance of the present invention provides a good foundation for future applications such as embedding into other interactive systems and dynamically adapting to user understanding capabilities.
[0072] The present invention has achieved a high accuracy in the recognition of the degree of understanding. The five deep neural networks used in the present invention, namely CNN, RNN, LSTM, BiLSTM, and Transformer, showed good prediction accuracy on the test set (MSE average value was 0.0625). The prediction accuracy of the present invention was further verified by experiments. It was found that the best performing Transformer model had a MAPE of 29.5% for the objective degree of understanding (meaning a prediction accuracy of 70.5%) and a MAPE of 14.9% for the subjective degree of understanding (meaning a prediction accuracy of 85.1%). The recognition accuracy of the present invention for the degree of understanding is much higher than random guessing (50%) and 41% of the existing prediction method based on eye movement data.
[0073] The present invention reduces the cost of understanding degree recognition, and has high cost performance and convenience of use. The present invention is developed as a network plug-in, which can be easily embedded in an existing human-computer interaction system. All models used in the system have been pre-trained, reducing the computing cost, data cost and energy cost of the user for training. In addition, the system of the present invention is lightweight in design, does not require any external equipment, and only requires a computer camera during use. Compared with existing understanding degree recognition methods based on wearable sensors, such as electroencephalograms, eye trackers, etc., the present invention significantly reduces the introduction cost, use cost and maintenance cost of the technology.
[0074] The recognition of the degree of understanding of the present invention takes into account the subjective degree of understanding and the objective degree of understanding respectively. Most of the existing methods for identifying the degree of understanding only consider the single dimension of understanding performance, lacking a comprehensive consideration of the subjective degree of understanding and the objective degree of understanding. However, for human-computer interaction systems, the subjective degree of understanding and the objective degree of understanding have the same important significance: the objective degree of understanding reflects the actual degree of understanding of the user, which is the most important purpose of the interaction task. The system can adaptively adjust the difficulty of the interaction task, auxiliary understanding information, etc. according to the objective degree of understanding; the subjective degree of understanding reflects the user's judgment on his or her own degree of understanding, which has an important impact on the quality of user experience in the interaction. Therefore, the comprehensive recognition of the degree of understanding from both subjective and objective perspectives of the present invention provides important technical support for the optimization and improvement of human-computer interaction systems.
[0075] Although some embodiments of the present disclosure have been shown and described, it will be appreciated by those skilled in the art that modifications may be made to the embodiments without departing from the principles and spirit of the present disclosure, the scope of which is defined by the claims and their equivalents.
Claims
1. A low-cost real-time understanding degree recognition system based on pupil data, characterized in that: It consists of two parts: front-end and back-end: The front end includes a display and interaction unit of a user interface, which captures the user's facial video stream data by calling a network camera, collects and generates the user's pupil data in real time, and transmits the pupil data to the back end via a Socket connection; Specifically, the front end uses a facial detection algorithm to identify the user's facial area and then locate the eye area. On this basis, a pupil tracking algorithm is used to identify the position and size of the pupil. First, the system obtains video stream data through a network camera and extracts the eye area in each frame. The pupil center and diameter are accurately calculated through edge detection and image segmentation technology. Secondly, the front end uses image processing technology to obtain the accurate position of the pupil using circle detection and contour extraction, tracks the movement of the user's pupil in real time, captures the dynamic changes of the pupil, including the pupil's dilation and contraction reactions, and updates the data; The pupil data of position, diameter, and change speed are transmitted to the backend in real time via a Socket connection, and the backend further analyzes the data; The back end receives the pupil data transmitted by the front end and performs preprocessing, and performs real-time analysis and prediction through multiple modules, the multiple modules including: The face detection module transfers the ratio to the pupil extraction module through the calibration module, and transfers the processed frame data to the pupil extraction module through the face recognition and feature extraction module; The pupil extraction module processes the frame data using the calibration ratio, performs pupil capture and pupil collection, obtains pupil sequence data and transmits it to the model recognition module; The model recognition module takes pupil sequence data as input and adopts the pupil extraction deep learning model. By embedding parallel convolutional neural networks, recurrent neural networks, long short-term memory networks, bidirectional long short-term memory networks and Transformer models, it integrates the prediction results of different network models through the soft voting strategy in ensemble learning to construct a complete analysis method.
2. A low-cost real-time understanding degree recognition system based on pupil data as claimed in claim 1, characterized in that: The specific implementation of the model recognition module is as follows: Firstly, receiving a pupil data sequence collected and preprocessed by a front end, wherein the pupil data includes information on the size, position and dynamic change of the pupil; It is then processed by a separately designed deep learning network, the structure of which includes: Convolutional neural network: A two-layer convolutional neural network structure is used, each layer contains 64 convolution kernels, the convolution kernel size is 3x3, and the ReLU activation function is used to enhance the nonlinear processing capability. The first convolution layer reduces the data dimension through maximum pooling, with a step size of 2, and the data dimension is halved; after convolution processing, the data is passed to the fully connected layer, and finally the comprehension score is generated through the Sigmoid activation function; Recurrent Neural Network: It uses a two-layer RNN structure, each layer contains 256 hidden units, and generates the user's understanding score through a fully connected layer at the final time point; Long short-term memory network: It consists of two layers, with 128 hidden units in each layer. In pupil data processing, a dropout rate of 0.5 is used to prevent overfitting. The output is mapped to between 0 and 1 through the Sigmoid function to generate the user's understanding score. Bidirectional long short-term memory network: Through the bidirectional connection of the forward and backward LSTM layers, the model can consider the past and future time step information at the same time. Each layer contains 128 hidden units and uses a dropout rate of 0.
5. The output results are converted into understanding scores through the Sigmoid function. Transformer model: Utilizes the self-attention mechanism, through 16 layers of encoders, 256 hidden units in each layer, and 8 attention heads, to process large-scale time series data in parallel. When generating comprehension scores, the final encoder output is mapped to the score space through a fully connected layer.
3. A low-cost real-time understanding degree recognition system based on pupil data as claimed in claim 2, characterized in that: The individually designed deep learning networks all use the cross entropy loss function during the training process, the optimizer adopts the Adam optimization algorithm, the learning setting is 0.00001, the batch size is 25, the number of training cycles is 100, and the performance of each neural network structure is evaluated by the mean square error and prediction accuracy indicators. The training process uses regularization and Dropout techniques to reduce overfitting.
4. A low-cost real-time understanding degree recognition system based on pupil data as claimed in claim 3, characterized in that: The predicted user understanding performance includes subjective understanding and objective understanding; subjective understanding is defined as the participants' self-reported understanding level, collected through questionnaires, and objective understanding is defined as the proportion of information that users can accurately repeat or reproduce, measured through tests or quizzes. Together, the two indicators provide a comprehensive picture of the user's understanding of the content; The calculation formula of subjective understanding degree is expressed as: Among them, the understanding score is the participant's self-reported understanding score, and the maximum possible score is the highest value of the score; The calculation formula of objective understanding degree is expressed as: The correct content is the part of the content that the user can accurately understand and repeat, and the total content is the entire content that needs to be repeated; Finally, through adaptive systems, when a low level of comprehension is detected, automated measures are taken to provide additional supplementary information and additional interpretation of the content to apply the level of comprehension to improve the user experience.
5. A low-cost real-time understanding degree recognition system based on pupil data as claimed in claim 4, characterized in that: The preprocessing methods include denoising and normalization.
Citation Information
Patent Citations
Reading ability assessment method, system and reading ability assessment auxiliary device
CN109567817A
Mouse pupil detection method, system, device, equipment and medium
CN117830386A
Artificial intelligence psychological assessment method and system based on multi-modal input
CN119108112A
Contaminated water advanced treatment system
KR102756888B1
Pupil dynamics, physiology, and performance for estimating competency in situational awareness
US20250005939A1