Artificial Intelligence-based Student Classroom Learning Status Assessment Method and System
By normalizing, augmenting and gamma correction processing of students' classroom learning images, combining low-level and advanced feature extraction, comprehensive feature sequences are constructed, and the problem of difficult to capture image quality differences and dynamic changes is solved, and a high-accurate learning state evaluation is achieved.
Patent Information
- Application Number
- CN202510505312.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The existing student classroom learning status evaluation methods have image quality differences, noise interference and rapid and complex actions that lead to blurred images and lack of details, resulting in low evaluation accuracy, difficulty in accurately capturing subtle dynamic changes and complex dynamic behaviors, and insufficient evaluation accuracy.
By normalizing each frame of image, primary and secondary enhancement, and gamma correction, low-level and advanced features are extracted, comprehensive feature sequences are constructed, and long-term memory networks are used for evaluation to enhance image details and learning status reflections.
It significantly improves image quality and evaluation accuracy, can accurately capture subtle dynamic changes and complex dynamic behaviors, provides refined classroom learning status scores, and enhances the accuracy and stability of evaluation.
Smart Images

Figure CN120032300B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and specifically refers to a method and system for evaluating the learning state of students in the classroom based on artificial intelligence. Background Art
[0002] The method for evaluating the learning state of students in the classroom is a method that uses image processing and artificial intelligence technologies to process and analyze the learning videos of students in the classroom, identify the key features reflecting the learning state of students, evaluate the learning state of students in the classroom in real time, provide a scientific basis for educators, thereby improving the learning effect of students, optimizing classroom management, and improving the overall teaching quality. However, in the existing methods for evaluating the learning state of students in the classroom, there are problems such as differences in image quality, noise interference, and blurring and lack of details in images caused by fast and complex movements, resulting in low accuracy of evaluating the learning state of students in the classroom; the existing methods for evaluating the learning state of students in the classroom have problems in accurately capturing the subtle dynamic changes and complex dynamic behaviors in the learning process of students, ignoring the detailed changes in the images, resulting in insufficient accuracy of evaluating the learning state of students. Summary of the Invention
[0003] In view of the above situation, to overcome the defects of the prior art, the present invention provides a method and system for evaluating the learning status of students in the classroom based on artificial intelligence. Aiming at the problems existing in the existing methods for evaluating the learning status of students in the classroom, such as differences in image quality, noise interference, blurring of images and lack of details caused by fast and complex actions, resulting in low accuracy of evaluating the learning status of students in the classroom, this solution performs normalization processing on each frame of image to eliminate equipment and environmental differences and ensure input consistency; by calculating the neighborhood pixel screening coefficient, local minimum divisible amount and pixel calibration factor, the image is initially enhanced to remove noise and retain learning-related features, enhancing the key details of the image; by calculating the mean and standard deviation of the image to obtain the enhancement upper limit value and enhancement lower limit value, the image is secondarily enhanced to reasonably control the image enhancement range and avoid distortion, making the details clearer; then gamma correction processing is performed to obtain a sequence of enhanced learning images, ensuring the consistency of image brightness, improving the stability of the evaluation model, significantly improving the quality of each frame of image, and thus increasing the accuracy of evaluating the learning status of students; aiming at the problems existing in the existing methods for evaluating the learning status of students in the classroom, such as difficulty in accurately capturing the subtle dynamic changes and complex dynamic behaviors in the learning process of students, and ignoring the detail changes in the image, resulting in insufficient accuracy of evaluating the learning status of students in the classroom, this solution extracts low-level features through standard convolution and uses depthwise separable convolution to extract high-level features, enhancing the comprehensive reflection of the learning status of students; by calculating global features to construct a cosine similarity matrix, and then calculating a quantization matrix and a two-dimensional quantization count matrix to obtain statistical features, the low-level features are enhanced, which can not only capture basic behavior features but also reflect the subtle changes in the learning status; by enhancing the high-level features through attention weights, the sensitivity to detail changes can be enhanced, improving the recognition accuracy of students' behavior status; constructing a comprehensive feature sequence and obtaining a status score through a long short-term memory network can comprehensively consider the behavior performance of students in the classroom environment, comprehensively reflect the learning status of students, provide a more refined score for the learning status in the classroom, and enhance the accuracy and stability of the evaluation.
[0004] The technical solution adopted by the present invention is as follows: The method for evaluating the learning status of students in the classroom based on artificial intelligence provided by the present invention includes the following steps:
[0005] Step S1: Collection of students' classroom learning data;
[0006] Step S2: Enhancement of students' classroom learning images;
[0007] Step S3: Construction of an evaluation model for students' classroom learning status;
[0008] Step S4: Generation of an evaluation report.
[0009] Further, in step S1, the collection of historical student classroom learning data is carried out, and the historical student classroom learning data includes student classroom learning videos and status scores. The status scores are used as data labels.
[0010] Further, in step S2, the enhancement of the student classroom learning images specifically includes the following steps:
[0011] Step S21: Construct a sequence of learning images; extract consecutive frames from the student classroom learning video to generate a sequence of learning images;
[0012] Step S22: Normalization; perform normalization processing on each frame image in the sequence of learning images;
[0013] Step S23: Initial image enhancement; perform the following initial enhancement processing on each normalized frame image in the sequence of learning images in turn:
[0014] Step S231: Design the neighborhood pixel screening coefficient; calculate the neighborhood pixel screening coefficient at each pixel point in the image according to the local maximum value, local minimum value and pixel value variance within the neighborhood of each pixel point; the formula used is as follows:
[0015] ;
[0016] In the formula, is the neighborhood pixel screening coefficient at the pixel point (x, y), and are respectively the local maximum value and local minimum value within the 5×5 neighborhood centered on the pixel point (x, y), is the variance of all pixel values within the 5×5 neighborhood centered on the pixel point (x, y), x is the abscissa index of the pixel point in the image, y is the ordinate index of the pixel point in the image, is the exponential function;
[0017] Step S232: Calculate the local minimum divisible amount; take the product of the neighborhood pixel screening coefficient and the local minimum value as the local minimum divisible amount at each pixel point;
[0018] Step S233: Design the pixel adjustment factor; calculate the pixel adjustment factor at each pixel point according to the local minimum divisible amount and the local maximum value; the formula used is as follows:
[0019] ;
[0020] In the formula, is the pixel adjustment factor at the pixel point (x, y), is the local minimum divisible amount at the pixel point (x, y);
[0021] Step S234: Initial enhancement; perform initial enhancement processing on each pixel point according to the pixel calibration factor, local minimum divisible amount, and local maximum value; the formula used is as follows:
[0022] ;
[0023] In the formula, is the normalized pixel value at pixel point (x, y), is the pixel value after initial enhancement;
[0024] Step S24: Image secondary enhancement; perform the following secondary enhancement processing on each frame of the image after initial enhancement in the learning image sequence in turn:
[0025] Step S241: Re - normalization; perform re - normalization processing on the image after initial enhancement according to the min - max normalization method;
[0026] Step S242: Calculate the enhancement upper limit value and enhancement lower limit value; calculate the enhancement upper limit value and enhancement lower limit value at each pixel point in the image according to the mean and standard deviation of all pixel values in the image after re - normalization; the formula used is as follows:
[0027] ;
[0028] ;
[0029] In the formula, and are the enhancement upper limit value and enhancement lower limit value at pixel point (x, y) respectively, is the pixel value processed according to the min - max normalization method, β is the adjustment factor, and μ and σ are the mean and standard deviation of all pixel values in the image after re - normalization respectively;
[0030] Step S243: Secondary enhancement; perform secondary enhancement processing on each pixel point according to the enhancement upper limit value and enhancement lower limit value; the formula used is as follows:
[0031] ;
[0032] In the formula, is the pixel value after secondary enhancement at pixel point (x, y);
[0033] Step S25: Gamma correction; perform gamma correction processing on each frame of the image after secondary enhancement in the learning image sequence to obtain the enhanced learning image sequence.
[0034] Further, in step S3, the construction of the student classroom learning status evaluation model specifically includes the following steps:
[0035] Step S31: Feature extraction; for each frame image in the learning image enhancement sequence, the following feature extraction processes are sequentially performed:
[0036] Step S311: Low-level feature extraction; input the image into the first convolutional layer for processing to extract the low-level feature map;
[0037] Step S312: High-level feature extraction; input the low-level feature map into the depthwise separable convolutional layer, the second convolutional layer, the average pooling layer, and the fully connected layer for processing to extract the high-level feature map;
[0038] Step S32: Low-level feature enhancement; for the low-level feature maps extracted from each frame image in the learning image enhancement sequence, the following enhancement processes are sequentially performed:
[0039] Step S321: Calculate the global feature; perform global average pooling on the low-level feature map to obtain the global feature Q;
[0040] Step S322: Construct the cosine similarity matrix; construct a cosine similarity matrix E based on the cosine similarity between the element values at each position in the low-level feature map and the global feature, and quantize the cosine similarity to obtain F quantization levels ; where S1, S f and S F are the 1st, the fth, and the Fth quantization levels, and f is the index;
[0041] Step S323: Construct the quantization matrix; perform quantization processing on the values at each position in the cosine similarity matrix E to obtain a quantization matrix H of size R×F; the formula used is as follows:
[0042] ;
[0043] In the formula, is the value at position (a, d) in the cosine similarity matrix E, r is the index, is the value at position (r, f) in the quantization matrix H, and R is the number of elements in the low-level feature map;
[0044] Step S324: Construct the two-dimensional quantization count matrix; generate the two-dimensional quantization count matrix L based on the quantization matrix H. The first dimension in the two-dimensional quantization count matrix represents each quantization level, and the second dimension represents the corresponding normalized count; the formula used is as follows:
[0045] ;
[0046] In the formula, Lf is the normalized count corresponding to the f-th quantization level, is a concatenation operation;
[0047] Step S325: Calculate statistical features; Process the two-dimensional quantization count matrix L through a multi-layer perceptron and concatenate it with the global feature Q to obtain the statistical feature Z;
[0048] Step S326: Calculate the enhanced low-level features; Calculate the enhanced low-level features according to the statistical feature Z and the quantization matrix H; The formula used is as follows:
[0049] ;
[0050] In the formula, V is the enhanced low-level feature, T is the matrix transpose operation, is a standard convolution of size 1×1, is the Softmax activation function;
[0051] Step S33: High-level feature enhancement; For the high-level feature map O extracted from each frame of the learning image enhancement sequence, perform enhancement processing in sequence; The enhancement processing is to first calculate the attention weight W of the high-level feature map; Then perform weighted summation on the high-level feature map through the attention weight to obtain the enhanced high-level feature P;
[0052] Step S34: Construct a comprehensive feature sequence; Merge the enhanced low-level feature V and the enhanced high-level feature P through a concatenation operation to obtain the comprehensive feature M of each frame of image, and integrate the comprehensive features of all frames of images in the order in the learning image enhancement sequence to obtain a comprehensive feature sequence;
[0053] Step S35: Evaluation; Input the comprehensive feature sequence into a long short-term memory network for evaluation to obtain the status score of the student's classroom learning.
[0054] Furthermore, in step S4, the generation of the evaluation report is to collect the real-time student classroom learning video and perform student classroom learning image enhancement processing, and then input it into the student classroom learning status evaluation model for evaluation, and generate an evaluation report on the student classroom learning according to the output status score.
[0055] The student classroom learning status evaluation system based on artificial intelligence provided by the present invention includes a student classroom learning data collection module, a student classroom learning image enhancement module, a module for constructing a student classroom learning status evaluation model, and a module for generating an evaluation report;
[0056] The student classroom learning data collection module collects historical student classroom learning data and sends the data to the student classroom learning image enhancement module;
[0057] The student classroom learning image enhancement module normalizes each frame of the image. By calculating the neighborhood pixel screening coefficient, local minimum divisible quantity, and pixel adjustment factor, the image is initially enhanced. By calculating the mean and standard deviation of the image, the enhancement upper limit value and enhancement lower limit value are obtained, and the image is enhanced for the second time. Then, gamma correction processing is performed to obtain the learning image enhancement sequence, and the data is sent to the module for constructing the student classroom learning status evaluation model;
[0058] The module for constructing the student classroom learning status evaluation model extracts low-level features using standard convolution, extracts high-level features using depthwise separable convolution, constructs a cosine similarity matrix through global features, calculates the quantization matrix and two-dimensional quantization count matrix to obtain statistical features, enhances the low-level features, enhances the high-level features through attention weights, constructs a comprehensive feature sequence, obtains the status score, and sends the data to the module for generating the evaluation report;
[0059] The module for generating the evaluation report collects the real-time student classroom learning video and performs student classroom learning image enhancement processing, and then inputs it into the student classroom learning status evaluation model for evaluation, and generates an evaluation report on the student classroom learning according to the output status score.
[0060] The beneficial effects achieved by the present invention using the above solution are as follows:
[0061] (1) Aiming at the problems in the existing student classroom learning status evaluation methods, such as image quality differences, noise interference, blurring of images due to fast and complex actions and lack of details, resulting in low evaluation accuracy of the student classroom learning status, this solution normalizes each frame of the image to eliminate equipment and environmental differences and ensure input consistency; by calculating the neighborhood pixel screening coefficient, local minimum divisible quantity, and pixel adjustment factor, the image is initially enhanced to remove noise and retain learning-related features, enhancing the key details of the image; by calculating the mean and standard deviation of the image, the enhancement upper limit value and enhancement lower limit value are obtained, and the image is enhanced for the second time to reasonably control the image enhancement range and avoid distortion, making the details clearer; then, gamma correction processing is performed to obtain the learning image enhancement sequence, ensuring image brightness consistency, improving the stability of the evaluation model, significantly improving the quality of each frame of the image, and thus improving the accuracy of the student learning status evaluation.
[0062] (2)In view of the problem that existing methods for evaluating students' in-class learning status are difficult to accurately capture the subtle dynamic changes and complex dynamic behaviors in the learning process of students, and ignore the detailed changes in images, resulting in insufficient accuracy in evaluating students' in-class learning status, this solution extracts low-level features through standard convolution and uses depthwise separable convolution to extract high-level features, enhancing the comprehensive reflection of students' learning status; by calculating global features to construct a cosine similarity matrix, and then calculating a quantization matrix and a two-dimensional quantization count matrix to obtain statistical features, enhancing the low-level features, which can not only capture basic behavioral features but also reflect the subtle changes in learning status; by enhancing the high-level features through attention weights, the sensitivity to detailed changes can be enhanced, improving the recognition accuracy of students' behavioral status; constructing a comprehensive feature sequence and obtaining a status score through a long short-term memory network can comprehensively consider the behavioral performance of students in the classroom environment, comprehensively reflect the learning status of students, provide a more refined in-class learning status score, and enhance the accuracy and stability of the evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is a schematic flowchart of the method for evaluating students' in-class learning status based on artificial intelligence provided by the present invention;
[0064] Figure 2 is a schematic diagram of the system for evaluating students' in-class learning status based on artificial intelligence provided by the present invention;
[0065] Figure 3 is a schematic flowchart of step S2;
[0066] Figure 4 is a schematic flowchart of step S3.
[0067] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0069] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. indicating the orientation or position relationship are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0070] Embodiment 1. Refer to Figure 1 , the method for evaluating the learning state of students in class based on artificial intelligence provided by the present invention includes the following steps:
[0071] Step S1: Collecting students' in-class learning data; collecting historical students' in-class learning data;
[0072] Step S2: Enhancing students' in-class learning images; performing normalization processing on each frame of the image, initially enhancing the image by calculating the neighborhood pixel screening coefficient, local minimum divisible amount, and pixel calibration factor, obtaining the upper and lower enhancement limit values by calculating the mean and standard deviation of the image, performing secondary enhancement on the image, and then performing gamma correction processing to obtain the learning image enhancement sequence;
[0073] Step S3: Constructing an evaluation model for students' in-class learning state; extracting low-level features using standard convolution, extracting high-level features using depthwise separable convolution, constructing a cosine similarity matrix through global features, calculating the quantization matrix and two-dimensional quantization count matrix to obtain statistical features, enhancing the low-level features, enhancing the high-level features through attention weights, constructing a comprehensive feature sequence, and obtaining the state score;
[0074] Step S4: Generating an evaluation report; collecting real-time students' in-class learning videos and performing in-class learning image enhancement processing on the students, and then inputting them into the evaluation model for students' in-class learning state to perform evaluation, and generating an evaluation report for students' in-class learning according to the output state score.
[0075] Embodiment 2. Refer to Figure 1 , based on the above embodiment, in step S1, the collection of students' in-class learning data is to collect historical students' in-class learning data, and the historical students' in-class learning data includes students' in-class learning videos and state scores, and the state score is used as the data label.
[0076] Embodiment 3. Refer to Figure 1 and Figure 3 , based on the above embodiment, in step S2, the enhancement of students' in-class learning images specifically includes the following steps:
[0077] Step S21: Construct a learning image sequence; extract 30 frames of images per second to ensure that every important moment in the student's learning process is captured, avoiding missing key information due to discontinuous data; extract consecutive frames from the student's in-class learning video at a rate of 30 frames per second to generate a learning image sequence;
[0078] Step S22: Normalization; the pixel values in video images can vary under the influence of different video devices, different lighting conditions, etc. Through normalization processing, the impact of device and environmental differences on image quality assessment is eliminated, making the subsequent image enhancement processing more reliable; perform normalization processing on each frame of the learning image sequence, scaling the pixel values of the image from [0, 255] to the range [0, 1]; the formula used is as follows:
[0079] ;
[0080] In the formula, and are the original pixel value and the normalized pixel value at the pixel point (x, y) respectively, and x and y are the abscissa index and ordinate index of the pixel point in the image;
[0081] Step S23: Initial image enhancement; in the student's in-class learning video, the student's attention and behavior may vary greatly in local areas. Through initial image enhancement, the local features of the image are strengthened, and noise interference is effectively suppressed, thus making the subsequent learning state recognition more accurate; for each normalized frame of the learning image sequence, perform the following initial enhancement processing in sequence:
[0082] Step S231: Design the neighborhood pixel screening coefficient; the images in the student's in-class learning video are often distorted locally due to background noise, light changes, or the student's rapid movements. These distortions will affect the correct recognition of the student's learning behavior. By calculating the neighborhood pixel screening coefficient, the noise in the local area can be effectively removed while retaining the image features related to the student's learning behavior; calculate the neighborhood pixel screening coefficient at each pixel point in the image according to the local maximum value, local minimum value, and pixel value variance within the neighborhood of each pixel point; the formula used is as follows:
[0083] ;
[0084] In the formula, is the neighborhood pixel screening coefficient at the pixel point (x, y), and are the local maximum value and local minimum value within the 5×5 neighborhood centered on the pixel point (x, y) respectively, , , is the normalized pixel value at pixel point (i, j), is the variance of all pixel values within a 5×5 neighborhood centered at pixel point (x, y), where x and i are the horizontal coordinates of pixel points in the image, and y and j are the vertical coordinates of pixel points in the image, 、 and are the exponential function, the maximum function, and the minimum function respectively;
[0085] Step S232: Calculate the local minimum divisible quantity; Each frame of the student's classroom learning video may have uneven brightness and contrast. By calculating the local minimum divisible quantity, the effectiveness and contribution degree of each pixel can be determined, and inefficient information can be suppressed; The product of the neighborhood pixel screening coefficient and the local minimum is used as the local minimum divisible quantity at each pixel point; The formula used is as follows:
[0086] ;
[0087] In the formula, is the local minimum divisible quantity at pixel point (x, y);
[0088] Step S233: Design the pixel adjustment factor; The inconsistent brightness and contrast of different video frames will cause the loss of some details and affect the judgment of the student's learning state. Through the pixel adjustment factor, fine brightness and contrast adjustments can be made in the local area of the image, making the key behavioral characteristics of the student more obvious; The pixel adjustment factor at each pixel point is calculated based on the local minimum divisible quantity and the local maximum; The formula used is as follows:
[0089] ;
[0090] In the formula, is the pixel adjustment factor at pixel point (x, y), is the local minimum divisible quantity at pixel point (x, y);
[0091] Step S234: Initial enhancement; The initial enhancement of the image highlights the key features and avoids the evaluation deviation caused by light or shooting angle; Each pixel point is subjected to initial enhancement processing based on the pixel adjustment factor, the local minimum divisible quantity, and the local maximum; The formula used is as follows:
[0092] ;
[0093] In the formula, is the normalized pixel value at pixel point (x, y), is the pixel value after initial enhancement;
[0094] Step S24: Secondary image enhancement; Some images may have low brightness or contrast, making it difficult to distinguish details and affecting the analysis of the learning state. Through secondary enhancement processing, not only is the contrast of the image improved, but also the details of each pixel become clearer through the calculation of the mean and standard deviation, thus helping the evaluation system to identify the subtle learning behaviors of students; For each frame of the image after primary enhancement in the learning image sequence, the following secondary enhancement processing is performed in sequence:
[0095] Step S241: Re-normalization; After primary enhancement, the pixel values of the image may fluctuate. Re-normalization can remap the pixel values of the image to a unified range, eliminating extreme value changes caused by primary enhancement, balancing the details of the image, and ensuring the stability of subsequent processing steps; Re-normalize the image after primary enhancement according to the min-max normalization method; The formula used is as follows:
[0096] ;
[0097] In the formula, is the pixel value processed according to the min-max normalization method, and are the minimum and maximum pixel values of the image after primary enhancement respectively;
[0098] Step S242: Calculate the enhancement upper limit value and the enhancement lower limit value; During the image enhancement process, it may cause over-enhancement or over-compression of some pixel values. By calculating the enhancement upper limit value and the enhancement lower limit value, it can ensure that the enhancement effect of the image is within a reasonable range, neither causing image distortion due to over-enhancement nor losing key information due to insufficient enhancement, and can more accurately reflect the subtle changes in the student's learning state; According to the mean and standard deviation of all pixel values in the image after re-normalization, calculate the enhancement upper limit value and the enhancement lower limit value at each pixel point in the image; The formula used is as follows:
[0099] ;
[0100] ;
[0101] In the formula, and are the enhancement upper limit value and the enhancement lower limit value at the pixel point (x, y) respectively, is the pixel value processed according to the min-max normalization method, β is an adjustment factor in the range of (0, 1), and μ and σ are the mean and standard deviation of all pixel values in the image after re-normalization respectively;
[0102] Step S243: Secondary enhancement; Secondary enhancement ensures that the details of the image remain in a natural and balanced state after enhancement, which helps to accurately identify the learning state of students, especially to precisely capture minute behaviors and emotional changes; perform secondary enhancement processing on each pixel point according to the enhancement upper limit value and the enhancement lower limit value; the formula used is as follows:
[0103] ;
[0104] In the formula, is the pixel value after secondary enhancement at pixel point (x, y);
[0105] Step S25: Gamma correction; Gamma correction adjusts the brightness of the image to an appropriate range, enabling the unified processing of images from different video sources and under different environmental conditions, and improving the stability and consistency of model analysis; perform gamma correction processing on each frame of the image after secondary enhancement in the learning image sequence to obtain the enhanced learning image sequence; the formula used is as follows:
[0106] ;
[0107] In the formula, is the pixel value after gamma correction at pixel point (x, y), γ is the gamma value, is the maximum pixel value of the image after secondary enhancement.
[0108] By performing the above operations, aiming at the problems in the existing student classroom learning state evaluation method, such as image quality differences, noise interference, blurring of images due to fast and complex actions and lack of details, resulting in low accuracy of evaluating students' classroom learning state, this solution performs normalization processing on each frame of the image, eliminates equipment and environmental differences, and ensures input consistency; by calculating the neighborhood pixel screening coefficient, local minimum divisible amount and pixel adjustment factor, the image is initially enhanced to remove noise and retain learning-related features, enhancing the key details of the image; by calculating the mean and standard deviation of the image to obtain the enhancement upper limit value and the enhancement lower limit value, the image is secondarily enhanced to reasonably control the image enhancement range, avoid distortion, and make the details clearer; then perform gamma correction processing to obtain the enhanced learning image sequence, ensure image brightness consistency, improve the stability of the evaluation model, significantly improve the quality of each frame of the image, and thus improve the accuracy of evaluating students' learning state.
[0109] Example 4, refer to Figure 1 and Figure 4 , based on the above example, in step S3, constructing a student classroom learning state evaluation model specifically includes the following steps:
[0110] Step S31: Feature extraction; For each frame of the image in the enhanced learning image sequence, perform the following feature extraction processing in sequence:
[0111] Step S311: Low-level feature extraction; The learning state of students in class is reflected by details such as expressions and postures. Traditional image processing methods are difficult to accurately capture these subtle changes. By extracting low-level features through a convolutional layer, the basic behavioral features of students can be effectively captured from the image. The image is input into the first convolutional layer for processing to obtain a low-level feature map. The first convolutional layer consists of a standard convolution with a size of 3×3.
[0112] Step S312: High-level feature extraction; Although low-level features can provide basic information, they cannot capture the more complex dynamic behaviors during the learning process of students. By extracting high-level features through a depthwise separable convolutional layer, more complex behavioral patterns can be captured, which reflects the comprehensive performance of students in class. The low-level feature map is input into the depthwise separable convolutional layer, the second convolutional layer, the average pooling layer, and the fully connected layer for processing to obtain a high-level feature map. The depthwise separable convolutional layer consists of a depthwise separable convolution with a size of 3×3 and an expansion factor of 1, a depthwise separable convolution with a size of 3×3 and an expansion factor of 6, a depthwise separable convolution with a size of 5×5 and an expansion factor of 6, two depthwise separable convolutions with a size of 3×3 and an expansion factor of 6, a depthwise separable convolution with a size of 5×5 and an expansion factor of 6, and a depthwise separable convolution with a size of 3×3 and an expansion factor of 6 in sequence. The second convolutional layer consists of a standard convolution with a size of 1×1.
[0113] Step S32: Low-level feature enhancement; Low-level features may not fully express the complex behavioral features related to the learning state of students, especially in dynamic scenarios. By enhancing low-level features, the sensitivity of low-level features to the assessment of students' learning state is improved. For each low-level feature map extracted from each frame image in the learning image enhancement sequence, the following enhancement processes are performed in sequence:
[0114] Step S321: Calculate global features; Individual local features are often easily affected by noise and are difficult to reflect the overall learning state of students, especially when dealing with complex classroom environments. By calculating global features through global average pooling, the global information of the image can be better captured, avoiding over-reliance on local features, so that the model can obtain a more comprehensive and stable assessment of the learning state. Global average pooling is performed on the low-level feature map to obtain the global feature Q. The formula used is as follows:
[0115] ;
[0116] In the formula, A and D are the width and height of the low-level feature map respectively, a and d are the row index and column index respectively, is the element value of the low-level feature map at the position (a, d);
[0117] Step S322: Construct a cosine similarity matrix; The features of different regions have different impacts on the learning status of students. Traditional methods tend to ignore the importance of local features, making it difficult for the model to capture subtle changes in the learning status. Through cosine similarity, the importance of each region can be quantified, which helps to focus on the regions that are most critical for evaluating the learning status of students, thereby improving accuracy and robustness; According to the cosine similarity between the element value at each position in the low-level feature map and the global feature, construct a cosine similarity matrix E of size A×D. The value of the cosine similarity matrix at position (a, d) is the cosine similarity between the element value at position (a, d) in the low-level feature map and the global feature, and the cosine similarity is quantified to obtain F quantization levels. The formula used is as follows:
[0118] ;
[0119] In the formula, S1, S f and S F are the 1st, f-th, and F-th quantization levels, f is the index, E min and E max are the minimum and maximum values of the cosine similarity respectively;
[0120] Step S323: Construct a quantization matrix; The numerical changes of features in the image are relatively complex. Directly using these values may lead to unstable or inconsistent results, affecting the accuracy of the evaluation. The process of quantifying the cosine similarity matrix transforms it into discrete numerical values, enabling the model to process more precise feature expressions and reducing the impact of differences between different images; Quantify the value at each position in the cosine similarity matrix E to obtain a quantization matrix H of size R×F, where R = A×D; The formula used is as follows:
[0121] ;
[0122] In the formula, is the value of the cosine similarity matrix E at position (a, d), r is the index, is the value of the quantization matrix H at position (r, f), and R is the number of elements in the low-level feature map;
[0123] Step S324: Construct a two-dimensional quantization counting matrix; traditional methods may not be able to effectively handle the relative importance and distribution differences between different features, resulting in information loss or misjudgment. Through the two-dimensional quantization counting matrix, the quantized features can be statistically analyzed to obtain the distribution and frequency of the features, enabling the model to better understand the distribution of the features, which helps to comprehensively consider the different behavioral performances of students in the classroom and provide a comprehensive evaluation basis; generate the two-dimensional quantization counting matrix L according to the quantization matrix H. The first dimension in the two-dimensional quantization counting matrix represents each quantization level, and the second dimension represents the corresponding normalized count; the formula used is as follows:
[0124] ;
[0125] In the formula, L f is the normalized count corresponding to the f-th quantization level, is the concatenation operation;
[0126] Step S325: Calculate statistical features; the multi-layer perceptron can effectively combine the two-dimensional quantization counting matrix with the global features to generate more abstract and comprehensive statistical features; process the two-dimensional quantization counting matrix L through the multi-layer perceptron and then concatenate it with the global feature Q to obtain the statistical feature Z; the formula used is as follows:
[0127] ;
[0128] In the formula, is the multi-layer perceptron;
[0129] Step S326: Calculate the enhanced low-level features; the enhanced low-level features can more accurately reflect the learning state of the students; calculate the enhanced low-level features according to the statistical feature Z and the quantization matrix H; the formula used is as follows:
[0130] ;
[0131] In the formula, V is the enhanced low-level feature, T is the matrix transpose operation, is a standard convolution of size 1×1, is the Softmax activation function;
[0132] Step S33: Advanced feature enhancement; Sometimes, advanced features may lack sensitivity to certain specific details. Especially when students' behaviors are complex or subtle, traditional methods are difficult to accurately capture. By introducing an attention mechanism, it can automatically focus on the feature regions that are most valuable for the evaluation results, avoiding interference from irrelevant information. Especially in a classroom environment, the subtle changes of students can be effectively identified through the attention mechanism; For the advanced feature maps O extracted from each frame of the learning image enhancement sequence, enhancement processing is performed in sequence; The enhancement processing is to first calculate the attention weights of the advanced feature maps ; Then, the advanced feature maps are weighted and summed through the attention weights to obtain the enhanced advanced features ; where d k is the feature dimension;
[0133] Step S34: Construct a comprehensive feature sequence; Using only low-level features or high-level features alone may not be able to comprehensively reflect the learning state of students. Especially in terms of the complexity and diversity of classroom behaviors, combining low-level features and high-level features can more comprehensively express the learning state of students, integrating different levels of information from details to the whole, enabling the model to more accurately evaluate the learning state of students; The enhanced low-level features V and the enhanced high-level features P are merged through a concatenation operation to obtain the comprehensive features of each frame of the image , and the comprehensive features of all frame images are integrated in the order in the learning image enhancement sequence to obtain a comprehensive feature sequence;
[0134] Step S35: Evaluation; The comprehensive feature sequence is input into a long short-term memory network for evaluation to obtain the status score of the student's classroom learning.
[0135] By performing the above operations, for the problem that the existing methods for evaluating the classroom learning state of students have difficulty accurately capturing the subtle dynamic changes and complex dynamic behaviors in the learning process of students, ignoring the detail changes in the images, resulting in insufficient accuracy in the evaluation of the classroom learning state of students, this solution extracts low-level features through standard convolution and extracts high-level features using depthwise separable convolution, enhancing the comprehensive reflection of the learning state of students; By calculating the global features to construct a cosine similarity matrix, and then calculating the quantization matrix and the two-dimensional quantization count matrix to obtain statistical features, the low-level features are enhanced, which can not only capture the basic behavior features but also reflect the subtle changes in the learning state; By enhancing the high-level features through attention weights, the sensitivity to detail changes can be enhanced, improving the recognition accuracy of the student behavior state; Constructing a comprehensive feature sequence and obtaining the status score through a long short-term memory network can comprehensively consider the behavior performance of students in the classroom environment, comprehensively reflect the learning state of students, provide a more refined classroom learning status score, and enhance the accuracy and stability of the evaluation.
[0136] Example 5, refer to Figure 2 , based on the above embodiments, the artificial intelligence-based student classroom learning status evaluation system provided by the present invention includes a student classroom learning data collection module, a student classroom learning image enhancement module, a module for constructing a student classroom learning status evaluation model, and a module for generating an evaluation report;
[0137] The student classroom learning data collection module collects historical student classroom learning data and sends the data to the student classroom learning image enhancement module;
[0138] The student classroom learning image enhancement module performs normalization processing on each frame of the image. By calculating the neighborhood pixel screening coefficient, the local minimum divisible amount, and the pixel calibration factor, the image is initially enhanced. By calculating the mean and standard deviation of the image, the enhancement upper limit value and the enhancement lower limit value are obtained, and the image is secondarily enhanced. Then, gamma correction processing is performed to obtain a learning image enhancement sequence, and the data is sent to the module for constructing a student classroom learning status evaluation model;
[0139] The module for constructing a student classroom learning status evaluation model uses standard convolution to extract low-level features, uses depthwise separable convolution to extract high-level features, constructs a cosine similarity matrix through global features, calculates a quantization matrix and a two-dimensional quantization count matrix to obtain statistical features, enhances the low-level features, enhances the high-level features through attention weights, constructs a comprehensive feature sequence, obtains a status score, and sends the data to the module for generating an evaluation report;
[0140] The module for generating an evaluation report collects real-time student classroom learning videos and performs student classroom learning image enhancement processing, and then inputs them into the student classroom learning status evaluation model for evaluation, and generates an evaluation report on student classroom learning according to the output status score.
[0141] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0142] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention.
[0143] The present invention and its implementation manners have been described above. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design, without creative efforts, structural manners and embodiments similar to the technical solution without departing from the gist of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. An artificial intelligence-based method for evaluating the learning status of students in the classroom, characterized in that: The method includes the following steps: Step S1: Acquisition of students' in-class learning data; acquisition of historical students' in-class learning data; Step S2: Enhancement of students' in-class learning images; Step S3: Construction of an evaluation model for students' in-class learning status; Step S4: Generation of an evaluation report; acquisition of real-time students' in-class learning videos and performing enhancement processing on students' in-class learning images, then inputting them into the evaluation model for students' in-class learning status for evaluation, and generating an evaluation report for students' in-class learning according to the output status scores; In step S2, it includes step S231: Design of neighborhood pixel screening coefficients; calculating the neighborhood pixel screening coefficients at each pixel point in the image according to the local maximum value, local minimum value, and pixel value variance within the neighborhood of each pixel point; the formula used is as follows: ; In the formula, is the neighborhood pixel screening coefficient at the pixel point (x, y), and are respectively the local maximum value and the local minimum value within the 5×5 neighborhood centered on the pixel point (x, y), is the variance of all pixel values within the 5×5 neighborhood centered on the pixel point (x, y), x is the abscissa index of the pixel point in the image, and y is the ordinate index of the pixel point in the image, is the exponential function; In step S2, the enhancement of students' in-class learning images specifically includes the following steps: Step S21: Construction of a learning image sequence; extracting consecutive frames from students' in-class learning videos to generate a learning image sequence; Step S22: Normalization; performing normalization processing on each frame image in the learning image sequence; Step S23: Initial image enhancement; Step S24: Secondary image enhancement; for each frame image after initial enhancement in the learning image sequence, perform the following secondary enhancement processing in sequence: Step S241: Re-normalization; performing re-normalization processing on the image after initial enhancement according to the min-max normalization method; Step S242: Calculation of enhancement upper limit value and enhancement lower limit value; calculating the enhancement upper limit value and enhancement lower limit value at each pixel point in the image according to the mean and standard deviation of all pixel values in the image after re-normalization; the formula used is as follows: ; ; Wherein, and are respectively the upper enhancement limit value and the lower enhancement limit value at the pixel point (x, y), is the pixel value processed according to the min-max normalization method, is the pixel value after primary enhancement, is the normalized pixel value at the pixel point (x, y), β is the adjustment factor, and μ and σ are respectively the mean and standard deviation of all pixel values in the image after re-normalization; Step S243: Secondary enhancement; performing secondary enhancement processing on each pixel point according to the enhancement upper limit value and enhancement lower limit value; the formula used is as follows: ; wherein, is the pixel value after secondary enhancement at the pixel point (x, y); Step S25: Gamma correction; performing gamma correction processing on each frame image after secondary enhancement in the learning image sequence to obtain a learning image enhancement sequence; In step S23, the initial image enhancement; for each frame image after normalization in the learning image sequence, perform the following initial enhancement processing in sequence: Step S231: Design of neighborhood pixel screening coefficients; Step S232: Calculation of local minimum divisible amount; taking the product of the neighborhood pixel screening coefficient and the local minimum value as the local minimum divisible amount at each pixel point; Step S233: Design of pixel adjustment factor; calculating the pixel adjustment factor at each pixel point according to the local minimum divisible amount and the local maximum value; the formula used is as follows: ; In the formula, is the pixel calibration factor at the pixel point (x, y), is the local minimum divisible amount at the pixel point (x, y); Step S234: Initial enhancement; performing initial enhancement processing on each pixel point according to the pixel adjustment factor, local minimum divisible amount, and local maximum value; the formula used is as follows: ; In the formula, is the normalized pixel value at the pixel point (x, y), is the pixel value after primary enhancement; In step S3, the construction of the evaluation model for students' in-class learning status specifically includes the following steps: Step S31: Feature extraction; for each frame image in the learning image enhancement sequence, perform the following feature extraction processing in sequence: Step S311: Low-level feature extraction; inputting the image into the first convolutional layer for processing to extract a low-level feature map; Step S312: Advanced feature extraction; input the low-level feature map into a depthwise separable convolutional layer, a second convolutional layer, an average pooling layer, and a fully connected layer for processing to extract the advanced feature map; Step S32: Low-level feature enhancement; Step S33: Advanced feature enhancement; for the advanced feature map O extracted from each frame of the learning image enhancement sequence, perform enhancement processing in sequence; the enhancement processing is to first calculate the attention weight W of the advanced feature map; then perform weighted summation on the advanced feature map through the attention weight to obtain the enhanced feature P; Step S34: Construct a comprehensive feature sequence; merge the enhanced low-level feature V and the enhanced advanced feature P through a concatenation operation to obtain the comprehensive feature M of each frame of the image, and integrate the comprehensive features of all frames of the image in the order in the learning image enhancement sequence to obtain the comprehensive feature sequence; Step S35: Evaluation; input the comprehensive feature sequence into a long short-term memory network for evaluation to obtain the state score of the student's classroom learning.
2. The method for evaluating the learning state of students in the classroom based on artificial intelligence according to claim 1, wherein: In step S32, the low-level feature enhancement; For the low-level feature map extracted from each frame of the learning image enhancement sequence, perform the following enhancement processing in sequence: Step S321: Calculate the global feature; Perform global average pooling on the low-level feature map to obtain the global feature Q; Step S322: Construct a cosine similarity matrix; construct a cosine similarity matrix E based on the cosine similarity between the element values at each position in the low-level feature map and the global feature, and quantize the cosine similarity to obtain F quantization levels ; where S1, S f and S F are the 1st, the f-th, and the F-th quantization levels, and f is the index; Step S323: Construct a quantization matrix; perform quantization processing on the value at each position in the cosine similarity matrix E to obtain a quantization matrix H with a size of R×F; the formula used is as follows: ; In the formula, is the value of the cosine similarity matrix E at the position (a, d), r is the index, is the value of the quantization matrix H at the position (r, f), and R is the number of elements of the low-level feature map; Step S324: Construct a two-dimensional quantization count matrix; generate a two-dimensional quantization count matrix L according to the quantization matrix H. The first dimension in the two-dimensional quantization count matrix represents each quantization level, and the second dimension represents the corresponding normalized count; the formula used is as follows: ; where L f is the normalized count corresponding to the f-th quantization level, is a concatenation operation; Step S325: Calculate the statistical feature; connect the two-dimensional quantization count matrix L processed by a multi-layer perceptron with the global feature Q to obtain the statistical feature Z; Step S326: Calculate the enhanced low-level feature; calculate the enhanced low-level feature according to the statistical feature Z and the quantization matrix H; the formula used is as follows: ; where V is the enhanced low-level feature, T is the matrix transpose operation, is a standard convolution of size 1×1, is the Softmax activation function.
3. The method for evaluating the learning status of students in class based on artificial intelligence according to claim 1, wherein: In step S4, the student classroom learning data collection is to collect historical student classroom learning data, and the historical student classroom learning data includes student classroom learning videos and state scores, and use the state score as the data label.
4. An artificial intelligence-based student classroom learning status evaluation system for implementing the artificial intelligence-based student classroom learning status evaluation method according to any one of claims 1-3, characterized in that: It includes a student classroom learning data collection module, a student classroom learning image enhancement module, a module for constructing a student classroom learning state evaluation model, and a module for generating an evaluation report; The student classroom learning data collection module collects historical student classroom learning data and sends the data to the student classroom learning image enhancement module; The student classroom learning image enhancement module performs normalization processing on each frame of the image, performs primary enhancement on the image by calculating the neighborhood pixel screening coefficient, the local minimum divisible amount, and the pixel calibration factor, obtains the enhancement upper limit value and the enhancement lower limit value by calculating the mean and standard deviation of the image, performs secondary enhancement on the image, and then performs gamma correction processing to obtain the learning image enhancement sequence and sends the data to the module for constructing a student classroom learning state evaluation model; The module for constructing the student classroom learning status evaluation model extracts low-level features using standard convolution, extracts high-level features using depthwise separable convolution, constructs a cosine similarity matrix through global features, calculates a quantization matrix and a two-dimensional quantization count matrix to obtain statistical features, enhances the low-level features, enhances the high-level features through attention weights, constructs a comprehensive feature sequence, obtains a status score, and sends the data to the module for generating an evaluation report; The module for generating an evaluation report collects real-time student classroom learning videos and performs image enhancement processing on student classroom learning images, then inputs them into the student classroom learning status evaluation model for evaluation, and generates an evaluation report on student classroom learning based on the output status score.
Citation Information
Patent Citations
Classroom teaching cognitive load measurement system
US20200098284A1
Multimodal data-based method and system for recognizing cognitive engagement in classroom
US20250022314A1