Online teaching management method and system based on AI
By enhancing the image dataset and extracting features, and using the improved YOLOv7 model and conditional generative adversarial network, the problem of feature extraction failure caused by low camera image resolution was solved, and real-time and accurate recognition of user learning status was achieved, thereby improving the recognition success rate of online teaching management.
Patent Information
- Application Number
- CN202510938755.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-07-08
AI Technical Summary
The image resolution captured by existing cameras is low, resulting in a high failure rate in facial feature extraction, making it impossible to accurately identify the user's learning status, and thus unable to effectively manage online teaching.
By obtaining the target user's image dataset, the image is enhanced, three-dimensional structural features and local features are extracted, and feature fusion is performed using the improved YOLOv7 model and conditional generative adversarial network, which are then input into the state recognition model for early warning.
It realizes real-time and accurate identification of user learning status, improves the recognition success rate, can dynamically adjust the collection frequency to ensure data richness and system resource occupancy, and enhances the effectiveness of online teaching management.
Smart Images

Figure CN120599684A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of teaching management technology, and specifically relates to an AI-based online teaching management method and system. Background Art
[0002] With the rapid advancement of Internet technology, new technologies such as 5G and artificial intelligence continue to emerge, social knowledge updates are accelerating, and the concept of lifelong learning is becoming popular. People's growing demand for convenience and personalization in education has laid a solid technical foundation for the development of online teaching technology.
[0003] Patent number: CN115936944A, discloses a virtual teaching management method and device based on artificial intelligence. It rationally utilizes the characteristics of different dimensions of both teachers and students in the classroom to integrate and achieve more accurate emotion recognition results than using a single feature. At the same time, based on the warning level corresponding to the classroom atmosphere evaluation value, corresponding measures are taken for teaching management. This can not only effectively promote student learning and help teachers reflect on their teaching, but also continuously optimize teaching plans by analyzing the characteristics of the video itself. It also provides good basic data support for improving teaching quality and determining the direction of education reform.
[0004] The image resolution captured by existing cameras is low, resulting in a high failure rate in facial feature extraction, making it impossible to accurately identify the user's learning status, and thus unable to effectively manage online teaching. Summary of the Invention
[0005] The purpose of the present invention is to solve the problem that the image resolution captured by existing cameras is low, resulting in a high failure rate in facial feature extraction, inability to accurately identify the user's learning status, and thus inability to effectively manage online teaching, and to propose an AI-based online teaching management method and system.
[0006] In a first aspect of the present invention, an AI-based online teaching management method is first proposed, the method comprising:
[0007] Acquire an image dataset of a target user, and enhance images in the image dataset to obtain a target image set; the target image set includes a plurality of target images;
[0008] Performing three-dimensional construction on the target images in the target image set to obtain structural features, and extracting features from the target images in the target image set to obtain a feature group; the feature group includes: global features and local features;
[0009] The feature group and the structural feature are input into a state recognition model to obtain a state score, and an early warning is issued to the target user according to the state score.
[0010] Optionally, enhancing the images in the image dataset to obtain a target image set includes:
[0011] Performing pixel recognition on each image in the image data set to obtain a blurred pixel atlas, and deblurring each blurred pixel image in the blurred pixel atlas by fusing a CNN to obtain a deblurred image set;
[0012] Determine the noise pixels of the target deblurred image, calculate an image matrix based on the noise pixels, and input the image matrix and the target deblurred image into a neural fuzzy system to obtain a noise-enhanced image; the noise-enhanced image is an image containing the positions of the noise pixels; and the target deblurred image is any one of the blurred images.
[0013] The noise-enhanced image is optimized through a conditional generative adversarial network to obtain a detail-enhanced image, the detail-enhanced image and the noise-enhanced image are fused to obtain a target image, and the target images corresponding to the deblurred images in the deblurred image set are combined into a target image set.
[0014] Optionally, before deblurring each blurred pixel image in the blurred pixel image set by fusing CNN to obtain a deblurred image set, the parameter optimization method of the fused CNN includes:
[0015] Step 1: Initialize the population and parameters to obtain an initial solution set, calculate the fitness of each initial solution in the initial solution set, and take the initial solution with the largest fitness value as the current optimal solution;
[0016] Step 2: Generate candidate optimal solutions by randomly perturbing the current optimal solution, calculate the Euclidean distance from each initial solution to the candidate optimal solution, sort the initial solutions in descending order according to the Euclidean distance, and filter the initial solutions according to the preset rules to obtain the tracking set and escape set;
[0017] Step 3: Update the initial solution in the escape set using the random offset escape formula to obtain a suboptimal solution, calculate the fitness of each suboptimal solution, and compare the fitness of the initial solution with the corresponding suboptimal solution. If the fitness of the initial solution is greater than the fitness of the corresponding suboptimal solution, replace the initial solution with the corresponding suboptimal solution; otherwise, retain the initial solution.
[0018] Step 4: Determine the optimal solution in the current population based on the fitness value, and repeat steps 2 to 3 until the maximum number of iterations is reached, or the fitness converges, then output the optimal solution.
[0019] Optionally, performing feature extraction on the target image in the target image set to obtain a feature group includes:
[0020] Extracting features of the target image in the target image set using a target YOLOv7 model to obtain a feature group;
[0021] The improvements to the target YOLOv7 include:
[0022] The CBS module in the backbone network is replaced with a separation Conv module, the CBL module in the YOLOv7-tiny model is replaced with a target CBL module, and the MDA module is added to the neck structure to obtain the target YOLOv7; the backbone network and the neck structure are components of the YOLOv7-tiny model;
[0023] The working principle of the separation Conv module includes:
[0024] Obtain an initial feature tensor, input the initial feature tensor into a batch normalization layer and a SiLU activation function in sequence to obtain a first feature tensor, input the initial feature tensor into a depthwise separable convolutional layer to obtain a second feature tensor, multiply the initial feature tensor, the first feature tensor, and the second feature tensor to obtain a third feature tensor, and use the third feature tensor as the output of the separation Conv module;
[0025] The working principle of the target CBL module includes:
[0026] Obtain an initial feature map, decompose the initial feature map into multiple blocks according to a preset size, splice the blocks according to the number of channels to obtain an intermediate feature map, input the intermediate feature map into a 1×1 Conv module to obtain an output feature map, and use the output feature map as the output of the target CBL module.
[0027] Optionally, feature extraction is performed on the target images in the target image set to obtain a feature group. The working principle of the MDA module includes:
[0028] Obtain a target feature map, input the target feature map into a 1×1 Conv module and a fully connected layer in sequence to obtain a Query, a Key, and a Value, input the Query and the Key into a separation Conv module to obtain a target Query and a target Key, perform a matrix multiplication operation on the target Query and the target Key to obtain a similarity matrix, and input the similarity matrix into a Softmax function to obtain an attention weight matrix;
[0029] The Value is input into the fully connected layer and the separation Conv module in sequence to obtain the target Value, and the attention weight matrix is multiplied by the target Value to obtain a fused feature map.
[0030] In a second aspect of the present invention, an AI-based online teaching management system is proposed, comprising: an image enhancement module, a feature extraction module, and a teaching management module:
[0031] The image enhancement module is used to obtain an image dataset of a target user and enhance images in the image dataset to obtain a target image set; the target image set includes multiple target images;
[0032] The feature extraction module is used to perform three-dimensional construction based on the target image in the target image set to obtain structural features, and perform feature extraction on the target image in the target image set to obtain a feature group; the feature group includes: global features and local features;
[0033] The teaching management module is used to input the feature group and the structural feature into a state recognition model to obtain a state score, and to issue an early warning to the target user according to the state score.
[0034] Optionally, the image enhancement module further includes: a deblurring module, a noise enhancement module and a feature fusion module:
[0035] The deblurring module is configured to perform pixel recognition on each image in the image data set to obtain a blurred pixel atlas, and deblur each blurred pixel image in the blurred pixel atlas by fusing a CNN to obtain a deblurred image set;
[0036] The noise enhancement module is configured to determine the noise pixels of the target deblurred image, calculate an image matrix based on the noise pixels, and input the image matrix and the target deblurred image into the neural fuzzy system to obtain a noise enhanced image; the noise enhanced image is an image containing the positions of the noise pixels; the target deblurred image is any one of the blurred images;
[0037] The feature fusion module is used to optimize the noise-enhanced image through a conditional generative adversarial network to obtain a detail-enhanced image, fuse the detail-enhanced image and the noise-enhanced image to obtain a target image, and form a target image set with the target images corresponding to each deblurred image in the deblurred image set.
[0038] Optionally, the system further includes:
[0039] The first operation module is used to initialize the population and parameters to obtain an initial solution set, calculate the fitness of each initial solution in the initial solution set, and take the initial solution with the largest fitness value as the current optimal solution;
[0040] The second operation module is configured to generate a candidate optimal solution by randomly perturbing the current optimal solution, calculate the Euclidean distance from each initial solution to the candidate optimal solution, sort the initial solutions in descending order according to the Euclidean distance, and filter the initial solutions according to a preset rule to obtain a tracking set and an escape set;
[0041] The third operation module is configured to update the initial solution in the escape set using a random offset escape formula to obtain a suboptimal solution, calculate the fitness of each suboptimal solution, compare the fitness of the initial solution with the fitness of the corresponding suboptimal solution, and if the fitness of the initial solution is greater than the fitness of the corresponding suboptimal solution, replace the initial solution with the corresponding suboptimal solution; otherwise, retain the initial solution;
[0042] The fourth operation module is used to determine the optimal solution in the current population according to the fitness value, repeatedly execute the second operation module and the third operation module until the maximum number of iterations is reached or the fitness converges, and then output the optimal solution.
[0043] Optionally, it is further used to extract features of the target image in the target image set using a target YOLOv7 model to obtain a feature group;
[0044] The improvements to the target YOLOv7 include:
[0045] The CBS module in the backbone network is replaced with a separation Conv module, the CBL module in the YOLOv7-tiny model is replaced with a target CBL module, and the MDA module is added to the neck structure to obtain the target YOLOv7; the backbone network and the neck structure are components of the YOLOv7-tiny model;
[0046] The working principle of the separation Conv module includes:
[0047] Obtain an initial feature tensor, input the initial feature tensor into a batch normalization layer and a SiLU activation function in sequence to obtain a first feature tensor, input the initial feature tensor into a depthwise separable convolutional layer to obtain a second feature tensor, multiply the initial feature tensor, the first feature tensor, and the second feature tensor to obtain a third feature tensor, and use the third feature tensor as the output of the separation Conv module;
[0048] The working principle of the target CBL module includes:
[0049] Obtain an initial feature map, decompose the initial feature map into multiple blocks according to a preset size, splice the blocks according to the number of channels to obtain an intermediate feature map, input the intermediate feature map into a 1×1 Conv module to obtain an output feature map, and use the output feature map as the output of the target CBL module.
[0050] Optionally, the working principle of the MDA module includes:
[0051] Obtain a target feature map, input the target feature map into a 1×1 Conv module and a fully connected layer in sequence to obtain a Query, a Key, and a Value, input the Query and the Key into a separation Conv module to obtain a target Query and a target Key, perform a matrix multiplication operation on the target Query and the target Key to obtain a similarity matrix, and input the similarity matrix into a Softmax function to obtain an attention weight matrix;
[0052] The Value is input into the fully connected layer and the separation Conv module in sequence to obtain the target Value, and the attention weight matrix is multiplied by the target Value to obtain a fused feature map.
[0053] Beneficial effects of the present invention:
[0054] This paper proposes an AI-based online teaching management method. The method involves acquiring an image dataset of a target user, enhancing the images in the dataset to obtain a target image set containing multiple target images, performing a three-dimensional construction based on the target images in the target image set to obtain structural features, and extracting features from the target images in the target image set to obtain a feature group. The feature group includes global features and local features. The feature group and structural features are input into a state recognition model to obtain a state score, and the target user is warned based on the state score. By acquiring and enhancing images, structural features and global and local features of facial motion can be accurately extracted. These features are then input into the state recognition model, enabling real-time and accurate identification of the user's learning state, improving the recognition success rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The present invention will be further described below with reference to the accompanying drawings.
[0056] Figure 1 A flowchart of an AI-based online teaching management method provided in an embodiment of the present invention;
[0057] Figure 2 A schematic diagram of the structure of an AI-based online teaching management method provided by an embodiment of the present invention;
[0058] Figure 3 A framework diagram of an AI-based online teaching management system provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0059] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments represent only a portion of the embodiments of the present invention, not all of them. The term "and / or" herein simply describes an association relationship between associated objects, indicating that three possible relationships exist. For example, "A" and "B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, references to "first," "second," and so on in the present invention are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include at least one of these features. Furthermore, the technical solutions of the various embodiments may be combined, but only if they are achievable by a person of ordinary skill in the art. If a combination of technical solutions contradicts or is unachievable, such combination shall be deemed non-existent and outside the scope of protection claimed by the present invention.
[0060] Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of the present invention.
[0061] The embodiment of the present invention provides an online teaching management method based on AI. Figure 1 , Figure 1 A flowchart of an AI-based online teaching management method provided in an embodiment of the present invention. The method includes the following steps:
[0062] S101, obtaining an image dataset of a target user, and enhancing images in the image dataset to obtain a target image set;
[0063] S102, performing three-dimensional construction on the target image in the target image set to obtain structural features, and performing feature extraction on the target image in the target image set to obtain a feature group;
[0064] S103: Input the feature group and the structural feature into the state recognition model to obtain a state score, and issue an early warning to the target user according to the state score.
[0065] The target image set includes multiple target images; the feature group includes: global features and local features;
[0066] An AI-based online teaching management method provided by an embodiment of the present invention can accurately extract the structural features and global and local features of facial movement by collecting and enhancing images, and input them into a state recognition model. It can then accurately identify the user's learning status in real time, thereby improving the recognition success rate.
[0067] In one implementation, when a user uses an electronic device (such as a tablet, mobile phone, or laptop computer) for online learning, the user's status is obtained through the camera of the electronic device. After the obtained data is processed locally, the user's learning status is evaluated (state recognition model) to obtain a status score, and the user's status score is uploaded to the cloud. The cloud issues an early warning based on the status score, such as voice prompts, text warnings, etc.; the judgment conditions for issuing early warnings to target users based on the status score are that the image data set has multiple images, and therefore also has multiple feature groups and structural features. The multiple feature groups and structural features are all input into the state recognition model to obtain multiple state scores. When a preset number of state scores are greater than a preset threshold, an early warning is issued.
[0068] In one implementation, the acquisition frequency for capturing images of the target user can be set to 10 to 30 frames per second. Dynamically setting the frequency range ensures both data richness and efficient system processing resources. If the acquisition frequency is too low, the data will lack real-time performance, while if it is too high, the data processing burden will increase.
[0069] In one implementation, structural features refer to lines used to characterize the motion state of a face. These lines include, but are not limited to, the contour lines of the eyebrows, the curvature of the mouth corners, and the contour lines of the eyes. These lines change dynamically during facial movement and can intuitively reflect the motion state of the face. For example, when a person smiles, the curvature of the mouth corners changes from a straight line to an upward curved arc; when frowning, the contour lines of the eyebrows become tighter and closer to the center. By extracting and analyzing these lines, the three-dimensional structural features of the face under different expressions and movements can be constructed, providing important geometric morphological basis for subsequent state recognition and helping the model to more accurately understand the motion patterns and state changes of the face.
[0070] In one implementation method, structural features can be extracted through a deep learning model to extract two-dimensional / three-dimensional key points, construct contour lines of areas such as eyebrows, eyes, and corners of the mouth, perform time series analysis on the key point coordinates of consecutive frames, and calculate the motion parameters of the lines (displacement, speed, and curvature change).
[0071] In one implementation, the feature group covers global and local features of the face. Global features focus on macro-level information such as the overall shape, size, and overall facial expression tendency of the face. For example, the overall outline shape of the face, the width-to-narrow ratio of the face, and whether the overall expression tends to be happy, sad, angry, or other emotional tendencies. These global features can provide a general direction and background information for state recognition. Local features focus on specific areas of the face, such as detailed features of the eye area (such as changes in pupil size and the degree of eyelid opening and closing), features of the nose area (such as the degree of expansion of the nostrils), and features of the mouth area (such as changes in lip thickness and the degree of tooth exposure). These local features can capture more subtle changes in expression and movement details. Combined with global features, they can more comprehensively and accurately describe the state of the face, providing rich feature information for the state recognition model, thereby improving the accuracy and reliability of state recognition.
[0072] In one implementation, the state recognition model can be: a convolutional neural network (CNN), a long short-term memory network (LSTM), etc.; the model training data includes: a manually annotated rating table (experts classify the user status in the image / video, such as "focused", "mildly distracted", "severely tired" and obtain a score based on a percentage), historical image data and historical structure data.
[0073] In one embodiment, enhancing images in an image dataset to obtain a target image set includes:
[0074] Pixel recognition is performed on each image in the image dataset to obtain a blurred pixel atlas, and each blurred pixel image in the blurred pixel atlas is deblurred by fusing CNN to obtain a deblurred image set;
[0075] Determine the noise pixels of the target deblurred image, calculate an image matrix based on the noise pixels, and input the image matrix and the target deblurred image into the neural fuzzy system to obtain a noise-enhanced image; the noise-enhanced image is an image containing the positions of the noise pixels; the target deblurred image is any one of the blurred images;
[0076] The noise-enhanced image is optimized through a conditional generative adversarial network to obtain a detail-enhanced image, the detail-enhanced image and the noise-enhanced image are fused to obtain a target image, and the target images corresponding to each deblurred image in the deblurred image set form a target image set.
[0077] In one implementation, a deep residual network (DRN) processes the surrounding area of an image to detect complex noise pixels and distinguish blurred from non-blurred pixels in the image. Blurred pixels do not refer to individual pixels being "blurry" per se, but rather to pixel values that deviate from ideal imaging, resulting in a smooth grayscale transition with adjacent pixels and a loss of edge information. In ideal imaging, sharp pixels should exhibit abrupt grayscale changes at their edges (e.g., from black to white), while blurred pixels exhibit smooth grayscale changes (e.g., gradual transitions), quantified by the pixel's gradient value or the grayscale variance of its neighborhood.
[0078] In one implementation, a CNN is fused to deblur each blurred pixel image in a set of blurred pixel images to obtain a set of deblurred images. The blurred images are then fed into the CNN encoder, where multi-layer convolution extracts multi-scale features (ranging from edges and textures to global structure). Skip connections or cross-layer fusion are used to preserve shallow details. The decoder restores resolution through deconvolution upsampling and incorporates attention mechanisms (such as channel / spatial attention) to focus on key areas and suppress redundant information introduced by blur. The network learns the nonlinear mapping from blurred to clear images, known as the deconvolution process, by minimizing the pixel-wise loss (MSE), the perceptual loss (VGG feature similarity), and the adversarial loss (GAN discrimination between real and clear images). By training on a large number of blurred-clear image pairs, the CNN implicitly models the point spread function (PSF) and noise distribution, directly predicting high-frequency details (such as edge sharpness and texture contrast) of the clear image.
[0079] In one implementation, the neuro-fuzzy system combines the explicit knowledge representation of the fuzzy inference system with the learning ability of the artificial neural network. The image matrix and the target deblurred image are used as inputs of the neuro-fuzzy system, and the nearest neighbors of the noise pixels are used to generate a new image matrix.
[0080] Image matrix calculation formula:
[0081]
[0082] Among them, X p Represents the original image matrix, X pAB Represents the generated new image matrix, AB represents the operation for all blurred pixels in the image, q and v represent the coordinates of the target pixels, r and n represent the sum index, represents the sum of neighborhood pixels, and 1 / 9 represents the mean calculation;
[0083] X p Represents the original image matrix, which represents the input noisy image, where each element X p (q,v) represents the pixel value at the coordinate (q,v); X pABRepresents the generated new image matrix, which is obtained by weighted averaging the nearest neighbor pixels of the noise pixel in the original image and is used for subsequent neural fuzzy processing; AB is the subscript suffix, indicating that the operation is for all blurred pixels in the image; (q, v) represents the coordinates of the target pixel, respectively representing the new matrix X p The row and column indices of the pixel to be calculated; (r,n) represents the sum index, which is used to traverse the neighborhood range of the target pixel in the original image, and is the 3×3 neighborhood of the target pixel (i.e., r,n∈{-1,0,1}); represents the sum of neighborhood pixels, which is the sum of the grayscale values of the target pixel (q, v) and its eight neighboring pixels (up, down, left, right, and diagonal); 1 / 9 is the mean calculation, which is the sum of the grayscale values of the nine neighboring pixels and the average of them as the pixel value at (q, v) in the new matrix to achieve noise smoothing.
[0084] In one implementation, the conditional generative adversarial network consists of a generator and a discriminator. The most important network is used to enhance contrast. When the generator is supported by the discriminator, the best enhancement results are achieved. During training, the generator is trained to fool the discriminator, making it unable to distinguish between the enhanced output and the labeled image.
[0085] In one implementation, the generative subnetwork: Image enhancement requires a model that can produce improved results in various applications and network performance. The cascaded features are restored to the original resolution through the deconvolution layer using the Tanh activation function. The generative network is as follows: GU-GSU-GSU-ρ-Tanh; where G represents the convolution layer, U represents the ReLU, s represents the batch normalization layer, ρ represents the deconvolution layer, and Tanh represents the hyperbolic tangent activation function;
[0086] G: Convolutional Layer, which extracts local image features (such as edges and textures) through convolution kernels of different sizes; U: ReLU activation function (RectifiedLinearUnit), which introduces nonlinear mapping capabilities to the network and enhances feature expression (the output is max(0,x)); S: Batch Normalization layer, which standardizes each batch of input data, stabilizes the training process and accelerates convergence; ρ: Deconvolution Layer, which achieves upsampling through transposed convolution operations and restores image resolution (mapping low-dimensional features to high-dimensional space); Tanh: Hyperbolic tangent activation function, which compresses the output value to the [-1,1] interval to adapt to the grayscale range of image pixels.
[0087] In one implementation, the discriminant subnetwork: The goal of the discriminant network is to identify label images from the enhanced results that help the generative subnetwork produce promising results. The structure of the discriminant network is as follows: Among them, G represents the convolutional layer, represents the LeakyReLU activation function, S represents the batch normalization layer, and ω represents the output layer;
[0088] G: Convolutional layer, similar to the generative sub-network, but more focused on extracting discriminative features of the image (such as real / forged texture differences); LeakyReLU activation function is a variant of ReLU that allows negative inputs to have a small slope (such as 0.2) to solve the problem of neuron "death"; S: Batch Normalization layer, as above, stabilizes the feature distribution of the discriminator; The output layer, usually a Sigmoid activation function, outputs a probability value between 0 and 1 (1 for "real image" and 0 for "generated image").
[0089] In one implementation, the weights and parameters of the conditional generative adversarial network are optimized using the FJBA algorithm, with the mean square error As the fitness function (where MSE fit represents the fitness function, represents the expected output, ξj represents the classification output of the conditional generative adversarial network, and P represents the sum of the data). Iterative optimization is performed until the termination condition is met to improve the image enhancement capability of the generative adversarial network.
[0090] FJBA: Fractional Jaya Bat Algorithm (Fractional Jaya Bat Algorithm) is a hybrid optimization algorithm that is a fusion of two basic algorithms: JBA (Jaya Bat Algorithm): a hybrid of Jaya algorithm (JOA) and Bat algorithm (BA); FC (Fractional Calculus): fractional calculus theory, used to enhance the algorithm's global search capability and convergence accuracy;
[0091] By optimizing the weights of the generative adversarial network for image enhancement conditions and enhancing the global optimization capability of the algorithm through non-integer-order differential operators, the model convergence is accelerated and the quality of image detail restoration is improved. Ultimately, the PSNR, SSIM and other indicators of image restoration are significantly improved, effectively solving the efficiency and accuracy bottlenecks of traditional algorithms in blurred pixel recognition and image enhancement.
[0092] In one embodiment, before deblurring each blurred pixel image in the blurred pixel image set by fusing a CNN to obtain a deblurred image set, a parameter optimization method for fusing the CNN includes:
[0093] Step 1: Initialize the population and parameters to obtain the initial solution set, calculate the fitness of each initial solution in the initial solution set, and take the initial solution with the largest fitness value as the current optimal solution;
[0094] Step 2: Generate candidate optimal solutions by randomly perturbing the current optimal solution, calculate the Euclidean distance from each initial solution to the candidate optimal solution, sort the initial solutions in descending order according to the Euclidean distance, and filter the initial solutions according to the preset rules to obtain the tracking set and escape set;
[0095] Step 3: Update the initial solution in the escape set using the random offset escape formula to obtain the suboptimal solution. Calculate the fitness of each suboptimal solution and compare the fitness of the initial solution with the corresponding suboptimal solution. If the fitness of the initial solution is greater than the fitness of the corresponding suboptimal solution, replace the initial solution with the corresponding suboptimal solution. Otherwise, retain the initial solution.
[0096] Step 4: Determine the optimal solution in the current population based on the fitness value, and repeat steps 2 to 3 until the maximum number of iterations is reached, or the fitness converges, then output the optimal solution.
[0097] In one implementation, CNN parameter optimization is achieved through a metaheuristic algorithm. By adjusting learnable weights / biases and hyperparameter configurations, the network optimizes the entire "feature extraction-expression-restoration" process in the deblurring task. For example, optimizing convolution kernel weights using the PSO algorithm (particle swarm optimization) can specifically enhance the recognition of blurred patterns, while adjusting the learning rate can balance convergence speed and accuracy, ultimately improving the structural authenticity and detail integrity of the deblurred image.
[0098] In one implementation, the optimization direction is determined by calculating the fitness of each solution in the initial solution set and selecting the maximum value as the current optimal solution, so that the algorithm can quickly locate a better solution in the initial stage, laying the foundation for subsequent iterative processes, ensuring that the algorithm can develop in the direction of improving the quality of the solution, avoiding blind search, and improving optimization efficiency.
[0099] In one implementation, the candidate optimal solution is generated by randomly perturbing the current optimal solution, w new =w best +rand*(α-β*w best ), where w new is the candidate optimal solution, w best is the current optimal solution, α is the global optimal eigenvector, β is the scaling factor, and rand is a random number in the range (0,1). This random perturbation introduces new solutions, balancing the algorithm's exploration (exploring new solution spaces) and exploitation (leveraging known optimal solutions). Furthermore, the Euclidean distance from the initial solution to the candidate optimal solutions is calculated and sorted, further strengthening the algorithm's ability to track high-quality solutions. This allows the algorithm to gradually approach the global optimal solution while maintaining population diversity.
[0100] In one implementation, for each initial solution, the Euclidean distance to the optimal solution is calculated:
[0101]
[0102] Where D is the Euclidean distance, w i For any initial solution, w new is the candidate optimal solution, m is the number of initial solutions;
[0103] In one implementation, the initial solution in the escape set is updated by using a random offset escape formula to obtain a suboptimal solution;
[0104] Random offset escape formula:
[0105]
[0106] in, is a suboptimal solution, P loc is the position of the candidate optimal solution, A is the balance parameter, iter now is the current number of iterations, iter max is the maximum number of iterations, rand is a random number in the interval (0,1), w i is any initial solution, and the initial solution corresponds to the suboptimal solution one by one;
[0107] The balance parameter A is a constant that controls the escape amplitude, typically set between 0.5 and 1.0. rand is a random number in the interval (0, 1), introducing directional randomness. The introduction of the balance parameter A and the random number rand allows the algorithm to dynamically adjust the escape amplitude and direction during the iteration process, preventing the algorithm from falling into a local optimum. As the number of iterations increases, the escape amplitude gradually decreases, allowing the algorithm to fully explore the solution space in the early stages and stably converge to the global optimal solution in the later stages. The adaptive escape mechanism enhances the algorithm's global search capabilities, while also improving its robustness and optimization accuracy.
[0108] In one implementation, a preset maximum number of iterations (e.g., 200) is reached; if the change in the fitness of the optimal solution is less than a threshold over T consecutive iterations, the fitness is considered converged. Setting the maximum number of iterations and the fitness convergence condition as the algorithm's termination conditions allows for precise control of the algorithm's execution. The maximum number of iterations provides a time constraint for the algorithm, preventing it from running for extended periods without convergence. The fitness convergence condition, however, is based on the quality of the solution. When the fitness of the optimal solution changes by less than a threshold over multiple consecutive iterations, the algorithm is deemed to have found a stable optimal solution, terminating the iterations. This ensures that the algorithm can find a high-quality solution within a limited timeframe, avoids excessive iterations, and improves its efficiency and reduces resource utilization.
[0109] In one embodiment, feature extraction is performed on a target image in a target image set to obtain a feature group, including:
[0110] The target image in the target image set is subjected to feature extraction by the target YOLOv7 model to obtain a feature group;
[0111] Improvements to target YOLOv7 include:
[0112] The CBS module in the backbone network is replaced with a separation Conv module, the CBL module in the YOLOv7-tiny model is replaced with the target CBL module, and the MDA module is added to the neck structure to obtain the target YOLOv7; the backbone network and neck structure are components of the YOLOv7-tiny model;
[0113] The working principle of the separation Conv module includes:
[0114] Get the initial feature tensor, input the initial feature tensor into the batch normalization layer and the SiLU activation function in sequence to obtain the first feature tensor, input the initial feature tensor into the depthwise separable convolution layer to obtain the second feature tensor, multiply the initial feature tensor, the first feature tensor and the second feature tensor to obtain the third feature tensor, and use the third feature tensor as the output of the separation Conv module;
[0115] The working principle of the target CBL module includes:
[0116] Obtain the initial feature map, decompose the initial feature map into multiple blocks according to the preset size, splice them according to the number of channels of each block to obtain the intermediate feature map, input the intermediate feature map into the 1×1 Conv module to obtain the output feature map, and use the output feature map as the output of the target CBL module.
[0117] In one implementation, see Figure 2 , Figure 2 A schematic diagram of the structure of an AI-based online teaching management method provided by an embodiment of the present invention; the working principle of the neck structure:
[0118] The output of the first CBL module and the output of the first MDA module are spliced to obtain a first output, the first output is sequentially input into the second CBL module and the third CBL module to obtain a second output, the second output is upsampled and then spliced with the output of the second MDA module to obtain a third output, the third output is sequentially input into the first ELAN module and the fourth CBL module to obtain a fourth output, the fourth output is upsampled and then spliced with the output of the third MDA module to obtain a fifth output, and the fifth output is passed through the fourth ELAN module as the output (first) of the neck structure;
[0119] Inputting the fifth output and the fourth output into a sixth CBL module to obtain a sixth output, concatenating the sixth output with the fourth output to obtain a target sixth output, inputting the target sixth output and the output of the first ELAN module into a third ELAN module to obtain a seventh output, and using the seventh output as the output (second) of the neck structure;
[0120] The seventh output and the second output are input into the fifth CBL module to obtain the eighth output, the eighth output and the second output are spliced together to obtain the ninth output, and the ninth output is sequentially passed through the second ELAN module and the CBL module as the output of the neck structure (third);
[0121] The numbers of the modules are only used to distinguish them (for example, "first" and "second" in the first CBL module and the second CBL module).
[0122] In one implementation, the Separate Conv module obtains feature tensors for different processing paths by inputting the initial feature tensor into a batch normalization layer, a SiLU activation function, and a depthwise separable convolution layer, respectively. These feature tensors are then multiplied together to obtain the final output. By leveraging the normalization and nonlinear enhancement effects of batch normalization and activation functions on features, and leveraging the efficient computational properties of depthwise separable convolution, the computational effort and number of parameters are reduced. The multiplication of feature tensors further integrates feature information from different paths, enhancing the expressive power of features. This allows the model to maintain efficient computation while reducing the number of parameters, lowering the computational burden and storage requirements of the model, and improving its operational efficiency.
[0123] In one implementation, the target CBL module decomposes the initial feature map into multiple blocks, concatenates them according to the number of channels in each block, and then processes them through a 1×1 convolution module to obtain the output feature map. This process recombines and optimizes the feature map in both spatial and channel dimensions. The decomposition and concatenation operations capture local feature information, and the concatenated features are further fused and compressed through the 1×1 convolution module, making the channel information in the feature map more compact and more semantically expressive. This helps improve the model's target feature extraction accuracy and enhances the model's detection capabilities for objects of different scales and shapes, thereby improving the accuracy and robustness of overall object detection.
[0124] In one embodiment, feature extraction is performed on target images in a target image set to obtain a feature group. The working principle of the MDA module includes:
[0125] Obtain the target feature map, input the target feature map into the 1×1 Conv module and the fully connected layer in sequence to obtain the Query, Key, and Value, input the Query and Key into the separation Conv module to obtain the target Query and target Key, perform matrix multiplication on the target Query and target Key to obtain the similarity matrix, and input the similarity matrix into the Softmax function to obtain the attention weight matrix;
[0126] Input the Value into the fully connected layer and the separation Conv module in sequence to obtain the target Value, and multiply the attention weight matrix by the target Value to obtain the fused feature map.
[0127] In one implementation, the target feature map is sequentially fed into a 1×1 Conv module and a fully connected layer to generate the Query (Q), Key (K), and Value (V). This allows different feature components to be generated from a single input data point, alleviating the limitation of the traditional self-attention mechanism of shared inputs for Q, K, and V. This allows the model to process these features independently, thus achieving data-parallel decoupling. This decoupling approach improves the model's flexibility, allowing each feature component to be transformed in an independent linear layer, enhancing the model's ability to express different features.
[0128] In one implementation, the query and key are input into the Separate Conv module to obtain the target query and target key. A similarity matrix is obtained through matrix multiplication, and then the attention weight matrix is obtained through the Softmax function. The introduction of the Separate Conv module further enhances the ability to fuse and extract features. Depthwise separable convolution reduces the number of parameters and computational complexity while maintaining feature independence and interpretability. The use of the Softmax function ensures the normalization of attention weights, enabling the model to more effectively aggregate features and improve the accuracy and robustness of feature fusion.
[0129] In one implementation, the value is sequentially fed into the fully connected layer and the separate Conv module to obtain the target value. The attention weight matrix is then multiplied by the target value to produce a fused feature map. This process, through independent decoupling operations and feature enhancement mechanisms, ensures a richer and more accurate representation of the value's features. The final fused feature map combines the feature information of the query, key, and value. Guided by the attention mechanism, it achieves precise extraction and fusion of target features, improving the model's recognition ability and overall performance.
[0130] Based on the same inventive concept, the embodiment of the present invention also provides an AI-based online teaching management system. Figure 3 , Figure 3A framework diagram of an AI-based online teaching management system provided in an embodiment of the present invention includes: an image enhancement module, a feature extraction module, and a teaching management module:
[0131] An image enhancement module is used to obtain an image dataset of a target user and enhance the images in the image dataset to obtain a target image set; the target image set contains multiple target images;
[0132] A feature extraction module is used to perform three-dimensional construction based on the target image in the target image set to obtain structural features, and extract features from the target image in the target image set to obtain a feature group; the feature group includes: global features and local features;
[0133] The teaching management module is used to input feature groups and structural features into the state recognition model to obtain state scores, and to issue early warnings to target users based on the state scores.
[0134] An AI-based online teaching management system provided by an embodiment of the present invention can accurately extract the structural features and global and local features of facial movements by collecting and enhancing images, and input them into a state recognition model. It can then accurately identify the user's learning status in real time, thereby improving the recognition success rate.
[0135] In one embodiment, the image enhancement module further includes: a deblurring module, a noise enhancement module, and a feature fusion module:
[0136] A deblurring module is used to perform pixel recognition on each image in the image dataset to obtain a blurred pixel atlas, and deblur each blurred pixel image in the blurred pixel atlas by fusing CNN to obtain a deblurred image set;
[0137] A noise enhancement module is used to determine the noise pixels of the target deblurred image, calculate an image matrix based on the noise pixels, and input the image matrix and the target deblurred image into the neural fuzzy system to obtain a noise enhanced image; the noise enhanced image is an image containing the positions of the noise pixels; the target deblurred image is any one of the blurred images;
[0138] The feature fusion module is used to optimize the noise-enhanced image through a conditional generative adversarial network to obtain a detail-enhanced image, fuse the detail-enhanced image and the noise-enhanced image to obtain a target image, and form a target image set with the target images corresponding to each deblurred image in the deblurred image set.
[0139] In one embodiment, the system further comprises:
[0140] The first operation module is used to initialize the population and parameters to obtain an initial solution set, calculate the fitness of each initial solution in the initial solution set, and take the initial solution with the largest fitness value as the current optimal solution;
[0141] The second operation module is used to generate candidate optimal solutions by randomly perturbing the current optimal solution, calculate the Euclidean distance from each initial solution to the candidate optimal solution, sort the initial solutions in descending order according to the Euclidean distance, and filter the initial solutions according to preset rules to obtain a tracking set and an escape set;
[0142] The third operation module is used to update the initial solution in the escape set by using the random offset escape formula to obtain a suboptimal solution, calculate the fitness of each suboptimal solution, and compare the fitness of the initial solution with the fitness of the corresponding suboptimal solution. If the fitness of the initial solution is greater than the fitness of the corresponding suboptimal solution, the initial solution is replaced by the corresponding suboptimal solution; otherwise, the initial solution is retained;
[0143] The fourth operation module is used to determine the optimal solution in the current population according to the fitness value, repeatedly execute the second operation module and the third operation module until the maximum number of iterations is reached or the fitness converges, and then output the optimal solution.
[0144] In one embodiment, the feature extraction module is further configured to extract features from the target image in the target image set using a target YOLOv7 model to obtain a feature group;
[0145] Improvements to target YOLOv7 include:
[0146] The CBS module in the backbone network is replaced with a separation Conv module, the CBL module in the YOLOv7-tiny model is replaced with the target CBL module, and the MDA module is added to the neck structure to obtain the target YOLOv7; the backbone network and neck structure are components of the YOLOv7-tiny model;
[0147] The working principle of the separation Conv module includes:
[0148] Get the initial feature tensor, input the initial feature tensor into the batch normalization layer and the SiLU activation function in sequence to obtain the first feature tensor, input the initial feature tensor into the depthwise separable convolution layer to obtain the second feature tensor, multiply the initial feature tensor, the first feature tensor and the second feature tensor to obtain the third feature tensor, and use the third feature tensor as the output of the separation Conv module;
[0149] The working principle of the target CBL module includes:
[0150] Obtain the initial feature map, decompose the initial feature map into multiple blocks according to the preset size, splice them according to the number of channels of each block to obtain the intermediate feature map, input the intermediate feature map into the 1×1 Conv module to obtain the output feature map, and use the output feature map as the output of the target CBL module.
[0151] In one embodiment, the working principle of the MDA module includes:
[0152] Obtain the target feature map, input the target feature map into the 1×1 Conv module and the fully connected layer in sequence to obtain the Query, Key, and Value, input the Query and Key into the separation Conv module to obtain the target Query and target Key, perform matrix multiplication on the target Query and target Key to obtain the similarity matrix, and input the similarity matrix into the Softmax function to obtain the attention weight matrix;
[0153] Input the Value into the fully connected layer and the separation Conv module in sequence to obtain the target Value, and multiply the attention weight matrix by the target Value to obtain the fused feature map.
[0154] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.
Claims
1. An AI-based online teaching management method, characterized in that: The method comprises: Acquire an image dataset of a target user, and enhance images in the image dataset to obtain a target image set; the target image set includes a plurality of target images; Performing three-dimensional construction on the target images in the target image set to obtain structural features, and extracting features from the target images in the target image set to obtain a feature group; the feature group includes: global features and local features; The feature group and the structural feature are input into a state recognition model to obtain a state score, and an early warning is issued to the target user according to the state score.
2. The AI-based online teaching management method according to claim 1, characterized in that: Enhancing the images in the image dataset to obtain a target image set includes: Performing pixel recognition on each image in the image data set to obtain a blurred pixel atlas, and deblurring each blurred pixel image in the blurred pixel atlas by fusing a CNN to obtain a deblurred image set; Determine the noise pixels of the target deblurred image, calculate an image matrix based on the noise pixels, and input the image matrix and the target deblurred image into a neural fuzzy system to obtain a noise-enhanced image; the noise-enhanced image is an image containing the positions of the noise pixels; and the target deblurred image is any one of the blurred images. The noise-enhanced image is optimized through a conditional generative adversarial network to obtain a detail-enhanced image, the detail-enhanced image and the noise-enhanced image are fused to obtain a target image, and the target images corresponding to the deblurred images in the deblurred image set are combined into a target image set.
3. The AI-based online teaching management method according to claim 2, characterized in that: Before deblurring each blurred pixel image in the blurred pixel image set by fusing CNN to obtain a deblurred image set, the parameter optimization method of the fused CNN includes: Step 1: Initialize the population and parameters to obtain an initial solution set, calculate the fitness of each initial solution in the initial solution set, and take the initial solution with the largest fitness value as the current optimal solution; Step 2: Generate candidate optimal solutions by randomly perturbing the current optimal solution, calculate the Euclidean distance from each initial solution to the candidate optimal solution, sort the initial solutions in descending order according to the Euclidean distance, and filter the initial solutions according to the preset rules to obtain the tracking set and escape set; Step 3: Update the initial solution in the escape set using the random offset escape formula to obtain a suboptimal solution, calculate the fitness of each suboptimal solution, and compare the fitness of the initial solution with the corresponding suboptimal solution. If the fitness of the initial solution is greater than the fitness of the corresponding suboptimal solution, replace the initial solution with the corresponding suboptimal solution; otherwise, retain the initial solution. Step 4: Determine the optimal solution in the current population based on the fitness value, and repeat steps 2 to 3 until the maximum number of iterations is reached, or the fitness converges, then output the optimal solution.
4. The AI-based online teaching management method according to claim 1, characterized in that: Extracting features from the target images in the target image set to obtain a feature group includes: Extracting features of the target image in the target image set using a target YOLOv7 model to obtain a feature group; The improvements to the target YOLOv7 include: The CBS module in the backbone network is replaced with a separation Conv module, the CBL module in the YOLOv7-tiny model is replaced with a target CBL module, and the MDA module is added to the neck structure to obtain the target YOLOv7; the backbone network and the neck structure are components of the YOLOv7-tiny model; The working principle of the separation Conv module includes: Obtain an initial feature tensor, input the initial feature tensor into a batch normalization layer and a SiLU activation function in sequence to obtain a first feature tensor, input the initial feature tensor into a depthwise separable convolutional layer to obtain a second feature tensor, multiply the initial feature tensor, the first feature tensor, and the second feature tensor to obtain a third feature tensor, and use the third feature tensor as the output of the separation Conv module; The working principle of the target CBL module includes: Obtain an initial feature map, decompose the initial feature map into multiple blocks according to a preset size, splice the blocks according to the number of channels to obtain an intermediate feature map, input the intermediate feature map into a 1×1 Conv module to obtain an output feature map, and use the output feature map as the output of the target CBL module.
5. The AI-based online teaching management method according to claim 1, characterized in that: Feature extraction is performed on the target image in the target image set to obtain a feature group. The working principle of the MDA module includes: Obtain a target feature map, input the target feature map into a 1×1 Conv module and a fully connected layer in sequence to obtain a Query, a Key, and a Value, input the Query and the Key into a separation Conv module to obtain a target Query and a target Key, perform a matrix multiplication operation on the target Query and the target Key to obtain a similarity matrix, and input the similarity matrix into a Softmax function to obtain an attention weight matrix; The Value is input into the fully connected layer and the separation Conv module in sequence to obtain the target Value, and the attention weight matrix is multiplied by the target Value to obtain a fused feature map.
6. An AI-based online teaching management system, characterized in that: The system includes: an image enhancement module, a feature extraction module and a teaching management module: The image enhancement module is used to obtain an image dataset of a target user and enhance images in the image dataset to obtain a target image set; the target image set includes multiple target images; The feature extraction module is used to perform three-dimensional construction based on the target image in the target image set to obtain structural features, and perform feature extraction on the target image in the target image set to obtain a feature group; the feature group includes: global features and local features; The teaching management module is used to input the feature group and the structural feature into a state recognition model to obtain a state score, and to issue an early warning to the target user according to the state score.
7. The AI-based online teaching management system according to claim 6, characterized in that: The image enhancement module also includes: a deblurring module, a noise enhancement module and a feature fusion module: The deblurring module is configured to perform pixel recognition on each image in the image data set to obtain a blurred pixel atlas, and deblur each blurred pixel image in the blurred pixel atlas by fusing a CNN to obtain a deblurred image set; The noise enhancement module is configured to determine the noise pixels of the target deblurred image, calculate an image matrix based on the noise pixels, and input the image matrix and the target deblurred image into the neural fuzzy system to obtain a noise enhanced image; the noise enhanced image is an image containing the positions of the noise pixels; the target deblurred image is any one of the blurred images; The feature fusion module is used to optimize the noise-enhanced image through a conditional generative adversarial network to obtain a detail-enhanced image, fuse the detail-enhanced image and the noise-enhanced image to obtain a target image, and form a target image set with the target images corresponding to each deblurred image in the deblurred image set.
8. The AI-based online teaching management system according to claim 7, characterized in that: The system further includes: a first operating module, a second operating module, a third operating module, a fourth operating module and a fifth operating module: The first operation module is used to initialize the population and parameters to obtain an initial solution set, calculate the fitness of each initial solution in the initial solution set, and take the initial solution with the largest fitness value as the current optimal solution; The second operation module is configured to generate a candidate optimal solution by randomly perturbing the current optimal solution, calculate the Euclidean distance from each initial solution to the candidate optimal solution, sort the initial solutions in descending order according to the Euclidean distance, and filter the initial solutions according to a preset rule to obtain a tracking set and an escape set; The third operation module is configured to update the initial solution in the escape set using a random offset escape formula to obtain a suboptimal solution, calculate the fitness of each suboptimal solution, compare the fitness of the initial solution with the fitness of the corresponding suboptimal solution, and if the fitness of the initial solution is greater than the fitness of the corresponding suboptimal solution, replace the initial solution with the corresponding suboptimal solution; otherwise, retain the initial solution; The fourth operation module is used to determine the optimal solution in the current population according to the fitness value, repeatedly execute the second operation module and the third operation module until the maximum number of iterations is reached or the fitness converges, and then output the optimal solution.
9. The AI-based online teaching management system according to claim 6, characterized in that: The feature extraction module is further configured to extract features of the target image in the target image set using a target YOLOv7 model to obtain a feature group; The improvements to the target YOLOv7 include: The CBS module in the backbone network is replaced with a separation Conv module, the CBL module in the YOLOv7-tiny model is replaced with a target CBL module, and the MDA module is added to the neck structure to obtain the target YOLOv7; the backbone network and the neck structure are components of the YOLOv7-tiny model; The working principle of the separation Conv module includes: Obtain an initial feature tensor, input the initial feature tensor into a batch normalization layer and a SiLU activation function in sequence to obtain a first feature tensor, input the initial feature tensor into a depthwise separable convolutional layer to obtain a second feature tensor, multiply the initial feature tensor, the first feature tensor, and the second feature tensor to obtain a third feature tensor, and use the third feature tensor as the output of the separation Conv module; The working principle of the target CBL module includes: Obtain an initial feature map, decompose the initial feature map into multiple blocks according to a preset size, splice the blocks according to the number of channels to obtain an intermediate feature map, input the intermediate feature map into a 1×1 Conv module to obtain an output feature map, and use the output feature map as the output of the target CBL module.
10. The AI-based online teaching management system according to claim 6, characterized in that: The working principle of the MDA module includes: Obtain a target feature map, input the target feature map into a 1×1 Conv module and a fully connected layer in sequence to obtain a Query, a Key, and a Value, input the Query and the Key into a separation Conv module to obtain a target Query and a target Key, perform a matrix multiplication operation on the target Query and the target Key to obtain a similarity matrix, and input the similarity matrix into a Softmax function to obtain an attention weight matrix; The Value is input into the fully connected layer and the separation Conv module in sequence to obtain the target Value, and the attention weight matrix is multiplied by the target Value to obtain a fused feature map.
Citation Information
Patent Citations
Virtual teaching management method and device based on artificial intelligence
CN115936944A
Online education scene student attention recognition method for CPU operation optimization
CN112597888A
Expression recognition method based on CNN model in teaching scene
CN113221683A
Marine target detection method under visible light based on improved YOLOv7 algorithm
CN116863293A
Method, system and device for accurately predicting human emotion based on facial expression recognition
CN117173767A