Online classroom concentration degree monitoring method for style disturbance privacy protection
By combining style transfer technology and models, the contradiction between privacy protection and attention monitoring in online education has been resolved, enabling accurate identification of student attention levels without compromising privacy and ensuring teaching quality.
Patent Information
- Application Number
- CN202511167200.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-25
AI Technical Summary
In the online education environment, there is a conflict between protecting student privacy and monitoring attention performance. Existing technologies struggle to achieve accurate attention recognition without compromising privacy.
Style perturbation was introduced through style transfer technology to establish an online classroom dataset. By combining visual privacy protection scores and attention prediction models, a balance was struck between privacy protection and attention recognition. Mamba-ST model, CatBoost regression model and association statistics model were used for feature extraction and prediction.
This approach enables the accurate identification of student focus levels while protecting student privacy, avoiding a decline in monitoring performance due to excessive privacy protection, and ensuring teaching quality.
Smart Images

Figure CN121010949A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision and pattern recognition, and particularly relates to a style disturbance privacy protection online classroom concentration monitoring method. BACKGROUND
[0002] At present, the traditional face-to-face education mode is severely impacted, in order to ensure the normal progress of education, the schools in various places are forced to quickly switch to the online education mode, and with the passage of time, the time is fragmented, and more classroom learning gradually changes from the traditional offline education mode to online education; however, this change of teaching mode faces many challenges. In the traditional teaching mode, the teacher can directly observe the facial expression, body movement, language and other real-time learning state and concentration degree of the student, and timely adjust the teaching strategy. In the online education environment, the direct teacher-student interaction mode is limited, many students choose to close the camera due to anxiety, fear of exposure, shyness, and worry about privacy leakage, which hinders the teacher's accurate grasp of the student's learning state, and closing the camera makes it difficult for the teacher to understand the student's learning state, and thus it is difficult to guarantee the teaching quality. On the contrary, if the students are forced to turn on the camera, it will further cause the students' concern about the exposure of personal privacy, and even may lead to the students' resistance, which will affect the learning enthusiasm.
[0003] In order to solve such technical defects, the prior art usually relies on directly collecting original videos, and evaluates the concentration level by analyzing the facial expression, eye movement, head posture and other behavior characteristics of the student, which inevitably involves the personal information of the student. In the private scene of home learning, this data collection mode will increase the risk of leakage of multi-dimensional privacy information such as student identity, behavior and environment, leading to the contradiction between privacy protection and concentration monitoring performance, and it is impossible to realize concentration recognition without leaking privacy. SUMMARY
[0004] In order to solve the technical problem that the existing concentration monitoring technology causes the contradiction between privacy protection and concentration monitoring performance, the purpose of the present application is to provide a style disturbance privacy protection online classroom concentration monitoring method, and the technical solution adopted is as follows:
[0005] Obtain a video data set, introduce style disturbance through style transfer technology, and establish an online classroom data set;
[0006] Obtain the corresponding visual privacy protection score through the online classroom data set;
[0007] Position the face region in the online classroom data set to obtain a face detection rate;
[0008] The face key point detection is performed according to the face region, and the behavior index and the facial high-level semantic feature are determined respectively;
[0009] The attention prediction model is constructed, the facial key point, the behavior index and the facial high-level semantic feature are input into the attention prediction model, and the attention recognition rate is output;
[0010] The association statistical model is established based on the visual privacy protection score, and the visual privacy protection and the attention recognition are balanced in combination with the face detection rate and the attention recognition rate.
[0011] Preferably, a video data set is acquired, a style disturbance is introduced through a style transfer technology, and an online classroom data set is established, specifically:
[0012] The effective video data is obtained based on the video data set, the features of the original image and the style image are extracted respectively through the Mamba-ST model, and the spatial feature map is output after fusion, the spatial feature map is mapped to the image space, and the online classroom data set is established.
[0013] Preferably, the online classroom data set is established, including:
[0014] The Mamba-ST model includes a parallel content encoder and a style encoder, and a Mamba decoder;
[0015] The original image and the style image are input into the content encoder and the style encoder respectively, and the features of the original image and the style image are obtained correspondingly;
[0016] The features of the original image and the style image are input into the Mamba decoder for feature fusion, and the fused features are rearranged through the Depatchify module to obtain the spatial feature map;
[0017] The spatial feature map is mapped to the image space by using the CNN decoder to obtain the stylized image, and the online classroom data set is established.
[0018] Preferably, the corresponding visual privacy protection score is acquired through the online classroom data set, including:
[0019] The visual saliency map is acquired based on the online classroom data set, the statistical features are extracted according to the visual saliency map, and the visual privacy protection score is obtained by analyzing the statistical features;
[0020] The CatBoost is trained through the statistical features of the visual saliency map and the corresponding visual privacy protection score, and the visual privacy protection score corresponding to each frame of the online classroom data set is obtained based on the CatBoost.
[0021] Preferably, the visual saliency map is acquired based on the online classroom data set, specifically:
[0022] The multi-channel features of the online classroom dataset are extracted and normalized, a graph weight matrix corresponding to each channel is constructed, the graph weight matrix is normalized to determine the balanced distribution corresponding to each channel, and a visual saliency map is obtained by fusing the balanced distribution.
[0023] Preferably, the statistical features include statistical quantity features, statistical histograms, local shape dimension distributions and Benford laws.
[0024] Preferably, the behavior indicators include eye width ratios, mouth width ratios, gaze direction vectors and head angles, and the facial high-level semantic features include eye states and head states.
[0025] Preferably, a concentration prediction model is constructed, the facial key points, the behavior indicators and the facial high-level semantic features are input into the concentration prediction model, and a concentration recognition rate is output, including:
[0026] The concentration prediction model includes at least three inputs, wherein the first input uses ST-GCN to extract features of the facial key points, the second input and the third input both use LSTM to process the behavior indicators and the facial high-level semantic features to form a complementary feature representation mechanism, and feature representations of each input are output.
[0027] The feature representations of each input are spliced and fused, weighted through a self-attention mechanism, and output through a fully connected layer to obtain the concentration recognition rate.
[0028] Preferably, the behavior indicators and the facial high-level semantic features are processed using LSTM to form a complementary feature representation mechanism, specifically:
[0029] The learning of the behavior indicators is guided by the facial high-level semantic features, and the facial high-level semantic features are supplemented with detailed information missing based on the behavior indicators.
[0030] Preferably, an association statistical model is established based on the visual privacy protection score, and the visual privacy protection and the concentration recognition are balanced by combining the face detection rate and the concentration recognition rate, including:
[0031] An association statistical model is established based on the visual privacy protection score and data availability, a fitting relationship is obtained, a fitting surface is obtained, the online classroom dataset to be monitored is predicted through the fitting surface, and the visual privacy protection and the concentration recognition are balanced by combining the face detection rate and the concentration recognition rate.
[0032] The present application has the following advantages:
[0033] By introducing non-real style disturbance to each frame image in the video data set through style transfer technology, the attention distribution of human eyes to the original sensitive area is reduced, the high-level semantic structure information of the image is preserved while realizing visual privacy protection; combined with the human eye visual perception mechanism, the statistical features of the visual saliency map are extracted and the CatBoost regression model is used to realize the quantitative evaluation of the visual privacy protection level; and a concentration prediction model is constructed to obtain the concentration recognition rate in the corresponding learning scene in the online classroom based on the facial key points, behavior indicators and facial high-level semantic features; finally, an association statistical model is established to effectively alleviate the contradiction between visual privacy protection and data usability, and to balance the visual privacy protection and concentration recognition of the face detection rate, that is, to select the appropriate visual privacy protection coding range to avoid the problem of serious decline in the performance of online classroom concentration monitoring application due to excessive privacy protection. BRIEF DESCRIPTION OF DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, and the advantages thereof, a brief introduction will be given to the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0035] Figure 1 A step flow chart of a style disturbance privacy protection online classroom concentration monitoring method provided by an embodiment of the present application;
[0036] Figure 2 An image visual privacy protection schematic diagram of a style disturbance privacy protection online classroom concentration monitoring method provided by an embodiment of the present application using style transfer technology;
[0037] Figure 3 A visual saliency map visualization result schematic diagram of a style disturbance privacy protection online classroom concentration monitoring method provided by an embodiment of the present application;
[0038] Figure 4 A fitting surface schematic diagram of an association statistical model of a style disturbance privacy protection online classroom concentration monitoring method provided by an embodiment of the present application. DETAILED DESCRIPTION
[0039] In order to further illustrate the technical means and effects taken by the present application to achieve the predetermined object of the application, the specific implementation, structure, features and effects of the style disturbance privacy protection online classroom concentration monitoring method according to the present application are described in detail as follows in combination with the drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0041] The specific scheme of the style disturbance privacy protection online classroom concentration monitoring method provided by the present application is specifically described below in combination with the drawings.
[0042] Please refer to Figure 1 , which shows the step flow chart of the style disturbance privacy protection online classroom concentration monitoring method provided by one embodiment of the present application, the method comprising:
[0043] Step S1: Obtain a video data set, introduce style disturbance through style transfer technology, and establish an online classroom data set;
[0044] Step S2: Obtain the corresponding visual privacy protection score through the online classroom data set;
[0045] Step S3: Locate the face region in the online classroom data set to obtain the face detection rate;
[0046] Step S4: Perform face key point detection according to the face region to determine the behavior index and the facial high-level semantic feature respectively;
[0047] Step S5: Construct a concentration prediction model, input the facial key point, the behavior index and the facial high-level semantic feature into the concentration prediction model, and output the concentration recognition rate;
[0048] Step S6: Establish an association statistical model based on the visual privacy protection score, and balance the visual privacy protection and the concentration recognition by combining the face detection rate and the concentration recognition rate.
[0049] To better illustrate, based on the online teaching situation, to ensure that teachers can monitor the concentration of students in real time, while avoiding the concern of students about the exposure of personal privacy, an online class concentration monitoring method based on style disturbance privacy protection is proposed to avoid the direct identification of the specific behavior data of students, thereby realizing privacy protection and ensuring that teachers can still obtain accurate concentration recognition rate of corresponding students for the online class data set with visual privacy protection effect, to improve the interaction of both parties for online class.
[0050] Please refer to Figure 2 Further, in step S1, specifically:
[0051] Based on the video data set, effective video data is obtained, the features of the original image and the style image are extracted respectively through the Mamba-ST model, and the spatial feature map is fused and output, the spatial feature map is mapped to the image space, and the online class data set is established.
[0052] It is explained that the video data set is the EngageNet video data set of public online classroom learning, which is a large-scale multi-oriented user participation prediction data set, containing the behavior data of multiple participants when watching educational videos, and covering computer laboratories, dormitories, open spaces and other diversified scenes, providing rich benchmark data resources for participation prediction problems in the fields of online education, human-computer interaction and user experience design.
[0053] It can be explained that the effective video data obtained by screening is the preliminary screening, which retains clear video frame data and avoids the influence of blurred data; the style transfer technology is to integrate the visual features of different artistic styles into the effective video data, mainly used in the field of non-realistic rendering; this technology can introduce the rendering of the target style while maintaining the high-level semantic features of the original image, generating new images with stylized features, i.e. stylized images; through the style transfer technology, the demand for visual privacy protection can be met, i.e. introducing non-realistic style disturbance to the original image while retaining the high-level semantic features of the original image to ensure subsequent video frame analysis.
[0054] Specifically, the Mamba-ST (Mamba Style Transfer) model is used as the basic framework of visual privacy protection, and the style transfer is performed on the whole original image through deep learning technology, i.e. the original image and the style image are processed to obtain a stylized image, and an online classroom video data with visual privacy protection effect is established.
[0055] Further, in step S1, it includes:
[0056] The Mamba-ST model includes a parallel content encoder and a style encoder, and a Mamba decoder.
[0057] It can be explained that in the embodiment, the Mamba-ST model adopts a double-branch symmetric structure, which is composed of parallel content encoders and style encoders, respectively processes the original image and the style image, and adopts the Mamba decoder to realize feature fusion, to generate video frame data that retains the original meaning and conforms to the specified style, that is, through the multi-level structure setting, the key features in the image are accurately captured, and the original details and clarity of the image are maximized while protecting privacy.
[0058] Step S11: input the original image and the style image into the content encoder and the style encoder respectively, and obtain the features of the original image and the style image correspondingly.
[0059] It is explained that the content encoder adopts the Patch Embedding technology, that is, the original image is divided into multiple small blocks for feature extraction, and each small block is converted into a fixed-dimensional vector through linear transformation; the style encoder is a Mamba encoder, and is composed of 3 Mamba encoding layers to capture style features at different levels.
[0060] Specifically, the corresponding calculation formula is:
[0061]
[0062] wherein, , represent the features of the original image and the style image respectively; , represent the processing of the content encoder and the style encoder respectively; , represent the original image and the style image respectively.
[0063] Step S12: input the features of the original image and the style image into the Mamba decoder for feature fusion, and rearrange the fused features through the Depatchify module to obtain the spatial feature map.
[0064] It is explained that in the embodiment, the ST-VSSM (Style Transfer-Vision State Space Model, i.e. vision state space model based on style transfer) technology is adopted for feature fusion of features, that is, by constructing a dynamic vision state space, the complex relationship between style features and content features is effectively captured and integrated, and accurate feature fusion is realized.
[0065] Specifically, cross-modal interaction is realized through a modified two-dimensional state space equation, and its core update process can be represented as:
[0066]
[0067] wherein, denotes the hidden state vector of the th step; , both denote the discretization of the basic matrix parameters of the state space model SSM, i.e. the matrices and obtained by discretizing the original continuous-time matrices denotes the embedding representation of the th patch in the feature sequence of the style image; denotes the feature representation vector after fusion of the current th step; denotes the output matrix generated by the feature mapping of the original image.
[0068] It can be explained that, , , The definition expression of
[0069]
[0070] wherein, denotes a linear layer operation. And , The definition expression of
[0071]
[0072] wherein, is an identity matrix; denotes the discretization step length parameter, and the definition expression is:
[0073]
[0074] Then, in the output reconstruction stage, the fused features are rearranged into spatial feature maps by the Depatchify module, i.e. the Depatchify module gradually restores the texture, edge and color information of the fused features through multiple levels of processing steps, achieving high-quality image reconstruction.
[0075] Step S13: mapping the spatial feature map to the image space by using the CNN decoder to obtain the stylized image, and establishing an online classroom dataset.
[0076] It is explained that the CNN (Convolutional Neural Network) decoder is used to dynamically adjust different style weight parameters , to map the spatial feature map to the image space, obtain a stylized image, realize continuous and controllable adjustment of the degree of visual privacy protection, and establish an online classroom dataset, i.e., an online classroom video dataset with visual privacy protection effect. Preferably, for valid video data, a style with good visual privacy protection is selected, and the style weight is set to be 0 to 1 in turn with a step of 0.2; under the visual privacy protection mechanism, recognition is performed based on the online classroom dataset.
[0077] Further, in step S2, the following steps are included:
[0078] Step S21: obtaining a visual saliency map based on the online classroom dataset, extracting statistical features from the visual saliency map, and analyzing the statistical features to obtain a visual privacy protection score.
[0079] It can be explained that the visual saliency map reflects the areas in the image that attract the most human attention, and is mainly divided into a bottom-up approach and a top-down approach, wherein the bottom-up approach relies on the saliency of visual stimuli to determine the region of interest, and the top-down approach is driven by tasks; in this embodiment, a graph-based visual saliency algorithm (GBVS, i.e., Graph-Based Visual Saliency) is used to reflect the attention change of the sensitive area before and after visual privacy protection.
[0080] Please refer to Figure 3 , wherein as the style weight increases, the original saliency region, i.e., the visual saliency map, gradually darkens, and the saliency region shifts to other regions; further, in step S21, the visual saliency map is obtained based on the online classroom dataset, specifically:
[0081] The multi-channel features of the online classroom dataset are extracted and normalized, a graph weight matrix corresponding to each channel is constructed, the normalized graph weight matrix is processed to determine the balanced distribution corresponding to each channel, and the balanced distribution is fused to obtain the visual saliency map.
[0082] Specifically, the multi-channel features of the video frames in the online classroom dataset are extracted and normalized, and the visual saliency map is denoted as , and the probability map obtained after normalization is For each channel, a graph weight matrix is constructed, and the corresponding calculation formula is:
[0083]
[0084] wherein, represents the pixel in channel and the pixel construct a graph weight matrix; , respectively represent the pixel with the pixel probability map; , represent the spatial position coordinates of the pixel with the pixel in the visual saliency map, i.e. the physical spatial coordinates of the pixel in the image; represent the standard deviation parameter of the spatial distance.
[0085] Then, the graph weight matrix is normalized to a Markov matrix, where the Markov matrix is a special matrix whose sum of elements in each row is equal to 1, representing that the sum of probabilities of transition from one state to another is 1; and the equilibrium distribution corresponding to each channel, the equilibrium distribution refers to in the Markov chain, after a long enough time of iteration, the state transition probability no longer changes, the system reaches a stable state, and the corresponding calculation formula is:
[0086]
[0087]
[0088] wherein, represents the Markov matrix after the graph weight matrix corresponding to the channel is normalized; represents the normalization processing; represents the graph weight matrix corresponding to the channel . represents the stationary distribution of the channel . represents the time step number of the Markov diffusion.
[0089] Finally, the equilibrium distributions of multiple channels are fused to obtain the visual saliency map, denoted as , and the corresponding calculation formula is:
[0090]
[0091] wherein, represents the edge weight of the pixel pair in the graph weight matrix corresponding to the channel .
[0092] It can be explained that in step S21, the statistical feature is used to more accurately identify and distinguish different visual objects, and improve the overall recognition performance. Further, the statistical feature includes statistical quantity feature, statistical histogram, local shape dimension distribution and Benford law.
[0093] It can be explained that the statistical features include mean, skewness, kurtosis and entropy, wherein the mean is used to reflect the overall level of the visual saliency map, and measures the pixel trend; the skewness is used to describe the symmetry of the visual saliency map, and when the distribution is completely symmetric, the skewness is zero; when the distribution is right-skewed, the skewness is positive, and when the distribution is left-skewed, the skewness is negative; the kurtosis is used to describe the sharpness of the visual saliency map, and when the distribution is relatively flat, the kurtosis is small, and when the distribution is relatively sharp, the kurtosis is large; the entropy is used to measure the uncertainty, and describes the complexity and randomness of the visual saliency map, and the greater the entropy value, the higher the uncertainty; otherwise, the smaller the entropy value, the higher the certainty.
[0094] Next, the statistical histogram is obtained, in this embodiment, the visual saliency map is normalized, divided into 32 bins, and the number of pixels in each bin is counted, denoted as , the histogram feature vector is calculated, and the corresponding calculation formula is:
[0095]
[0096] wherein, denotes the histogram feature vector, i.e. the histogram feature vector with a length of 32, denoted as ; denotes the total number of pixels in the visual saliency map.
[0097] Then, the Benford law is obtained, which involves the leading digit of each pixel value in the visual saliency map, i.e. describes the distribution rule of the first digit in the natural data set, which proves that the leading digit , and conforms to the corresponding probability function:
[0098]
[0099] It is explained that the Benford law is used in the wavelet domain and the gradient domain of the visual saliency map to reflect the change of the image space attention distribution, i.e. by counting the probability distribution of each region in the image, the attention concentration point of the human visual system when observing the image is effectively reflected.
[0100] Finally, the local fractal dimension distribution is obtained, which reflects the relationship between the image structure complexity and the scale change; in this embodiment, the box counting method is used to calculate the local fractal dimension feature of the visual saliency map to quantify the structure complexity of the local region of the visual saliency map. Specifically, each pixel of the visual saliency map is taken as the center, a 7x7 local neighborhood region is extracted, and the fractal dimension of the neighborhood region is estimated by using the box counting method to generate a fractal dimension map of the entire visual saliency map.
[0101] Step S23: training CatBoost through the statistical features of the visual saliency map and the corresponding visual privacy protection score, and obtaining the visual privacy protection score corresponding to each frame of the online classroom dataset based on the CatBoost.
[0102] It is explained that CatBoost, as a high-efficiency gradient boosting decision tree regression algorithm, realizes excellent prediction performance in regression tasks through technologies such as ordered boosting and target statistical encoding; specifically, based on the online classroom dataset, the training set and the test set are determined, the visual saliency map and the corresponding visual privacy protection score are obtained for the training set, and are input to the CatBoost for training, the trained CatBoost is used to predict the test set, and the corresponding predicted visual privacy protection score in each frame in the test set is obtained.
[0103] It can be explained that in step S3, the face region in the online classroom dataset is located to obtain the face detection rate; specifically, the RetinaFace method with strong robustness is used to locate the face region in each frame for the online classroom dataset, wherein the RetinaFace method uses deep learning technology, through multi-scale feature fusion and multi-task learning framework, the accuracy and robustness of face detection are improved, so that in various complex backgrounds and occlusion conditions, a higher detection rate can still be maintained; then the number of faces in the current frame and the number of faces in the total number of frames of the online classroom dataset are counted to determine the ratio between the two to obtain the face detection rate, which is used to obtain the facial key points and measure the data effectiveness.
[0104] Please refer to Figure 4 Wherein, as the degree of visual privacy protection increases, the face detection rate gradually decreases, indicating that the availability of data is gradually decreasing.
[0105] It can be explained that in step S4, the face key point detection is performed according to the face region, and the behavior indicators and the facial high-level semantic features are determined respectively; that is, a general key point detection model considering artistic style is used to detect the key points of the face region, which to some extent preserves and embodies the uniqueness of artistic style, and in the key point detection process, the accuracy and robustness of the key points are guaranteed, and specific artistic requirements can also be met in visual effect.
[0106] Further, in step S4, the behavior indicators include eye width ratio, mouth width ratio, gaze direction vector and head angle; the facial high-level semantic features include eye state and head state.
[0107] Specifically, the corresponding behavior indicators are calculated based on the key points, and the corresponding calculation formula is:
[0108]
[0109]
[0110] wherein, represents an eye width ratio; represents a mouth width ratio; represents the i-th key point of the face; represents the i-th key point of the face; represents the Euclidean norm.
[0111] It can be understood that the eye width ratio refers to the ratio of the vertical distance of the eyes to the horizontal distance of the eyes; the mouth width ratio refers to the ratio of the vertical distance of the mouth to the horizontal distance of the mouth; in the embodiment, 68 key points of the face are detected based on the key points, which have a certain corresponding relationship with the organs of the face, wherein the 37th-42nd key points correspond to the eye part; the 60th-66th key points correspond to the mouth part.
[0112] The head pose refers to the direction and angle of the head in space, i.e. the three-dimensional rotation angle, including up-down tilt (pitch), left-right rotation (yaw) and forward-back swing (roll); in the embodiment, the three-dimensional rotation angle of the head is accurately calculated by the PnP (Perspective-n-Point) algorithm to reflect the head pose change trend of the student in the online classroom in time sequence in detail, which is beneficial to analyze the gaze direction, the attention concentration point and the head movement mode.
[0113] The gaze direction is used to extract the visual behavior clues of the student; in the embodiment, a gaze direction estimation method based on eye key points is adopted, i.e. the average coordinates of the left and right eye center points are calculated as the eye center distance, and the offset vector of the image center and the eye center is calculated to estimate the gaze direction, which reflects the eye movement trajectory of the student, and can maintain stable performance under different light conditions.
[0114] As an optional implementation, the eye state includes the open-closed state of the eyes, the blinking frequency, the eye rotation direction and the pupil dilation degree, etc.; the face state includes the change of facial expressions, such as smiling, frowning, surprise, anger, etc.
[0115] Further, in step S5, the following steps are included:
[0116] Step S51: The concentration prediction model includes at least three inputs, wherein the first input uses ST-GCN to extract the features of the face key points; the second input and the third input both use LSTM to process the behavior indicators and the high-level semantic features of the face to form a complementary feature representation mechanism; and the feature representation of each path is output.
[0117] To make an explanation, the ST-GCN (Spatial-Temporal Graph Convolutional Network) is used to effectively capture the complex relationship between facial expressions and facial key points.
[0118] Further, in step S51, the behavior indicators and the facial high-level semantic features are processed using LSTM to form a complementary feature representation mechanism, specifically:
[0119] The learning of the behavior indicators is guided by the facial high-level semantic features, and the facial high-level semantic features are supplemented with detailed information missing based on the behavior indicators.
[0120] To make an explanation, LSTM (Long Short-Term Memory) is used to capture long-term dependencies in the time series data of eye and head features; specifically, by using high-level state semantic features to guide the learning of low-level behavior indicators, a feature representation from coarse to fine is achieved, and the behavior indicators supplement the state features, i.e., the detailed information missing in the high-level state semantic features, forming a complementary feature representation mechanism.
[0121] Step S52: The feature representations of each path are spliced and fused, and after being processed by a self-attention mechanism, the attention recognition rate is output through a fully connected layer.
[0122] Specifically, the features output by the three paths are spliced and fused, and the fused features are processed by a self-attention mechanism module, which can dynamically learn the important weights of different feature types to ensure that the contribution of each feature in the final result is accurately evaluated; then, the attention recognition rate is output through a fully connected layer to realize accurate judgment of the student's attention state.
[0123] Further, in step S6, it includes:
[0124] According to the visual privacy protection score and the data availability, a correlation statistical model is established to obtain a fitting relationship, and a fitting surface is obtained, and the to-be-monitored online classroom data set is predicted through the fitting surface, and the face detection rate and the attention recognition rate are combined to balance the visual privacy protection and the attention recognition.
[0125] To make an explanation, in order to alleviate the contradictory relationship between visual privacy protection and attention recognition performance, a correlation statistical model is established according to the visual privacy protection score and the data availability, i.e., by analyzing the influence of different privacy protection levels on image clarity and the availability and accuracy of data under various privacy protection strategies, the best balance point is found.
[0126] Specifically, by using a bicubic polynomial to establish a fitting relationship between the visual privacy protection score and the data availability, a fitting surface is obtained as a prediction of other data points, providing an effective basis for the selection of privacy protection strength, that is, by using the smoothing property of the bicubic polynomial, the subtle balance between privacy protection demand and data utility is accurately captured; and by using the face detection rate and the concentration recognition rate, the data availability is evaluated, wherein the face detection rate reflects the performance retention degree of the data after visual privacy protection processing on the basic visual task; and the concentration recognition rate reflects the utility of the data after visual privacy protection processing in specific application scenarios; that is, the trade-off between privacy protection and data utilization is optimized through the correlation statistical model.
[0127] Understandably, by introducing non-real style disturbance to each frame of image in the video data set through the style transfer technology, the attention distribution of the human eye to the original sensitive area is reduced, the high-level semantic structure information of the image is retained while the visual privacy protection is realized; combined with the human eye visual perception mechanism, the statistical features of the visual saliency map are extracted and the CatBoost regression model is used to realize the quantitative evaluation of the visual privacy protection level; and a concentration prediction model is constructed, the concentration recognition rate in the corresponding learning scene of the online class is obtained based on the facial key points, the behavior indicators and the facial high-level semantic features; finally, a correlation statistical model is established, which effectively alleviates the contradictory relationship between visual privacy protection and data availability, and balances the visual privacy protection and the concentration recognition rate through the face detection rate and the concentration recognition rate, that is, the appropriate visual privacy protection coding range is selected to avoid the problem of serious decline in the performance of the online class concentration monitoring application due to excessive privacy protection.
[0128] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or can be advantageous.
[0129] Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other, and each embodiment mainly describes the difference from other embodiments.
Claims
1. A style perturbation privacy-preserving online class attentiveness monitoring method, characterized in that, The method comprises: Obtaining a video data set, introducing style disturbance through a style transfer technology, and establishing an online classroom data set; Obtaining a corresponding visual privacy protection score through the online classroom data set; Positioning a face region in the online classroom data set to obtain a face detection rate; Performing face key point detection according to the face region to determine a behavior index and a facial high-level semantic feature respectively; Building a concentration prediction model, inputting the face key point, the behavior index and the facial high-level semantic feature into the concentration prediction model, and outputting a concentration recognition rate; Based on the visual privacy protection score, a correlation statistical model is established, and the face detection rate and the concentration recognition rate are combined to balance the visual privacy protection and the concentration recognition.
2. The style perturbation privacy-preserving online class attentiveness monitoring method of claim 1, wherein, Obtaining a video data set, introducing style disturbance through a style transfer technology, and establishing an online classroom data set, specifically: Based on the video data set, effective video data is obtained, the features of the original image and the style image are extracted respectively through the Mamba-ST model, and are fused to output a spatial feature map, the spatial feature map is mapped to an image space, and the online classroom data set is established.
3. The style perturbation privacy-preserving online class attentiveness monitoring method of claim 2, wherein, Establishing an online classroom data set includes: The Mamba-ST model comprises a parallel content encoder and a style encoder, and a Mamba decoder; The original image and the style image are input into the content encoder and the style encoder respectively, and the features of the original image and the style image are obtained correspondingly; The features of the original image and the style image are input into the Mamba decoder for feature fusion, and the fused features are rearranged through the Depatchify module to obtain a spatial feature map; The spatial feature map is mapped to an image space by using a CNN decoder to obtain a stylized image, and the online classroom data set is established.
4. The style perturbation privacy-preserving online class attentiveness monitoring method of claim 1, wherein, Obtaining a corresponding visual privacy protection score through the online classroom data set includes: Based on the online classroom data set, a visual saliency map is obtained, statistical features are extracted according to the visual saliency map, and a visual privacy protection score is obtained by analyzing the statistical features; CatBoost is trained based on the statistical features of the visual saliency map and the corresponding visual privacy protection score, and the visual privacy protection score corresponding to each frame of the online classroom data set is obtained based on CatBoost.
5. The style perturbation privacy-preserving online classroom attentiveness monitoring method of claim 4, wherein, Based on the online classroom data set, a visual saliency map is obtained, specifically: Multi-channel features of the online classroom data set are extracted and normalized, a graph weight matrix corresponding to each channel is constructed, a balanced distribution corresponding to each channel is determined by normalizing the graph weight matrix, and a visual saliency map is obtained by fusing the balanced distribution.
6. The style perturbation privacy-preserving online classroom attentiveness monitoring method of claim 4, wherein, The statistical features include statistical quantity features, statistical histogram, local shape dimension distribution and Benford law.
7. The style perturbation privacy-preserving online classroom attentiveness monitoring method of claim 1, wherein, The behavior index includes eye width ratio, mouth width ratio, gaze direction vector and head angle; and the facial high-level semantic feature includes eye state and head state.
8. The style perturbation privacy-preserving online classroom attentiveness monitoring method of claim 7, wherein, Building a concentration prediction model, inputting the face key point, the behavior index and the facial high-level semantic feature into the concentration prediction model, and outputting a concentration recognition rate includes: The concentration prediction model comprises at least three inputs, wherein the first input uses ST-GCN to extract the features of facial key points; the second input and the third input both use LSTM to process the behavior indicators and the facial advanced semantic features, forming a complementary feature representation mechanism; and the features of each path are outputted; The feature representations of each path are spliced and fused, are processed through a self-attention mechanism after being weighted, and are outputted through a fully connected layer to obtain the concentration recognition rate.
9. The style perturbation privacy-preserving online classroom attentiveness monitoring method of claim 8, wherein, The behavior indicators and the facial advanced semantic features are processed through LSTM to form a complementary feature representation mechanism, specifically: The learning of the behavior indicators is guided by the facial advanced semantic features, and the facial advanced semantic features are supplemented with detailed information missing based on the behavior indicators.
10. The style perturbation privacy-preserving online classroom attentiveness monitoring method of claim 1, wherein, An association statistical model is established based on the visual privacy protection score, and the visual privacy protection and the concentration recognition are balanced in combination with the face detection rate and the concentration recognition rate, including: An association statistical model is established based on the visual privacy protection score and the data availability, a fitting relationship is obtained, a fitting curved surface is obtained, the online classroom data set to be monitored is predicted through the fitting curved surface, and the visual privacy protection and the concentration recognition are balanced in combination with the face detection rate and the concentration recognition rate.
Citation Information
Cited By
Vehicle end track data desensitization method and device, storage medium and electronic device
CN122065348A