Classroom face recognition system based on deep learning

Through a deep learning-based classroom face recognition system, the CNN model and SVM classifier are used, combined with adaptive filtering processing, the recognition bias caused by the differences in facial features of different races is solved, and the recognition accuracy and image quality are improved.

CN120299076AInactive Publication Date: 2025-07-11ZHEJIANG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510411128.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, due to differences in facial features between different races, deviations or errors occur in facial recognition, especially in classroom environments.

Method used

A classroom facial recognition system based on deep learning is adopted, including a collection module, a preprocessing module, a face recognition and analysis module, a feedback module, a data storage and management module, a system configuration and management module and a permission management module, a CNN model is used to extract face features, combine an SVM classifier for identity and status recognition, and select a one-to-one or many-to-many SVM classification strategy based on race differences, and combine it with adaptive filtering to reduce noise interference.

Benefits of technology

It improves the accuracy of face recognition, reduces recognition errors caused by ethnic differences, and can finely process images under different lighting conditions, retain detailed information, reduce noise interference, and improve image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299076A_ABST
    Figure CN120299076A_ABST
Patent Text Reader

Abstract

The invention discloses a classroom face recognition system based on deep learning. An acquisition module is responsible for acquiring face images of students in a classroom; the preprocessing module is used for carrying out image enhancement, contrast adjustment and self-adaptive threshold filtering processing on the collected face image to obtain a preprocessed image, and the self-adaptive threshold filtering processing is to adjust filtering parameters, reserve image details and reduce noise interference according to image statistical information and illumination conditions; the face recognition and analysis module is used for learning the preprocessed image by using a CNN model, extracting face features and student features, extracting the face features by using convolution kernels of different scales, judging different races according to the face features, classifying the extracted face features and student features by using an SVM classifier, and outputting the classified face features and student features; according to the invention, identification of student identities and student states is realized, and a one-to-one or many-to-many SVM classification strategy is adopted for different types of people existing in a classroom environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face recognition, and in particular, to a classroom face recognition system based on deep learning. Background Art

[0002] Face recognition technology is an important branch in the field of artificial intelligence, involving multiple links such as face image capture, processing, feature extraction, and comparison. Machine learning algorithms play a core role in face recognition. Through learning from a training dataset, machine learning algorithms can automatically extract face features and build a classification model that can distinguish different individuals. With the continuous deepening of educational reform, personalized teaching has become an important trend in the field of education. A classroom face recognition system can capture features such as students' expressions and head postures in real time, providing teachers with information about students' classroom states, thereby helping teachers better understand students' learning situations and formulate personalized teaching plans.

[0003] In the prior art, face recognition technology, especially expression recognition, may encounter challenges among different ethnic groups. This is mainly because there are differences in facial features among different ethnic groups, such as skin color, facial contour, eye shape, etc. These differences may lead to biases or errors in recognition. Therefore, a classroom face recognition system based on deep learning is proposed. Summary of the Invention

[0004] The purpose of the present invention is to solve the drawbacks in the prior art that due to differences in facial features among different ethnic groups, biases or errors occur in face recognition, and to propose a classroom face recognition system based on deep learning.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A classroom face recognition system based on deep learning, comprising: A collection module: responsible for collecting face images of students in the classroom; A preprocessing module: including an image enhancement unit, a contrast adjustment unit, and an adaptive threshold filtering processing unit, which respectively perform image enhancement, contrast adjustment, and adaptive threshold filtering processing on the collected face images to obtain preprocessed images. The adaptive threshold filtering processing adjusts the filtering parameters according to image statistical information and lighting conditions, retains image details, and reduces noise interference; Face recognition and analysis module: Use a CNN model to learn and extract the features of the face and student features from the preprocessed image. The student features include features such as the student's expression and head pose. Use convolutional kernels of different scales to extract face features. At the same time, distinguish different ethnic groups according to the features of the face. Use an SVM classifier to classify the extracted face features and student features to achieve the recognition of student identities and the states of students. For different ethnic groups existing in the classroom environment, adopt a one-to-one or many-to-many SVM classification strategy. If there are multiple ethnic groups in the classroom environment, adopt a many-to-many SVM classification strategy. If there is mainly only one or two ethnic groups in the classroom environment, then adopt a one-to-one classification strategy; Feedback module: Real-time feedback the recognition results to the teacher; Data storage and management module: Responsible for storing the collected face images, preprocessed images, and analysis results, providing data backup and recovery functions to ensure data security; System configuration and management module: Provide a system configuration interface, allowing teachers to adjust system parameters according to actual needs, such as camera angles, recognition thresholds, etc., providing system log and error reporting functions for system maintenance and troubleshooting; Permission management module: Set different permissions for different users (such as teachers, administrators, etc.) to ensure the security and controllability of the system, support functions such as role assignment and permission inheritance, and improve the flexibility and usability of the system.

[0006] The above technical solution further includes: Further, the image enhancement unit adjusts the gray histogram of the image through histogram equalization technology, making the gray value distribution of the output image uniform and enhancing the contrast of the image. Let the input image be , and its gray level range is , represents the total number of gray levels in the image, the number of pixels with gray level is , the total number of pixels is , then the probability density function of gray level is , and the transformation function of histogram equalization is where, is the gray level of the output image, is the transformation function.

[0007] Further, the contrast adjustment unit calculates the local contrast of each pixel and adjusts the gray value of the pixel according to the size of the contrast. Let the input image be , the local contrast is , then the output image of adaptive contrast enhancement is Among them, is the gain function, which is adjusted according to the local contrast When the local contrast is low, the gain function increases the gray value of the output image. When the local contrast is high, the gain function keeps or decreases the gray value of the output image.

[0008] Furthermore, the specific steps of the adaptive threshold filtering processing unit for adaptive threshold filtering processing are as follows: Image input: The image to be processed is used as the input, denoted as , where x and y respectively represent the pixel positions in the image; Statistical information calculation: Statistical analysis is performed on the input image to calculate statistical information such as the gray histogram, mean, and variance of the image; Illumination condition evaluation: According to the statistical information of the image, the illumination condition of the image is evaluated by calculating indicators such as the brightness mean and brightness standard deviation of the image. The evaluation result of the illumination condition is used to adjust the filtering parameters to adapt to different illumination environments; Filtering parameter calculation: According to the statistical information of the image and the evaluation result of the illumination condition, the filtering parameters are calculated; Filtering operation: The input image is filtered using the calculated filtering parameters. The filtering operation is expressed as , where represents the output image after filtering, represents the filtering function, represents the filtering parameters; Output image: After the filtering operation, the processed image is obtained. While retaining the details of the image, this image effectively reduces the interference of noise.

[0009] Furthermore, the specific steps of the filtering parameter calculation are as follows: Determine the filter type: Select the Gaussian filter; Calculate the filter size: The size of the filter (i.e., the window size) is determined according to the resolution of the image and the degree of detail retention. Set the size of the filter to . A larger filter can smooth more details, but may also cause image blurring; a smaller filter can better retain the details of the image, but the smoothing effect may not be obvious; Set the filter weights: The Gaussian filter weight calculation formula is , where represents the pixel position, represents the standard deviation of the Gaussian function, which determines the smoothing degree of the filter; Adjust parameters in combination with lighting conditions: Fine-tune the filtering parameters according to the evaluation result of the lighting conditions of the image. For example, in darker areas, it may be necessary to increase the weight or window size of the filter to enhance the contrast of the image; in brighter areas, it may be necessary to reduce these parameters to avoid loss of details due to over-smoothing.

[0010] Further, the face recognition and analysis module includes a feature extraction unit, a race discrimination unit, a feature classification unit, and a strategy selection unit. The feature extraction unit uses a CNN model to extract features from the preprocessed image. The features include the deep features of the face (such as the shapes and positions of parts like eyes, nose, mouth, etc.) and features such as the expressions and head postures of the students. The race discrimination unit discriminates different races according to the extracted face features. The feature classification unit uses an SVM classifier to classify the extracted face features and student features (such as expressions, head postures, etc.) to achieve the recognition of student identities and student states. The strategy selection unit selects an SVM classification strategy, including one-versus-one or many-versus-many, for different races present in the classroom environment. The feature extraction unit passes the face features to the race discrimination unit, the feature extraction unit passes the face features and student features to the feature classification unit, the race discrimination unit passes the discrimination result to the strategy selection unit, and the strategy selection unit selects an SVM classification strategy according to the discrimination result and passes the selected SVM classification strategy to the feature classification unit.

[0011] Further, the specific steps for the feature extraction unit to use a CNN model to extract features from the preprocessed image are as follows: Image preprocessing: Preprocess the input image, including operations such as grayscale conversion, denoising, and normalization, to improve the accuracy and efficiency of subsequent feature extraction; Convolution operation: Use convolution kernels of different scales to perform convolution operations on the preprocessed image. The size and number of convolution kernels will affect the effect of feature extraction. Generally, smaller convolution kernels can capture the detailed features in the image, while larger convolution kernels can capture more global features. The formula for the convolution operation is , where represents the pixel value of the convolved image at the position, represents the input image, represents the convolution kernel, and respectively represent the row and column indices of the convolution kernel;

[0012] Activation function: After the convolution operation, use ReLU to increase the nonlinearity of the network and improve the expressive ability of the model. The ReLU formula is where Denotes the output after the convolution operation, Denotes the output after passing through the ReLU activation function; Pooling operation: Max-pooling is used to further reduce the dimension of the features, reduce the computational load, and at the same time improve the robustness of the model. The formula for max-pooling is , where Denotes the pixel value of the image after max-pooling at the position, Denotes the pooling window corresponding to the position, Denotes the input image; Multi-layer convolution and pooling: Through multi-layer convolution and pooling operations, high-level feature representations are gradually abstracted. Each layer will extract more abstract and complex features, which are crucial for subsequent classification and recognition tasks; Feature extraction: After multi-layer convolution and pooling operations, features of the human face (such as the shapes and positions of parts like eyes, nose, mouth, etc.) and features such as the expression and head pose of the student are obtained.

[0013] Further, the strategy selection unit selects the SVM classification strategy based on the output result of the ethnic discrimination unit. If there are multiple ethnic groups in the classroom environment, a one-versus-one SVM classification strategy is adopted, and an SVM classifier is trained for each pair of ethnic groups. A total of classifiers need to be trained, where is the number of ethnic groups, Denotes the number of ways to choose any two from ethnic groups for combination. In the classification stage, for each face to be classified, all trained SVM classifiers are used for classification, and the final classification result is determined through a voting mechanism; if there is mainly only one or a few ethnic groups in the classroom environment, a one-versus-one classification strategy is adopted, and an SVM classifier is trained for each ethnic group against other ethnic groups respectively. A total of n classifiers need to be trained, where n is the number of ethnic groups. In the classification stage, for each face to be classified, all trained SVM classifiers are used for classification, and the ethnic group corresponding to the classifier with the highest score is selected as the final classification result.

[0014] The present invention has the following beneficial effects: In the present invention, a CNN model is used to learn and extract the features of human faces and student features from the preprocessed images. According to the features of human faces, different ethnic groups are discriminated. If there are multiple ethnic groups in the classroom environment, a many-to-many SVM classification strategy is adopted. If there is mainly only one or two ethnic groups in the classroom environment, a one-to-one classification strategy is adopted. It can select a suitable classification method according to the facial feature differences between different ethnic groups, thereby improving the recognition accuracy and avoiding recognition errors caused by ethnic differences. At the same time, by intelligently adjusting the filtering parameters, the image can be refined according to different lighting conditions and image statistical information. This filtering method can effectively retain the detailed information in the image, such as edges, textures, etc., while reducing the interference of noise, thereby improving the overall quality of the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a system block diagram of a classroom face recognition system based on deep learning proposed by the present invention; Figure 2 It is a system block diagram of the preprocessing module in the present invention; Figure 3 It is a system block diagram of the face recognition and analysis module in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present invention.

[0017] Please refer to Figures 1-3 As shown, the present invention is a classroom face recognition system based on deep learning, including: Collection module: responsible for collecting the face images of students in the classroom; Preprocessing module: including an image enhancement unit, a contrast adjustment unit, and an adaptive threshold filtering processing unit, which respectively perform image enhancement, contrast adjustment, and adaptive threshold filtering processing on the collected face images to obtain preprocessed images. The adaptive threshold filtering processing is to adjust the filtering parameters according to the image statistical information and lighting conditions, retain the image details, and reduce the noise interference; Face recognition and analysis module: Use a CNN model to learn and extract the features of the face and student features from the preprocessed image. The student features include features such as the student's expression and head pose. Use convolutional kernels of different scales to extract face features. At the same time, distinguish different ethnic groups according to the features of the face. Use an SVM classifier to classify the extracted face features and student features to achieve the recognition of student identities and student statuses. For different ethnic groups in the classroom environment, adopt a one-to-one or many-to-many SVM classification strategy. If there are multiple ethnic groups in the classroom environment, adopt a many-to-many SVM classification strategy. If there is mainly only one or two ethnic groups in the classroom environment, then adopt a one-to-one classification strategy; Feedback module: Real-time feedback the recognition results to the teacher; Data storage and management module: Responsible for storing the collected face images, preprocessed images, and analysis results, providing data backup and recovery functions to ensure data security; System configuration and management module: Provide a system configuration interface that allows teachers to adjust system parameters according to actual needs, such as camera angle, recognition threshold, etc., and provide system log and error reporting functions for system maintenance and troubleshooting; Permission management module: Set different permissions for different users (such as teachers, administrators, etc.) to ensure the security and controllability of the system, support functions such as role assignment and permission inheritance, and improve the flexibility and usability of the system.

[0018] In one embodiment, for the above image enhancement, the image enhancement unit adjusts the gray histogram of the image through histogram equalization technology, making the gray value distribution of the output image uniform and enhancing the contrast of the image. Let the input image be , and its gray level range is , represents the total number of gray levels in the image, and the number of pixels with gray level is , and the total number of pixels is , then the probability density function of gray level is , and the transformation function of histogram equalization is where, is the gray level of the output image, is the transformation function.

[0019] In one embodiment, for the above contrast adjustment, the contrast adjustment unit calculates the local contrast of each pixel and adjusts the gray value of the pixel according to the size of the contrast. Let the input image be , and the local contrast is , then the output image of adaptive contrast enhancement is Among them, is the gain function, which is adjusted according to the local contrast When the local contrast is low, the gain function increases the gray value of the output image. When the local contrast is high, the gain function keeps or reduces the gray value of the output image.

[0020] In one embodiment, for the above adaptive threshold filtering process, the specific steps of the adaptive threshold filtering process by the adaptive threshold filtering unit are as follows: Image input: The image to be processed is taken as the input, denoted as , where x and y respectively represent the pixel positions in the image; Statistical information calculation: Perform statistical analysis on the input image to calculate statistical information such as the gray histogram, mean, and variance of the image; Illumination condition evaluation: According to the statistical information of the image, evaluate the illumination condition of the image by calculating indicators such as the brightness mean and brightness standard deviation of the image. The evaluation result of the illumination condition is used to adjust the filtering parameters to adapt to different illumination environments; Filtering parameter calculation: Calculate the filtering parameters according to the statistical information of the image and the evaluation result of the illumination condition; Filtering operation: Use the calculated filtering parameters to perform a filtering operation on the input image. The filtering operation is expressed as , where represents the filtered output image, represents the filtering function, represents the filtering parameters; Output image: After the filtering operation, the processed image is obtained. While retaining the details of the image, this image effectively reduces the interference of noise.

[0021] In one embodiment, for the above filtering parameter calculation, the specific steps of the filtering parameter calculation are as follows: Determine the filter type: Select a Gaussian filter; Calculate the filter size: The size of the filter (i.e., the window size) is determined according to the resolution of the image and the degree of detail retention. Set the size of the filter to , A larger filter can smooth more details, but may also cause image blurring; a smaller filter can better retain the details of the image, but the smoothing effect may not be obvious; Set the filter weights: The Gaussian filter weight calculation formula is , where, represents the pixel position, represents the standard deviation of the Gaussian function, which determines the smoothing degree of the filter; Adjust parameters in combination with lighting conditions: Based on the evaluation result of the lighting conditions of the image, fine-tune the filtering parameters. For example, in areas with darker lighting, it may be necessary to increase the weight or window size of the filter to enhance the contrast of the image; in areas with brighter lighting, these parameters may need to be reduced to avoid loss of details due to excessive smoothing.

[0022] Suppose there is an image taken under uneven lighting conditions that needs to be processed by adaptive threshold filtering. The following is the specific parameter calculation process:

[0023] Select the filter type: According to the characteristics of the image, select the Gaussian filter as the filtering tool because it can retain better edge information while smoothing the image.

[0024] Calculate the filter size: Based on the resolution of the image and the requirement for retaining details, set the size of the filter to 3×3. This size can smooth small noises in the image and retain the details of the image well.

[0025] Set the filter weights: According to the Gaussian filter weight calculation formula, calculate the weight value at each pixel position. For example, for the center point (0,0), its weight value is G(0,0,σ); for other positions (x,y), calculate the corresponding weight value according to the formula.

[0026] Adjust parameters in combination with lighting conditions: Evaluate the lighting conditions of the image and find that the lighting is darker in the upper left corner area and brighter in the lower right corner area of the image. Therefore, increase the weight or window size of the filter in the upper left corner area (for example, adjust the filter size to 5×5) to enhance the contrast of the image; in the lower right corner area, reduce these parameters (for example, keep the filter size at 3×3) to avoid loss of details due to excessive smoothing.

[0027] In one embodiment, for the above-mentioned face recognition and analysis module, the face recognition and analysis module includes a feature extraction unit, a race discrimination unit, a feature classification unit, and a strategy selection unit. The feature extraction unit uses a CNN model to extract features from the preprocessed image. The features include the deep features of the face (such as the shapes and positions of parts like eyes, nose, mouth, etc.) and features such as the expressions and head postures of the students. The race discrimination unit discriminates different races based on the extracted face features. The feature classification unit uses an SVM classifier to classify the extracted face features and student features (such as expressions, head postures, etc.) to achieve the recognition of student identities and student states. The strategy selection unit selects an SVM classification strategy, including one-versus-one or many-versus-many, for different races existing in the classroom environment. The feature extraction unit transfers the face features to the race discrimination unit, the feature extraction unit transfers the face features and student features to the feature classification unit, the race discrimination unit transfers the discrimination result to the strategy selection unit, and the strategy selection unit selects an SVM classification strategy according to the discrimination result and transfers the selected SVM classification strategy to the feature classification unit.

[0028] In one embodiment, for the above-mentioned feature extraction unit, the specific steps for the feature extraction unit to use a CNN model to extract features from the preprocessed image are as follows: Image preprocessing: Preprocess the input image, including operations such as grayscale conversion, denoising, and normalization, to improve the accuracy and efficiency of subsequent feature extraction; Convolution operation: Use convolution kernels of different scales to perform convolution operations on the preprocessed image. The size and number of convolution kernels will affect the effect of feature extraction. Generally, smaller convolution kernels can capture the detailed features in the image, while larger convolution kernels can capture more global features. The formula for the convolution operation is , where represents the pixel value of the convolved image at the position, represents the input image, represents the convolution kernel, and respectively represent the row and column indices of the convolution kernel;

[0029] Activation function: After the convolution operation, use ReLU to increase the non-linearity of the network and improve the expression ability of the model. The formula for ReLU is where represents the output after the convolution operation, represents the output after passing through the ReLU activation function; Pooling operation: Use max pooling to further reduce the dimension of the features, reduce the computational amount, and improve the robustness of the model at the same time. The formula for max pooling is , where represents the pixel value of the image after max - pooling at the position, represents the pooling window corresponding to the position; Multi - layer convolution and pooling: Through multi - layer convolution and pooling operations, high - level feature representations are gradually abstracted. Each layer extracts more abstract and complex features, which are crucial for subsequent classification and recognition tasks; Feature extraction: After multi - layer convolution and pooling operations, features of the face (such as the shapes and positions of parts like eyes, nose, mouth, etc.) and features of the student's expression, head pose, etc. are obtained.

[0030] Suppose there is a pre - processed student face image, and a feature extraction unit is used to extract its face features and expression features.

[0031] Convolution operation: Select 3 convolution kernels of different scales (such as 3x3, 5x5, 7x7) to perform convolution operations on the image. Each convolution kernel generates a feature map, and these feature maps contain feature information of different scales and positions in the image.

[0032] Activation function: After the convolution operation, the ReLU activation function is used to increase the non - linearity of the network. The ReLU activation function sets all negative values to 0 and retains all positive values, thus increasing the sparsity of the network.

[0033] Pooling operation: Select the max - pooling operation to reduce the dimension of the features. Within each pooling window, the maximum value is selected as the output of the window, thus retaining the most important feature information.

[0034] Multi - layer convolution and pooling: Repeat the convolution and pooling operations multiple times to gradually abstract higher - level feature representations. Each layer extracts more abstract and complex features, which are crucial for subsequent classification and recognition tasks.

[0035] Feature extraction: After multi - layer convolution and pooling operations, deep features of the face (such as the shapes and positions of parts like eyes, nose, mouth, etc.) and expression features of the student (such as smiling, being surprised, etc.) can be obtained. These features can be used for subsequent classification and recognition tasks, such as student identity recognition, expression recognition, etc.

[0036] In one embodiment, for the above-mentioned policy selection unit, the policy selection unit selects the SVM classification policy based on the output result of the race discrimination unit. If there are multiple races in the classroom environment, the one-against-one SVM classification policy is adopted, and an SVM classifier is trained for each pair of races. A total of classifiers are required, where is the number of races, represents the number of ways to select any two from races for combination. In the classification stage, for each face to be classified, all the trained SVM classifiers are used for classification, and the final classification result is determined through a voting mechanism; if there is mainly only one or a few races in the classroom environment, the one-against-one classification policy is adopted, and an SVM classifier is trained for each race and the other races respectively. A total of n classifiers are required, where n is the number of races. In the classification stage, for each face to be classified, all the trained SVM classifiers are used for classification, and the race corresponding to the classifier with the highest score is selected as the final classification result.

[0037] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A classroom face recognition system based on deep learning, characterized in that, Including: Collection module: Responsible for collecting the face images of students in the classroom; Preprocessing module: Includes an image enhancement unit, a contrast adjustment unit, and an adaptive threshold filtering processing unit, which are respectively responsible for performing image enhancement, contrast adjustment, and adaptive threshold filtering processing on the collected face images to obtain preprocessed images. The adaptive threshold filtering processing adjusts the filtering parameters according to the image statistical information and the lighting conditions; Face recognition and analysis module: Uses a CNN model to learn and extract the features of the face and student features from the preprocessed images, extracts face features using convolutional kernels of different scales. At the same time, discriminates different ethnic groups according to the features of the face, and uses an SVM classifier to classify the extracted face features and student features to achieve the recognition of student identities and student states. For different ethnic groups existing in the classroom environment, an one-to-one or many-to-many SVM classification strategy is adopted. If there are multiple ethnic groups in the classroom environment, a many-to-many SVM classification strategy is adopted. If there is mainly only one or two ethnic groups in the classroom environment, an one-to-one classification strategy is adopted; Feedback module: Real-time feedbacks the recognition results to the teacher; Data storage and management module: Responsible for storing the collected face images, preprocessed images, and analysis results; System configuration and management module: Provides a system configuration interface, allows the teacher to adjust system parameters according to actual needs, and provides system log and error reporting functions; Permission management module: Sets different permissions for different users.

2. The classroom face recognition system based on deep learning according to claim 1, wherein, The image enhancement unit adjusts the grayscale histogram of the image through histogram equalization technology, making the grayscale value distribution of the output image uniform and enhancing the contrast of the image. Let the input image be , and its grayscale range is , represents the total number of grayscale levels in the image. The number of pixels with grayscale level is , and the total number of pixels is . Then the probability density function of grayscale level is , and the transformation function of histogram equalization is Among them, is the gray level of the output image, is the transformation function.

3. The classroom face recognition system based on deep learning according to claim 1, characterized in that, The contrast adjustment unit calculates the local contrast of each pixel and adjusts the gray value of the pixel according to the magnitude of the contrast. Let the input image be , the local contrast be , then the output image of the adaptive contrast enhancement is , where is a gain function that is adjusted according to the magnitude of the local contrast . When the local contrast is low, the gain function increases the gray value of the output image. When the local contrast is high, the gain function maintains or decreases the gray value of the output image.

4. The classroom face recognition system based on deep learning according to claim 1, characterized in that, The specific steps of the adaptive threshold filtering processing unit for adaptive threshold filtering processing are as follows: Image input: The image to be processed is used as the input, denoted as , where x and y respectively represent the pixel positions in the image; Statistical information calculation: Performs statistical analysis on the input image and calculates the statistical information of the image; Lighting condition evaluation: Evaluates the lighting condition of the image by calculating the metrics of the image according to the statistical information of the image. The evaluation result of the lighting condition is used to adjust the filtering parameters; Filtering parameter calculation: Calculates the filtering parameters according to the statistical information of the image and the evaluation result of the lighting condition; Filtering operation: Perform a filtering operation on the input image using the calculated filtering parameters. The filtering operation is expressed as , where represents the filtered output image, represents the filtering function, represents the filtering parameters; Output image: After the filtering operation, the processed image is obtained .

5. The classroom face recognition system based on deep learning according to claim 4, characterized in that, The specific steps of the filtering parameter calculation are as follows: Determine the filter type: Select a Gaussian filter; Calculate the filter size: The size of the filter is determined according to the resolution of the image and the degree of detail retention, and the size of the filter is set to be ; Set filter weights: The calculation formula for Gaussian filter weights is , where represents the pixel position, represents the standard deviation of the Gaussian function, which determines the smoothness of the filter; Adjust the parameters in combination with the lighting condition: Fine-tune the filtering parameters according to the evaluation result of the lighting condition of the image.

6. The classroom face recognition system based on deep learning according to claim 1, characterized in that, The face recognition and analysis module includes a feature extraction unit, a racial discrimination unit, a feature classification unit, and a strategy selection unit. The feature extraction unit uses a CNN model to extract features from the preprocessed image. The racial discrimination unit discriminates different races according to the extracted face features. The feature classification unit uses an SVM classifier to classify the extracted face features and student features to achieve the recognition of student identities and student statuses. The strategy selection unit selects an SVM classification strategy, including one-to-one or many-to-many, for different races existing in the classroom environment. The feature extraction unit transmits the face features to the racial discrimination unit. The feature extraction unit transmits the face features and student features to the feature classification unit. The racial discrimination unit transmits the discrimination result to the strategy selection unit. The strategy selection unit selects an SVM classification strategy according to the discrimination result and transmits the selected SVM classification strategy to the feature classification unit.

7. The classroom face recognition system based on deep learning according to claim 6, characterized in that, The specific steps for the feature extraction unit to extract features from the preprocessed image using a CNN model are as follows: Image preprocessing: Preprocess the input image. Convolution operation: Use convolution kernels of different scales to perform convolution operations on the preprocessed image. The formula for the convolution operation is , where represents the pixel value of the convolved image at the position, represents the input image, represents the convolution kernel, and represent the row and column indices of the convolution kernel respectively; Activation function: After the convolution operation, ReLU is used to increase the non-linearity of the network. The ReLU formula is where represents the output after the convolution operation, represents the output after passing through the ReLU activation function; Pooling operation: Use max pooling to further reduce the dimension of the features. The formula for max pooling is , where represents the pixel value of the image after max pooling at the position, represents the pooling window corresponding to the position, and represents the input image; Multi-layer convolution and pooling: Gradually abstract high-level feature representations through multi-layer convolution and pooling operations. Feature extraction: After the multi-layer convolution and pooling operations, obtain the face features and student features.

8. The classroom face recognition system based on deep learning according to claim 6, characterized in that, The strategy selection unit selects the SVM classification strategy based on the output result of the race discrimination unit. If there are multiple races in the classroom environment, the one - against - one SVM classification strategy is adopted, and an SVM classifier is trained for each pair of races, and a total of classifiers are required, where is the number of races, represents the number of ways to choose any two from races for combination. In the classification stage, for each face to be classified, all the trained SVM classifiers are used for classification, and the final classification result is determined through a voting mechanism; if there is mainly only one or a few races in the classroom environment, the one - against - one classification strategy is adopted, and an SVM classifier is trained for each race and the other races respectively, and a total of n classifiers are required, where n is the number of races. In the classification stage, for each face to be classified, all the trained SVM classifiers are used for classification, and the race corresponding to the classifier with the highest score is selected as the final classification result.