A multi-perspective teaching expression recognition method and system in a classroom environment
By collecting multi-view image sequences and temporal attention networks, combined with long-wave infrared images, the accuracy of teacher expression recognition in classroom environments is solved, and adaptability to light and angle changes are enhanced and the meticulousness of expression recognition is improved.
Patent Information
- Application Number
- CN202111359859.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-17
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-11-17
AI Technical Summary
In the prior art, in the recognition of facial expressions by teachers in classroom environments, facial images are unclear, light changes affect, expression uncertainty and differences, resulting in insufficient recognition accuracy.
Multi-view image sequences are collected, and multi-view expression recognition is performed through self-similarity matrix extraction module, feature extractor, time attention network and classification network, combined with long-wave infrared images, multi-view expression recognition is carried out, shallow and deep features are integrated, and the time attention mechanism is used to deal with expression changes.
It improves the accuracy and robustness of teacher expression recognition, can effectively deal with the influence of light changes and angle changes, and enhances the meticulousness and accuracy of expression recognition.
Smart Images

Figure CN114120403B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pattern recognition, and more specifically, relates to a multi-view teaching expression recognition method and system in a classroom environment. Background Art
[0002] The teaching emotional state of a teacher is an important means to measure the enthusiasm of the teacher in the classroom, and can provide more intelligent services for the intelligentization of classroom teaching evaluation. Facial expression is one of the most common signals for humans to express inner emotions and intentions, and the facial expression of a teacher is an effective manifestation form of the teacher's teaching emotional state. Applying facial expression recognition technology to the detection of the teacher's teaching state ensures the fairness and accuracy of the evaluation to a certain extent. The difficulties in facial expression recognition are as follows: ① The facial images are not clear, and the image pixels are low due to the influence of other factors (such as lighting, shooting angle, etc.); ② The facial expressions in the facial images have great uncertainty; ③ The differences between the same expressions and the similarities between different expressions greatly affect the recognition performance of the model, etc.
[0003] Most traditional expression recognition methods are based on the convolutional neural network model of 2D images. The basic process of this model method is as follows: ① Directly input the picture for 2D convolution processing, and optimize the parameters in the convolution through continuous training; ② After passing through the convolutional layer, enter the max pooling layer and the global normalization layer; ③ Obtain the final expression recognition result, calculate the loss between the predicted value and the true value, and perform backpropagation.
[0004] However, the limitations of this type of traditional method are mainly reflected in two aspects. First, the image is directly input into the convolution module, which destroys the internal information of the image and cannot accurately capture the shallow information of the image. Second, the lighting changes in the learning environment, whether the lighting is too strong or too weak, will cause the loss of facial details and sometimes generate shadows, affecting the recognition result. Summary of the Invention
[0005] Aiming at at least one defect or improvement requirement of the prior art, the present invention provides a multi-view teaching expression recognition method and system in a classroom environment, which can improve the accuracy of teacher expression recognition in a classroom environment.
[0006] To achieve the above object, according to the first aspect of the present invention, there is provided a multi-view teaching expression recognition method in a classroom environment, including the steps of:
[0007] S1, collecting a multi-view image sequence of the user's expression, where the multi-view image sequence includes image sequences of multiple types of views;
[0008] S2. Input the multi-view image sequence into an expression recognition model to obtain the user expression recognition result. The expression recognition model includes a self-similarity matrix extraction module, a feature extractor, a temporal attention network, a fusion and splicing module, and a classification network. The self-similarity matrix extraction module is used to extract the self-similarity matrix of each type of view from the multi-view image sequence. The feature extractor is used to obtain the first type of feature vector of each type of view according to the self-similarity matrix of each type of view. The temporal attention network is used to obtain the second type of feature vector of each type of view according to the self-similarity matrix of each type of view. The fusion and splicing module is used to fuse and splice the two types of feature vectors and then input them into the classification network.
[0009] Preferably, before the S2, the method further includes the step of: preprocessing the multi-view image sequence, and the preprocessing includes:
[0010] Convert the images in the multi-view image sequence into PIL images, then adjust the size of the PIL images to 224×224 pixels, and then convert the PIL images into tensor images, and normalize the tensor images with the mean value and standard deviation.
[0011] Preferably, after the S2, the method further includes the step of:
[0012] S3. Obtain the total number of user expression recognition results and the number of expression changes within a preset time period, calculate the frequency of user classroom expression changes, and draw a line chart of user emotional changes.
[0013] Preferably, the multi-view image sequence includes three types, namely: the left-view image sequence, the front-view image sequence, and the right-view image sequence.
[0014] Preferably, the multi-view image sequence is a long-wave infrared image sequence. The left-view image sequence includes multiple image sequences at multiple angles belonging to the left view, and the right-view image sequence includes multiple image sequences at multiple angles belonging to the right view.
[0015] Preferably, the S2 includes sub-steps:
[0016] Extract the first self-similarity matrix of the left-view image sequence, the second self-similarity matrix of the front-view image sequence, and the third self-similarity matrix of the right-view image sequence respectively;
[0017] Input the first self-similarity matrix, the second self-similarity matrix, and the third self-similarity matrix into the feature extractor respectively to obtain the first feature vector, the second feature vector, and the third feature vector respectively;
[0018] Input the first self-similar matrix, the second self-similar matrix, and the third self-similar matrix into the temporal attention network respectively to obtain a fourth eigenvector, a fifth eigenvector, and a sixth eigenvector respectively;
[0019] Calculate the tensor product of the first eigenvector and the fourth eigenvector to obtain a left-view fusion eigenvector, calculate the tensor product of the second eigenvector and the fifth eigenvector to obtain a front-view fusion eigenvector, and calculate the tensor product of the third eigenvector and the sixth eigenvector to obtain a right-view fusion eigenvector;
[0020] Concatenate the left-view fusion eigenvector, the front-view fusion eigenvector, and the right-view fusion eigenvector and input them into the classification network to obtain the user expression recognition result.
[0021] Preferably, the feature extractor is a 3D feature extractor.
[0022] Preferably, the 3D feature extractor includes two convolutional modules, which are connected by a 3D average pooling layer. Each convolutional module includes a 3D convolutional layer, a Rule layer, two 3D convolutional layers, a Rule layer, a 3D max pooling layer, two 3D convolutional layers, and a 3D max pooling layer connected in sequence.
[0023] Preferably, the temporal attention network includes a fully connected layer, a Rule layer, a fully connected layer, a Rule layer, a fully connected layer, a sigmoid layer, and a repeat layer connected in sequence.
[0024] According to the second aspect of the present invention, a multi-view teaching expression recognition system in a classroom environment is provided, including:
[0025] A data acquisition module for acquiring a multi-view image sequence of a user's expression, where the multi-view image sequence includes image sequences of multiple types of views;
[0026] An output module for inputting the multi-view image sequence into an expression recognition model to obtain a user expression recognition result; the expression recognition model includes a self-similar matrix extraction module, a feature extractor, a temporal attention network, a fusion and concatenation module, and a classification network; the self-similar matrix extraction module is used to extract the self-similar matrix of each type of view from the multi-view image sequence; the feature extractor is used to obtain the first type of eigenvector of each type of view; the temporal attention network is used to obtain the second type of eigenvector of each type of view; the fusion and concatenation module is used to fuse and concatenate the two types of eigenvectors and then input them into the classification network.
[0027] Generally speaking, compared with the prior art, the present invention has beneficial effects:
[0028] (1) The present invention collects expression image sequences from multiple different perspectives. By extracting the features of image sequences from different perspectives and fusing auxiliary information such as the features at various positions of the same person, more detailed feature information is obtained to assist in expression recognition, greatly improving the accuracy of expression recognition.
[0029] (2) In order to effectively address the problems caused by the changing teacher expressions over time in the teacher teaching emotion detection task, a temporal attention mechanism is incorporated. At the same time, the image is subjected to two feature extractions using a self-similarity matrix and a feature extractor, and the shallow feature extraction and deep feature extraction are fused to overcome the problem that traditional convolutional neural networks cannot simultaneously balance the depth of convolution and the accuracy of recognition.
[0030] (3) The collected images are long-wave infrared images, which can obtain more expression detail information and effectively cope with the influence of different light changes in the classroom and the angular changes caused by teacher activities. Brief Description of the Drawings
[0031] Figure 1 is a schematic diagram of the principle of the multi-perspective teaching expression recognition method in the classroom environment according to an embodiment of the present invention;
[0032] Figure 2 is a simulation diagram of the teacher's classroom scene according to an embodiment of the present invention;
[0033] Figure 3 is a network diagram of the expression recognition model according to an embodiment of the present invention. Detailed Embodiments
[0034] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0035] As Figure 1 shown, the long-wave infrared multi-perspective teaching expression recognition method in the classroom environment according to an embodiment of the present invention includes the steps:
[0036] S1, collecting multi-perspective image sequences of the user's expression, where the multi-perspective image sequences include image sequences of multiple types of perspectives.
[0037] Further, the user can be a teacher.
[0038] Further, the multi-perspective image sequences are long-wave infrared image sequences.
[0039] Further, the multi-view image sequence includes three categories, namely: the left-view image sequence, the front-view image sequence, and the right-view image sequence.
[0040] Further, the left-view image sequence includes multiple image sequences at multiple angles belonging to the left view, and the right-view image sequence includes multiple image sequences at multiple angles belonging to the right view.
[0041] In one embodiment, the multi-view image sequence teacher includes: (1) the left-view image sequence, including the teacher's upper left 45° (L_45) and the teacher's left-side infrared image (L_90); (2) the teacher's front infrared image (Fro); (3) the right-view image sequence, including the right-side infrared image (R_90) and the teacher's upper right 45° infrared image (R_45).
[0042] As Figure 2 shown, long-wave infrared cameras are installed at various positions in the classroom. A total of 10 cameras are placed around the classroom and in the first row of students facing the podium. During class, each camera collects the teacher's classroom status in real time and transmits the data to the background server for storage.
[0043] The background server selects clear facial images of the teacher in the classroom from the transmitted image data according to time, and respectively selects the teacher's right-side infrared image (R_90), the teacher's upper right 45° infrared image (R_45), the teacher's front infrared image (Fro), the teacher's upper left 45° infrared image (L_45), and the teacher's left-side infrared image (L_90), and saves these multi-view data of the classroom teacher to the database.
[0044] S2. Input the multi-view image sequence into the facial expression recognition model to obtain the user facial expression recognition result; the facial expression recognition model includes a self-similarity matrix extraction module, a feature extractor, a temporal attention network, a fusion and splicing module, and a classification network; the self-similarity matrix extraction module is used to extract the self-similarity matrix of each type of view from the multi-view image sequence; the feature extractor is used to obtain the first type of feature vector of each type of view according to the self-similarity matrix of each type of view; the temporal attention network is used to obtain the second type of feature vector of each type of view according to the self-similarity matrix of each type of view; the fusion and splicing module is used to fuse and splice the two types of feature vectors and then input them into the classification network.
[0045] The self-similarity matrix is a matrix representing the difference between two images.
[0046] Further, before S2, it further includes the steps of: preprocessing the multi-view image sequence, and the preprocessing includes:
[0047] Convert the images of the multi-view image sequence into PIL images, resize the input PIL images to a size of 224×224, and convert them into tensor images (RImg_90, RImg_45, FroImg, LImg_45, LImg_90). Normalize the tensor images with the mean (mean = [0.485, 0.456, 0.406]) and standard deviation (std = [0.229, 0.224, 0.225]) calculated by randomly sampling millions of images from the dataset to obtain the processed facial images of each view (R_90, R_45, Fro, L_45, L_90).
[0048] The formula description is as follows:
[0049] R_90 = (RImg_90 – mean) / std
[0050] R_45 = (RImg_45 – mean) / std
[0051] Fro = (Fro Img – mean) / std
[0052] L_45 = (LImg_45 – mean) / std
[0053] L_90 = (LImg_90 – mean) / std
[0054] Input the preprocessed multi-view images into the facial expression recognition model for facial expression recognition.
[0055] The training process of the facial expression recognition model is as follows.
[0056] Preparation of samples. Locate the facial area in the image, crop the infrared image of the facial area, perform horizontal and vertical flips on the obtained facial images, and add Gaussian noise to the images, which can increase the effective samples and enhance the learning ability of the model. Then perform the same actions as the above preprocessing on the samples.
[0057] Divide the sample data into a training set and a test set in a ratio of 7:3.
[0058] Select the Adam optimizer and use cross-entropy as the loss function to assist the model in better learning image features.
[0059] Perform cyclic training, input the data into the learning model for forward propagation, calculate the loss, perform backpropagation, and continuously update the model parameters.
[0060] The cross-entropy loss calculation process is described as follows:
[0061]
[0062] Where \(L\) represents the loss, \(N\) is the number of samples, and \(M\) represents the number of expressions. In one embodiment, \(M = 20\), and \(y\) ic is the sign function of the true label. If the true class of sample \(i\) is equal to \(c\), then \(y\) ic = 1; otherwise, it takes 0. represents the predicted probability that the observed sample \(i\) belongs to class \(c\).
[0063] In one embodiment, the expression recognition model can recognize 20 expressions: Happy: label \(Ha\), Surprised: label \(Su\), Neutral: label \(Ne\), Sad: label \(Sa\), Fear: label \(Fe\), Disgust: label \(Di\), Angry: label \(An\), Happy and Sad: label \(Ha\_sa\), Happy and Surprised: label \(Ha\_su\), Happy and Disgusted: label \(Ha\_di\), Sad and Fearful: label \(Sa\_fe\), Sad and Angry: label \(Sa\_an\), Sad with Surprise: label \(Sa\_su\), Sad and Disgusted: label \(Sa\_di\), Fearful and Angry: label \(Fe\_an\), Fearful and Surprised: label \(Fe\_su\), Angry and Surprised: label \(An\_su\), Happy and Fearful: label \(Ha\_fe\), Angry and Disgusted: label \(An\_di\), Awe: label \(Aw\).
[0064] Further, when the input multi-view image sequence of the expression recognition model includes a left-view image sequence, a front-view image sequence, and a right-view image sequence, S2 includes sub-steps:
[0065] Extract the first self-similarity matrix of the left-view image sequence, the second self-similarity matrix of the front-view image sequence, and the third self-similarity matrix of the right-view image sequence respectively;
[0066] Input the first self-similarity matrix, the second self-similarity matrix, and the third self-similarity matrix into the feature extractor respectively to obtain the first feature vector, the second feature vector, and the third feature vector;
[0067] Input the first self-similarity matrix, the second self-similarity matrix, and the third self-similarity matrix into the temporal attention network respectively to obtain the fourth feature vector, the fifth feature vector, and the sixth feature vector;
[0068] Calculate the tensor product of the first feature vector and the fourth feature vector to obtain the left-view fusion feature vector, calculate the tensor product of the second feature vector and the fifth feature vector to obtain the front-view fusion feature vector, and calculate the tensor product of the third feature vector and the sixth feature vector to obtain the right-view fusion feature vector;
[0069] Concatenate the left-view fusion feature vector, the front-view fusion feature vector, and the right-view fusion feature vector and input them into the classification network to obtain the user expression recognition result.
[0070] Such as Figure 3As shown, taking the input of the facial expression recognition model as the pre - processed facial images from each perspective (R_90, R_45, Fro, L_45, L_90) and the output as the teacher's classroom emotional state at a certain moment (one of 20 expressions) as an example for illustration.
[0071] Take the left - hand facial image (L_90) and the upper - left 45 - degree facial image (L_45) as the first group, the frontal facial image (Fro) as the second group, and the right - hand facial image (R_90) and the upper - right 45 - degree facial image (R_45) as the third group.
[0072] Input the first group and the third group into the pairwise joint graph respectively, and input the second group into the self - joint graph. The joint graph measures the difference of facial images according to the arbitrary distance of relative pixels in the image, and uses the standard method in metric learning to learn the Mahalanobis metric. The Mahalanobis metric method has good generalization ability and flexibility, and the Mahalanobis distance is invariant to linear transformation, and constructs three pairwise joint matrices (L_D, F_D, R_D).
[0073] L_D, F_D, and R_D calculate their self - similarity matrices L_Similarity, F_Similarity, and R_Similarity respectively to characterize the shallow features of the left, frontal, and right sides of the teacher's classroom facial expression images at the same moment.
[0074] The feature extractor uses a 3D feature extractor. The self - similarity matrices L_Similarity, F_Similarity, and R_Similarity are input into the 3D feature extractor. The 3D feature extractor contains a total of 10 3×3×3 3D convolutional layers, 4 Relu layers, and 4 3D max - pooling layers. The extractor contains two convolutional modules, which are connected by a 3D average - pooling layer. Each convolutional module includes a 3D convolutional layer, a Rule layer, two 3D convolutional layers, a Rule layer, a 3D max - pooling layer, two 3D convolutional layers, and a 3D max - pooling layer connected in sequence, and outputs features LV of a fixed - scale size. M 、FV M 、RV M 。
[0075] At the same time, the self - similarity matrices L_Similarity, F_Similarity, and R_Similarity are input into the temporal attention network. The temporal attention network contains 3 fully - connected layers, 2 Relu layers, 1 sigmoid, and 1 repeat layer. Its structure is 1 fully - connected layer, a Rule layer, 1 fully - connected layer, a Rule layer, 1 fully - connected layer, a sigmoid, and 1 repeat layer, and outputs three feature vectors LV of a fixed - scale size. A(Feature vector of the left - hand side view image sequence), FV A (Feature vector of the front - hand side view image sequence), RV A (Feature vector of the right - hand side view image sequence).
[0076] Fuse the output of the temporal attention network with the output of the 3D feature extractor. The outer product of each output is used to obtain LV (left - hand side view fusion feature vector), FV (front - hand side view fusion feature vector) and RV (right - hand side view fusion feature vector). The three features incorporating temporal attention are concatenated, followed by connecting a fully - connected layer. Finally, the probability output of the class is obtained through the classification network (softmax), and the result with the maximum probability is used as the prediction result.
[0077] The calculation process of the pairwise joint matrix of the right - hand side view image sequence is as follows:
[0078]
[0079] where L_D ij is used to measure the difference between the i - th key position on one image and the j - th key position on another image, p i represents the i - th key position in the facial key area of an image in a certain angular image sequence in the right - hand side view image sequence (one of the images in R_90 or R_45), p j represents the j - th key position in the facial key area of an image in another angular image sequence in the right - hand side view image sequence. M is a decomposable symmetric positive - semi - definite matrix M = L T L. The elements of the pairwise joint matrix can be calculated using the second - norm by learning the linear transformation L on each image in the training set data. The formula is as follows:
[0080]
[0081] The self - similarity matrix is composed of the elements of the joint matrix, with a size of N l ×N l , N l represents the number of key positions in the facial area. The formula is as follows;
[0082]
[0083] The calculation methods of R_Similarity and F_Similarity are the same as that of L_Similarity. It should be noted that the input of F_Similarity is a self - connected graph, so the diagonal elements of F_Similarity are formed by comparing a key position with itself and are therefore zero.
[0084] The calculation process of the self - joint matrix is as follows:
[0085]
[0086] Among them, F_D ij can be used to measure any distance in the frontal image, representing the distance difference between the i - th key position and the j - th key position in the frontal image, and p i represents the i - th key position in the facial key area of the frontal - view image, and p j represents the j - th key position in the facial key area of the frontal - view image. When i = j, it represents the distance of the same key position in the frontal image. Therefore, F_D ii = 0, M is a decomposable symmetric positive - semi - definite matrix M = L T L. The elements of the pairwise joint matrix can be calculated using the second - norm by learning the linear transformation L on the data. The formula is as follows:
[0087]
[0088] The self - similarity matrix is composed of the elements of the joint matrix, with a size of N l ×N l , where N l represents the number of key positions in the facial area. The formula is as follows;
[0089]
[0090] The eigenvector LV A (another type of eigenvector of the left - side - view image sequence) is calculated using the temporal attention network as follows:
[0091]
[0092] Among them, W1 and W2 represent the weights of the fully - connected layers, D represents the number of channels of the final network, L_Similarity is the self - similarity matrix of the first group, θ and σ represent the Relu and sigmoid functions respectively, and repeat(.) represents repeating elements.
[0093] RV A (another type of eigenvector of the right - side - view image sequence) and FV A (another type of eigenvector of the frontal - view image sequence) are calculated in the same way as LV A is consistent.
[0094] Then, the output eigenvector of the 3D feature extractor is fused with the output eigenvector of the temporal attention network.
[0095]
[0096] Among them, LV represents the left - view fusion feature vector, and LV M represents the output of the 3D feature extractor. LV A represents the output of the temporal attention mechanism, and it is the tensor product.
[0097] Furthermore, after S2, it also includes the steps of:
[0098] S3: Obtain the total number of user facial expression recognition results and the number of facial expression changes within a preset time period, calculate the frequency of user facial expression changes in the classroom, and draw a line chart of user emotional changes.
[0099] That is, determine the teaching emotional state of the teacher according to the teaching facial expression changes in the classroom, and count the teaching emotional state of the whole class, so as to provide a reference for the intelligent evaluation of classroom teaching.
[0100] In one embodiment, multi-directional long-wave infrared cameras are placed around the teacher's position in the classroom as the core. A total of 10 cameras are respectively placed around the classroom and in the first row of students facing the podium. The 10 long-wave infrared cameras simultaneously collect images of the teacher in the classroom. Five images of the teacher at different angles (90° to the left, 45° to the left, front, 45° to the right, and 90° to the right) are selected from multiple images at the same moment. The above images are subjected to facial area localization, cropped into 224×224 size, and data augmentation and data preprocessing are performed. The Mahalanobis distance at the element level is calculated, and a pairwise joint graph (90° to the left and 45° to the left images, 45° to the right and 90° to the right images) or a self-joint graph (front image) with a size of 3×224×224 is constructed to form a self-similarity matrix, which is respectively input into a 3D feature extractor and a temporal attention network. In the 3D feature extractor, it sequentially passes through a first convolutional module with a 3×3×3, 4 3D convolution, a Relu layer, two 3D convolutions (3×3×3, 8), a Relu layer, a 2×1×1 max pooling layer, two 3D convolutions (3×3×3, 16), a 2×2×2 max pooling layer, passes through a 2×2×2 average pooling layer, and then passes through a second convolutional module with a 3D convolution (3×3×3, 32), a Relu layer, two 3D convolutions (3×3×3, 64), a Relu layer, a 2×2×2 max pooling layer, two 3D convolutions (3×3×3, 128), and a 2×2×2 max pooling layer, and outputs a tensor with a size of 8×8. The output of the temporal attention network is also an 8×8 tensor. The outer product of the two tensors is obtained to get an 8×8 tensor. The tensors of the three branches are concatenated to get a 24×4 tensor, and then through a fully connected layer and softmax, the probabilities of 20 expressions are obtained, and finally the recognition result is output. The model selects the Adam optimizer with an initial learning rate of 0.01. Starting from the 20th iteration, the learning rate is modified (the current learning rate is multiplied by 0.8), and the learning rate is modified after each iteration to optimize the model. According to the teacher's facial recognition results per minute in the classroom, a line graph of the teacher's classroom emotional changes is drawn to represent the changes in the teacher's teaching expressions, and the frequency of the teacher's classroom expression changes is statistically calculated (the number of classroom expression changes / the total number of statistical expressions).
[0101] A multi-view teaching expression recognition system in a classroom environment according to an embodiment of the present invention includes:
[0102] A data acquisition module for collecting a multi-view image sequence of a user's expression, where the multi-view image sequence includes image sequences of multiple types of views;
[0103] An output module for inputting a multi-view image sequence into an expression recognition model to obtain a user expression recognition result; the expression recognition model includes a self-similar matrix extraction module, a feature extractor, a temporal attention network, a fusion splicing module, and a classification network; the self-similar matrix extraction module is used to extract the self-similar matrix of each type of view from the multi-view image sequence; the feature extractor is used to obtain the first type of feature vector of each type of view; the temporal attention network is used to obtain the second type of feature vector of each type of view; the fusion splicing module is used to fuse and splice the two types of feature vectors and then input them into the classification network.
[0104] The implementation principle and technical effect of the system are similar to those of the above method, which will not be elaborated here.
[0105] It must be noted that in any of the above embodiments, the methods do not necessarily need to be executed in the order of the serial numbers. As long as it cannot be inferred from the execution logic that they must be executed in a certain order, it means that they can be executed in any other possible order.
[0106] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A multi - perspective teaching expression recognition method in a classroom environment, characterized in that, Including the steps: S1. Collect a multi-view image sequence of the user's expression, where the multi-view image sequence includes image sequences of multiple types of views; S2. Input the multi-view image sequence into an expression recognition model to obtain a user expression recognition result; the expression recognition model includes a self-similarity matrix extraction module, a feature extractor, a temporal attention network, a fusion and splicing module, and a classification network; the self-similarity matrix extraction module is used to extract the self-similarity matrix of each type of view from the multi-view image sequence respectively; the feature extractor is used to obtain the first type of feature vectors of each type of view according to the self-similarity matrix of each type of view; the temporal attention network is used to obtain the second type of feature vectors of each type of view according to the self-similarity matrix of each type of view; the fusion and splicing module is used to fuse and splice the two types of feature vectors and then input them into the classification network.
2. The multi - perspective teaching expression recognition method in a classroom environment according to claim 1, wherein, Before the S2, it also includes the steps: preprocess the multi-view image sequence, and the preprocessing includes: Convert the images of the multi-view image sequence into PIL images, then adjust the size of the PIL images to 224×224 pixels, and then convert the PIL images into tensor images, and normalize the tensor images with the mean value and standard deviation.
3. The multi-view teaching expression recognition method in a classroom environment according to claim 1, wherein After the S2, it also includes the steps: S3. Obtain the total number of user expression recognition results and the number of expression changes within a preset time period, calculate the frequency of the user's classroom expression changes, and draw a line chart of the user's emotional changes.
4. The multi-perspective teaching expression recognition method according to claim 1, wherein The multi-view image sequence includes three types, namely: the left-view image sequence, the front-view image sequence, and the right-view image sequence.
5. The multi-view teaching expression recognition method according to claim 4, wherein The multi-view image sequence is a long-wave infrared image sequence, the left-view image sequence includes multiple image sequences of multiple angles belonging to the left view, and the right-view image sequence includes multiple image sequences of multiple angles belonging to the right view.
6. The multi-view teaching expression recognition method according to claim 4, characterized in that, The S2 includes sub-steps: Extract the first self-similarity matrix of the left-view image sequence, the second self-similarity matrix of the front-view image sequence, and the third self-similarity matrix of the right-view image sequence respectively; Input the first self-similarity matrix, the second self-similarity matrix, and the third self-similarity matrix into the feature extractor respectively to obtain the first feature vector, the second feature vector, and the third feature vector respectively; Input the first self-similarity matrix, the second self-similarity matrix, and the third self-similarity matrix into the temporal attention network respectively to obtain the fourth feature vector, the fifth feature vector, and the sixth feature vector respectively; Calculate the tensor product of the first feature vector and the fourth feature vector to obtain the left-view fusion feature vector, calculate the tensor product of the second feature vector and the fifth feature vector to obtain the front-view fusion feature vector, calculate the tensor product of the third feature vector and the sixth feature vector to obtain the right-view fusion feature vector; Splice the left-view fusion feature vector, the front-view fusion feature vector, and the right-view fusion feature vector and then input them into the classification network to obtain the user expression recognition result.
7. The multi-view teaching expression recognition method according to claim 1, wherein The feature extractor is a 3D feature extractor.
8. The multi-perspective teaching expression recognition method according to claim 7, wherein, The 3D feature extractor includes two convolutional modules, which are connected by a 3D average pooling layer. Each convolutional module includes a 3D convolutional layer, a Rule layer, two 3D convolutional layers, a Rule layer, a 3D max pooling layer, two 3D convolutional layers, and a 3D max pooling layer, which are connected in sequence.
9. A multi - perspective teaching expression recognition system in a classroom environment, characterized in that, It includes: A data acquisition module for acquiring a multi-view image sequence of a user's expression, where the multi-view image sequence includes image sequences of multiple types of views; An output module for inputting the multi-view image sequence into an expression recognition model to obtain a user expression recognition result; the expression recognition model includes a self-similarity matrix extraction module, a feature extractor, a temporal attention network, a fusion splicing module, and a classification network; the self-similarity matrix extraction module is used to respectively extract the self-similarity matrix of each type of view from the multi-view image sequence; the feature extractor is used to obtain the first type of feature vector of each type of view according to the self-similarity matrix of each type of view; the temporal attention network is used to obtain the second type of feature vector of each type of view according to the self-similarity matrix of each type of view; the fusion splicing module is used to fuse and splice the two types of feature vectors and then input them into the classification network.
Citation Information
Patent Citations
Multi-angle facial expression recognition method based on generation of countermeasure network
CN108446609A
Classroom student expression recognition and classroom state evaluation method and device
CN113239914A