Student behavior detection method based on course learning and cross-domain Vision Transform
By adopting cross-domain Vision Transformer and course learning strategies in student behavior analysis, behavior detection knowledge is migrated from the source domain with sufficient label samples to the target domain with insufficient label samples, which solves the problem of insufficient label data and improves the adaptability and accuracy of the model.
Patent Information
- Application Number
- CN202510380416.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-05-13
AI Technical Summary
The existing student behavior analysis algorithms have bottlenecks in obtaining high-quality labeled data, resulting in high labeling costs, low efficiency and high data noise, affecting the accuracy and reliability of model training.
Student behavior detection methods based on course learning and cross-domain Vision Transformer are adopted to migrate behavior detection knowledge from the source domain public data set with sufficient label samples to the target domain local data set with insufficient label samples, and sort training samples through course learning strategies to improve the knowledge transfer effect.
It effectively solves the problem of insufficient local labeling samples, improves the adaptability and accuracy of the behavioral detection model, and reduces the labeling cost and noise impact.
Smart Images

Figure CN119992666A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of behavior detection technology, and in particular to a student behavior detection method based on course learning and cross-domain Vision Transformer. Background Art
[0002] With the in-depth integration of artificial intelligence in educational technology, student behavior analysis has become a key tool for optimizing teaching strategies. By capturing students' concentration, interaction level, and emotional response, artificial intelligence systems can provide instant feedback, allowing educators to flexibly adjust teaching methods to improve learning outcomes. However, the current advancement of student behavior analysis algorithms is severely limited by the acquisition of high-quality labeled data, and this bottleneck has caused multiple problems.
[0003] On the one hand, creating a detailed and accurate manually annotated database is a time-consuming and labor-intensive task. This not only requires expert intervention to ensure the accuracy of the data, but also comes with high time and economic costs. Especially when faced with massive data, maintaining the efficiency and consistency of annotation becomes a major challenge.
[0004] On the other hand, subjective judgment and human error inevitably introduce data noise, which may lead to diversity and inaccuracy of labels, thereby weakening the accuracy and reliability of model training and may produce misleading results in practical applications. In short, although artificial intelligence provides unprecedented insights into student behavior analysis, the difficulty of building reliable annotated datasets limits the full release of its potential, and innovative solutions are urgently needed to overcome these obstacles. Summary of the invention
[0005] The purpose of the present invention is to provide a student behavior detection method based on curriculum learning and cross-domain Vision Transformer. The behavior detection knowledge is transferred from a source domain public dataset with sufficient labeled samples to a target domain local dataset with insufficient labeled samples through a cross-domain Vision Transformer. The training samples are sorted based on the curriculum learning strategy to improve the knowledge transfer effect, and a behavior detection model adapted to the local data distribution is obtained, which effectively solves the problem of insufficient local labeled samples.
[0006] To achieve the above objectives, the present invention provides a student behavior detection method based on course learning and cross-domain Vision Transformer, comprising the following steps:
[0007] S1, collect and preprocess student behavior image data to build a local data set;
[0008] S2. Construct a student behavior detection model based on Vision Transformer, wherein the student behavior detection model includes an encoder, a classifier, and an attention map alignment module;
[0009] S3. Take the Stanford40 dataset as the source domain dataset and the local dataset as the target domain dataset, divide the target domain dataset into batches, and obtain a target domain batch set; for each target domain batch, adjust the order and divide the source domain dataset into batches based on the curriculum learning idea, and obtain a sorted source domain batch set;
[0010] S4. Based on each target domain batch and its corresponding sorted source domain batch set, the encoder and classifier are updated according to the loss function to obtain a trained student behavior detection model.
[0011] S5. Use the trained student behavior detection model to realize real-time student behavior detection.
[0012] Preferably, the student behavior data in step S1 includes student movements, postures, and scene image data, and the preprocessing includes image scaling, denoising, and standardization.
[0013] Preferably, the step of constructing the student behavior detection model in step S2 includes:
[0014] Construct an encoder, in which the student behavior image data x passes through the Patch Embedding layer to obtain the image feature X∈R N×D ; Image feature X∈R N×D After being processed by the cascaded L-layer Transformer Block, the encoder output feature Z is obtained. L ∈R (N+1)×D ; Where N is the number of image blocks segmented by the Patch Embedding layer, and D is the feature dimension of each image block;
[0015] Construct a classifier, in which the output feature Z of the encoder is converted based on the fully connected layer L Map it into the probability distribution of behavior categories and obtain the behavior classification result.
[0016] Preferably, the step of constructing the student behavior detection model in step S2 further includes:
[0017] Image features X and classification labels cls∈R 1×D Perform splicing to obtain the spliced feature X'∈R (N+1)×D ; The concatenated feature X' is processed by the cascaded L-layer Transformer Block to obtain the encoder output feature Z L ∈R (N+1)×D .
[0018] Preferably, the step of constructing the student behavior detection model in step S2 further includes:
[0019] Image features X and classification labels cls∈R 1×D Perform splicing to obtain the spliced feature X'∈R (N+1)×D ; The concatenated feature X' and position code P∈R (N+1)×D Add together to get the input feature Z0∈R (N+1)×D ; After the input feature Z0 is processed by the cascaded L-layer Transformer Block, the encoder output feature Z is obtained L ∈R (N+1)×D ;
[0020] The calculation method of the elements in the position code P is:
[0021] ,
[0022] ,
[0023] Where, the position number is pos∈[0,N+1) and the feature dimension index is i∈[0,D / 2].
[0024] Preferably, the specific steps of step S3 include:
[0025] In preset batch size Sequentially Perform batch division to obtain the target domain batch set ,in, is the sample j in the target domain dataset, N t is the number of samples in the target domain dataset, is the target domain batch b in the target domain batch set;
[0026] For source domain dataset Samples and target domain datasets in Extract features from samples in the source domain to obtain the source domain feature set for sorting and the target domain feature set for sorting ,in is the sample i in the source domain dataset, N s is the number of samples in the source domain dataset, is the feature of sample i in the source domain dataset, is the feature of sample j in the target domain dataset;
[0027] For each target domain batch, the mean square error between it and each feature in the source domain feature set for sorting is calculated, and the samples in the source domain dataset are reordered in ascending order according to the obtained mean square error, and then the batch size is set to the preset batch size. The source domain data set after the order adjustment is divided into batches in sequence to obtain a sorted source domain batch set.
[0028] Preferably, a pre-trained encoder model E based on ResNet-18 p Feature extraction is performed on samples in the source domain dataset and samples in the target domain dataset.
[0029] Preferably, for each target domain batch, the mean square error between it and each feature in the source domain feature set for sorting is calculated:
[0030] ,
[0031] in, The mean square error between feature i in the source domain ranking feature set and batch b in the target domain.
[0032] Preferably, in step S4, the encoder is updated according to the attention graph alignment loss function, and then the encoder and the classifier are jointly updated according to the cross entropy classification loss function, wherein:
[0033] Attention graph alignment loss function L ATT The expression is:
[0034] ,
[0035] Among them, B is the preset batch size, A l,i src and A l,i tgt are the attention maps of sample i in the source domain and target domain batches at layer l, respectively, and H is the number of attention heads;
[0036] Cross entropy classification loss function L CE The expression is:
[0037] ,
[0038] In the formula, is the sample j' in the target domain batch b, is the sample j' in the sorted source domain batch b', and for and Corresponding category labels, is the number of behavior categories; is the indicator function; and The tth round and The features obtained by the encoder are and They are and Probability distribution of behavior categories output after inputting the classifier and Corresponding to the category The probability value of .
[0039] Preferably, the step of using the trained student behavior detection model to implement real-time student behavior detection in step S4 includes:
[0040] The student behavior image data collected in real time is trained into a student behavior detection model to obtain a probability distribution vector of the behavior detection results, where the category with the largest probability is the real-time student behavior detection result.
[0041] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned student behavior detection and abnormal behavior identification method are implemented.
[0042] The present invention also provides an electronic device, comprising:
[0043] Memory for storing computer programs;
[0044] A processor is used to implement the steps of the above-mentioned student behavior detection method when executing the computer program.
[0045] Therefore, the present invention adopts the above-mentioned student behavior detection method based on curriculum learning and cross-domain Vision Transformer, and transfers the behavior detection knowledge from the source domain public dataset with sufficient labeled samples to the target domain local dataset with insufficient labeled samples through the cross-domain Vision Transformer, and sorts the training samples based on the curriculum learning strategy to improve the knowledge transfer effect. This method can effectively solve the problem of insufficient local labeled samples.
[0046] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 The present invention is a flowchart of an embodiment of a student behavior detection method based on course learning and cross-domain Vision Transformer.
[0048] Figure 2It is a structural schematic diagram of the student behavior detection model based on Vision Transformer of the present invention.
[0049] Figure 3 It is a schematic diagram of the structure of the Transformer block.
[0050] Figure 4 This is a schematic diagram of the structure of the pre-trained encoder based on ResNet-18.
[0051] Figure 5 This is a comparison chart of experimental results of a student behavior detection method based on curriculum learning and cross-domain Vision Transformer in the present invention and a comparison method that removes the source domain dataset sorting and attention map alignment modules based on curriculum learning. DETAILED DESCRIPTION
[0052] The technical solution of the present invention is further described below through the accompanying drawings and embodiments.
[0053] Unless otherwise defined, the technical terms or scientific terms used in the present invention should be understood by people with ordinary skills in the field to which the present invention belongs. The words "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0054] In one embodiment, Figure 1 As shown, the present invention provides a student behavior detection method based on curriculum learning and cross-domain VisionTransformer, comprising the following steps:
[0055] S1. Collect student behavior data and perform preprocessing to build a local data set. The student behavior data includes student movements, postures, and scene image data. The preprocessing includes image scaling, denoising, and standardization.
[0056] S2. Build a student behavior detection model based on Vision Transformer, including encoder, classifier and attention map alignment module.
[0057] S3. Take the Stanford40 dataset as the source domain dataset and the local dataset as the target domain dataset, divide the target domain dataset into batches, and obtain a target domain batch set; for each target domain batch, adjust the order and divide the source domain dataset into batches based on the course learning idea, and obtain a sorted source domain batch set; based on each target domain batch and its corresponding sorted source domain batch set, update the encoder and classifier according to the loss function to obtain a trained student behavior detection model.
[0058] S4. Use the trained student behavior detection model to achieve real-time student behavior detection.
[0059] In one embodiment, Figure 2 As shown in the figure, the encoder is based on the image feature extraction module of Vision Transformer. The input image x is segmented into N image blocks through the Patch Embedding layer. The feature dimension of each image block is D, and the image block feature X∈R is obtained. N×D , image features X and classification labels cls∈R 1×D Concatenate and obtain the concatenated feature X'=Concat(cls,X), where X'∈R (N+1)×D , the concatenated feature X' and the position code P∈R (N+1)×D Add together to get the input feature Z0=X'+P, where Z0∈R (N+1)×D The input feature Z0 is processed by L layers of Transformer Block, such as Figure 3 As shown in Figure 1, each layer of Transformer Block includes Multi-Head Self-Attention (MHSA) and Feed-Forward Network (FFN). The output Z of the lth layer is l The calculation method is:
[0060] ,
[0061] ,
[0062] Among them, LayerNorm(·) is the layer normalization operation, and the output feature of the final encoder is Z L ∈R (N+1)×D , where L is the total number of layers of Transformer Block.
[0063] Specifically, the position code P is calculated as follows:
[0064] ,
[0065] ,
[0066] The position number is pos∈[0,N+1) and the feature dimension index is i∈[0,D / 2].
[0067] In one embodiment, the classifier is based on a fully connected layer, which transforms the feature Z output by the encoder into L Mapped to behavior category probability distribution:
[0068] y=Softmax(W·Z L [:,0]),
[0069] Among them, Z L [:,0] is the classification mark feature, W∈R (K×D) is the classifier weight, is the number of behavior categories.
[0070] In one embodiment, the attention map alignment module is used to calculate the alignment loss of the source domain and the target domain attention maps. For the multi-head attention of the l-th layer Transformer Block, its attention map A l ∈R (H×N×N) The calculation formula is:
[0071] ,
[0072] Where Q l , K l ∈R (N×D⁄H) are the query matrix and the key matrix respectively, and H is the number of attention heads. The attention graph alignment loss function L ATT Defined as the mean squared error between the source domain and the target domain attention map:
[0073] ,
[0074] Where B is the preset batch size, and The source domain and the target domain correspond to the labels The sample is in Attention map of the layer.
[0075] In one embodiment, the specific steps of S3 include:
[0076] Using the pre-trained model E based on ResNet-18 p For source domain dataset Samples and target domain datasets in Extract features from samples in the source domain to obtain the source domain feature set for sorting and the target domain feature set for sorting , and are the total number of samples in the source domain and the target domain, respectively.
[0077] like Figure 4 As shown, E p It contains an initial convolution layer and four convolution blocks. The initial convolution layer consists of a The convolutional layer and a maximum pooling layer are composed of two residual units. The convolutional layer is composed of several layers, and finally the feature vector is output through the global average pooling layer.
[0078] First, with a preset batch size Sequentially Perform batch division to obtain the target domain batch set ,in For any target domain batch , calculate its feature set for sorting with the source domain The mean square error of each feature in :
[0079] ,
[0080] According to the mean square error from small to large, the source domain data set Adjust the order to get the batch size for the target domain The sorted source domain dataset , and then divide it into batches to obtain batches for the target domain The source domain batch collection ,in According to this rule, for the target domain batch set Each target domain batch in , we all calculate a batch for the target domain The source domain batch collection .
[0081] After order adjustment, in each training round t, for the target domain batch set Each target domain sample batch in , are all related to the source domain batch set for the target domain sample batch Each batch in Perform the following calculations:
[0082] For the convenience of explanation, and Expressed as and , and Each sample in the training process is passed through the encoder of each training round t (here E tRepresents) to obtain the source domain feature batch and the target domain feature batch , and then through the classifier Get the source domain output result Output results with target domain , and then calculate the cross entropy classification loss function :
[0083] ,
[0084] in, and For sample and Corresponding category label; is the number of behavior categories; is the indicator function; and The probability distribution vector of the result output by the classifier and Corresponding to the category The probability value of .
[0085] At the same time, the source domain attention map of each layer of each sample is extracted from the encoder Attention map with target domain , calculate the attention map alignment loss function :
[0086] .
[0087] Finally, the loss function is aligned according to the attention map Update encoder , and then according to the cross entropy loss function Joint Update Encoder With classifier , the updates within a single batch are as follows:
[0088] ,
[0089] ,
[0090] in, is the learning rate, , are the updated encoder and classifier. Repeat the above training process until the encoder and classifier converge.
[0091] In one embodiment, the real-time sample x' is trained by the encoder Get the source domain feature E(x'), the source domain feature is trained by the classifier The probability distribution vector W(E(x')) of the behavior detection result can be obtained, where the category with the highest probability is the detection result.
[0092] like Figure 5 As shown, there is a comparison chart of experimental results of a student behavior detection method based on curriculum learning and cross-domain Vision Transformer in the present invention and a comparison method in which the source domain dataset sorting and attention map alignment modules based on curriculum learning are removed.
[0093] Therefore, the present invention adopts the above-mentioned student behavior detection method based on curriculum learning and cross-domain Vision Transformer, and transfers the behavior detection knowledge from the source domain public dataset with sufficient labeled samples to the target domain local dataset with insufficient labeled samples through the cross-domain Vision Transformer, and sorts the training samples based on the curriculum learning strategy to improve the knowledge transfer effect, which can effectively solve the problem of insufficient local labeled samples.
[0094] Based on the same technical solution, the present invention also provides an electronic device, including:
[0095] Memory for storing computer programs;
[0096] A processor is used to implement the steps of the above-mentioned student behavior detection and abnormal behavior identification method when executing the computer program.
[0097] Based on the same technical solution, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned student behavior detection and abnormal behavior identification method are implemented. The computer-readable storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0098] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A student behavior detection method based on curriculum learning and cross-domain Vision Transformer, characterized by: The following steps are involved: S1, collect and preprocess student behavior image data to build a local data set; S2. Construct a student behavior detection model based on Vision Transformer, wherein the student behavior detection model includes an encoder and a classifier; S3, taking the Stanford40 dataset as the source domain dataset and the local dataset as the target domain dataset, dividing the target domain dataset into batches to obtain the target domain batch set; For each target domain batch, the source domain dataset is reordered and divided into batches based on the curriculum learning idea to obtain a sorted source domain batch set; S4. Based on each target domain batch and its corresponding sorted source domain batch set, the encoder and classifier are updated according to the loss function to obtain a trained student behavior detection model. S5. Use the trained student behavior detection model to realize real-time student behavior detection.
2. The method according to claim 1, characterized in that The student behavior data in step S1 includes student movements, postures, and scene image data, and the preprocessing includes image scaling, denoising, and standardization.
3. The method according to claim 1, characterized in that The steps of constructing the student behavior detection model in step S2 include: Construct an encoder, in which the student behavior image data x passes through the Patch Embedding layer to obtain the image feature X∈R N×D ; Image feature X∈R N×D After being processed by the cascaded L-layer Transformer Block, the encoder output feature Z is obtained. L ∈R (N+1)×D ; Where N is the number of image blocks segmented by the Patch Embedding layer, and D is the feature dimension of each image block; Construct a classifier, in which the output feature Z of the encoder is converted based on the fully connected layer L Map it into the probability distribution of behavior categories and obtain the behavior classification result.
4. The method according to claim 3, characterized in that: The step of constructing the student behavior detection model in step S2 also includes: Image features X and classification labels cls∈R 1×D Perform splicing to obtain the spliced feature X'∈R (N+1)×D ; The concatenated feature X' is processed by the cascaded L-layer Transformer Block to obtain the encoder output feature Z L ∈R (N+1)×D .
5. The method according to claim 3, characterized in that: The step of constructing the student behavior detection model in step S2 also includes: Image features X and classification labels cls∈R 1×D Perform splicing to obtain the spliced feature X'∈R (N+1)×D ; The concatenated feature X' and position code P∈R (N+1)×D Add together to get the input feature Z0∈R (N+1)×D ; After the input feature Z0 is processed by the cascaded L-layer Transformer Block, the encoder output feature Z is obtained L ∈R (N+1)×D ; The calculation method of the elements in the position code P is: , , Where, the position number is pos∈[0,N+1) and the feature dimension index is i∈[0,D / 2].
6. The method according to claim 1, characterized in that The specific steps of step S3 include: In preset batch size Sequentially Perform batch division to obtain the target domain batch set ,in, is the sample j in the target domain dataset, N t is the number of samples in the target domain dataset, is the target domain batch b in the target domain batch set; For source domain dataset Samples and target domain datasets in Extract features from samples in the source domain to obtain the source domain feature set for sorting and the target domain feature set for sorting ,in is the sample i in the source domain dataset, N s is the number of samples in the source domain dataset, is the feature of sample i in the source domain dataset, is the feature of sample j in the target domain dataset; For each target domain batch, the mean square error between it and each feature in the source domain feature set for sorting is calculated, and the samples in the source domain dataset are reordered in ascending order according to the obtained mean square error, and then the batch size is set to the preset batch size. The source domain data set after the order adjustment is divided into batches in sequence to obtain a sorted source domain batch set.
7. The method according to claim 6, characterized in that Pre-trained encoder model E based on ResNet-18 p Feature extraction is performed on samples in the source domain dataset and samples in the target domain dataset.
8. The method according to claim 6, characterized in that For each target domain batch, calculate the mean squared error between it and each feature in the source domain feature set used for sorting: , in, The mean square error between feature i in the source domain ranking feature set and batch b in the target domain.
9. The method according to claim 1, characterized in that: In step S4, the encoder is updated according to the attention graph alignment loss function, and then the encoder and classifier are jointly updated according to the cross entropy classification loss function, wherein: Attention graph alignment loss function L ATT The expression is: , Among them, B is the preset batch size, A l,i src and A l,i tgt are the attention maps of sample i in the source domain and target domain batches at layer l, respectively, and H is the number of attention heads; Cross entropy classification loss function L CE The expression is: , In the formula, is the sample j' in the target domain batch b, is the sample j' in the sorted source domain batch b', and for and Corresponding category labels, is the number of behavior categories; is the indicator function; and The tth round and The features obtained by the encoder are and They are and Probability distribution of behavior categories output after inputting the classifier and Corresponding to the category The probability value of .
10. The method for detecting student behavior based on course learning and cross-domain Vision Transformer according to claim 1, characterized in that: The step of using the trained student behavior detection model in step S4 to implement real-time student behavior detection includes: The student behavior image data collected in real time is trained into a student behavior detection model to obtain a probability distribution vector of the behavior detection results, where the category with the largest probability is the real-time student behavior detection result.