Online learning behavior recognition method and system based on multi-view depth behavior recognition
By adopting a multi-view deep behavior recognition method in online learning behavior monitoring, combining the circular generation adversarial network between views and modals, the problems of false learning state recognition, perspective limitations and lack of non-verbal interaction paths are solved, and more accurate learning behavior recognition and teaching effect evaluation are achieved.
Patent Information
- Application Number
- CN202510233864.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
AI Technical Summary
Currently, online learning behavior monitoring has problems such as algorithms that are difficult to identify false learning states, single-camera visual detection has limitations in perspective, and lack of non-verbal interaction paths between teachers and students.
The online learning behavior recognition method based on multi-view deep behavior recognition is adopted. By obtaining students' front RGB video data, frontal skeleton data, side RGB video data and side skeleton data, a multi-layer perceptron is built for classification, combining inter-view loop generation adversarial networks and inter-modal loop generation adversarial networks to achieve cross-view and cross-modal features interactive fusion.
It improves the accuracy of online learning behavior recognition, can more accurately identify students' learning behavior, provide abnormal behavior prompts, and help teachers more accurately understand students' learning behavior status, thereby evaluating teaching effectiveness and taking improvement measures.
Smart Images

Figure CN120164258A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of educational informatization, and in particular, to an online learning behavior recognition method and system based on multi-view depth behavior recognition. Background Art
[0002] With the rapid development of artificial intelligence technology and the intelligent transformation of society, lifelong learning has become the core driving force for personal development. As an important carrier of lifelong learning, online learning, with its advantages of breaking through time and space limitations and rich and flexible resources, has strongly promoted educational equity and innovation. To ensure learning quality, current mainstream learning platforms have built a multi-dimensional monitoring system: verifying through pop-up questions and detecting split screens to enforce learning coherence; constructing a concentration model by using facial expression analysis (concentration heat map) and physiological behavior monitoring (blinking frequency, head posture); adopting interactive mechanisms such as timed expression verification and manual resume to improve participation. These technical means have significantly optimized the supervision of the learning process and formed a digital teaching management paradigm different from traditional classrooms.
[0003] However, there are still three major shortcomings in current online learning behavior monitoring: First, the algorithm is difficult to identify false learning states, such as "hanging up learning" (only returning to the screen when answering questions), and other utilitarian behaviors lead to distorted concentration data; Second, single-camera visual detection has a perspective limitation and cannot accurately distinguish the essential differences between learning behaviors (such as taking notes and playing with mobile phones); Third, there is a lack of non-verbal interaction paths between teachers and students, making it difficult to intuitively judge the degree of knowledge mastery through body language, micro-expressions, etc., which affects the timeliness of teaching adjustment.
[0004] For most current multi-view based behavior recognition algorithms, only RGB video data is used as input. However, with the gradual maturity of human keypoint detection algorithms based on deep learning, using both human skeleton data and RGB data as model input helps capture the spatio-temporal complementary information of behavior representations. In addition, in multi-view based behavior recognition algorithms, the traditional early feature-level fusion method is prone to heterogeneous feature interference, while late decision-level fusion is difficult to mine the deep semantic information between cross-views and cross-modalities, resulting in limited behavior recognition accuracy. Based on these research results, this paper proposes a behavior recognition algorithm based on interactive fusion of multi-view and multi-modal features (MMFINet). The model is divided into two sub-networks, each sub-network consists of two branches, and view-intercycle generative adversarial network and modality-intercycle generative adversarial network are introduced to achieve cross-view and cross-modal feature interaction. MMFINet solves the problems of existing methods being limited to single-modal input, not effectively fusing cross-modal features, and conventional fusion strategies (early / late) being difficult to capture the deep semantic interaction and complementary details of multi-view and multi-modal. Through the heterogeneous multi-branch design, it realizes the feature co-optimization of cross-view and cross-modal, and improves the accuracy of behavior recognition. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides an online learning behavior recognition method based on multi-view depth behavior recognition, including the following steps:
[0006] S1. Obtain the frontal RGB video data, frontal skeleton data, lateral RGB video data, and lateral skeleton data of students during online learning, and perform preprocessing on each of them;
[0007] S2. Construct a branch network for extracting RGB video data features, and extract video fusion feature vectors based on the frontal RGB video data and the lateral RGB video data;
[0008] S3. Construct a branch network for extracting skeleton data features, and extract skeleton fusion feature vectors based on the frontal skeleton data and the lateral skeleton data;
[0009] S4. Input the video fusion feature vectors and the skeleton fusion feature vectors into the modality-intercycle generative adversarial network for feature alignment to obtain total fusion feature vectors;
[0010] S5. Input the total fusion feature vectors into a multi-layer perceptron for classification to obtain the classification result of the students' online learning behaviors.
[0011] Further, step S2 includes the following steps:
[0012] S201. Perform boundary cropping on the frontal RGB video data according to the frontal skeleton data to obtain frontal learning behavior video data, and perform boundary cropping on the side RGB video data according to the side skeleton data to obtain side learning behavior video data;
[0013] S202. Input the frontal learning behavior video data and the side learning behavior video data into the motion time information extraction model to obtain the frontal spatio-temporal feature map and the side spatio-temporal feature map of the student, and generate a frontal motion map and a side motion map according to the frontal spatio-temporal feature and the side spatio-temporal feature respectively;
[0014] S203. Based on the ResNet model, perform feature extraction on the frontal motion map and the side motion map respectively to obtain a frontal motion feature vector and a side motion feature vector, and perform feature fusion on the frontal motion feature vector and the side motion feature vector through an inter-view cyclic generative adversarial network to obtain a video fusion feature vector.
[0015] Further, the boundary box range of the boundary clipping in step S201 is determined based on the maximum and minimum values of the skeleton coordinates in the skeleton data.
[0016] Further, step S202 includes the following steps:
[0017] S202-1. Perform dense frame sampling on the frontal learning behavior video data and the side learning behavior video data respectively, and compress the sampled frontal learning behavior video data and side learning behavior video data to obtain a frontal learning behavior video sparse set and a side learning behavior video sparse set respectively;
[0018] S202-2. Calculate the frame difference between adjacent frames in the frontal learning behavior video sparse set to obtain a first frame difference, and calculate the frame difference between adjacent frames in the side learning behavior video sparse set to obtain a second frame difference;
[0019] S202-3. Perform feature extraction on the first frame difference and the frontal learning behavior video sparse set through a convolutional kernel to obtain a frontal spatio-temporal feature map, and perform feature extraction on the second frame difference and the side learning behavior video sparse set through a convolutional kernel to obtain a side spatio-temporal feature map;
[0020] S202-4. Perform max pooling processing on the first frame difference and the second frame difference respectively, and generate a first time attention mask and a second time attention mask through a convolutional layer and a sigmoid function respectively;
[0021] S202-5. Multiply the frontal spatio-temporal feature map by the first temporal attention mask to obtain the frontal motion map, and multiply the lateral spatio-temporal feature map by the second temporal attention mask to obtain the lateral motion map.
[0022] Further, the frontal skeleton data and the lateral skeleton data are extracted based on the HRNet network.
[0023] Further, step S3 includes the following steps:
[0024] S301. Input the frontal skeleton data and the lateral skeleton data into the Prune MSSTNet network to model the temporal and spatial dimensions of the frontal skeleton data and the temporal and spatial dimensions of the lateral skeleton data;
[0025] S302. Calculate the temporal and spatial dimension models of the frontal skeleton data to obtain the frontal skeleton data feature vector, and calculate the temporal and spatial dimension models of the lateral skeleton data to obtain the lateral skeleton data feature vector;
[0026] S303. Fuse the frontal skeleton data feature vector and the lateral skeleton data feature vector through an inter-view cyclic generative adversarial network to obtain the skeleton fusion feature vector.
[0027] The present invention also provides an online learning behavior recognition system based on multi-view depth behavior recognition. This system is implemented based on the online learning behavior recognition method described in any one of the foregoing, and includes a user layer, a gateway layer, a service layer, and a data layer that are sequentially signal-connected; the user layer is used for students to log in to the system through single sign-on and perform identity authentication; the gateway layer is used to process requests from the user layer; the service layer is used to process business logic; the data layer is used to store structured data, memory data, and video file data.
[0028] Further, the user layer includes unified authentication. Students perform identity authentication through single sign-on to the system, enabling students to access at least 2 systems using one account and password.
[0029] Further, the business logic includes file upload, behavior recognition, statistical evaluation, and anomaly warning.
[0030] Further, the data layer includes a MySQL relational database, a Redis database, and a file storage database; the MySQL relational database is used to store structured data; the Redis database is used to store memory data; the file storage database is used to store video data.
[0031] The beneficial effects of the present invention are:
[0032] (1) Process the videos of the students' learning process from two different perspectives through the multi-view modal feature interaction and fusion algorithm to identify the students' online learning behaviors. If there are abnormal learning behaviors such as sleeping, eating, playing with mobile phones, etc., the system can give prompts for abnormal behaviors and, combined with the students' after-class homework scores and exam scores, etc., give reasonable feedback, such as whether they need to learn again. This method not only improves the teacher's evaluation of the learners' learning process but also helps the teacher understand the learners' learning behavior status more accurately, so as to evaluate the teaching effect and take targeted improvement measures to enhance the learners' learning effect.
[0033] (2) In the multi-view behavior recognition algorithm, introducing the skeleton modality provides an effective solution. Human skeleton data is usually represented by key points with 2D or 3D coordinates, and these key points are connected to form a wireframe representing the human skeleton. This data form has high anti-interference ability against environmental changes (such as lighting conditions and background complexity) due to its abstraction, and focuses on capturing the structural dynamics of the human body. And with the gradual maturity of the deep learning-based human key point detection algorithm, obtaining human skeleton data becomes simpler and faster, and this technology usually does not depend on a specific type of camera or device and can adapt to a variety of shooting devices and environments. Introducing skeleton data can make up for the deficiencies of single RGB data, and can give full play to the complementary advantages of both, so as to maintain a high recognition effect in complex or harsh environments. Using this multi-modal method can perform behavior recognition and analysis more accurately and stably in complex scenarios.
[0034] In the algorithm, by introducing the Inter-View CycleGAN and Inter-Modal CycleGAN, the potential semantic information between different views and different modalities can be more fully utilized to improve the accuracy of behavior recognition. Description of the Drawings
[0035] Figure 1 Flowchart of an online learning behavior recognition method based on multi-view deep behavior recognition of the present invention.
[0036] Figure 2 Structural schematic diagram of an online learning behavior recognition system based on multi-view deep behavior recognition of the present invention.
[0037] Figure 3 Schematic diagram of the behavior recognition algorithm based on multi-view multi-modal feature interaction and fusion in an embodiment of the present invention.
[0038] Figure 4 Schematic diagram of the residual network in an embodiment of the present invention.
[0039] Figure 5 、Schematic diagram of the time shift module structure according to an embodiment of the present invention.
[0040] Figure 6 、Schematic diagram of the structure of the cropped multi-scale spatio-temporal convolutional neural network according to an embodiment of the present invention.
[0041] Figure 7 、Schematic diagram of boundary clipping according to an embodiment of the present invention.
[0042] Figure 8 、Schematic diagram of the MOTIEM structure according to an embodiment of the present invention.
[0043] Figure 9 、Schematic diagram of the inter-view and inter-modal recurrent generative adversarial network according to an embodiment of the present invention.
[0044] Figure 10 、Schematic diagram of the HRNet network structure according to an embodiment of the present invention. Detailed implementation manners
[0045] For those skilled in the art to better understand the content of the present invention, make the objectives, technical solutions and advantages of the present invention clearer, the following further details the present invention in conjunction with embodiments and drawings. The illustrative embodiments of the present invention and their descriptions are only used to explain the present invention and are not used as a further limitation of the present invention.
[0046] Embodiment 1:
[0047] As Figure 1 shown, Embodiment 1 of the present invention provides an online learning behavior recognition method based on multi-view depth behavior recognition, including the following steps:
[0048] S1. Obtain the frontal RGB video data, frontal skeleton data, side RGB video data and side skeleton data of students during online learning, and perform preprocessing on them respectively;
[0049] Specifically, first, RGB video data and skeleton data are input, which come from the front-facing camera and the side-facing camera respectively. These data are all selected from a subset that conforms to online learning behavior classification in the NTU RGB+D dataset (RGB-Skeleton Online Learning Behavior Subset, RS-OLBR). This subset includes 5 behavior classifications: playing with mobile phone, watching video, taking notes, sleeping, and eating. The number of videos in each behavior category is roughly the same, and each view contains 1550 videos. All data are divided into a training set and a validation set according to the ratio of 8:2. Each video data is a video clip of 2-5s, and these video clips are processed into frames. The input format of the video frames is k×t×c×h×w, and the skeleton data input is k×m×o, where k represents the number of video clips, t represents the number of frames contained in each video clip, c represents the number of channels, h×w represents the resolution of the video frames, m represents the number of human joint points, and o represents the x, y coordinates and confidence.
[0050] Specifically, in practical applications, the skeleton data can be extracted based on a Figure 10 High-Resolution Network (HRNet) as shown. The feature of HRNet is to fuse high-resolution and low-resolution features at different stages of the network. This network cleverly uses two 3×3 convolutional layers with a stride of 2 to perform 4-fold downsampling on the input image. After this downsampling process, batch normalization (BN) and ReLU activation functions are connected to ensure the effectiveness of the data when input into the network and increase the non-linear expression to improve the quality of feature extraction.
[0051] S2. Construct a feature extraction branch network for RGB video data, and extract a video fusion feature vector according to the front RGB video data and the side RGB video data;
[0052] Specifically, it includes the following steps:
[0053] S201. Perform boundary cropping on the front RGB video data according to the front skeleton data to obtain the front learning behavior video data C a , and perform boundary cropping on the side RGB video data according to the side skeleton data to obtain the side learning behavior video data C b ;
[0054] Specifically, as Figure 7As shown, during training, this embodiment uses skeleton data corresponding to RGB data to guide the cropping of RGB frames. The input of the RGB video data feature extraction branch network is R:k×t×c×h×w, where k represents the number of video segments, t represents the number of frames in each video segment, c represents the number of channels, and h×w represents the resolution of the video. This module needs to extract motion information from RGB video frames, locate the dynamic human body area through skeleton, and crop the redundant background to improve the focus of the subsequent MOTIEM module. First, use the human object detection algorithm MMDection and the human pose estimation algorithm MMPose to detect the human joint points (such as elbows, etc.) in each frame, and record the set of horizontal coordinates of all joint points as X = {x1, x2, x3, …, x n} and the set of vertical coordinates as Y = {y1, y2, y3, …, y n}, where the minimum coordinate value and the maximum coordinate value of the joint points are min(X), max(X), min(Y), max(Y);
[0055] Secondly, determine the initial active area. The input is the extreme values of the joint point coordinates, and the width w i of the initial bounding box is calculated through the extreme values, which is w = max(X) - min(X), and the height h j = max(Y) - min(Y). Then, expand the initial bounding box. Determine the filling direction according to the size relationship between w i and h j . If h j > w i , then expand the horizontal short side by 30%, and the expanded part is Δ x = 0.3w i , expand the vertical long side by 20%, and the expanded part is Δ y = 0.2h j . The adjusted horizontal coordinate range and vertical coordinate range can be expressed as:
[0056] [max(min(X) - Δ x , 1), min(max(X) + Δ x , w)]
[0057] [max(min(Y) - Δ y , 1), min(max(Y) + Δ y , h)]
[0058] Δ x = 0.3w i , w i = max(X) - min(X), Δ y = 0.2h j , h j= max(Y) - min(Y)
[0059] where X is the horizontal coordinate of all joint points, Y is the vertical coordinate of all joint points, c is the number of channels, and h×w is the video resolution; the constant 1 represents the initial pixel value. If the extended coordinates exceed the image boundary, for example, the left boundary is negative or the right boundary exceeds the width w i , then the forced cropping boundary value is [1, w]. If the lower boundary is negative or the upper boundary exceeds the height h j , then the forced cropping boundary value is [1, h], which serves as the fixed area for cropping all frames;
[0060] After that, unified cropping is performed on all sampled RGB frames to ensure that the cropped frames have the same background area. The cropped frames are c×h′×w′, and the RGB frames of the front view and side view obtained after cropping are r a and r b . On the one hand, this module can eliminate the interference of background changes and force the network to focus on dynamic regions. On the other hand, it can reduce the computational amount and improve the computational efficiency of the model.
[0061] S202. Input the front learning behavior video data and side learning behavior video data into the motion time information extraction model to obtain the front spatio-temporal feature map and side spatio-temporal feature map of the student, and generate the front motion map M a and the side motion map M b ;
[0062] Specifically, the video frames are input into the motion time information extraction module MOTIEM as shown in Figure 8 . Directly inputting dense RGB frames into the backbone network will result in high computational costs, and the traditional frame difference method has two defects. One is that the time scale is too long, resulting in overlapping of key context information, and the other is that the manually designed preprocessing is difficult to adapt to different action types. This MOTIEM module compresses the dense RGB frames in the video segment into sparse motion maps, and combines short-term spatio-temporal features and an adaptive attention mechanism to retain key action regions and suppress irrelevant backgrounds. The RGB video cropped by the Cropping module in the previous section is [k, t, c, h′, w′], which is input into the MOTIEM module. First, the inter-frame difference F diffs is calculated. The original RGB frame sequence is [R1, R2, R3,.., R t, the shape of each frame is [c, h′, w′]. The input shape is [t×c, h′, w′], which means that the frames of each segment are unfolded and processed independently. The frames of each segment are arranged in channel order, and the last c×(t - 1) channels and the first c×(t - 1) channels of all segments are taken to calculate t - 1 groups of adjacent frame differences. The frame difference sequence calculated by Equation 2 - 3 is with the shape of [k, (t - 1)×c, h′, w′]; then the cropped RGB frames and the frame differences are concatenated along the channel dimension, and the shape after concatenation is [k, (2t - 1)×c, h′, w′].
[0063] F diffs =R i+1 -R i (i = 1, 2, 3,... t - 1)
[0064] Then, the spatio - temporal features of the video frames are extracted through the m1 convolutional layer, and the shape of the obtained feature map is [k, c, h′, w′], obtaining the spatio - temporal feature map F st , which can be expressed as:
[0065] F st =Conv3×3(Concat(cropped RGB frames, F diffs ))
[0066] To further extract effective motion information, this paper uses F diffs to generate a temporal attention mask. First, the shape of F diffs is reshaped into [k, (t - 1), c, h′, w′]. Then, the most significant motion features are aggregated along the temporal dimension through max - pooling operation, and the shape of the obtained feature map is [k, c, h′, w′]. Then, it is processed through the m2 convolutional layer to keep the number of channels unchanged, and then the temporal attention mask Mask is obtained through the sigmoid activation function, which is used to emphasize the motion regions. Finally, Mask and F st are multiplied element - by - element to obtain Motion maps, and the final output shape is [k, c, h′, w′], which can be expressed as:
[0067] Motion maps = σ(Conv3×3(maxpool(F diffs )))⊙F st
[0068] The obtained Motion maps are input into ResNet to obtain the feature map and This module can enhance the model's attention to the motion regions and ignore unimportant background information.
[0069] S203. Based on the ResNet model, respectively perform feature extraction on the frontal motion map M a and the side motion map M b to obtain the frontal motion feature vector and the side motion feature vector and perform feature fusion on the frontal motion feature vector and the side motion feature vector through an inter-view cycle generative adversarial network to obtain the RGB video fusion feature vector g1.
[0070] Specifically, the motion maps are input into the residual network ResNet to obtain the feature vectors and Then use the Inter-view CycleGAN network shown in the left figure as Figure 9 to perform feature fusion. CycleGAN was first used in the field of image generation. It consists of two generators and two discriminators. The generator is used to generate data in the target domain, and the role of the discriminator is to judge whether the fake data generated by the generator is real target domain data. In the RGB video data feature extraction branch network, although the video data of the front view and the side view express the same semantic information, their expression forms are different. After these feature vectors are preliminarily extracted, they are input into the Inter-view CycleGAN. This network contains multiple convolutional layers and residual blocks. During the mutual generation process, the feature vectors of the front view will gradually approach the feature expression of the side view, and the feature vectors of the side view will gradually approach the feature expression of the front view. Through this mutual generation method, the feature vectors of the two views are concatenated together to form the final fusion feature vector g1. This fusion process ensures the consistency and integrity of the feature representation, thereby enhancing the model's ability to understand and analyze video content.
[0071] Specifically, when processing RGB video data, this embodiment adopts a temporal shift module (Temporal Shift Module, TSM), and the backbone network is the temporal shift-residual network TSM-ResNet50. This network structure is mainly composed of a residual network and a TSM structure. The residual network solves the problem that the performance does not improve with the increase in the number of layers in the deep network by introducing residual blocks.
[0072] Specifically, the residual network is as Figure 4 shown. A residual block contains two paths, where F(x) represents the residual path, and x serves as the identity mapping, also known as the shortcut. The TSM structure can be effectively embedded into a two-dimensional convolutional neural network for efficient temporal dimension modeling. Specifically, we can use A ∈ R N×C×T×H×WRepresents a video clip, where N represents the batch size, C represents the number of channels of the video, H×W represents the resolution of the video frame, and T represents the time dimension. Usually, 2D convolution does not involve the time dimension. As Figure 5 shown, the squares of the same color represent the convolution maps at the same time point. TSM realizes temporal modeling by shifting the channels of the convolution maps generated by 2D convolution forward or backward in the time dimension. For example, some channels are shifted -1 or +1 in the time dimension, so as to enhance the model's ability to capture the temporal changes of video data.
[0073] S3. Construct a backbone data feature extraction branch network, and extract a backbone fusion feature vector according to the frontal backbone data and the lateral backbone data;
[0074] Specifically, in the backbone data feature extraction branch network, the backbone data is first input into the Figure 6 Prune MSSTNet (Pruned Multi-Scale Spatio-Temporal Network) shown as follows. The difference between it and MSSTNet is that it has no large convolution kernels of 11×1 and 1×11. Prune MSSTNet contains 7 convolution modules and 2 pooling modules, and uses one-dimensional convolutions of multiple scales to model the time and space dimensions of the backbone data. The stride in the MSST module is controlled by the 3×3, 5×1, 7×1, and the first 1×1 convolution layers to control the size of the time dimension, and the stride of other convolution layers is 1. The size of the space dimension is controlled by the stride of the pooling module. The settings of different convolution kernel sizes and channel numbers help the network capture features at different scales. Through this structure setting, the network can effectively capture and process time and space information, and enhance the accuracy and efficiency of the model for action recognition. The obtained feature vector and are input into the Inter-View CycleGAN, and the concatenated feature vector g2 is obtained.
[0075] S4. Input the video fusion feature vector and the backbone fusion feature vector into an inter-modal cycle generative adversarial network for feature alignment to obtain a total fusion feature vector;
[0076] Specifically, g1 and g2 are input into the inter-modal cycle generative adversarial network Inter-ModalCycleGAN shown in the Figure 9 right figure. CycleGAN is used between modalities to align the two modalities and make full use of the information of different modalities to obtain the finally concatenated feature vector g. Thus, the recognition accuracy is improved.
[0077] Step 7. Input g into the MLP (Multi-Layer Perceptron) to obtain the final classification of the student's online learning behavior.
[0078] To enhance the representational differences between different view features and heterogeneous modal features and improve the semantic consistency of cross-domain features, this paper designs Inter-View CycleGAN and Inter-Modal CycleGAN for feature fusion between views and modalities respectively. The architecture of this module is as Figure 9 shown.
[0079] When fusing different view features and modal features, the feature maps and are input into CycleGAN, which consists of generators G1 and G2 and discriminators D1 and D2. Specifically, as Figure 9 shown, given different view or modal feature representations and learn the mapping function between the two feature representations. Among them, the forward transformation generator is used to learn the implicit mapping from the source domain to the target domain, and the reverse transformation generator is used to implement the inverse transformation from the target domain to the source domain. At the same time, two adversarial discriminators D1 and D2 are introduced. Among them, D1 is used to discriminate and the generated by G2. Similarly, D2 is used to discriminate and the generated by G1. The objective loss consists of two types of losses: adversarial loss and cycle consistency loss. The adversarial loss is used to match the generated sample distribution with the view data or modal data distribution representation in the target domain; the cycle consistency loss is used to prevent the mappings G1 and G2 learned by the generator from falling into local optima and causing mode collapse problems.
[0080] CycleGAN uses adversarial loss for both mapping functions. For the mapping function and its discriminator D2, in this subsection, the adversarial loss is expressed as:
[0081]
[0082] where G1 attempts to generate samples similar to while D2 is used to discriminate the generated samples from the real samples where
[0083] Similarly, for the mapping function and its discriminator D1, in this subsection, the adversarial loss is expressed as:
[0084]
[0085] Among them, G2 attempts to generate samples similar to while D1 is used to identify the generated samples and real samples Among them
[0086] In the cyclic generative adversarial network, not only can the source domain generate the target domain, but the generated target domain also needs to be able to return to the source domain using another generator. Here, the cyclic consistency loss is used to ensure the effectiveness of the transformation. For each sample from , the view representation transformation should be able to bring the sample back to the original view representation sample, and the modality representation transformation should be able to bring the sample back to the original modality sample, that is This equation is called forward cyclic consistency. Similarly, for each sample from , the view representation transformation should be able to bring the sample back to the original view representation sample, and the modality representation transformation should be able to bring the sample back to the original modality sample, that is This equation is called backward cyclic consistency. As shown in Equation 2-8, in this subsection, the forward consistency loss and the backward consistency loss are used for excitation:
[0087]
[0088] Among them, the first term on the right side of the equal sign is the forward cyclic consistency loss, and the second term on the right side of the equal sign is the backward cyclic consistency loss.
[0089] The complete loss function for this part is:
[0090]
[0091] The goal of this method is to solve:
[0092]
[0093] Extract the intermediate features in the view or modality transformation between the source domain and the target domain and splice them to obtain the dual-view or dual-modality representation g r :
[0094]
[0095] Among them, and are and the generated feature map representations.
[0096] Embodiment 2
[0097] Based on Embodiment 1, as Figure 2 shown, the present invention further provides an online learning behavior recognition system based on multi-view depth behavior recognition, including a user layer, a gateway layer, a service layer and a data layer that are sequentially connected by signal; the user layer is used for students to log in to the system through single sign-on and perform identity authentication; the gateway layer is used to process requests from the user layer; the service layer is used to process business logic; the data layer is used to store structured data, memory data and video file data.
[0098] Specifically, the user layer includes unified authentication. Students perform identity authentication through single sign-on to the system, which allows users to access multiple related systems with one account and password. The user interface is built using the Vue.js framework combined with the ElementUI component library to provide a responsive and highly interactive front-end user interface.
[0099] Specifically, the gateway layer serves as the API gateway of the system, processing all requests from the client. It can perform functions such as routing forwarding, load balancing, and full authentication.
[0100] Specifically, the service layer is responsible for processing specific business logic, including services such as file upload, behavior recognition, statistical evaluation, and anomaly warning.
[0101] Specifically, the data layer includes a MySQL relational database, a Redis database, and a file storage database; the MySQL relational database is used to store structured data; the Redis database is used to store memory data to cache for performance improvement; the file storage database is used to store video data.
[0102] Specifically, PyTorch is used on the algorithm side, and the Flask framework is used for algorithm embedding.
[0103] Its working process is as follows: When students are watching course videos for learning, they need to turn on the computer camera facing the front position, and a camera needs to be set up at the side and rear. The two cameras are used together to recognize the learning behavior of learners during the online learning process, such as when watching course learning videos. The video data collected by the cameras is saved to the backend, and the backend forwards it to the algorithm side. After the algorithm recognizes, it will return the result. The system will perform visual display based on the result of behavior recognition for students and teachers to view. Combining the students' after-class homework and exam situations, teachers can take targeted measures, such as re-watching the video for learning.
[0104] The present invention is described from the viewpoints of purpose of use, efficacy, progressiveness and novelty. It has practical progressiveness and meets the functional enhancement and use requirements emphasized by the patent law. The above description and drawings of this application are only preferred embodiments of this application and do not limit this application. Therefore, all those that are similar or identical to the structure, device, features, etc. of this application, that is, all equivalent substitutions or modifications made according to the scope of the patent application of this application, shall fall within the scope of protection of the patent application of this application.
[0105] The specific implementation manners described above have further elaborated on the purpose, technical solution and beneficial effects of the present invention. It should be understood that the above are only specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An online learning behavior recognition method based on multi-view deep behavior recognition, characterized in that: The following steps are involved: S1. Obtain the front RGB video data, front skeleton data, side RGB video data and side skeleton data of students during online learning, and pre-process them respectively; S2. Construct a branch network for extracting features from RGB video data, and extract a video fusion feature vector based on the front RGB video data and the side RGB video data; S3. Construct a skeleton data feature extraction branch network, and extract a skeleton fusion feature vector based on the front skeleton data and the side skeleton data; S4. Input the video fusion feature vector and the skeleton fusion feature vector into the inter-modality recurrent generative adversarial network for feature alignment to obtain a total fusion feature vector; S5. Input the total fused feature vector into the multi-layer perceptron for classification to obtain the classification results of students’ online learning behavior.
2. The online learning behavior recognition method based on multi-view deep behavior recognition according to claim 1, characterized in that: Step S2 includes the following steps: S201. Crop the front RGB video data according to the front skeleton data to obtain the front learning behavior video data, and crop the side RGB video data according to the side skeleton data to obtain the side learning behavior video data; S202. Input the front learning behavior video data and the side learning behavior video data into the motion time information extraction model to obtain the front spatiotemporal feature map and the side spatiotemporal feature map of the student, and generate the front motion map and the side motion map according to the front spatiotemporal feature and the side spatiotemporal feature respectively; S203. Based on the ResNet model, feature extraction is performed on the front motion map and the side motion map respectively to obtain a front motion feature vector and a side motion feature vector, and feature fusion is performed on the front motion feature vector and the side motion feature vector through an inter-view cyclic generative adversarial network to obtain a video fusion feature vector.
3. The online learning behavior recognition method based on multi-view deep behavior recognition according to claim 2 is characterized in that: In step S201 , the bounding box range of the boundary clipping is determined based on the maximum and minimum values of the skeleton coordinates in the skeleton data.
4. The online learning behavior recognition method based on multi-view deep behavior recognition according to claim 2, characterized in that: Step S202 includes the following steps: S202-1. Densely sampling the front learning behavior video data and the side learning behavior video data, and compressing the sampled front learning behavior video data and the side learning behavior video data to obtain a sparse set of front learning behavior videos and a sparse set of side learning behavior videos, respectively; S202-2. Calculate the frame difference between adjacent frames in the sparse set of front learning behavior videos to obtain a first frame difference, and calculate the frame difference between adjacent frames in the sparse set of side learning behavior videos to obtain a second frame difference; S202-3. Extract features from the first frame difference and the sparse set of front learning behavior videos by convolution kernel to obtain a front spatiotemporal feature map, and extract features from the second frame difference and the sparse set of side learning behavior videos by convolution kernel to obtain a side spatiotemporal feature map; S202-4. Perform maximum pooling processing on the first frame difference and the second frame difference respectively, and generate a first time attention mask and a second time attention mask respectively through a convolution layer and a sigmoid function; S202-5. Multiply the front spatiotemporal feature map with the first time attention mask to obtain a front motion map, and multiply the side spatiotemporal feature map with the second time attention mask to obtain a side motion map.
5. The online learning behavior recognition method based on multi-view deep behavior recognition according to claim 1, characterized in that: The front skeleton data and the side skeleton data are extracted based on the HRNet network.
6. The online learning behavior recognition method based on multi-view deep behavior recognition according to claim 1, characterized in that: Step S3 includes the following steps: S301. Input the front skeleton data and the side skeleton data into the Prune MSSTNet network, model the time and space dimensions of the front skeleton data, and model the time and space dimensions of the side skeleton data; S302. Calculate the time and space dimension model of the front skeleton data to obtain the feature vector of the front skeleton data, and calculate the time and space dimension model of the side skeleton data to obtain the feature vector of the side skeleton data; S303. The front skeleton data feature vector and the side skeleton data feature vector are fused through an inter-view cyclic generative adversarial network to obtain a skeleton fusion feature vector.
7. An online learning behavior recognition system based on multi-view deep behavior recognition, the system is implemented based on the online learning behavior recognition method based on multi-view deep behavior recognition as claimed in any one of claims 1 to 6, characterized in that: It includes a user layer, a gateway layer, a service layer and a data layer which are connected in sequence by signals; the user layer is used for students to log in to the system through a single point and perform identity authentication; the gateway layer is used to process user layer requests; and the service layer is used to process business logic; The data layer is used to store structured data, memory data and video file data.
8. The online learning behavior recognition system based on multi-view deep behavior recognition according to claim 7, characterized in that: The user layer includes unified authentication, and students perform identity authentication through a single sign-on system, enabling students to access at least two systems using one account and password.
9. The online learning behavior recognition system based on multi-view deep behavior recognition according to claim 7, characterized in that: The business logic includes file uploading, behavior recognition, statistical evaluation and abnormal warning.
10. The online learning behavior recognition system based on multi-view deep behavior recognition according to claim 7, characterized in that: The data layer includes a MySQL relational database, a Redis database and a file storage database; the MySQL relational database is used to store structured data; the Redis database is used to store memory data; and the file storage database is used to store video data.
Citation Information
Cited By
Bone tumor classification model training method, classification method and system based on multi-view fusion
CN120783137A
Teaching plan content generation method based on generative adversarial network and hierarchical teaching target analysis
CN121071193A
Method for generating teaching plan content based on generative adversarial network and hierarchical teaching objectives
CN121071193B
Student classroom behavior intelligent analysis method based on video semantic understanding
CN121963058A