Student behavior data analysis method and system based on transformer model
Patent Information
- Application Number
- CN202610837568.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-08-28
AI Technical Summary
具体而言,数据采集环节多依赖单一维度感知设备,缺乏对课堂场景中多元行为的全面捕捉;数据处理过程以基础统计分析为主,仅能实现行为数据的表面归类;分析结果呈现多为静态数据报表,缺乏对行为与学情关联的深度解读,整体流程侧重于数据的初步处理与展示,未形成从采集到研判的全流程智能闭环
[0007] Beneficial Effects: This invention proposes a student behavior data analysis method and system based on the Transformer model. It comprehensively captures diverse classroom behavior data through multimodal acquisition devices, mines the temporal dimension features of the behavior data through a temporal correlation identification mechanism, deeply integrates multidimensional behavioral features using a specific fusion model, and accurately classifies student learning status through hierarchical correlation reasoning. Finally, it completes data storage, analysis, and report output through an intelligent judgment platform, with each unit collaboratively constructing a full-process intelligent analysis system. Through multi-dimensional data acquisition and temporal feature capture, it overcomes the shortcomings of traditional technologies in data acquisition—namely, the one-sidedness of data collection and the difficulty in capturing dynamic correlations—achieving comprehensive coverage and in-depth mining of behavioral data. Furthermore, through feature fusion and hierarchical reasoning mechanisms, it addresses the insufficient accuracy of existing technologies in interpreting student learning status, enabling a refined presentation of student learning status levels. Simultaneously, the heterogeneous computing architecture and high-speed communication design at the hardware level, combined with optimized parameter settings at the software level, significantly improves data processing efficiency and the reliability of analysis results, providing support for personalized teaching decisions.
Smart Images

Figure CN122654546A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of student behavior data analysis technology, and in particular to a method and system for student behavior data analysis based on the Transformer model. Background Technology
[0002] In the process of digital transformation in education, classroom student behavior data contains rich information about student learning. Accurately mining this data plays a crucial supporting role in optimizing teaching strategies and improving teaching quality. In current classroom teaching scenarios, student behavior exhibits multidimensional and dynamic characteristics. Traditional data collection and analysis methods struggle to fully capture the underlying learning relationships behind these behaviors. There is an urgent need to build an efficient and intelligent data analysis method and system to achieve a deep transformation from behavioral data to learning insights, meeting the dual needs of personalized teaching and large-scale teaching management. Currently, student behavior data is mainly collected through manual observation and recording or single-sensor devices. Information is processed using simple statistical and classification methods, and then the analysis results are presented through conventional software platforms. Specifically, the data collection stage relies heavily on single-dimensional sensing devices, lacking comprehensive capture of diverse behaviors in the classroom scenario; the data processing stage focuses on basic statistical analysis, only achieving superficial classification of behavioral data; and the analysis results are mostly presented as static data reports, lacking in-depth interpretation of the relationship between behavior and learning. The overall process emphasizes preliminary data processing and display, failing to form a fully intelligent closed loop from collection to analysis.
[0003] Existing technologies have two significant drawbacks: First, the comprehensiveness of data collection and processing is insufficient. Limited by single collection methods and basic processing means, they cannot effectively capture the temporal correlation and multi-dimensional feature fusion information of behavioral data, making it difficult for the analysis results to reflect the true learning orientation of students' behavior. Second, the accuracy of learning interpretation is lacking. There is a lack of in-depth mining mechanism for the intrinsic relationship between behavioral data and learning status. It can only provide general behavioral statistics results and cannot achieve refined stratification and accurate judgment of learning status, making it difficult to support the formulation and implementation of personalized teaching decisions. Summary of the Invention
[0004] In order to overcome the shortcomings and deficiencies of existing technologies, this invention provides a method and system for analyzing student behavior data based on the Transformer model.
[0005] The technical solution adopted in this invention is a student behavior data analysis method based on the Transformer model, comprising the following steps: S1, collecting multi-dimensional behavioral data of students in classroom scenarios through a multi-dimensional intelligent judgment and analysis platform for classroom behavior, wherein the multi-dimensional behavioral data includes action behavior data, interaction behavior data, attention-related data, and expression behavior data; S2, calling a time-series-aware student behavior identification model to identify the time-series correlation of the collected multi-dimensional behavioral data, and generating a time-series behavioral feature sequence by capturing the dynamic correlation features of different behavioral data in the time dimension; S3, inputting the time-series behavioral feature sequence into the multi-dimensional behavioral feature fusion Transformer model. The Mer model uses a multi-head attention mechanism to weight and deeply fuse behavioral features across different dimensions, generating a fused feature matrix. S4 employs a hierarchical learning status association reasoning algorithm to perform hierarchical reasoning on the fused feature matrix, determining the learning status level for each student based on preset learning status classification standards. S5 utilizes a multi-dimensional intelligent judgment and analysis platform for classroom behavior to associate and store learning status levels and corresponding behavioral feature data, generating multi-dimensional behavioral analysis results. S6, based on various parameters of student behavioral data analysis, features are extracted and associated with the intelligent judgment results, outputting a behavioral analysis report corresponding to the learning status.
[0006] This student behavior data analysis system based on the Transformer model comprises: a multi-dimensional behavior data acquisition unit, a temporal behavior feature identification unit, a multi-dimensional feature fusion processing unit, a hierarchical reasoning unit for student learning status, an intelligent judgment and analysis unit, and a result output and storage unit. The multi-dimensional behavior data acquisition unit and the temporal behavior feature identification unit are connected via a high-speed data bus to transmit the acquired multi-dimensional behavior data to the temporal behavior feature identification unit. The temporal behavior feature identification unit is communicatively connected to the multi-dimensional feature fusion processing unit, processing the data through a temporal-aware student behavior identification model and outputting a temporal behavior feature sequence to the multi-dimensional feature fusion processing unit. The multi-dimensional feature fusion processing unit is equipped with multiple... The system integrates a Transformer model with multidimensional behavioral features and establishes a data interaction channel with the hierarchical reasoning unit for learning status, transmitting the integrated feature matrix to the learning status hierarchical reasoning unit. The learning status hierarchical reasoning unit employs a hierarchical association reasoning algorithm for learning status, communicating bidirectionally with the intelligent judgment and analysis unit to output the learning status hierarchy to the intelligent judgment and analysis unit. The intelligent judgment and analysis unit embeds a setting module of the multidimensional intelligent judgment and analysis platform for classroom behavior, connecting with the result output and storage unit to transmit multidimensional behavioral analysis results to the result output and storage unit. Based on various parameters of student behavioral data analysis, the result output and storage unit structures and stores the analysis results and outputs a behavioral analysis report in a preset format. All units synchronize data and collaborate through a distributed communication protocol.
[0007] Beneficial Effects: This invention proposes a student behavior data analysis method and system based on the Transformer model. It comprehensively captures diverse classroom behavior data through multimodal acquisition devices, mines the temporal dimension features of the behavior data through a temporal correlation identification mechanism, deeply integrates multidimensional behavioral features using a specific fusion model, and accurately classifies student learning status through hierarchical correlation reasoning. Finally, it completes data storage, analysis, and report output through an intelligent judgment platform, with each unit collaboratively constructing a full-process intelligent analysis system. Through multi-dimensional data acquisition and temporal feature capture, it overcomes the shortcomings of traditional technologies in data acquisition—namely, the one-sidedness of data collection and the difficulty in capturing dynamic correlations—achieving comprehensive coverage and in-depth mining of behavioral data. Furthermore, through feature fusion and hierarchical reasoning mechanisms, it addresses the insufficient accuracy of existing technologies in interpreting student learning status, enabling a refined presentation of student learning status levels. Simultaneously, the heterogeneous computing architecture and high-speed communication design at the hardware level, combined with optimized parameter settings at the software level, significantly improves data processing efficiency and the reliability of analysis results, providing support for personalized teaching decisions. Attached Figure Description
[0008] Figure 1 This is a flowchart illustrating the overall process of the method of the present invention. Figure 2 This is a flowchart of method step S2 of the present invention; Figure 3 This is a flowchart of method step S3 of the present invention; Figure 4 This is a flowchart of method step S4 of the present invention; Figure 5 This is a flowchart of step S5 of the method of the present invention; Figure 6 This is a diagram showing the system unit composition of the present invention. Detailed Implementation
[0009] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0010] like Figure 1As shown, the student behavior data analysis method based on the Transformer model includes the following steps: S1, collecting multi-dimensional behavioral data of students in classroom scenarios through a multi-dimensional intelligent judgment and analysis platform for classroom behavior, wherein the multi-dimensional behavioral data includes action behavior data, interaction behavior data, attention-related data, and expression behavior data; S2, calling the time-series-aware student behavior identification model to identify the time-series correlation of the collected multi-dimensional behavioral data, and generating a time-series behavioral feature sequence by capturing the dynamic correlation features of different behavioral data in the time dimension; S3, inputting the time-series behavioral feature sequence into the multi-dimensional behavioral feature fusion Transformer model. The model employs a multi-head attention mechanism to weight and deeply fuse behavioral features across different dimensions, generating a fused feature matrix. In step S4, a hierarchical reasoning algorithm for learning status is used to perform hierarchical reasoning on the fused feature matrix, determining the learning status level for each student based on preset learning status classification standards. In step S5, a multi-dimensional intelligent judgment and analysis platform for classroom behavior is used to associate and store the learning status levels and corresponding behavioral feature data, generating multi-dimensional behavioral analysis results. Finally, based on various parameters from the student behavioral data analysis, feature extraction and association mapping are performed on the intelligent judgment results, outputting a behavioral analysis report corresponding to the learning status.
[0011] Step S1 involves the comprehensive collection of multi-dimensional behavioral data of students in the classroom through a multi-dimensional intelligent analysis platform for classroom behavior. The platform is equipped with multi-modal data acquisition terminals, including high-definition visual sensors, sound acquisition modules, motion capture sensors, and interaction status monitoring units. Each terminal establishes a high-speed communication connection with the edge computing node cluster via 5G gigabit Ethernet to ensure the real-time performance and stability of data transmission. The edge computing node cluster is configured with a 16-core heterogeneous processor and a dedicated neural network acceleration chip, supporting parallel processing of 128 channels of multi-dimensional behavioral data. During the collection process, four core data categories are acquired simultaneously: action behavior data, interaction behavior data, attention-related data, and expressive behavior data. Action behavior data includes information such as changes in posture and the amplitude of limb movements; interaction behavior data includes the frequency of teacher-student interaction and the duration of student-student interaction; attention-related data involves indicators such as the duration of eye contact and the frequency of distraction; and expressive behavior data includes information such as the number of times a student speaks and the duration of their speech. The quantification of these four core data categories is based on a classroom unit time (default 45 minutes), combined with specific standards formulated from objectively collectable dimensions: In the action behavior data, posture changes are quantified by identifying the number of times "standard sitting posture → non-standard sitting posture" is switched through posture sensors (e.g., each obvious posture deviation is counted as 1 time), and the amplitude of limb movements is assigned a quantification value of 1 / 2 / 3 according to three levels of "slight / moderate / violent" (based on the action coverage and speed threshold); In the interaction behavior data, the frequency of teacher-student interaction is directly counted by counting the number of times students actively raise their hands and answer teachers' questions, and the duration of student-student interaction is recorded by audio devices and seat sensors in collaboration with each other to record the cumulative number of seconds of effective communication between students; In the attention-related data, the duration of eye focus is counted by counting the cumulative number of seconds that students' eyes stay on the blackboard, screen, or book through visual tracking technology, and the frequency of inattentiveness is judged by the criterion of "eyes deviating from the effective area for more than 3 seconds", and the number of times it occurs per unit time is counted; In the expression behavior data, the number of times students speak is directly counted by counting the number of times students actively speak in class, and the speaking time is the cumulative effective speech time of each speech (excluding pauses and invalid noise). Data normalization was performed using the Min-Max standardization method. First, the minimum (Min) and maximum (Max) values in the entire student sample were determined for each type of indicator. Then, the normalized value was mapped to the [0,1] interval using the formula "normalized value = (original data - Min) / (Max - Min)". For inverse indicators such as the number of times sitting posture changes and the frequency of distraction (the higher the value, the more negative the tendency), the "Max - original data" inverse conversion was performed before normalization to ensure that all indicators were consistent in direction. This ultimately achieved the comparability and applicability of data from different dimensions and scales for integrated analysis.The data collection process strictly adheres to the principle of data integrity. The data collection frequency for each channel is set to 30 frames per second, and the number of bytes stored in a single data entry is controlled within a preset range. Through multi-terminal collaborative collection, the data achieves comprehensive coverage of student behavior data in the classroom setting, providing comprehensive and accurate raw data support for subsequent behavior analysis. After the data collection is completed, it is initially screened and then transmitted to the platform's temporary storage module, awaiting subsequent processing.
[0012] Step S2 invokes the time-series-aware student behavior identification model to identify the temporal correlation of the collected multi-dimensional behavior data. The model runs on the heterogeneous computing unit FPGA logic processing module of the classroom behavior multi-dimensional intelligent judgment and analysis platform to ensure the efficiency of feature extraction. During implementation, the collected multi-dimensional behavior data is first sorted according to timestamp order to establish a time-behavior data mapping relationship. Then, the time interval between different behavior data is calculated through the model's built-in time-series detection module to accurately identify the correlation strength of each dimension of data in the time dimension. The correlation strength calculation uses a 10-second time window, and the data correlation threshold within the window is set to a preset standard value. Based on the calculated time correlation strength, dynamic weights are assigned to each dimension of behavior data. The weight coefficient ranges from 0.1 to 0.9; the higher the correlation strength, the larger the corresponding weight coefficient, achieving focused attention on behavior data at key time nodes. Subsequently, the model's feature generation module extracts temporal features from the weighted behavioral data. The extraction process uses a sliding window method with a window size of 20 seconds and a step size of 5 seconds. By aggregating and filtering the features of the data within the window, a temporal behavioral feature sequence including time dimension information is generated. The sequence length is set to 1024 dimensions to ensure that the feature sequence can fully reflect the temporal dynamic changes of student behavior and provide high-quality temporal feature data for subsequent feature fusion.
[0013] Step S3 inputs the temporal behavioral feature sequence into the multi-dimensional behavioral feature fusion Transformer model. The model is deployed on the GPU computing module of the classroom behavior multi-dimensional intelligent judgment and analysis platform. This module supports multi-threaded parallel computing, significantly improving the model's processing speed. First, the temporal behavioral feature sequence is converted into a vector format recognizable by the model, constructing a feature input matrix with a dimension of 512×1024. Each element in the matrix corresponds to the quantized value of a single behavioral feature. The model's multi-head attention mechanism is then activated, with 8 attention heads. Each attention head independently performs correlation calculations on different dimensions of features in the feature input matrix. During the calculation, the cosine similarity between features is used as the basis for correlation determination, with similarity values ranging from -1 to 1. The weight coefficients of each dimension feature are dynamically adjusted based on the correlation calculation results. The sum of the weight coefficients is normalized to 1. Key behavioral features with correlation higher than a preset threshold are assigned high weights of 0.6 to 0.9, while irrelevant features with correlation lower than a preset threshold are assigned low weights of 0.1 to 0.3, thereby strengthening key features and weakening irrelevant features. Finally, the weighted feature matrix is nonlinearly transformed through the fully connected layer of the model. The fully connected layer includes two hidden layers. The number of neurons in the first hidden layer is set to 1024, and the number of neurons in the second hidden layer is set to 512. The activation function adopts a preset nonlinear function. Through layer-by-layer feature transformation, a 512-dimensional deep fusion feature matrix is generated. This matrix comprehensively integrates the key information and correlation of behavioral features in various dimensions, providing strong feature support for subsequent learning status inference.
[0014] Step S4 employs a hierarchical association reasoning algorithm based on learning status to perform hierarchical reasoning operations on the fused feature matrix. The algorithm runs on the edge computing node cluster of the multi-dimensional intelligent judgment and analysis platform for classroom behavior, leveraging the parallel computing capabilities of a 16-core heterogeneous processor to improve reasoning efficiency. First, the 512-dimensional fused feature matrix is divided into feature dimensions, categorized by behavior type into four reasoning channels: action features, interaction features, attention features, and expression features. Each channel corresponds to a 256-dimensional feature sub-matrix. The algorithm's reasoning engine is then invoked to perform hierarchical matching calculations on the feature sub-matrices of each reasoning channel. The matching process is based on preset learning status feature templates, which include feature standard values corresponding to four learning status levels: excellent, good, average, and needing improvement. The Euclidean distance between the feature sub-matrix and each level template is used as the matching degree index; a smaller distance indicates a higher matching degree. Based on the matching degree results of each channel and combined with the preset learning status judgment rules, the rules stipulate that when the matching degree of all four channels is higher than 0.8, it is judged as an excellent level; when the matching degree of three channels is higher than 0.7, it is judged as a good level; when the matching degree of two channels is higher than 0.6, it is judged as a medium level; and the rest are judged as levels that need improvement. This initially determines the candidate learning status levels for each student. The conflict resolution module of the algorithm is then activated to perform consistency checks on the initially determined candidate levels. If two or more conflicting level judgment results exist during the check, the level with the highest total matching degree is taken as the final result, and contradictory level information is eliminated. Finally, a unique learning status level corresponding to each student is output. The accuracy of level judgment is controlled above the preset standard, providing accurate learning status data for subsequent intelligent analysis.
[0015] Step S5 utilizes a multi-dimensional intelligent analysis platform for classroom behavior to correlate and intelligently analyze the learning status levels and corresponding behavioral characteristic data. The platform's distributed storage server employs a storage architecture combining a distributed file system and a time-series database, with data read / write response latency controlled within 50 milliseconds. First, an index is established linking the learning status levels and corresponding behavioral characteristic data, using the student's unique identifier as the key and the learning status level and behavioral characteristic data of each dimension as the value, forming a key-value pair storage structure. A hash algorithm is used during index creation to ensure efficient index queries. The indexed dataset is then transmitted to the platform's storage module via 5G gigabit Ethernet, partitioned according to learning status levels and data types. Each storage partition has independent access permissions and a data backup mechanism to ensure data storage security and integrity. The platform's intelligent analysis engine is then activated. The engine employs a multi-threaded parallel analysis mechanism, simultaneously performing multi-dimensional correlation analysis on 128 data streams. Analysis dimensions include the correlation between behavioral characteristics and learning status, the distribution of behavioral characteristics among students at different learning status levels, and the behavioral change trends of the same student over different time periods. By calculating the correlation coefficients between behavioral characteristics of various dimensions and learning status levels, key behavioral factors affecting learning status are identified. Statistics such as the mean and variance of students at different learning status levels on each behavioral characteristic are statistically analyzed. The patterns of student behavior changes over time are analyzed. Based on these analysis results, multi-dimensional behavioral analysis results are generated, including behavioral trend curves, feature distribution histograms, and level proportion pie charts. The amount of data in the analysis results is controlled within a preset range to ensure the efficiency of subsequent output.
[0016] Step S6, based on the parameters of student behavior data analysis, performs feature extraction and correlation mapping on the intelligent assessment results, ultimately outputting a behavior analysis report corresponding to the student's learning status. The parameters for student behavior data analysis include feature extraction threshold, correlation mapping coefficient, and report generation format parameters. The feature extraction threshold is set to 0.7, retaining only key behavioral features with a correlation higher than this threshold. The correlation mapping coefficient is differentiated according to different learning status levels: 0.9 for excellent levels, 0.8 for good levels, 0.7 for average levels, and 0.6 for levels needing improvement. During implementation, key behavioral features are first extracted from the intelligent assessment results, including core behavioral indicators for each learning status level, key nodes in behavioral change trends, and behavioral differences between students at different levels. The extraction process uses a feature importance ranking algorithm, sorting features from highest to lowest according to their impact on the learning status, selecting the top 50 key features. Subsequently, the extracted key features are mapped to their corresponding learning status levels, establishing a feature-learning status association table. This clarifies the interpretation of learning status corresponding to different feature combinations. The mapping process strictly adheres to preset association rules to ensure the accuracy of the mapping results. Finally, according to preset report generation format parameters, the mapping results are organized into a structured behavioral analysis report. The report includes three parts: individual student behavior analysis, overall class learning status analysis, and key behavior improvement suggestions. The individual analysis section details each student's learning status level, core behavioral characteristics, and behavioral change trends. The overall analysis section presents the proportion of each learning status level and the distribution of behavioral characteristics within the class. The improvement suggestions section provides targeted optimization directions based on key behavioral characteristics. The report output format supports two common formats: PDF and Excel. The output process is completed through the platform's result output unit, with an output speed controlled at 10 reports per second. Simultaneously, the reports are stored on a distributed storage server, forming a complete archive of analysis results, providing comprehensive and accurate reference for teaching decisions.
[0017] Preferably, the expression of the time-aware student behavior identification model includes: in, This represents the temporal behavior characteristic value at time t. Indicates the number of dimensions in the behavioral data. This represents the time-series weighting coefficient of the i-th dimension of the data. This represents the characteristic function of the i-th dimension of the data at time t. This represents the time-series decay coefficient of the i-th dimension of the data. This represents the time interval of the i-th dimension of the data. Represents a temporal behavioral feature sequence, where Conv1d represents a one-dimensional convolution operation. Indicates the kernel size. This indicates the step size, and MaxPool1d represents the one-dimensional max pooling operation. Indicates the pooling window size. This indicates a feature concatenation operation.
[0018] Specifically, the temporal awareness student behavior identification model is based on the characteristics of temporal continuity and dimensional differences in classroom student behavior. It first quantifies single-moment features and then constructs a complete logical chain through sequence building. The model aims to accurately capture the feature contributions of behavioral data across various dimensions at specific moments. Considering the varying importance of different behavioral data in reflecting learning progress, a temporal weight coefficient is introduced. Simultaneously, the effectiveness of behavioral data decays with increasing time intervals; therefore, a temporal decay coefficient and an S-shaped function are incorporated to construct a decay mechanism, achieving dynamic quantification of single-moment features. The number of behavioral data dimensions is determined to be four based on the four categories of data collected: actions, interactions, focus, and expression. The temporal weight coefficient ranges from 0.1 to 0.9, and its allocation is optimized through sample training based on the degree of influence of behavioral data on learning progress. The temporal decay coefficient ranges from 0.05 to 0.2 to ensure the reasonable impact of time intervals on features. Based on single-moment feature values, a model is established to extract temporally related features. One-dimensional convolution operations capture local temporal features, combined with max pooling operations to reduce data dimensionality while retaining key information. Finally, feature concatenation integrates the results of the two operations to form a complete temporal behavioral feature sequence. Regarding parameters, the convolution kernel size was set to 3, the stride to 1, and the pooling window size to 2. Extensive experiments verified that this parameter combination effectively balances feature extraction performance and computational efficiency. In implementation, the feature values at each time point are first calculated. Then, all single-time-point feature values are arranged chronologically, and convolution, pooling, and concatenation operations are performed to ultimately generate a temporal behavioral feature sequence, providing foundational data including temporal dimension information for subsequent feature fusion.
[0019] Preferably, the expression for the multidimensional behavioral feature fusion Transformer model includes: in, This represents the attention weight between the i-th query vector and the j-th key vector. This represents the i-th query vector. Let j represent the j-th key vector. This represents the dimension of the query vector and the key vector. Indicates the total number of vectors. The fusion feature matrix is represented by LayerNorm, the layer normalization operation is represented by MultiHead, and the multi-head attention computation function is represented by MultiHead. Represents the query vector matrix. Represents the key vector matrix, Represents a value vector matrix. This represents the output weight matrix. This represents the output bias vector.
[0020] Specifically, a multi-dimensional behavioral feature fusion Transformer model is constructed to meet the correlation requirements of multi-dimensional student behavioral features. The attention weight calculation formula is based on the strength of the correlation between behavioral features of different dimensions, which determines their fusion weight. The similarity between the query vector and the key vector is calculated, and then normalized using the softmax function to obtain the attention weights between each feature, achieving focused attention on key correlated features. In the parameters, the total number of vectors is set to 1024, consistent with the length of the temporal behavioral feature sequence, and the dimensions of the query vector and key vector are set to 512. This dimension value ensures computational accuracy while controlling the computational load. The feature fusion formula, based on the attention weight calculation results, processes feature correlation in parallel through a multi-head attention mechanism, then adds residual connections to retain the original temporal feature information. After layer normalization to stabilize the training process, feature mapping is performed through a fully connected layer to generate a deeply fused feature matrix. The output weight matrix dimension is set to 512×512, and the output bias vector dimension is consistent with the output weight matrix. Layer normalization uses a standardization method with a mean of 0 and a variance of 1. In implementation, the temporal behavioral feature sequence is first converted into three types of vector matrices: query, key, and value. The attention weights are then calculated by substituting them into the first formula. The weights are then combined with the value vector matrix, and the second formula is used to complete operations such as multi-head attention calculation, residual connection, layer normalization, and fully connected mapping. This achieves deep fusion of multi-dimensional behavioral features, solves the problem of one-sided information in a single feature dimension, and improves the feature representation capability.
[0021] Preferably, the expression for the hierarchical association reasoning algorithm for learning status is: ;in, This indicates the final determined level of learning progress. This represents the set of all learning status levels. This indicates the number of features in the fused feature matrix. This represents the m-th eigenvalue in the fused feature matrix. This represents the conditional probability of obtaining a learning status level of I based on the inference of the m-th feature value. The prior weight coefficient represents the learning status level I.
[0022] Specifically, the learning status hierarchical correlation reasoning algorithm is constructed based on the hierarchical characteristics of learning status and the correlation with behavioral features, using a combination of probabilistic reasoning and prior knowledge. The learning status level is jointly determined by the conditional probabilities of each feature in the fusion feature matrix. Simultaneously, prior weight coefficients are introduced to reflect the statistical distribution characteristics of different learning status levels. By weighting the product of the conditional probabilities of each feature with the prior weight coefficients, the level corresponding to the maximum value is taken as the final judgment result, ensuring the accuracy and rationality of the reasoning. The number of features in the fusion feature matrix is set to 512, consistent with the feature fusion dimension. The set of all learning status levels includes four levels: excellent, good, average, and need improvement. The prior weight coefficients are determined through statistical analysis of a large amount of sample data: 0.3 for excellent, 0.35 for good, 0.25 for average, and 0.1 for need improvement. These values conform to the general distribution pattern of classroom learning status. During implementation, each feature value is first extracted from the fusion feature matrix. Based on the preset learning status feature template, the conditional probability of each feature value corresponding to each learning level is calculated. Then, the conditional probabilities of each feature are multiplied together and multiplied by the prior weight coefficient of the corresponding learning level to obtain the comprehensive score of each level. Finally, the level with the highest comprehensive score is selected as the final learning status level, realizing the accurate mapping from fusion features to learning status and providing core data support for subsequent intelligent judgment.
[0023] Preferably, the multi-dimensional intelligent judgment and analysis platform for classroom behavior includes an edge computing node cluster, a multimodal data acquisition terminal, a distributed storage server, and a heterogeneous computing unit. The multimodal data acquisition terminal includes a high-definition visual sensor, a sound acquisition module, a motion capture sensor, and an interactive status monitoring unit. Each terminal communicates with the edge computing node cluster via 5G gigabit Ethernet. The edge computing node cluster is configured with a 16-core heterogeneous processor and a dedicated neural network acceleration chip, supporting parallel processing of 128 channels of multi-dimensional behavioral data. The distributed storage server adopts a storage architecture combining a distributed file system and a time-series database, with data read / write response latency controlled within a preset range. The heterogeneous computing unit includes a GPU computing module and an FPGA logic processing module, wherein the GPU computing module is used for model inference calculation, the FPGA logic processing module is used for feature extraction calculation, the number of model training iterations at the platform software level is set to a preset threshold, the number of heads in the attention mechanism is set to 8, and the feature fusion dimension is set to 512.
[0024] Preferred, such as Figure 2As shown, S2 includes the following sub-steps: S21, calling the input interface of the time-series-aware student behavior identification model to sort the collected multi-dimensional behavior data according to the timestamp order and establish a time-behavior data mapping relationship; S22, using the time-series detection module built into the model to calculate the time interval of the sorted behavior data and identify the time correlation strength between different behavior data; S23, dynamically assigning weights to each dimension of behavior data based on the time correlation strength, with the weight assignment result being positively correlated with the time correlation strength; S24, using the model's feature generation module to extract time-series features from the weighted behavior data and generate a time-series behavior feature sequence including time dimension information.
[0025] Specifically, step S2 generates the temporal behavior feature sequence through a step-by-step process. S21 calls the input interface of the temporal-aware student behavior recognition model, rigorously sorting the collected multi-dimensional behavior data (action, interaction, focus, and expression) according to their timestamps, with sorting accuracy down to the millisecond level to ensure the accuracy of the data's temporal dimension. Simultaneously, a one-to-one time-behavior data mapping relationship is established, stored in a key-value pair structure for easy subsequent retrieval and retrieval. S22 uses the model's built-in temporal detection module to calculate the time intervals of the sorted behavior data, with a calculation step size set to 10 milliseconds. The time interval of each behavior data is obtained by the difference between two consecutive timestamps. Based on the magnitude of the time interval, the temporal correlation strength between different behavior data is identified, categorized into three levels: strong, medium, and weak. The threshold for this classification is determined through extensive sample training. S23 dynamically assigns weights to behavioral data across dimensions based on the identified temporal correlation strength. The weight values range from 0.1 to 0.9, with strong correlations having a weight coefficient of 0.7 to 0.9, medium correlations 0.4 to 0.6, and weak correlations 0.1 to 0.3, ensuring a strict positive correlation between the weight assignment and the temporal correlation strength. S24 uses the model's feature generation module to extract temporal features from the weighted behavioral data. The extraction process employs a sliding window method with a window size of 20 seconds and a step size of 5 seconds. Features are aggregated and filtered within each window, ultimately generating a fixed-dimensional temporal behavioral feature sequence. This provides structurally sound and informationally complete foundational data for subsequent multi-dimensional feature fusion.
[0026] Preferred, such as Figure 3As shown, step S3 includes the following sub-steps: S31, converting the temporal behavioral feature sequence into a vector format recognizable by the Transformer model to construct a feature input matrix; S32, activating the multi-head attention mechanism of the multi-dimensional behavioral feature fusion Transformer model to perform correlation calculations on features of different dimensions in the feature input matrix; S33, adjusting the weight coefficients of each dimension feature based on the correlation calculation results to strengthen the influence of the calibrated behavioral features and weaken the interference of irrelevant features; S34, performing a nonlinear transformation on the weighted feature matrix through the fully connected layer of the model to generate a deeply fused feature matrix.
[0027] Specifically, step S3 achieves deep fusion of multi-dimensional behavioral features. S31 first converts the generated temporal behavioral feature sequence into a fixed-dimensional vector format according to the input requirements of the Transformer model. The vector dimension is set to 512 dimensions. Then, a feature input matrix is constructed based on the converted vector. The number of rows in the matrix is the same as the length of the temporal behavioral feature sequence, and the number of columns is 512. The value range of each element in the matrix is controlled within a preset range after normalization to ensure efficient model processing. S32 activates the multi-head attention mechanism of the Transformer model for multi-dimensional behavioral feature fusion. The number of attention heads is set to 8. Each attention head independently performs correlation calculations on different dimensions of features in the feature input matrix. The similarity between features is the core indicator during the calculation. The similarity calculation uses the cosine similarity algorithm, and large-scale feature fast correlation analysis is achieved through matrix operations. S33 dynamically adjusts the weight coefficients of each dimension of features based on the correlation calculation results. The sum of the weight coefficients is normalized to 1. Key behavioral features with correlation higher than a preset threshold (threshold set to 0.7) are assigned high weights of 0.6 to 0.9, while irrelevant features with correlation lower than the preset threshold are assigned low weights of 0.1 to 0.3. This weight adjustment strengthens the influence of key features and weakens the interference of irrelevant features. S34 performs a nonlinear transformation on the weighted feature matrix through the fully connected layer of the model. The fully connected layer includes two hidden layers. The number of neurons in the first hidden layer is set to 1024, and the number of neurons in the second hidden layer is set to 512. A specific nonlinear function is used as the activation function. Through layer-by-layer feature transformation and information integration, a 512-dimensional deep fusion feature matrix is finally generated. This matrix comprehensively integrates the key information and intrinsic correlations of behavioral features in each dimension, providing high-quality feature support for learning state reasoning.
[0028] Preferred, such as Figure 4As shown, S4 includes the following sub-steps: S41, dividing the fused feature matrix into feature dimensions and assigning different types of behavioral features to the corresponding inference channels; S42, calling the inference engine of the learning status hierarchical association inference algorithm to perform hierarchical matching calculation on the feature data of each inference channel; S43, based on the hierarchical matching calculation results and combined with the preset learning status judgment rules, initially determining the candidate level of each student's learning status; S44, performing consistency verification on the candidate levels through the algorithm's conflict resolution module, eliminating contradictory level information, and outputting the final learning status level.
[0029] Specifically, step S4 includes S41, which first divides the generated 512-dimensional fused feature matrix into four independent inference channels according to behavior type: action features, interaction features, attention features, and expression features. Each inference channel corresponds to a 128-dimensional feature sub-matrix. The division process strictly follows the semantic relevance of features to ensure the independence and specificity of features in each channel. S42, the inference engine of the learning status hierarchical association inference algorithm is invoked to perform hierarchical matching calculation on the feature sub-matrix of each inference channel. The matching process is based on a preset learning status feature template, which includes feature standard values corresponding to four learning status levels: excellent, good, average, and needing improvement. The distance between the feature sub-matrix and each level template is calculated as the matching degree index. The distance calculation uses the Euclidean distance algorithm to ensure the accuracy of the matching results. S43 determines candidate levels based on the matching degree results of each channel and the preset learning status judgment rules. The rules clearly stipulate that when the matching degree of all four channels is higher than 0.8, it is judged as an excellent level; when the matching degree of three channels is higher than 0.7, it is judged as a good level; when the matching degree of two channels is higher than 0.6, it is judged as an average level; and the rest are judged as levels to be improved. This rule is used to initially determine the candidate levels of each student's learning status. S44 uses the algorithm's conflict resolution module to perform consistency verification on the initially determined candidate levels. If there are two or more conflicting level judgment results during the verification process, the level with the highest sum of matching degrees of all channels is taken as the final result, effectively eliminating contradictory level information and ensuring that the output learning status level is unique and accurate, providing reliable core data for subsequent intelligent judgment.
[0030] Preferred, such as Figure 5As shown, step S5 includes the following sub-steps: S51, establishing an association index between the learning status level and the corresponding behavioral feature data to form a key-value pair storage structure; S52, transmitting the dataset after association indexing to the storage module of the classroom behavior multi-dimensional intelligent judgment and analysis platform, and storing it in partitions according to data type; S53, starting the platform's intelligent judgment engine to perform multi-dimensional association analysis on the stored dataset and identify the inherent association between behavioral features and learning status; S54, generating multi-dimensional behavioral analysis results including behavioral trends, feature distribution, and level proportions based on the association analysis results.
[0031] Specifically, step S5 involves the associated storage and intelligent analysis of learning status levels and behavioral characteristic data. S51 first establishes an index linking learning status levels with corresponding behavioral characteristic data, using the student's unique identifier as the key and the learning status level and behavioral characteristic data of each dimension as the value, forming a key-value pair storage structure. The index creation process employs a hash algorithm, and the hash function selection is optimized to ensure efficient index queries, with query response time controlled within a preset range. S52 transmits the indexed dataset to the storage module of the classroom behavior multi-dimensional intelligent analysis platform via 5G gigabit Ethernet. The transmission rate is stable at gigabit levels to ensure real-time data transmission. Storage is partitioned according to learning status levels and data types, with each storage partition having independent access permissions and a data backup mechanism. The backup frequency is set to once per hour to ensure the security and integrity of data storage. The S53 platform's intelligent analysis engine employs a multi-threaded parallel analysis mechanism, supporting simultaneous multi-dimensional correlation analysis of 128 data streams. Analysis dimensions include the correlation between behavioral characteristics and learning status, the distribution of behavioral characteristics among students at different learning levels, and the behavioral trends of the same student over different time periods. A sliding window statistical method is used during the analysis, with the window size set to one hour to ensure the analysis results reflect dynamic behavioral changes. Based on the correlation analysis results, S54 generates multi-dimensional behavioral analysis results, including behavioral trend curves, feature distribution histograms, and level percentage pie charts. The behavioral trend curves are displayed as line charts with a time granularity of minutes. The feature distribution histogram shows the value distribution of each behavioral characteristic, and the level percentage pie chart visually presents the proportion of students at each learning status level. All analysis results are compressed, with the data volume controlled within a preset range to ensure efficient subsequent output.
[0032] like Figure 6As shown, a student behavior data analysis system based on the Transformer model is implemented. This system includes: a multi-dimensional behavior data acquisition unit, a temporal behavior feature identification unit, a multi-dimensional feature fusion processing unit, a hierarchical reasoning unit for learning status, an intelligent judgment and analysis unit, and a result output and storage unit. The multi-dimensional behavior data acquisition unit and the temporal behavior feature identification unit are connected via a high-speed data bus to transmit the acquired multi-dimensional behavior data to the temporal behavior feature identification unit. The temporal behavior feature identification unit is communicatively connected to the multi-dimensional feature fusion processing unit, which processes the data using a temporal-aware student behavior identification model and outputs a temporal behavior feature sequence to the multi-dimensional feature fusion processing unit. The multi-dimensional feature fusion processing unit is equipped with... A multi-dimensional behavioral feature fusion Transformer model is established to create a data interaction channel with the learning status hierarchical reasoning unit, transmitting the fused feature matrix to the learning status hierarchical reasoning unit. The learning status hierarchical reasoning unit uses a learning status hierarchical association reasoning algorithm to communicate bidirectionally with the intelligent judgment and analysis unit, outputting the learning status hierarchy to the intelligent judgment and analysis unit. The intelligent judgment and analysis unit embeds the setting module of the multi-dimensional intelligent judgment and analysis platform for classroom behavior, connecting with the result output and storage unit to transmit the multi-dimensional behavioral analysis results to the result output and storage unit. Based on the various parameters of student behavioral data analysis, the result output and storage unit stores the analysis results in a structured manner and outputs a behavioral analysis report in a preset format. All units synchronize data and collaborate through a distributed communication protocol.
[0033] The student behavior data analysis method and system based on the Transformer model captures diverse behavioral data through a multimodal acquisition terminal, and achieves efficient collaboration between data acquisition and processing by combining high-speed communication and heterogeneous computing architecture. It utilizes a time-series feature identification mechanism to mine the dynamic correlation of behavioral data, integrates multi-dimensional information through feature fusion technology, and then achieves refined classification of learning status through a hierarchical reasoning mechanism, forming a full-process intelligent processing system from data acquisition to result output, thereby improving the depth and reliability of behavioral data analysis.
[0034] This invention addresses the shortcomings of traditional technologies in data collection, such as incomplete data acquisition and insufficient dynamic correlation capture. It employs multimodal acquisition devices and a temporal feature identification mechanism to comprehensively cover diverse classroom behavioral data. Simultaneously, it uncovers the inherent correlations between behaviors over time, achieving complete data acquisition and effective capture of dynamic features. Furthermore, addressing the lack of accuracy in interpreting student learning, it establishes a precise mapping between behavioral data and student learning status through deep fusion of multidimensional features and hierarchical correlation reasoning. This enables refined stratification and scientific assessment of student learning status, providing precise and reliable support for personalized teaching decisions and completely overcoming the limitations of existing technologies that can only perform superficial data statistics.
[0035] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0036] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A student behavior data analysis method based on the Transformer model, characterized in that, Includes the following steps: S1. Collect multi-dimensional behavioral data of students in classroom scenarios through a multi-dimensional intelligent analysis platform for classroom behavior. This multi-dimensional behavioral data includes action behavior data, interaction behavior data, attention-related data, and expressive behavior data. S2. Call a time-series-aware student behavior identification model to identify the time-series correlation of the collected multi-dimensional behavioral data. By capturing the dynamic correlation features of different behavioral data in the time dimension, a time-series behavioral feature sequence is generated. S3. Input the time-series behavioral feature sequence into a multi-dimensional behavioral feature fusion Transformer model. The model's multi-head attention mechanism is used to analyze different dimensions. S4. Behavioral features are weighted and deeply integrated to generate a fused feature matrix; S5. A hierarchical association reasoning algorithm for learning status is used to perform hierarchical reasoning operations on the fused feature matrix, and the learning status level corresponding to each student is determined according to the preset learning status classification standard; S6. The learning status level and corresponding behavioral feature data are associated, stored and intelligently judged through a multi-dimensional intelligent judgment and analysis platform for classroom behavior, generating multi-dimensional behavioral analysis results; S7. Based on the various parameters of student behavioral data analysis, features are extracted and associated with the intelligent judgment results, and a behavioral analysis report corresponding to the learning status is output.
2. The student behavior data analysis method based on the Transformer model according to claim 1, characterized in that, The expression of the time-series-aware student behavior identification model includes: in, This represents the temporal behavior characteristic value at time t. Indicates the number of dimensions in the behavioral data. This represents the time-series weighting coefficient of the i-th dimension of the data. This represents the characteristic function of the i-th dimension of the data at time t. This represents the time-series decay coefficient of the i-th dimension of the data. This represents the time interval of the i-th dimension of the data. Represents a temporal behavioral feature sequence, where Conv1d represents a one-dimensional convolution operation. Indicates the kernel size. This indicates the step size, and MaxPool1d represents the one-dimensional max pooling operation. Indicates the pooling window size. This indicates a feature concatenation operation.
3. The student behavior data analysis method based on the Transformer model according to claim 1, characterized in that, The expression for the multidimensional behavioral feature fusion Transformer model includes: in, This represents the attention weight between the i-th query vector and the j-th key vector. This represents the i-th query vector. Let j represent the j-th key vector. This represents the dimension of the query vector and the key vector. Indicates the total number of vectors. The fusion feature matrix is represented by LayerNorm, the layer normalization operation is represented by MultiHead, and the multi-head attention computation function is represented by MultiHead. Represents the query vector matrix. Represents the key vector matrix, Represents a value vector matrix. This represents the output weight matrix. This represents the output bias vector.
4. The student behavior data analysis method based on the Transformer model according to claim 1, characterized in that, The expression for the hierarchical association reasoning algorithm for learning status is: ;in, This indicates the final determined level of learning progress. This represents the set of all learning status levels. This indicates the number of features in the fused feature matrix. This represents the m-th eigenvalue in the fused feature matrix. This represents the conditional probability of obtaining a learning status level of I based on the inference of the m-th feature value. The prior weight coefficient represents the learning status level I.
5. The student behavior data analysis method based on the Transformer model according to claim 1, characterized in that, The classroom behavior multidimensional intelligent judgment and analysis platform includes an edge computing node cluster, a multimodal data acquisition terminal, a distributed storage server, and a heterogeneous computing unit; the multimodal data acquisition terminal includes a high-definition visual sensor, a sound acquisition module, a motion capture sensor, and an interactive status monitoring unit, and each terminal communicates with the edge computing node cluster via 5G gigabit Ethernet; The edge computing node cluster is configured with a 16-core heterogeneous processor and a dedicated neural network acceleration chip, supporting parallel processing of 128 channels of multi-dimensional behavioral data; the distributed storage server adopts a storage architecture that combines a distributed file system and a time-series database, and the data read and write response latency is controlled within a preset range; the heterogeneous computing unit includes a GPU computing module and an FPGA logic processing module, where the GPU computing module is used for model inference calculation, and the FPGA logic processing module is used for feature extraction calculation. The number of model training iterations at the platform software level is set to a preset threshold, the number of heads in the attention mechanism is set to 8, and the feature fusion dimension is set to 512.
6. The student behavior data analysis method based on the Transformer model according to claim 1, characterized in that, S2 includes the following sub-steps: S21, calling the input interface of the time-series-aware student behavior identification model to sort the collected multi-dimensional behavior data according to the timestamp order and establish a time-behavior data mapping relationship; S22, using the time-series detection module built into the model to calculate the time interval of the sorted behavior data and identify the time correlation strength between different behavior data; S23, dynamically assigning weights to each dimension of behavior data based on the time correlation strength, with the weight assignment result being positively correlated with the time correlation strength; S24, using the model's feature generation module to extract time-series features from the weighted behavior data and generate a time-series behavior feature sequence including time dimension information.
7. The student behavior data analysis method based on the Transformer model according to claim 1, characterized in that, S3 includes the following sub-steps: S31, converting the temporal behavioral feature sequence into a vector format recognizable by the Transformer model to construct a feature input matrix; S32, activating the multi-head attention mechanism of the multi-dimensional behavioral feature fusion Transformer model to perform correlation calculations on features of different dimensions in the feature input matrix; S33, adjusting the weight coefficients of each dimension feature based on the correlation calculation results to strengthen the influence of the calibrated behavioral features and weaken the interference of irrelevant features; S34, performing a nonlinear transformation on the weighted feature matrix through the fully connected layer of the model to generate a deeply fused feature matrix.
8. The student behavior data analysis method based on the Transformer model according to claim 1, characterized in that, S4 includes the following sub-steps: S41, dividing the fused feature matrix into feature dimensions and assigning different types of behavioral features to the corresponding inference channels; S42, calling the inference engine of the learning status hierarchical association inference algorithm to perform hierarchical matching calculation on the feature data of each inference channel; S43, based on the hierarchical matching calculation results and combined with the preset learning status judgment rules, initially determining the candidate level of each student's learning status; S44, performing consistency verification on the candidate levels through the algorithm's conflict resolution module, eliminating contradictory level information, and outputting the final learning status level.
9. The student behavior data analysis method based on the Transformer model according to claim 1, characterized in that, S5 includes the following steps: S51, establishing an association index between the learning status level and the corresponding behavioral feature data to form a key-value pair storage structure; S52, transmitting the dataset after association indexing to the storage module of the classroom behavior multi-dimensional intelligent judgment and analysis platform, and storing it in partitions according to data type; S53, starting the platform's intelligent judgment engine to perform multi-dimensional association analysis on the stored dataset and identify the inherent association between behavioral features and learning status. S54 generates multi-dimensional behavioral analysis results based on the correlation analysis results, including behavioral trends, feature distribution, and hierarchical proportions.
10. A student behavior data analysis system based on the Transformer model, characterized in that, This system is applied to the student behavior data analysis method based on the Transformer model as described in claim 1, comprising: a multi-dimensional behavior data acquisition unit, a time-series behavior feature identification unit, a multi-dimensional feature fusion processing unit, a hierarchical reasoning unit for learning status, an intelligent judgment and analysis unit, and a result output and storage unit; the multi-dimensional behavior data acquisition unit and the time-series behavior feature identification unit are connected via a high-speed data bus for transmitting the acquired multi-dimensional behavior data to the time-series behavior feature identification unit; the time-series behavior feature identification unit is communicatively connected to the multi-dimensional feature fusion processing unit, and outputs a time-series behavior feature sequence to the multi-dimensional feature fusion processing unit after processing by the time-series perception student behavior identification model; the multi-dimensional feature fusion processing unit is equipped with a multi-dimensional behavior feature fusion Tr The Ansformer model establishes a data interaction channel with the learning status hierarchical reasoning unit, transmitting the fused feature matrix to the learning status hierarchical reasoning unit. The learning status hierarchical reasoning unit adopts the learning status hierarchical association reasoning algorithm, communicating bidirectionally with the intelligent judgment and analysis unit, and outputting the learning status hierarchy to the intelligent judgment and analysis unit. The intelligent judgment and analysis unit embeds the setting module of the classroom behavior multi-dimensional intelligent judgment and analysis platform, connects with the result output and storage unit, and transmits the multi-dimensional behavior analysis results to the result output and storage unit. Based on the various parameters of student behavior data analysis, the result output and storage unit stores the analysis results in a structured manner and outputs a behavior analysis report in a preset format. Each unit synchronizes data and works collaboratively through a distributed communication protocol.