Intelligent analysis system for class learning situation based on deep learning
Patent Information
- Application Number
- CN202610726569.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-08-18
AI Technical Summary
本申请的实施例提供了基于深度学习的课堂学情智能分析系统,其有效解决了单一计算算子导致的局部感知机制僵化的缺陷,从而提高了系统在复杂多变的课堂环境下提取特征的准确度以及最终学情评估的可靠性
[0015]与现有技术相比,采用根据本申请实施例的基于深度学习的课堂学情智能分析系统,通过发散状态实时重构池化阶数,克服了常规特征融合与降维机制在模型构建后即被静态固化的局限性,使得计算模型的底层算子具备了与输入数据动态演化趋势相耦合的干预能力,允许模型根据课堂交互数据在高度同步与严重分散之间的客观变化,在平滑提取与异常聚焦两种感知逻辑之间进行连续切换,避免了单一计算算子导致的局部感知机制僵化,从而提高了系统在复杂多变的课堂环境下提取特征的准确度以及最终学情评估的可靠性。
Smart Images

Figure CN122595200A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of classroom learning analysis technology, and in particular to an intelligent classroom learning analysis system based on deep learning. Background Technology
[0002] With the deepening development of smart education, using deep learning networks to extract and align features from and with time-series data of teachers' teaching characteristics and students' multimodal interactions in the classroom has become an important technical means to achieve accurate assessment of individual learning and adaptive learning path recommendation.
[0003] Existing deep learning-based learning analysis systems suffer from significant technical shortcomings: their underlying feature fusion and dimensionality reduction mechanisms are typically statically fixed after model construction; real classroom teaching scenarios are highly dynamic and exhibit complex temporal fluctuations, with the distribution of student interaction data often changing drastically between high synchronization and severe dispersion; existing analysis models lack underlying dynamic intervention capabilities coupled with these data evolution trends, and cannot adaptively reconstruct their internal data aggregation logic when faced with abrupt changes in the distribution manifold of classroom input data; this makes it difficult for existing systems to flexibly switch between smoothly extracting global background features and precisely focusing on local anomalies based on the actual degree of data dispersion, resulting in a rigid local perception mechanism in the computational model, ultimately leading to limited feature capture capabilities and distorted learning assessment in complex and ever-changing classroom environments. Summary of the Invention
[0004] In view of the above-mentioned prior art, this application is hereby proposed. Embodiments of this application provide a deep learning-based intelligent classroom learning analysis system, which effectively solves the problem of rigid local perception mechanisms caused by single computational operators, thereby improving the accuracy of feature extraction in complex and ever-changing classroom environments and the reliability of the final learning assessment.
[0005] According to one aspect of this application, a deep learning-based intelligent classroom learning analysis system is provided, comprising:
[0006] The feature acquisition module is used to acquire teacher query features corresponding to teacher teaching data and class response features corresponding to student interaction data.
[0007] The cross-modal association module is used to input the teacher query features and the class response features into a preset cross-modal network to calculate the feature association distribution that reflects the teacher-student association status.
[0008] The divergence calculation module is used to calculate the divergence index of the feature correlation distribution and calculate the difference between the divergence index and the preset dissociation critical threshold to obtain the dissociation gradient.
[0009] The dynamic order adjustment module is used to obtain the order adjustment parameter for adjusting the pooling processing layer in the cross-modal network using the disjointed gradient mapping, and to update the initial pooling order in the pooling processing layer to the dynamic pooling order according to the order adjustment parameter.
[0010] The feature aggregation module is used to perform feature aggregation operations based on the dynamic pooling order on the feature association distribution using the pooling processing layer, and output fused features; wherein, the feature aggregation operation specifically includes: when the divergence index indicated by the disjoint gradient does not exceed the disjoint critical threshold, controlling the dynamic pooling order to tend to a first value, so that the pooling processing layer performs average pooling processing on the feature association distribution; and
[0011] When the dissociation gradient indicates that the divergence index exceeds the dissociation critical threshold, the dynamic pooling order is controlled to be changed to a second value greater than the first value, so that the pooling processing layer performs maximum pooling processing on the feature correlation distribution focusing on extreme features.
[0012] The classification output module is used to input the fused features into a classification network cascaded at the output end of the cross-modal network for probability mapping and output the corresponding learning assessment results.
[0013] According to another aspect of this application, an electronic device is provided, including a memory and a processor, the memory being used to store computer-executable instructions, and the processor being used to execute the computer-executable instructions, which, when executed by the processor, implement the functions of the system as described above.
[0014] According to another aspect of this application, a computer storage medium is provided that stores computer-executable instructions thereon, which, when executed by a processor, implement the functions of the system described above.
[0015] Compared with existing technologies, the deep learning-based intelligent classroom learning analysis system according to the embodiments of this application overcomes the limitations of conventional feature fusion and dimensionality reduction mechanisms, which are statically fixed after model construction, by reconstructing the pooling order in real time through divergent states. This enables the underlying operators of the computational model to have the intervention capability coupled with the dynamic evolution trend of the input data. It allows the model to continuously switch between two perception logics, smooth extraction and anomaly focusing, based on the objective changes in classroom interaction data between high synchronization and severe dispersion. This avoids the rigidity of local perception mechanisms caused by a single computational operator, thereby improving the accuracy of feature extraction in complex and ever-changing classroom environments and the reliability of the final learning assessment. Attached Figure Description
[0016] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0017] Figure 1 This is a flowchart illustrating the intelligent classroom learning analysis system based on deep learning, as described in this invention.
[0018] Figure 2 This is a schematic diagram of the logical framework of the deep learning-based intelligent classroom learning analysis system of the present invention.
[0019] Figure 3 This is an extended schematic diagram of the intelligent classroom learning analysis system based on deep learning, as described in this invention. Detailed Implementation
[0020] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.
[0021] Example 1:
[0022] Existing deep learning-based classroom learning analysis models generally employ static, fixed feature fusion mechanisms, which cannot adapt to the drastic fluctuations in the cognitive state of classroom groups between "highly synchronous" and "severely dispersed" states. When the distribution manifold of classroom interaction data undergoes abrupt changes, static operators lack the underlying intervention capability coupled with the data evolution trend, making it difficult to flexibly switch between extracting globally smooth features and precisely focusing on local anomalies. This leads to a rigid perception mechanism in the computational model and ultimately distorted evaluation. To address the technical challenge of fixed network topologies being unable to adapt to highly dynamic distributed data, this paper proposes a deep learning-based intelligent classroom learning analysis system. This system aims to overcome the limitations of static operators by quantifying the manifold evolution state of the data and constructing a low-level closed-loop feedback mechanism.
[0023] This application obtains the feature association distribution reflecting the teacher-student relationship through a cross-modal network, calculates its divergence index, and extracts the disjoint gradient by combining it with a preset threshold. Subsequently, the disjoint gradient is mapped to an order adjustment parameter, and the initial pooling order of the pooling layer is updated in real time to a dynamic pooling order. Based on this, the model adaptively reconstructs the data aggregation logic: in the stationary state where the disjoint gradient does not exceed the threshold, an average pooling operation tending to the first value is performed; in the disjoint state where the disjoint gradient exceeds the threshold, a maximum pooling operation climbing to a higher order to focus on extreme features is performed. This application reconstructs the pooling order in real time based on the data divergence state, endowing the underlying operators with morphological variation capabilities, realizing continuous adaptive switching between global background extraction and local anomaly focusing, effectively improving the accuracy and robustness of learning assessment in complex classroom environments.
[0024] Reference Figures 1-3 As an embodiment of the present invention, a classroom learning intelligence analysis system based on deep learning is provided, including: a feature acquisition module, a cross-modal association module, a divergent calculation module, a dynamic order adjustment module, a feature aggregation module, and a classification output module.
[0025] To facilitate understanding of the processing flow of subsequent modules, a 40-minute math class in an eighth-grade class at a middle school is used as an example. The class has 45 students, the teacher is Ms. Li, and the lesson topic is the translation transformation of quadratic function graphs. The classroom is equipped with: one front-facing wide-angle camera to capture the teacher's lecture and the overall view of the student seating area; one ceiling-mounted microphone array to capture the teacher's voice and ambient sound from student discussions; 45 response terminals placed on student desks to collect student clicks, answers, and interactive data; and one server deployed at the front of the classroom to run the processing modules of this embodiment. All raw data is synchronized and aligned by the front-end server according to a unified timestamp, with a data acquisition frame rate of 2 frames per second. The preset sliding time window length is... Take 10 seconds as the time interval between adjacent historical time points and the current time point. The system triggers a learning assessment process based on data from the most recent 10 seconds every 5 seconds. The processing flow of subsequent modules will follow this example for detailed explanation.
[0026] Figure 1 and Figure 2 The illustration shows a deep learning-based intelligent classroom learning analysis system according to an embodiment of this application, specifically including:
[0027] like Figure 1 As shown, in the feature acquisition module, teacher query features corresponding to teacher teaching data and class response features corresponding to student interaction data are acquired.
[0028] In this module, teacher teaching data includes, but is not limited to, visual data, audio data, and teaching interaction data of the teacher during the teaching period. Visual data originates from the teacher's teaching footage captured by the front-facing camera, and after preprocessing such as human posture estimation and facial expression recognition, it yields sequences of the teacher's body movements and facial expressions. Audio data originates from the teacher's voice captured by the ceiling-mounted microphone array, and after automatic speech recognition and acoustic feature extraction, it yields sequences of the teacher's semantic text and tone intensity. Teaching interaction data originates from the teacher's actions on the teaching terminal, such as writing on the blackboard, turning pages, and asking questions. Student interaction data includes, but is not limited to, visual data, audio data, and response data of the student group. Student visual data originates from the student seating area captured by the camera, and after preprocessing such as human posture estimation and gaze direction estimation, it yields student posture features and gaze direction features. Student audio data originates from the ambient sound of students speaking and discussing, captured by the microphone array, and after sound source localization and acoustic feature extraction, it yields group response intensity and rhythm features. Response data originates from the student response terminal, including but not limited to click response time, response content, and interaction frequency.
[0029] Teacher query features refer to the feature sequence obtained by multimodal encoding of the aforementioned teacher teaching data, used to characterize the teacher's current teaching intention and direction of instruction. Class response features refer to the feature sequence obtained by multimodal encoding of the aforementioned student interaction data, used to characterize the current cognitive response state of the student group. The aforementioned multimodal encoding can be implemented based on convolutional neural networks, recurrent neural networks, or encoders based on self-attention mechanisms. This embodiment does not limit the specific network structure of the encoder, as long as it can map the original multimodal data to a feature vector space of a unified dimension. It should be noted that the channel dimension of both teacher query features and class response features is uniformly set to 1. The query feature vector and response feature vector used in subsequent cross-modal association operations are both located in this... In 3D space.
[0030] Using the aforementioned classroom example, the system triggers a feature acquisition process every 5 seconds, extracting the teacher query features corresponding to this assessment and the class response features consisting of all 45 students in the class from the multi-source data of the most recent 10 seconds, as well as the channel dimension constants of the query feature vector and the response feature vector. The value is 64, meaning that each query feature vector and each response feature vector are both 64-dimensional vectors.
[0031] After obtaining the teacher query features and class response features, these two heterogeneous features need to be correlated to reflect the relationship between teachers and students within the current time window, and then returned. Figure 1In the cross-modal association module, teacher query features and class response features are input into a preset cross-modal network to calculate the feature association distribution reflecting the teacher-student relationship status. The specific formula is as follows:
[0032] ;
[0033] in, This represents the dimensionless feature association value between the feature sub-item corresponding to the u-th teacher query and the feature sub-item corresponding to the v-th class response in the feature association distribution. Let represent the u-th query feature vector (dimensionless) in the teacher query features. This represents the v-th response feature vector (dimensionless) in the class response features. This represents the channel dimension constant between the query feature vector and the response feature vector. It is a normalized exponential function.
[0034] In this module, the cross-modal network employs a cross-modal fusion network based on scaled dot product attention. Specifically, teacher query features are used as the query end, and class response features are used as the key end. The correlation strength between the query feature vector and the response feature vector is measured by the dot product operation. The dot product value is then scaled by the square root of the channel dimension constant to avoid excessive expansion of the dot product value as the channel dimension increases. Subsequently, it is normalized by an exponential function. The mapping yields feature association values within the interval (0, 1). The feature association values between all teacher query feature items and all class response feature items together constitute the feature association distribution. It should be noted that the specific structure of the preset cross-modal network is not limited to single-layer scaled dot product attention; it can also adopt extended forms such as multi-head attention stacking and bidirectional cross-modal attention. This embodiment does not limit this.
[0035] Channel dimension constant of query feature vector and response feature vector The network hyperparameters are pre-configured and their values usually fall within the range of 16 to 512. Values that are too small will limit the feature representation ability, while values that are too large will increase the computational cost and make the distribution of softmax input values tend to be flat, affecting the discriminativeness of feature association distribution.
[0036] Using the aforementioned classroom example, The cross-modal fusion network takes 64 as an example. It performs association operations on the teacher query features and class response features triggered every 5 seconds and outputs a feature association distribution in the form of a matrix. The value of each element in the matrix is located in the interval (0, 1) and the sum of the elements in each row is 1. This matrix reflects the association strength distribution between each teacher query feature item and all response feature items of the class.
[0037] The feature correlation distribution characterizes the instantaneous state of teacher-student interaction within the current time window. However, classroom teaching is a dynamic evolutionary process, and the instantaneous state alone is insufficient to determine whether the teacher-student interaction is in a stationary or disjointed state. Therefore, the evolution rate of the feature correlation distribution between adjacent time nodes is introduced into the divergent calculation module as an indicator to measure the degree to which the teacher-student interaction deviates from a stationary state. Specifically, returning... Figure 1 In the divergence calculation module, the divergence index of the feature correlation distribution is calculated, and the difference between the divergence index and the preset dissociation critical threshold is calculated to obtain the dissociation gradient, including:
[0038] Extract the first distribution state parameter of the feature association at the current time node, and the second distribution state parameter at adjacent historical time nodes;
[0039] Specifically, the first distribution state parameter refers to the state representation vector obtained by vectorizing the feature association distribution at the current time node according to a preset element arrangement order; the second distribution state parameter refers to the historical state representation vector obtained by vectorizing the feature association distribution calculated at adjacent historical time nodes according to the same element arrangement order; the dimension of both is equal to the total number of features N contained in the distribution state parameter. Specific extraction methods include, but are not limited to: flattening each feature association value in the feature association distribution matrix into a one-dimensional sequence according to row priority or column priority, and using the value at each position in this sequence as the i-th feature association value in the distribution state parameter. During system operation, the feature association distribution output at the current time node is extracted as the first distribution state parameter in the above manner and stored in the system's internal historical cache queue. Simultaneously, the second distribution state parameter corresponding to adjacent historical time nodes is read from this historical cache queue for use in subsequent state perturbation parameter calculations.
[0040] Calculate the relative distance between the first and second distribution state parameters to obtain the state perturbation parameter used to characterize the magnitude of the associated state drift. The specific formula is as follows:
[0041] ;
[0042] in, This represents the state perturbation parameter (dimensionless) corresponding to the current time node t. This represents the dimensionless value of the i-th feature in the first distribution state parameter. Let be the dimensionless value associated with the i-th historical feature in the second distribution state parameter, and N be the total number of features contained in the distribution state parameter. The time interval (s) between adjacent historical time nodes and the current time node. For adjacent historical time nodes (s);
[0043] The logarithmic rate of change of the state perturbation parameter with respect to the time evolution interval is calculated as the divergence exponent, and the specific formula is as follows:
[0044] ;
[0045] in, This represents the divergence exponent (1 / s) at the current time point t. This represents the state perturbation parameter corresponding to the current time node t. Indicates adjacent historical time nodes Corresponding historical state perturbation parameters;
[0046] It should be noted that at the first time point after system startup, since there are no historical state perturbation parameters corresponding to adjacent historical time points, in order to ensure the normal calculation of the divergence index, this embodiment adopts one of the following two processing methods: First, the historical state perturbation parameter is set to the state perturbation parameter of the current time point, so that the output of the divergence index calculation result of the first time point is zero, that is, it is assumed that the teacher-student relationship state at the first time point has not evolved and is judged to be in a stationary state; Second, the divergence index calculation process of the first time point is skipped, and the divergence index calculation is started from the second time point. The learning assessment result corresponding to the first time point is output by the subsequent modules according to the stationary state by default. This embodiment does not limit the above two processing methods;
[0047] It should be noted that the reason why this application uses the logarithmic rate of change as the divergence index is mainly based on the following factors:
[0048] First, if the first difference of the state perturbation parameter is directly used as the divergence index, the range of the divergence index will drift with the absolute magnitude of the state perturbation parameter, making it incomparable between the early and late stages of the class, and between different classes or different subjects; while the logarithmic rate of change normalizes the state perturbation parameter, so that the evolution rate reflected by the divergence index no longer depends on the absolute magnitude of the state perturbation parameter.
[0049] Secondly, the evolution of teacher-student relationship in the classroom exhibits a typical exponential change characteristic. Logarithmic operations can map this exponential change to a linear change, making it easier to set subsequent threshold comparisons and gradient calculations. Alternatively, KL divergence, Wasserstein distance, and other distribution difference measures can be used to characterize the evolution rate of teacher-student relationship. However, the former requires prior assumption that the characteristic relationship distribution follows a specific probability distribution form, while the latter will bring higher computational overhead. Therefore, this embodiment uses the logarithmic rate of change as the specific form of the divergence exponent.
[0050] The difference between the divergence index and the dissociation threshold is used to obtain the evolutionary bias value with positive and negative signs. This evolutionary bias value is then used as the dissociation gradient, as shown in the following formula:
[0051] ;
[0052] in, This represents the evolutionary bias value with positive and negative signs, i.e., the disjoint gradient (1 / s). This represents the divergence exponent (1 / s) at the current time point t. This represents the critical threshold for disconnection, characterizing the maximum tolerable boundary divergence rate (1 / s) of classroom cognitive state.
[0053] It should be noted that the critical threshold for disconnection The specific value is obtained based on historical classroom data statistics during the system deployment phase: In several sample classes of the same type that have been manually annotated, the divergence index corresponding to the time window marked as stationary is statistically analyzed, and the upper quantile (e.g., the 95th percentile) of this statistical distribution is taken as the initial value of the disconnection critical threshold; subsequently, the initial value can be adjusted within a small range based on the early warning accuracy feedback of the online operation of the system. It should be noted that the disconnection critical threshold can be separately calibrated for different grade levels, different subjects or different classes, and this embodiment does not limit this.
[0054] The time interval between adjacent historical time points and the current time point Based on the pace of classroom development, typical values fall within the range of 1 to 30 seconds; A value that is too small will make the divergence exponent overly sensitive to frame-level noise. If the value is too large, the divergence index will lag in responding to the rapid evolution in the classroom.
[0055] Using the aforementioned classroom example, Taking 5 seconds, the critical threshold for disconnection was obtained by statistical analysis of the history class data of this class. Approximately 0.15 (1 / s), within a certain time window during the explanation of the quadratic function graph transformation at the 18th minute, the system extracts the first distribution state parameter of the current time node and the second distribution state parameter of the adjacent historical time nodes. After the above calculation, the state perturbation parameter is approximately 0.32, and thus the divergence index corresponding to the current time node t is obtained. The dissociation gradient is approximately 0.21 (1 / s) and is obtained by subtracting it from the dissociation critical threshold. (1 / s), with a positive sign, indicates that the current teacher-student relationship has deviated from a steady state and entered a disconnected state.
[0056] Disjoint gradients characterize the direction and magnitude of the current teacher-student relationship state deviating from a stationary state, but disjoint gradients themselves cannot directly affect the feature aggregation behavior of cross-modal networks; it is necessary to map disjoint gradients as intrinsic parameters that can be used to adjust the operational behavior of pooling layers. Specifically: Return Figure 1 In the dynamic order adjustment module, the disjointed gradient mapping is used to obtain the adjustment parameter for adjusting the order of the pooling layer in the cross-modal network. The specific formula is as follows:
[0057] ;
[0058] in, This represents the order adjustment parameter at the current time point t. is the smoothing inertia constant, used as the time scale scaling factor (s) to convert the evolution rate signal into the system intrinsic parameter adjustment signal;
[0059] In this module, the pooling layer is a parameterizable pooling layer cascaded at the feature association output of the cross-modal network. Its operational behavior is controlled by the dynamic pooling order. A continuous transition between different pooling modes can be achieved through the same pooling layer. The specific transition logic will be further explained in the subsequent feature aggregation module. The smoothing inertia constant τ is a time scale scaling factor configured by the system, numerically related to... Keep the order of magnitude in the same range, with a typical value range of 1s to 10s. If the value is too small, the order adjustment parameter will be overly sensitive to the disjointed gradient and prone to frequent fluctuations. If the value is too large, the order adjustment parameter will be slow to respond and will not keep up with the pace of classroom evolution.
[0060] The initial pooling order in the pooling layer is updated to a dynamic pooling order based on the order adjustment parameter, including:
[0061] By inputting the order adjustment parameter into the saturated activation function and performing constraint mapping, the target order offset value, which is limited to a bounded numerical range, is obtained. The specific formula is as follows:
[0062] ;
[0063] in, This represents the target value of the order offset, which is limited to a bounded numerical range. This represents the maximum order offset extreme value constant used to limit the amplitude of a single adjustment. This represents the order adjustment parameter at the current time point t. The hyperbolic tangent function is the saturated activation function.
[0064] The reason why this application uses the hyperbolic tangent function to constrain the order adjustment parameter is based on the following factors:
[0065] First, the value of the order adjustment parameter may be sharp due to the occasional large value of the decoupled gradient. If it is directly superimposed on the initial pooling order, it will cause the dynamic pooling order to jump to an ill-conditioned value, thereby making the operation behavior of the pooling layer unstable. The function smoothly constrains the input value to the interval (−1, 1), and then multiplies it by... You can get the location located at ( The order offset target value within the bounded interval is used to limit the single adjustment range.
[0066] Secondly, The function is an odd function, which can symmetrically represent two cases: order-up adjustment (corresponding to a positive disjoint gradient) and order-down adjustment (corresponding to a negative disjoint gradient). This conforms to the symmetric semantics of switching between disjoint and stationary states. Alternatively, a sigmoid function with a constant bias or a hard truncation function can be used to achieve a similar constraint effect; however, the former destroys the symmetry of the upward and downward adjustment, and the latter introduces a non-differentiable hard boundary and affects gradient backpropagation. Therefore, this embodiment adopts... The function is used as a saturation activation function.
[0067] The bounded numerical interval refers to the bilaterally restricted interval for the target value of the order offset. This interval is determined by the maximum order offset extremum constant, and numerically it is ( ), Used to limit the magnitude of a single adjustment, its value is in principle 0.5 to 2 times the initial pooling order, to ensure that the pooling order after superimposing the offset target value is always within the physically meaningful pooling order range. The initial pooling order is the baseline order of the pooling processing layer, with a typical value of 1, corresponding to the baseline behavior of average pooling in a steady state.
[0068] Extract the historical dynamic pooling order of the pooling layer within the previous adjacent time window;
[0069] Specifically, the previous adjacent time window refers to the time window corresponding to one time evolution interval preceding the current time node; the historical dynamic pooling order refers to the dynamic pooling order output and stored by the system when executing the dynamic order adjustment module within the previous adjacent time window. During system operation, after each dynamic order adjustment is completed, the output dynamic pooling order is cached in the order history queue; when the current time node executes this step, the most recently cached dynamic pooling order is read from the order history queue as the historical dynamic pooling order. It should be noted that, at the first time window after system startup, since there is no historical dynamic pooling order corresponding to the previous adjacent time window, in order to ensure the normal operation of subsequent exponential moving average calculations, this embodiment initializes the historical dynamic pooling order to the initial pooling order. Even if the order evolution residual under the first time window is obtained by subtracting the initial pooling order after superimposing the offset target value from the initial pooling order itself, this embodiment does not limit the above initialization method. The historical dynamic pooling order can also be initialized to a preset order value calibrated based on historical classroom data experience.
[0070] The time-shifting sliding factor is invoked to perform a weighted iterative calculation on the historical dynamic pooling order and the initial pooling order superimposed with the order offset target value, to obtain the dynamic pooling order, including:
[0071] Based on the initial pooling order that superimposes the target value of the order offset and the historical dynamic pooling order, the order evolution residual is obtained, and the specific formula is as follows:
[0072] ;
[0073] in, For order evolution residuals, This represents the initial pooling order in the pooling processing layer. This represents the target value of the order offset, which is limited to a bounded numerical range. Indicates adjacent historical time nodes The historical dynamic pooling order within;
[0074] When the order evolution residual is greater than or equal to the baseline zero value, the time translation momentum factor is assigned the first decay coefficient;
[0075] When the order evolution residual is less than the baseline zero value, the time translation momentum factor is assigned a second attenuation coefficient that is greater than the first attenuation coefficient;
[0076] The reason this application assigns different decay coefficients to the positive and negative residuals of the order evolution is that when the classroom switches from a steady state to a disjointed state, the system needs to quickly follow up and raise the dynamic pooling order to a level that can focus on extreme features as soon as possible to avoid missing the capture of sudden cognitive deviations. Conversely, when the classroom returns from a disjointed state to a steady state, it should not fall back too quickly to avoid prematurely relaxing attention to abnormal features due to a brief illusion of synchronization. Specifically, when the residual of the order evolution is greater than or equal to the baseline zero value (i.e., the pooling order after superimposing the offset target value is higher than the historical dynamic pooling order and is on an upward trend), the time smoothing momentum factor is assigned a relatively small first decay coefficient, so that the historical order has a lower weight and the new order has a higher weight, in order to accelerate the upward response. Conversely, when the residual of the order evolution is less than the baseline zero (i.e., on a downward trend), the time smoothing momentum factor is assigned a second decay coefficient greater than the first decay coefficient, so that the historical order has a higher weight and the new order has a lower weight, in order to smooth the downward response. Alternatively, a symmetrical fixed momentum factor can be used; however, this will bring about two problems: hysteresis in the rising response and overshoot in the falling response. Therefore, this embodiment adopts an asymmetric attenuation coefficient assignment method based on the positive and negative signs of the order evolution residual. It should be noted that the reference zero value is the reference value for scalar comparison, which is always equal to 0 in numerical value, and is used to determine the positive and negative directions of the order evolution residual.
[0077] Using the assigned time-smooth factor as the weight allocation ratio, an exponential moving average is performed on the historical dynamic pooling order and the initial pooling order with the added order offset target value to obtain the dynamic pooling order. The specific formula is as follows:
[0078] ;
[0079] in, Let t be the dynamic pooling order at the current time point. This represents the time smoothing factor called after assignment, which is used as the first weight allocation ratio. This represents the auxiliary smoothing coefficient used as the second weighting allocation ratio. This represents the initial pooling order in the pooling processing layer. This represents the order offset target value that is limited to a bounded numerical range.
[0080] Following the previous classroom example, the smoothing inertia constant τ is set to 5 seconds, which is used to limit the maximum order of deviation extreme constant of a single adjustment amplitude. We set the pooling order to 2, the initial pooling order to 1, the first decay coefficient to 0.3, and the second decay coefficient to 0.7; within the 18-minute time window, the gradient becomes disjointed. (1 / s), order adjustment parameter ;Will After inputting the tanh function, multiply by ,get Let the historical dynamic pooling order of the previous time window be... The initial pooling order after superimposing the target offset value is 1.58, and the order evolution residual is 0.58, which is greater than the baseline zero value. Therefore, the first decay coefficient of 0.3 is used as the time smoothing momentum factor. After exponential moving average calculation, the dynamic pooling order at the current time node is obtained. .
[0081] Once the dynamic pooling order is determined, the pooling layer can perform feature aggregation operations on the feature association distribution based on that order and return the result. Figure 1 In the feature aggregation module, the pooling processing layer performs feature aggregation operations based on the dynamic pooling order on the feature association distribution and outputs fused features.
[0082] This application achieves continuous switching between average pooling and max pooling through a single dynamic pooling order. The mathematical principle behind this is the limiting property of p-norm pooling: when the pooling layer performs p-norm pooling, its operation can be uniformly represented as taking the p-th power mean of the input feature association distribution and then taking the p-th root. Under this uniform form, when p approaches 1, the operation result is the arithmetic mean of the input elements, which corresponds to average pooling and is suitable for extracting global background features from the whole under stationary conditions. When p approaches a large value, the operation result is dominated by the extreme value element with the largest value, which approaches max pooling and is suitable for focusing on individual abnormal elements under disjointed conditions.
[0083] Based on the above characteristics, a continuous transition between the two pooling modes can be formed by dynamically adjusting a single order p, without having to deploy two independent pooling operators in parallel in the network and switch them with a switch. Alternatively, a scheme of deploying average pooling and max pooling in parallel with a gated switch can be adopted, or an attention pooling scheme based on additional trainable parameters can be adopted. However, the former will introduce a step-like feature jump at the moment of switching, and the latter requires the introduction of additional trainable parameters and increases training overhead. Therefore, this embodiment adopts a scheme based on continuous switching of a single order.
[0084] The feature aggregation operation specifically includes the following two cases:
[0085] Case 1: When the divergence exponent of the disjoint gradient indicator does not exceed the disjoint critical threshold, the dynamic pooling order is controlled to tend to the first value so that the pooling layer performs average pooling on the feature correlation distribution.
[0086] Case 2: When the divergence index of the disjoint gradient indicator exceeds the disjoint critical threshold, the dynamic pooling order is changed to a second value greater than the first value, so that the pooling processing layer performs maximum pooling processing on the feature correlation distribution focusing on the extreme features.
[0087] Specifically, the first value refers to the baseline order corresponding to average pooling, which is numerically close to 1. In engineering implementation, it approximates average pooling behavior with a finite value close to 1 (e.g., 1.0). The second value refers to a larger order corresponding to the tendency of max pooling, which is usually a finite value between 6 and 10. In engineering implementation, it approximates max pooling behavior with this upper bound (e.g., 10). The transition between the first and second values is continuously controlled by the dynamic pooling order generated by the aforementioned dynamic order adjustment module, thereby avoiding feature jumps caused by hard switching between the two sets of pooling operators.
[0088] Using the aforementioned classroom example, the dynamic pooling order within this time window Between the first and second values (6 to 10), the pooling processing layer shifts the operation behavior of the feature association distribution from average pooling to max pooling. The output fused features retain global background information while increasing the response weight of individual abnormal items, which is beneficial to reflect the local cognitive deviation that has appeared in the current time window in the subsequent classification output stage.
[0089] The fusion feature reflects the teacher-student relationship status after dynamic pooling integration within the current time window. Based on this, the fusion feature needs to be mapped into a readable learning assessment result. This process is completed in the classification output module. The classification output module simultaneously monitors the dynamic pooling order corresponding to the current fusion feature to ensure that the final probability mapping result matches the current pooling logic tendency. Specifically, it returns... Figure 1 In the classification output module, the fused features are input to the classification network cascaded at the output of the cross-modal network for probability mapping, and the corresponding learning assessment results are output, including:
[0090] The dynamic pooling order corresponding to the generation of fused features is extracted synchronously, and the dynamic pooling order is input into a preset routing control function to map and generate a pair of gated weights with normalization properties. The specific formula is as follows:
[0091] ;
[0092] in, Indicates the first gating weight. This indicates the dynamic pooling order corresponding to the synchronously extracted fused features. This is a pre-defined median point for the order boundary used to distinguish different pooling logic tendencies. This is the mapping sensitivity scaling factor;
[0093] ;
[0094] Indicates the second gating weight;
[0095] In this module, the routing control function adopts a sigmoid-based normalized form, which maps the dynamic pooling order to a first gate weight in the (0, 1) interval, and then subtracts the first gate weight from 1 to obtain a second gate weight, thus forming a pair of gate weights with normalized properties (the sum of the two is always 1); where, Numerically, the value is taken between the first and second values, typically ranging from 3 to 5; the mapping sensitivity scaling factor κ is used to control the deviation of the gating weights from the dynamic pooling order. The steepness of the transition during the transition typically falls between 0.5 and 2; a value that is too high will cause the gating weight to... The nearby transition is close to a step, and if the value is too small, the transition of the gating weight between the two branches will be slow.
[0096] The fused features are input into the first and second parallel mapping branches of the classification network, respectively, and the first and second candidate probability vectors are calculated respectively.
[0097] Specifically, both the first and second mapping branches are fully connected subnetworks cascaded at the output of the cross-modal network, and their respective outputs are passed through a normalized exponential function. The two branches are mapped to the probability space where the learning situation classification labels are located. The main difference between the two branches lies in the feature distribution used in their training phase: the first mapping branch is trained on samples whose feature aggregation output corresponds to a lower pooling order (i.e., average pooling tendency, stationary state), and is good at making learning situation judgments based on global background features; the second mapping branch is trained on samples whose feature aggregation output corresponds to a higher pooling order (i.e., max pooling tendency, disjointed state), and is good at making learning situation judgments based on extreme value focusing features. It should be noted that the specific network structure of the two mapping branches is not limited to fully connected sub-networks, and can also adopt other forms such as residual connections and lightweight Transformers. This embodiment does not limit this.
[0098] By using a pair of gating weights to perform cross-weighted calculations on the first candidate probability vector and the second candidate probability vector respectively, the target probability distribution vector is obtained, as shown in the following formula:
[0099] ;
[0100] in, This represents the target probability value corresponding to the c-th learning classification label in the obtained target probability distribution vector. Indicates the first gating weight. Indicates the second gating weight. Let be the candidate probability value corresponding to the c-th learning classification label in the first candidate probability vector. This represents the candidate probability value corresponding to the c-th learning situation classification label in the second candidate probability vector;
[0101] Extract the maximum probability index value from the numerical space of the target probability distribution vector, and use the learning assessment label corresponding to the maximum probability index value as the learning assessment result.
[0102] In this module, the learning situation classification tags can be configured according to the needs of the application scenario. Typical tag sets include, but are not limited to, four categories: "focus", "participation", "distraction" and "confusion". Further refinement can be made on this basis, such as adding subcategories such as "high focus", "passive following" and "active thinking". This embodiment does not limit the specific set of learning situation classification tags.
[0103] Using the aforementioned classroom example, the midpoint of the order boundary is used to distinguish different pooling logic tendencies. Set to 4, mapping sensitivity ratio coefficient Taking 1, the learning assessment labels adopt the aforementioned four-category configuration. The dynamic pooling order under the current time window is approximately 1.41. After mapping by the routing control function, the first gate weight is approximately 0.07 and the second gate weight is approximately 0.93. The fusion features are synchronously input into the first and second mapping branches to obtain the first and second candidate probability vectors. Assuming that the first candidate probability vector takes the values (0.55, 0.25, 0.10, 0.10) for the four labels "focused, engaged, distracted, confused", and the second candidate probability vector takes the values (0.10, 0.20, 0.55, 0.15) for the four labels, then the target probability distribution vector after cross-weighting by the gate weights takes the values (0.13, 0.20, 0.52, 0.15) for the four labels. The maximum probability index value corresponds to the "distracted" label, and the learning assessment result for this time window is "distracted", which is consistent with the disjointed state judgment indicated by the aforementioned disjointed gradient.
[0104] The pooling layer uses a globally shared dynamic pooling order during feature aggregation, which is suitable for situations where the overall cognitive state of the class tends to be consistent. However, in real classrooms, situations such as "students in the front row keeping up with the teacher's explanation while some students in the back row start to lose focus" or "students near the window experiencing decreased attention due to external distractions" often occur, leading to sub-region differentiation. In these cases, using a single global order would dilute the coverage of local anomalies. To address these shortcomings, this application further proposes introducing a sub-region refinement scheme based on spatial region segmentation into the dynamic order adjustment module. Specifically, for example... Figure 3 As shown, before updating the initial pooling order in the pooling processing layer to the dynamic pooling order based on the order adjustment parameter, the following steps are also included:
[0105] Spatial region segmentation is performed on the feature association distribution to extract features from multiple sub-regions;
[0106] In this step, the feature association distribution is spatially segmented. This can be done using a regular grid based on feature map rows / columns, a physical segmentation based on class seating space (such as a combination of front / middle / back area and left / middle / right column), or an irregular segmentation based on clustering. This embodiment does not limit the segmentation method. The number of features M in the sub-regions obtained after segmentation is usually between 4 and 16. If the value is too small, it will not have a refining effect, and if the value is too large, it will increase the computational cost and reduce the statistical reliability of samples in a single sub-region.
[0107] Calculate the local information entropy of the features in each sub-region, including:
[0108] Extract the feature response evolution sequence of sub-region features within a preset sliding time window;
[0109] It should be noted that the preset sliding time window length The window length can be set according to the pace of classroom evolution, with typical values falling between 10 and 60 seconds. If the window length is too short, it will lead to insufficient statistical samples of subsequent cognitive state stability and consistency and excessive noise. If the window length is too long, it will make the local information entropy overly sensitive to slow evolution.
[0110] Analyze the numerical fluctuation variance of the characteristic response evolution sequence on the time axis to quantify the stability of cognitive states within sub-regions;
[0111] Specifically, the feature response evolution sequence refers to a time series consisting of multiple feature response values arranged chronologically within a preset sliding time window for a sub-region. The variance of the numerical fluctuation of the feature response evolution sequence on the time axis is quantified in the following way: first, the arithmetic mean of the sequence is calculated in the time dimension; then, the squares of the differences between each feature response value and the arithmetic mean in the sequence are summed and averaged. The resulting variance value is used as a numerical index of cognitive state stability. The closer the value of this index is to zero, the more stable the feature response in the sub-region is in the time dimension; conversely, the larger the value of this index, the more drastic the fluctuation of the feature response in the sub-region is in the time dimension.
[0112] The numerical distribution dispersion of the statistical characteristic response evolution sequence in the spatial dimension quantifies the consistency of cognitive states within sub-regions;
[0113] Specifically, the spatial dispersion of the feature response evolution sequence is quantified as follows: the time mean of the feature response values at each spatial location within a sub-region is calculated within a preset sliding time window to obtain the spatial distribution of the feature response in that sub-region; then, the standard deviation or range of this feature response distribution is calculated in the spatial dimension, and the obtained standard deviation or range is used as a numerical index of cognitive state consistency; the closer the index value is to zero, the more consistent the feature responses at each spatial location within the sub-region tend to be; conversely, the larger the index value, the more obvious the differentiation of feature responses at each spatial location within the sub-region.
[0114] By performing multi-dimensional feature mapping on cognitive state stability and cognitive state consistency, the local information entropy is obtained, as shown in the following formula:
[0115] ;
[0116] in, The local information entropy corresponding to the feature of the i-th sub-region is... This is a numerical index for the stability of cognitive states within the quantified i-th sub-region. This is a numerical index of cognitive state consistency obtained by quantification within the i-th sub-region;
[0117] Specifically, the cognitive state stability index is quantified as the variance of the sub-region feature response evolution sequence over time, while the cognitive state consistency index is quantified as the dispersion of the sequence's numerical distribution in the spatial dimension. When the cognitive state within a sub-region exhibits both temporal stability and spatial consistency, both indices approach 0, and the local information entropy correspondingly approaches 0. Conversely, when the cognitive state within a sub-region exhibits temporal fluctuations or spatial differentiation, both indices increase, and the local information entropy increases accordingly. It should be noted that the specific form of the multidimensional feature mapping can employ logarithmic weighted summation, Euclidean norm, or other mapping methods; this embodiment does not limit this approach.
[0118] The spatial distribution weights corresponding to the features of each sub-region are generated based on local information entropy, and the specific formula is as follows:
[0119] ;
[0120] in, This represents the spatial distribution weight corresponding to the feature of the i-th sub-region. The local information entropy corresponding to the feature of the i-th sub-region is... For the local information entropy corresponding to the j-th sub-region feature, M represents the total number of sub-region features obtained by spatial region segmentation, and j represents the index number of all sub-region features participating in the summation;
[0121] The reason why this application adopts spatial distribution weights based on local information entropy is based on the following factors:
[0122] First, the cognitive state of different sub-regions within a class may exhibit a pattern of local stability and local fluctuation. Directly sharing the same order adjustment parameter among the sub-regions with uniform weights will average out local abnormal signals. By characterizing the fluctuation and dispersion of the cognitive state in each sub-region through local information entropy, sub-regions with relatively fluctuating cognitive states can obtain higher weights and thus obtain stronger order shifts, while sub-regions with relatively stable cognitive states can obtain lower weights and maintain stable behavior close to the globally shared order.
[0123] Secondly, the spatial distribution weights are obtained by normalizing the local information entropy over all sub-regions, which ensures that the local order adjustment parameters of each sub-region remain comparable to the global order adjustment parameters in terms of numerical magnitude. Alternatively, a fixed grid block plus weighting method can be used, or an additional spatial attention module can be introduced to generate weights. However, the former cannot reflect local differences, and the latter will introduce additional trainable parameters and increase training overhead. Therefore, this embodiment adopts the spatial distribution weight generation method based on local information entropy.
[0124] The spatial distribution weights are fused with the order adjustment parameter to obtain the local order adjustment parameter for each sub-region feature. The specific formula is as follows:
[0125] ;
[0126] in, This represents the local order adjustment parameter corresponding to the feature of the i-th sub-region. This represents the spatial distribution weight corresponding to the feature of the i-th sub-region. This represents the global order adjustment parameter for the current time node t, where t is the current time node (s).
[0127] The parameters are adjusted according to each local order to update the initial pooling order to the local dynamic pooling order corresponding to the features of each sub-region, which is then used as the dynamic pooling order. The specific formula is as follows:
[0128] ;
[0129] in, This represents the local dynamic pooling order corresponding to the updated feature of the i-th sub-region. This represents the initial pooling order that is shared globally by the pooling processing layer. This represents the maximum order offset extreme value constant. This represents the local order adjustment parameter corresponding to the feature of the i-th sub-region. This represents the hyperbolic tangent saturated activation function.
[0130] Using the previous classroom example, the spatial area is divided into 9 sub-regions by combining "front / middle / back zone" and "left / middle / right column". Preset sliding time window length Take 10 seconds. Within the 18-minute time window, statistics show that the cognitive stability and consistency indices of the sub-region in the 3rd row, 1st column (corresponding to the left-hand seats in the back row of the class) are both at a high level, with relatively large local information entropy. After normalization, the spatial distribution weight is approximately 0.22 (higher than the uniform weight 1 / 9≈0.11). In contrast, the local information entropy of the sub-region in the 1st row, 2nd column (corresponding to the middle seats in the front row of the class) is relatively small, with a spatial distribution weight of approximately 0.06. Based on the above spatial distribution weights, the global order adjustment parameters are fused and calculated to obtain the local order adjustment parameters corresponding to each sub-region. This leads to the local dynamic pooling order corresponding to each sub-region, resulting in a larger pooling order offset for the left-hand sub-region in the back row and a smaller offset for the middle sub-region in the front row.
[0131] The aforementioned divergence calculation process directly uses the original feature association distribution as the input data for calculating the divergence index. However, the feature association distribution may be disturbed by occasional transient noise (such as a student briefly turning around or a short environmental sound) at certain time points. The local anomalies caused by these transient disturbances do not truly reflect the evolution of the teacher-student relationship, but will be misinterpreted by the divergence index as an overall disconnect. Therefore, this application further proposes to perform a stability pre-correction on the feature association distribution based on multi-scale feature deconstruction before calculating the divergence index, specifically including:
[0132] Multi-scale feature deconstruction is performed on the feature correlation distribution to obtain a multi-scale smooth feature set;
[0133] In this step, the feature association distribution is deconstructed at multiple scales. This can be achieved by applying multiple smoothing filters with different kernel sizes (such as moving averages with different window lengths, Gaussian smoothing at different scales, etc.) to the feature association distribution, thereby obtaining a set of smoothed features corresponding to different smooth receptive fields. The total number of scales Q of the different smooth receptive fields obtained by multi-scale feature deconstruction is usually between 3 and 5, with a typical value of Q=3, corresponding to small, medium, and large scales. Too few scales are insufficient to represent cross-scale differences, while too many scales will increase computational overhead and reduce benefits. Among them, the smoothed feature value corresponding to the largest deconstruction scale has the strongest low-pass filtering effect and represents the macro trend of the feature association distribution. In this embodiment, it is selected as the benchmark feature value for subsequent cross-scale comparisons.
[0134] The specific formula for calculating the cross-scale distribution difference between features of different scales in a multi-scale smooth feature set is as follows:
[0135] ;
[0136] in, This represents the cross-scale distribution difference of the m-th feature value element in the multi-scale smooth feature set, and Q represents the total number of different scales of the smooth receptive fields obtained from the multi-scale feature deconstruction. This represents the value of the m-th smoothed feature extracted at the q-th deconstruction scale. This indicates the m-th benchmark feature value selected as the macro trend benchmark at the largest deconstruction scale, where m represents the index number of the feature value element and q represents the index number of the deconstruction scale.
[0137] Furthermore, a stability confidence factor for suppressing anomalous feature interference is obtained based on cross-scale distribution difference mapping, and the specific formula is as follows:
[0138] ;
[0139] in, This represents the stability confidence factor for the corresponding m-th feature element obtained based on the cross-scale distribution difference mapping. This represents the difference penalty ratio, used to control the degree of confidence decay. This represents the cross-scale distribution difference of the numerical element corresponding to the m-th feature in a multi-scale smooth feature set;
[0140] It should be noted that the difference penalty ratio coefficient The differential penalty coefficient is used to control the attenuation of the stability confidence factor as the cross-scale distribution difference increases. Its value usually falls within the range of 0.5 to 5. If the value is too small, the confidence factor will not attenuate sufficiently and the noise suppression will be weak. If the value is too large, the confidence factor will attenuate too quickly and may falsely affect the actual evolution region. The specific value of the differential penalty coefficient can be adjusted based on the feedback of the false alarm rate and false alarm rate during the online operation of the system.
[0141] The feature association distribution is updated by weighting and correcting it using a stability confidence factor, as shown in the following formula:
[0142] ;
[0143] in, This represents the m-th final feature association value in the weighted and corrected feature association distribution. This value will be used as the input data for calculating the divergence index. This represents the stability confidence factor for the corresponding m-th feature element obtained based on the cross-scale distribution difference mapping. This represents the m-th original feature association value in the feature association distribution before correction. This represents the value of the m-th reference feature used for calibration and anchoring.
[0144] The updated feature association distribution is used as the input data for calculating the divergence index.
[0145] Specifically, using the updated feature correlation distribution as input data for the divergence index calculation means replacing the first and second distribution state parameters extracted from the original feature correlation distribution in the divergence calculation module with the weighted and corrected updated feature correlation distribution. This ensures that subsequent calculations of state perturbation parameters, logarithmic rate of change, and the difference operation with the dissociation threshold are all performed based on the corrected and updated feature correlation distribution. As a result, the impact of instantaneous noise is suppressed before the data enters the divergence index calculation process, rather than applying smoothing filtering after the divergence index has been formed, thus avoiding the lag of filtering operations on the real rapid evolution response.
[0146] The reason why this application uses multi-scale feature deconstruction plus stability confidence factor is based on the following factors:
[0147] First, the evolution of the true teacher-student relationship should exhibit a consistent trend across different temporal / spatial scales; that is, it should be evident at small scales and roughly in the same direction at large scales to be considered a true evolution. Instantaneous noise, on the other hand, is stronger at small scales but is naturally averaged out at large scales, exhibiting cross-scale differences that are large at small scales and small at large scales. By comparing the differences between the smoothed feature values at different scales and the baseline feature values at the largest deconstruction scale (which serves as the macro-trend benchmark), we can identify which feature deviations primarily originate from instantaneous noise and which deviations originate from true evolution.
[0148] Secondly, based on the aforementioned cross-scale differences, a stability confidence factor is generated. Positions that deviate mainly from instantaneous noise are assigned lower confidence weights, and weighted corrections are made towards the reference feature values. This allows the interference of instantaneous noise on the divergence index to be suppressed without relying on external labels.
[0149] In contrast, time-domain median filtering and time-domain low-pass filtering can be used to smooth the feature correlation distribution; however, the former will blur the real state switching boundary, and the latter will cause the divergence index to lag in response to the real rapid evolution. Therefore, this embodiment adopts a stability pre-correction method based on multi-scale feature deconstruction.
[0150] By combining the above modules and two further extension schemes, the deep learning-based intelligent classroom learning analysis system provided in this embodiment no longer fixes the feature aggregation behavior at the bottom layer of the cross-modal network to either average pooling or max pooling. Instead, it forms a continuous transition between the two pooling behaviors based on the divergence of the current teacher-student relationship. In cases where cognitive differentiation occurs in sub-regions within the class, the pooling order can be further differentiated spatially based on the local information entropy of each sub-region. Furthermore, when the feature association distribution is disturbed by occasional transient noise, the pooling order adjustment signal can avoid being misled by transient noise through multi-scale pre-correction. These three levels together form a bottom-layer closed-loop feedback mechanism coupled with the dynamic evolution trend of classroom data, which helps improve the reliability of learning assessment results in complex and ever-changing classroom environments.
[0151] Example 2:
[0152] In one embodiment of the present invention, which differs from the previous embodiment, the electronic device includes one or more processors and a memory.
[0153] A processor can be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and can control other components in an electronic device to perform desired functions.
[0154] The memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0155] In one example, the electronic device may also include input devices and output devices, which are interconnected via a bus system and / or other forms of connection mechanisms (not shown). In addition, depending on the specific application, the electronic device may include any other suitable components.
[0156] Example 3:
[0157] Embodiments of this application may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps described in the "Embodiment 1" section of this specification according to the various embodiments of this application.
[0158] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0159] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not restrict the application from being implemented using the specific details described above.
[0160] The block diagrams of devices, apparatuses, devices, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0161] It should also be noted that in the apparatus, equipment, and methods of this application, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of this application.
[0162] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0163] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A classroom learning analysis system based on deep learning, characterized in that: include: The feature acquisition module is used to acquire teacher query features corresponding to teacher teaching data and class response features corresponding to student interaction data. The cross-modal association module is used to input the teacher query features and the class response features into a preset cross-modal network to calculate the feature association distribution that reflects the teacher-student association status. The divergence calculation module is used to calculate the divergence index of the feature correlation distribution and calculate the difference between the divergence index and the preset dissociation critical threshold to obtain the dissociation gradient. The dynamic order adjustment module is used to obtain the order adjustment parameter for adjusting the pooling processing layer in the cross-modal network using the disjointed gradient mapping, and to update the initial pooling order in the pooling processing layer to the dynamic pooling order according to the order adjustment parameter. The feature aggregation module is used to perform feature aggregation operations based on the dynamic pooling order on the feature association distribution using the pooling processing layer, and output fused features; wherein, the feature aggregation operation specifically includes: when the divergence index indicated by the disjoint gradient does not exceed the disjoint critical threshold, controlling the dynamic pooling order to tend to a first value, so that the pooling processing layer performs average pooling processing on the feature association distribution; and When the dissociation gradient indicates that the divergence index exceeds the dissociation critical threshold, the dynamic pooling order is controlled to be changed to a second value greater than the first value, so that the pooling processing layer performs maximum pooling processing on the feature correlation distribution focusing on extreme features. The classification output module is used to input the fused features into a classification network cascaded at the output end of the cross-modal network for probability mapping and output the corresponding learning assessment results.
2. The classroom learning intelligent analysis system based on deep learning according to claim 1, characterized in that, Before updating the initial pooling order in the pooling processing layer to the dynamic pooling order based on the order adjustment parameter, the method further includes: The feature association distribution is spatially segmented to extract multiple sub-region features; Calculate the local information entropy of each of the sub-region features, and generate spatial distribution weights corresponding to each of the sub-region features based on the local information entropy; The spatial distribution weights are fused with the order adjustment parameters to obtain the local order adjustment parameters corresponding to the features of each sub-region. The initial pooling order is updated to the local dynamic pooling order corresponding to the features of each sub-region by adjusting the parameters according to each local order.
3. The classroom learning intelligent analysis system based on deep learning according to claim 2, characterized in that, The calculation of the local information entropy of the features of each of the sub-regions includes: Extract the feature response evolution sequence of the sub-region features within a preset sliding time window; Analyze the numerical fluctuation variance of the feature response evolution sequence over time to quantify the stability of the cognitive state within the sub-region; The numerical distribution dispersion of the feature response evolution sequence in the spatial dimension is statistically analyzed to quantify the consistency of cognitive state within the sub-region; The local information entropy is obtained by performing multidimensional feature mapping on the stability and consistency of the cognitive state.
4. The classroom learning analysis system based on deep learning according to claim 1, characterized in that, Before obtaining the disjoint gradient, the following is also included: The feature association distribution is deconstructed at multiple scales to obtain a set of smooth features at multiple scales. Calculate the cross-scale distribution difference between features of different scales in the multi-scale smooth feature set, and obtain a stability confidence factor for suppressing the interference of anomalous features based on the cross-scale distribution difference. The feature association distribution is weighted and corrected using the stability confidence factor to update the feature association distribution, and the updated feature association distribution is used as the input data for calculating the divergence index.
5. The classroom learning intelligent analysis system based on deep learning according to claim 1, characterized in that, The update to the dynamic pooling order includes: The order adjustment parameter is input into the saturated activation function for constraint mapping to obtain the order offset target value limited to a bounded numerical range. Extract the historical dynamic pooling order of the pooling layer within the previous adjacent time window; The time smoothing sliding factor is invoked to perform a weighted iterative calculation on the historical dynamic pooling order and the initial pooling order superimposed with the order offset target value, so as to obtain the dynamic pooling order.
6. The classroom learning intelligent analysis system based on deep learning according to claim 5, characterized in that, Obtaining the dynamic pooling order includes: Based on the initial pooling order that superimposes the target value of the order offset and the historical dynamic pooling order, the order evolution residual is obtained; When the order evolution residual is greater than or equal to the reference zero value, the time translation momentum factor is assigned the first decay coefficient; When the order evolution residual is less than the reference zero value, the time translation momentum factor is assigned a second attenuation coefficient that is greater than the first attenuation coefficient; Using the assigned time smoothing factor as the weight allocation ratio, an exponential moving average is calculated on the historical dynamic pooling order and the initial pooling order superimposed with the order offset target value to obtain the dynamic pooling order.
7. The classroom learning intelligent analysis system based on deep learning according to claim 1, characterized in that, The process of obtaining the disjoint gradient includes: Extract the first distribution state parameter of the feature association at the current time node, and the second distribution state parameter at adjacent historical time nodes; Calculate the relative distance between the first distribution state parameter and the second distribution state parameter to obtain the state perturbation parameter used to characterize the amplitude of the associated state drift; The logarithmic rate of change of the state perturbation parameter with respect to the time evolution interval is calculated as the divergence index; The difference between the divergence index and the dissociation critical threshold is calculated to obtain an evolutionary deviation value with positive and negative signs, and the evolutionary deviation value is used as the dissociation gradient.
8. The classroom learning intelligent analysis system based on deep learning according to claim 1, characterized in that, The corresponding learning assessment results for the output include: The dynamic pooling order corresponding to the generation of the fusion feature is extracted synchronously, and the dynamic pooling order is input into a preset routing control function to map and generate a pair of gated weights with normalization properties. The fused features are respectively input into the first and second mapping branches in the parallel classification network to calculate the first candidate probability vector and the second candidate probability vector respectively. By using a pair of gating weights to perform cross-weighted calculations on the first candidate probability vector and the second candidate probability vector respectively, the target probability distribution vector is obtained. Extract the maximum probability index value from the numerical space of the target probability distribution vector, and use the learning assessment label corresponding to the maximum probability index value as the learning assessment result.
9. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the functions of the system as described in any one of claims 1 to 8.
10. A computer storage medium storing computer-executable instructions thereon, characterized in that: When the computer-executable instructions are executed by the processor, they implement the functions of the system as described in any one of claims 1 to 8.