Classroom concentration intelligent analysis method based on multi-mode student behavior data
By constructing a dynamic model library of multimodal student behavior data and combining it with teaching scenarios and personalized adjustments, the accuracy problem of existing classroom attention assessment methods has been solved, achieving a more efficient attention assessment.
Patent Information
- Application Number
- CN202511672770.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-02-13
AI Technical Summary
Existing methods for assessing classroom attention rely on single or limited-dimensional behavioral characteristics, ignoring the influence of different teaching scenarios, leading to inaccurate assessment results.
We construct a standard dynamic model library based on multimodal student behavior data, dynamically adjust the model according to the teaching scenario, combine students' personalized behavioral habits to generate real-time dynamic models, and evaluate students' concentration through conformity analysis.
It improves the accuracy of classroom attention assessment, adapts to different teaching scenarios and individual differences, and enhances the adaptability and accuracy of the assessment.
Smart Images

Figure CN121526413A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of classroom attention analysis technology, specifically to an intelligent analysis method for classroom attention based on multimodal student behavior data. Background Technology
[0002] In education, accurately assessing students' classroom focus is crucial for improving teaching quality and learning outcomes. Traditional methods for assessing classroom focus rely primarily on teachers' direct observation and subjective experience. This approach is not only inefficient but also prone to inaccurate and incomplete results due to teachers' limited personal energy and subjective biases. With the development of educational informatization, student behavior analysis systems based on computer vision or sensor technologies have emerged.
[0003] Existing technological solutions typically attempt to determine a student's focus by monitoring single or limited-dimensional behavioral characteristics (such as the number of times they look up or the direction of their face). However, these methods have significant limitations: most rely on a static, uniform assessment standard, ignoring the impact of different teaching scenarios on students' expected behavioral patterns. Due to the lack of scenario diversity, existing assessment systems are prone to misjudgment, thus affecting the accuracy of student focus assessment.
[0004] To address these issues, we propose an intelligent analysis method for classroom focus based on multimodal student behavior data. Summary of the Invention
[0005] The purpose of this invention is to provide an intelligent analysis method for classroom attention based on multimodal student behavior data, so as to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: an intelligent analysis method for classroom attention based on multimodal student behavior data, the method comprising the following steps: Based on different classroom scenarios, a standard dynamic model library for assessing classroom focus is pre-built. Each standard dynamic model corresponds to a classroom scenario and defines the range of standard behavioral parameters that characterize students' focus in that scenario. Acquire multimodal behavioral data of target students during real-time teaching, and generate real-time dynamic models corresponding to target students based on multimodal behavioral data; Identify the current classroom scenario, match the corresponding standard dynamic model from the standard dynamic model library, and adjust it to obtain the target standard dynamic model; compare the real-time dynamic model with the target standard dynamic model to determine the degree of conformity between the two. At least two compliance thresholds are preset, and the compliance is compared with the compliance thresholds to determine the real-time focus level of the target student; wherein, the focus level includes a level that meets the focus requirements and at least one level that does not meet the focus requirements.
[0007] Preferably, the step of pre-constructing a standard dynamic model library for assessing classroom focus based on different classroom scenarios includes: Classroom scenarios are categorized according to multiple pre-defined independent dimensions, and a unique scenario identifier is created for each specific classroom scenario determined by a combination of dimensions. Collect multimodal behavior data sequences of standard student groups for each specific classroom scenario, and label the data sequences with corresponding scenario identifiers and teaching progress timestamps; Based on the behavioral characteristics corresponding to the scene identifiers, a set of behavioral features related to the attention assessment of that scene is extracted from the labeled data. A dynamic benchmark range of standard behavioral parameters is established for each specific classroom scenario based on time series data. The dynamic benchmark range is used to define the normal fluctuation range of behavioral parameters under focused conditions. The scene identifier, the corresponding set of behavioral features, and the corresponding dynamic benchmark range are integrated to obtain a standard dynamic model, which is then stored in the standard dynamic model library.
[0008] Preferably, the step of establishing a dynamic benchmark range for standard behavioral parameters for each specific classroom scenario based on time series, wherein the dynamic benchmark range is used to define the normal fluctuation range of behavioral parameters under a state of focus, includes: For at least one specific spatiotemporal behavior related to attention, a reference anchor point is set at multiple key moments in its standard execution process, where each reference anchor point contains a set of multi-dimensional spatial pose parameters; Multiple reference anchor points are sequentially connected according to their corresponding key moments to form a standard behavior trajectory template for this specific spatiotemporal behavior. The primary standard dynamic reference range is obtained by associating the spatial attitude parameter fluctuation range with each reference anchor point in the standard behavior trajectory template; Data on target students performing specific spatiotemporal behaviors multiple times while in a state of perceived focus is obtained, and personalized behavioral trajectory templates for target students are extracted and generated. The time axis of the baseline anchor point is scaled proportionally based on the duration of the personalized behavior trajectory template. Based on the movement amplitude of the personalized behavior trajectory template at the key reference anchor points, the fluctuation range of the spatial attitude parameters of the corresponding reference anchor points is adjusted, and the adjusted primary standard dynamic reference range is used as the final standard dynamic reference range.
[0009] Preferably, the step of acquiring multimodal behavioral data of the target student during real-time teaching and generating a real-time dynamic model corresponding to the target student based on the multimodal behavioral data includes: Multiple sensors are deployed in the teaching environment to simultaneously collect multimodal behavioral data of target students within a preset time window. Based on the behavioral data of each modality, the corresponding behavioral feature vectors are extracted; the behavioral feature vectors from different modalities are aligned and fused on the time axis to form a multimodal temporal feature sequence; The real-time dynamic model is constructed based on multimodal temporal feature sequences.
[0010] Preferably, the step of identifying the current classroom scenario, matching the standard dynamic model corresponding to the current classroom scenario from the standard dynamic model library, and adjusting it to obtain the target standard dynamic model includes: Identify the current classroom scenario and match the corresponding standard dynamic model from the standard dynamic model library as the primary standard dynamic model; Identify the target student, obtain the duration and range of motion of the target student's personalized movement trajectory, adjust the dynamic reference range to obtain the personalized dynamic reference range, and use the standard dynamic model corresponding to the personalized dynamic reference range as the target standard dynamic model.
[0011] Preferably, the step of comparing the real-time dynamic model with the target standard dynamic model to determine the degree of conformity between the two includes: Obtain the scene-relative time axis corresponding to the real-time dynamic model and the target standard dynamic model respectively; establish a synchronization chain based on the same time node on the scene-relative time axis; Traverse each time node in the scene relative to the time axis corresponding to the real-time dynamic model, and extract the first multi-dimensional pose parameters of the key points in the real-time dynamic model at that time node. Based on the synchronization chain, the second multi-dimensional attitude parameters and their personalized dynamic reference range of the key points corresponding to the target standard dynamic model at the time node of the scene relative to the time axis are called; Calculate the difference information between the personalized dynamic reference range of the first multi-dimensional attitude parameters and the second multi-dimensional attitude parameters, and generate the comprehensive conformity between the real-time dynamic model and the target standard dynamic model based on the difference information.
[0012] Preferably, the step of calculating the difference information between the personalized dynamic reference range of the first multi-dimensional attitude parameters and the second multi-dimensional attitude parameters, and generating the comprehensive conformity between the real-time dynamic model and the target standard dynamic model based on the difference information includes: Iterate through each time point in the scene relative to the time axis; for each traversed time point, determine whether the first multi-dimensional attitude parameter of the real-time dynamic model falls within the personalized dynamic reference range of the second multi-dimensional attitude parameter. If the value of any dimension of the first multi-dimensional attitude parameter does not fall within the reference range of its corresponding dimension, it is determined that the attitude parameter of the node does not completely fall within the range of the personalized dynamic reference. The number of nodes that are determined to fall completely within the range of the personalized dynamic benchmark is counted among all the time nodes that are traversed. Calculate the ratio of the number of nodes to the total number of time nodes in the synchronization chain, and output this ratio as the degree of conformity between the real-time dynamic model and the target standard dynamic model.
[0013] Preferably, the step of presetting at least two compliance thresholds and comparing the compliance score with the compliance thresholds to determine the real-time focus level of the target student includes: At least two compliance thresholds are preset, namely a first compliance threshold and a second compliance threshold, wherein the first compliance threshold is less than the second compliance threshold, and the second compliance threshold is the minimum compliance standard for meeting the focus requirement; The compliance rate is compared with the preset compliance rate threshold. When the compliance rate is greater than or equal to the second compliance rate threshold, the real-time focus level of the target student is determined to meet the focus requirement level. When the compliance rate is less than the second compliance rate threshold but greater than or equal to the first compliance rate threshold, the target student's real-time focus level is determined to be the first non-compliance level. When the compliance rate is less than the first compliance rate threshold, the target student's real-time focus level is determined to be the second non-compliance level, where the focus level of the second non-compliance level is lower than that of the first non-compliance level.
[0014] Compared with the prior art, the beneficial effects of the present invention are: Based on the teaching scenario, standard dynamic models corresponding to different student concentration states are constructed, and the dynamic benchmark range of the standard dynamic models is adjusted according to the unique behavioral habits of students to increase the adaptability of the standard dynamic models to the target students. The real-time dynamic models of the target students are compared with the standard dynamic models to analyze students' concentration in specific scenarios and improve the accuracy of student concentration assessment. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram of the method architecture of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example
[0018] Please see Figures 1 to 2 This invention provides a technical solution for an intelligent analysis method of classroom attention based on multimodal student behavior data: the intelligent analysis method of classroom attention based on multimodal student behavior data includes the following steps: S1: Based on different classroom scenarios, a standard dynamic model library for evaluating classroom focus is pre-built. Each standard dynamic model corresponds to a classroom scenario and defines the range of standard behavioral parameters that represent students' focused state in that scenario. The steps for pre-constructing a standard dynamic model library for assessing classroom focus, based on different classroom scenarios, include: classifying classroom scenarios according to multiple preset independent dimensions, creating a unique scenario identifier for each specific classroom scenario determined by a combination of dimensions; collecting multimodal behavioral data sequences of a standard student group for each specific classroom scenario, and labeling the data sequences with corresponding scenario identifiers and teaching progress timestamps; extracting a set of behavioral features related to the focus assessment of that scenario from the labeled data based on the behavioral characteristics corresponding to the scenario identifiers; establishing a dynamic benchmark range for standard behavioral parameters for each specific classroom scenario based on the time series, the dynamic benchmark range being used to define the normal fluctuation range of behavioral parameters under focused conditions; and integrating the scenario identifiers, the corresponding set of behavioral features, and their corresponding dynamic benchmark ranges to obtain a standard dynamic model, which is then stored in the standard dynamic model library. Specifically, multiple independent dimensions include teaching interaction modes, knowledge transmission media, and classroom activity organization formats. The teaching interaction mode dimension includes teacher-led lectures, teacher-student Q&A, group collaboration, and individual study; the knowledge transmission media dimension includes blackboard teaching, slide presentations, video playback, and physical demonstrations; the classroom activity organization format dimension includes new knowledge instruction, in-class exercises, project-based learning, and summary and review. Multimodal behavioral data sequences generated by standard student groups in each specific classroom scenario are collected. These sequences are generated by standard student groups confirmed to be in a focused state. The criteria for confirming a "standard student group in a focused state" are the combined results of post-class knowledge mastery test scores and teacher / expert annotations during classroom video playback. The collected annotated data is then analyzed based on the behavioral characteristics corresponding to its scenario identifiers, from the perspective of multimodal... The most relevant set of behavioral features for the attention assessment in a given scenario is extracted from the behavioral data and parameterized. This is a scenario-adaptive feature extraction method. Specifically, for scenarios characterized by "teacher lecturing," features related to the duration of audiovisual attention are extracted first, including facial orientation stability and the proportion of time the gaze is focused on the blackboard or main display screen. For scenarios characterized by "group collaboration," features related to interactive participation are extracted first, including the frequency of eye contact with group members, specific gesture semantic features, and the number of voice dialogue rounds. The dynamic baseline range is a tolerance interval that changes over time. It is achieved by calculating the quantile interval of standard student group behavioral parameters at the same time or the same teaching stage. The tolerance interval is defined by the α-quantile and the β-quantile, where α and β are preset values based on the required strictness of the model. The steps for establishing a dynamic baseline range of standard behavioral parameters for each specific classroom scenario based on time series, and defining the normal fluctuation range of behavioral parameters under focused conditions, include: setting a baseline anchor point at multiple key moments during the standard execution of at least one specific spatiotemporal behavior related to focus, where each baseline anchor point contains a set of multi-dimensional spatial posture parameters; sequentially connecting multiple baseline anchor points according to their corresponding key moments to form a standard behavioral trajectory template for that specific spatiotemporal behavior; associating the fluctuation range of spatial posture parameters with each baseline anchor point in the standard behavioral trajectory template to obtain a primary standard dynamic baseline range; acquiring data on the target student's multiple executions of the specific spatiotemporal behavior under a state of focus, extracting and generating a personalized behavioral trajectory template for the target student; scaling the time axis of the baseline anchor points proportionally according to the duration of the personalized behavioral trajectory template; adjusting the fluctuation range of spatial posture parameters of the corresponding baseline anchor points according to the movement amplitude of the personalized behavioral trajectory template at the key baseline anchor points, and using the adjusted primary standard dynamic baseline range as the final standard dynamic baseline range. It should be noted that scaling the time axis involves: calculating the ratio of the total duration of a specific spatiotemporal behavior completed by the personalized behavior trajectory template to the total duration of the standard behavior trajectory template. This ratio is used as the time axis scaling factor, and the timestamp intervals of all reference anchor points are uniformly scaled. Here, the time axis is the behavior trajectory template time axis, with the start time of a specific behavior as the relative farthest point. It is used to describe and compare the execution rhythm and sequence within an independent behavior, and can be personalized. Depending on student habits, the execution duration of the action itself is stretched or compressed (multiplied by the time scaling factor), which is a local time axis. Scaling the range of spatial posture parameter fluctuations involves: dynamically adjusting the tolerance of the fluctuation range based on the statistical dispersion of the spatial posture parameters at each reference anchor point in the personalized behavior trajectory template; the greater the dispersion, the wider the scaled fluctuation range. Multi-dimensional spatial posture parameters are a set of parameters used to quantitatively describe the posture and motion state of the human body or its parts in three-dimensional space. This set of parameters collectively constitutes a feature vector, which may include, but is not limited to: joint angles, three-dimensional spatial coordinates of limb extremities, Euler angles of the human trunk orientation, and relative positional relationships with external interacting objects; a reference anchor point is a data structure used to record the nominal values of a set of multi-dimensional spatial posture parameters representing the behavioral state at a critical moment; a behavioral trajectory template refers to a reference trajectory representing the standard execution mode of a behavior, constructed by sequentially connecting multiple reference anchor points within a complete execution cycle of a specific spatiotemporal behavior according to their corresponding critical moments. It defines the ideal dynamic sequence of the behavior in time and space; the spatial posture parameter fluctuation range refers to a numerical range set at a specific reference anchor point for each dimension of the multi-dimensional spatial posture parameters that allows fluctuation; a personalized behavioral trajectory template refers to a benchmark model suitable for a specific student, obtained by deforming the standard behavioral trajectory template. Deformation mainly includes time axis scaling (changing the speed of behavior) and spatial range adjustment (changing the tolerance of the range of motion), thereby adapting the general standard to the individual habits of students. The adjustment usually includes scaling the template time axis and / or scaling the range of fluctuations in spatial posture parameters. The dynamic reference range corresponding to the standard dynamic model is adjusted according to the individual habits of students, so that it can better adapt to the unique behavioral habits of each student, thereby improving the accuracy of the assessment of students' concentration.
[0019] S2: Acquire multimodal behavioral data of the target student during real-time teaching, and generate a real-time dynamic model corresponding to the target student based on the multimodal behavioral data; The steps for acquiring multimodal behavioral data of target students during real-time teaching and generating a real-time dynamic model corresponding to the target students based on the multimodal behavioral data include: deploying multiple sensors in the teaching environment and synchronously collecting multimodal behavioral data of target students within a preset time window based on the multiple sensors; extracting corresponding behavioral feature vectors based on the behavioral data of each modality; aligning and fusing the behavioral feature vectors from different modalities on the time axis to form a multimodal temporal feature sequence; and constructing a real-time dynamic model based on the multimodal temporal feature sequence. Specifically, facial expressions (such as blinking frequency and smile intensity) and head posture (such as angle changes) are captured through student terminal cameras to generate behavioral data streams. The real-time dynamic model is represented as a temporal sequence or statistical distribution of one or more attention-related behavioral parameters within a preset time window. Multimodal behavioral data includes at least two of the following: visual modality data, auditory modality data, and environmental modality data. For visual modality data, the extracted behavioral feature vectors include at least one of the following: head horizontal rotation angle, head vertical pitch angle, gaze point coordinates, eye closure state index, mouth state (such as whether speaking), and upper body posture key point coordinates. Auditory modality behavioral data includes speech activity detection and non-speech event detection (recognizing sounds with behavioral indications such as yawning, sighing, tapping on the table, turning pages in a book), and student-related sound signal data acquired through one or more audio acquisition devices (such as microphone arrays). The preset time window is a sliding, continuous time window, allowing the real-time attention dynamic model to be updated periodically or continuously to reflect the dynamic changes in student attention.
[0020] S3: Identify the current classroom scenario, match the corresponding standard dynamic model from the standard dynamic model library, and adjust it to obtain the target standard dynamic model; compare the real-time dynamic model with the target standard dynamic model to determine the degree of conformity between the two; The steps of identifying the current classroom scenario, matching the standard dynamic model corresponding to the current classroom scenario from the standard dynamic model library, and adjusting it to obtain the target standard dynamic model include: identifying the current classroom scenario, matching the standard dynamic model corresponding to the current classroom scenario from the standard dynamic model library as the initial standard dynamic model; identifying the target student, obtaining the duration and amplitude of the target student's personalized movement trajectory, adjusting the dynamic reference range to obtain the personalized dynamic reference range, and using the standard dynamic model corresponding to the personalized dynamic reference range as the target standard dynamic model; Specifically, identifying the current classroom scenario is achieved by analyzing at least one of the teacher's voice content, teaching materials content, or classroom activity instructions. The time axis of the standard dynamic model is scaled proportionally based on the duration of the personalized motion trajectory. The range of spatial posture parameters at key points in the standard dynamic model is scaled based on the stability of the personalized motion trajectory's amplitude. The personalized dynamic reference range here is specific to the target student; each student's range is different. Since each student has different habits when listening in class, these small habits do not necessarily indicate decreased concentration. For example, student A might constantly twirl their pen while listening. This could easily lead to misjudgment: based on the general standard model that "focused students should maintain a still posture," the system might classify this behavior as a "small movement," indicating low concentration. However, in reality, student A's gaze consistently follows the teacher and the blackboard, they take notes on key points, and they accurately answer classroom questions. His pen-spinning is an accompanying habit that helps him concentrate. Through learning phases, it was discovered that pen-spinning occurred frequently during student A's historical periods of focus (such as after answering a question correctly). Therefore, it was necessary to generate a personalized behavioral trajectory template for student A corresponding to the "pen-spinning" behavior, including his habitual hand movement range, rhythm, and posture. Real-time assessment revealed that student A's hand movements conformed to his personal focus template, and core dimensions such as gaze and head posture also met the standards. Therefore, he was determined to be in a focused state. A dynamic benchmark range was generated based on the personalized behavioral trajectory template to adapt to each student's differences, thereby generating a personalized standard dynamic benchmark range, which can improve the accuracy of focus assessment for each student. The steps to compare the real-time dynamic model with the target standard dynamic model and determine their conformity include: obtaining the scene-relative time axis corresponding to the real-time dynamic model and the target standard dynamic model respectively; and establishing a synchronization chain based on the same time node on the scene-relative time axis. Traverse each time node in the scene relative to the time axis corresponding to the real-time dynamic model, and extract the first multi-dimensional pose parameters of the key points in the real-time dynamic model at that time node. Based on the synchronization chain, the second multi-dimensional attitude parameters and their personalized dynamic reference ranges of the corresponding key points in the target standard dynamic model at the time node of the scene relative to the target standard dynamic model are called; the difference information between the personalized dynamic reference ranges of the first multi-dimensional attitude parameters and the second multi-dimensional attitude parameters is calculated, and the comprehensive conformity between the real-time dynamic model and the target standard dynamic model is generated based on the difference information. The steps for calculating the difference information between the personalized dynamic reference ranges of the first multi-dimensional attitude parameters and the second multi-dimensional attitude parameters, and generating the comprehensive conformity between the real-time dynamic model and the target standard dynamic model based on the difference information, include: traversing each time node in the synchronization chain; for each traversed time node, determining whether the first multi-dimensional attitude parameters of the real-time dynamic model fall within the personalized dynamic reference range of the second multi-dimensional attitude parameters; if the value of any dimension of the first multi-dimensional attitude parameter does not fall within the reference range of its corresponding dimension, it is determined that the attitude parameter of that node does not completely fall within the personalized dynamic reference range; counting the number of nodes determined to completely fall within the personalized dynamic reference range among all traversed time nodes; calculating the ratio of the number of nodes to the total number of time nodes in the synchronization chain, and outputting this ratio as the conformity between the real-time dynamic model and the target standard dynamic model.
[0021] It should be noted that the method for calculating the difference is as follows: It determines whether the first multi-dimensional attitude parameter falls within the personalized dynamic baseline range of the second multi-dimensional attitude parameter, and then calculates the proportion of nodes falling within this range across all time points in the synchronization chain. For example, in a classroom scenario where a teacher explains a new formula on the blackboard, the specific spatiotemporal behavior of "taking notes" is selected for analysis. A personalized dynamic baseline range for "taking notes" is generated for this student in this scenario, dividing the 3-second action of "taking notes" evenly into 3 key nodes (t1, t2, t3) on the timeline. Each node defines the personalized dynamic baseline range for the student's note-taking while focused. Assuming each node's baseline range contains only two dimensions: Dimension A: vertical distance between the head and the desktop (unit: cm), and Dimension B: vertical distance between the fingertips of the right hand holding the pen and the desktop (unit: cm), in the real-time dynamic model, a student's actual "note-taking" action is captured, lasting 3 seconds, and the first multi-dimensional pose parameters are extracted at the same three time nodes (t1, t2, t3). The synchronization chain is: [time node t1, time node t2, time node t3]. The total number of time nodes is 3. Assuming the baseline range and real-time data for each node are as shown in the table below:
[0022] Judgment Process: At node t1: Both dimension A (35cm) and dimension B (7cm) of the real-time data fall within their respective personalized baseline ranges. Therefore, node t1 is judged as "fully within the baseline range". At node t2: Dimension A (25cm) of the real-time data is below the lower limit of the baseline range (28cm), even though dimension B (5cm) falls within the range. According to the rule that "if any dimension does not fall within the baseline range, the node is not fully within the baseline range", node t2 is judged as "not fully within the baseline range". At node t3: Dimension B (12cm) of the real-time data is above the upper limit of the baseline range (10cm), even though dimension A (38cm) falls within the baseline range. Similarly, node t3 is judged as "not fully included". The number of nodes that are fully included is counted as 1, and the total number of event nodes in the synchronization chain is 3. The formula for calculating the overall compliance is: Overall Compliance = (Number of nodes that are fully included) / (Total number of nodes) = 1 / 3 ≈ 0.33. The overall compliance of this "note-taking" behavior is output as 0.33. This compliance indicates that the student's spatial posture of the head and hands during this note-taking process deviates significantly from the habitual pattern formed when he is focused. For example, at time t2, the head is too low (possibly lying on the table), and at time t3, the hands are raised too high (possibly playing with the pen rather than writing). These detailed deviations suggest that the student is likely to have insufficient focus during this note-taking behavior. At each time point, a strict "range hit" test is performed on each behavioral dimension. If any dimension does not comply, the entire behavior at that time point is judged as non-compliant. The proportion of compliant time points to the total number of time points is used as the measure of overall compliance, thereby improving the accuracy of the judgment.
[0023] The specific steps for determining whether the first multi-dimensional attitude parameter falls within the personalized dynamic reference range of the second multi-dimensional attitude parameter are as follows: Deconstruct the first multi-dimensional attitude parameter into N independent dimensional parameters, forming a set of dimensional parameters. Simultaneously, the personalized dynamic benchmark range is deconstructed into N dimensional benchmark ranges that correspond one-to-one with the dimensional parameters, forming a set of dimensional ranges. The baseline range for each dimension Defined dimension parameters Normal fluctuation range ; Iterate through each dimension parameter in the dimension parameter set Execute the judgment: This dimension parameter Is the value within its corresponding dimensional baseline range? interval Within; if and only if, in the dimension-by-dimensional comparison step, all dimension parameters All were determined to be within their corresponding dimensional baseline range. If the first multi-dimensional attitude parameter falls completely within the range of the personalized dynamic benchmark, then it is determined that the first multi-dimensional attitude parameter falls completely within the range of the personalized dynamic benchmark; if any dimension parameter... Not within its corresponding dimensional benchmark range If the position is within the specified range, it is determined that the position has not been fully entered. The multi-dimensional posture parameters include at least two of the following dimensions: X-axis coordinates, Y-axis coordinates, and Z-axis coordinates of the hand key points; rotation angle of the head around the X-axis and rotation angle around the Y-axis; and elevation angle of the pen relative to the hand.
[0024] Specifically, a relative timeline is defined for a specific classroom scenario. This relative timeline uses the start time of the specific classroom scenario as a common time origin (t=0) to measure the relative execution progress of the teaching process or action sequence within the scenario, rather than actual absolute clock time. During comparison, the real-time action data at time T on the relative timeline of the scenario in the real-time dynamic model is compared with the standard dynamic benchmark range at the same time T on the relative timeline of the scenario in the standard dynamic model to assess focus. The standard dynamic benchmark range at time T on the relative timeline of the scenario in the standard dynamic model is pre-calibrated based on a personalized behavioral trajectory template. The personalized dynamic benchmark range generated after adjusting the quasi-benchmark ensures that the comparison is made on the same teaching progress point, reducing interference caused by different start and end times of the class or different paces of progress, and ensuring the accuracy of the selected standard action benchmark. Based on the dynamic benchmark range of the scene relative to the time axis, it can be further adjusted according to the student's personalized behavior trajectory template (such as time axis scaling and spatial range scaling), thereby ensuring the consistency of the real-time dynamic model and the standard dynamic model on the time axis. At the same time, the dynamic benchmark range corresponding to the standard dynamic model can be adjusted according to individualization, thereby improving the accuracy of the analysis and evaluation of focus.
[0025] S4: Preset at least two compliance thresholds, compare compliance with the compliance thresholds to determine the real-time focus level of the target student; wherein, the focus level includes a level that meets the focus requirements and at least one level that does not meet the focus requirements; The steps for determining the real-time focus level of a target student by presetting at least two compliance thresholds and comparing the compliance score with these thresholds include: presetting at least two compliance thresholds, namely a first compliance threshold and a second compliance threshold, wherein the first compliance threshold is less than the second compliance threshold, and the second compliance threshold is the minimum compliance standard for meeting the focus requirement; comparing the compliance score with the preset compliance thresholds; when the compliance score is greater than or equal to the second compliance threshold, the target student's real-time focus level is determined to be a level that meets the focus requirement; when the compliance score is less than the second compliance threshold but greater than or equal to the first compliance threshold, the target student's real-time focus level is determined to be a first level that does not meet the focus requirement; when the compliance score is less than the first compliance threshold, the target student's real-time focus level is determined to be a second level that does not meet the focus requirement, wherein the focus level of the second level that does not meet the focus requirement is lower than that of the first level that does not meet the focus requirement.
[0026] Specifically, standard dynamic models corresponding to different student attention states are constructed based on teaching scenarios. The dynamic baseline range of the standard dynamic models is adjusted according to the unique behavioral habits of students to increase the adaptability of the standard dynamic models to the target students. The real-time dynamic models of the target students are compared with the standard dynamic models to analyze students' attention in specific scenarios and improve the accuracy of students' attention assessment.
[0027] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0028] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for intelligent analysis of classroom attentiveness based on multi-modal student behavior data, characterized in that, The method comprises the following steps: According to different classroom scenes, a standard dynamic model library for evaluating classroom concentration is pre-built, wherein each standard dynamic model corresponds to a classroom scene and defines the standard behavior parameter range representing the student's concentration state under the scene; Obtain the multi-modal behavior data of the target student in the real-time teaching process, and generate a real-time dynamic model corresponding to the target student based on the multi-modal behavior data; Identify the current classroom scene, match the standard dynamic model corresponding to the current classroom scene from the standard dynamic model library and adjust it to obtain a target standard dynamic model; compare the real-time dynamic model with the target standard dynamic model to determine the degree of coincidence between them; Pre-set at least two coincidence thresholds, compare the coincidence with the coincidence thresholds, and determine the real-time concentration level of the target student; wherein the concentration level includes a concentration requirement level and at least one non-concentration requirement level. 2.The method of claim 1, wherein: The step of pre-building a standard dynamic model library for evaluating classroom concentration according to different classroom scenes comprises: Classify the classroom scenes according to pre-set multiple independent dimensions, and create a unique scene identifier for each specific classroom scene determined by the dimension combination; Collect multi-modal behavior data sequences of a standard student group in each specific classroom scene, and label the data sequences with corresponding scene identifiers and teaching progress time stamps; Based on the behavior characteristics corresponding to the scene identifier, extract a set of behavior characteristics related to the scene concentration evaluation from the labeled data; Based on the time sequence, establish a dynamic reference range of standard behavior parameters for each specific classroom scene, which is used to define the normal fluctuation interval of behavior parameters in the concentration state; Integrate the scene identifier, the corresponding behavior characteristic set and the corresponding dynamic reference range to obtain a standard dynamic model, and store it in the standard dynamic model library. 3.The method of claim 2, wherein: The step of establishing a dynamic reference range of standard behavior parameters for each specific classroom scene based on the time sequence, which is used to define the normal fluctuation interval of behavior parameters in the concentration state, comprises: For at least one specific spatio-temporal behavior related to concentration, set a reference anchor point at each key moment in its standard execution process, wherein each reference anchor point contains a set of multi-dimensional spatial posture parameters; Connect the multiple reference anchor points in sequence according to their corresponding key moments to form a standard behavior trajectory template of the specific spatio-temporal behavior; Associate the spatial posture parameter fluctuation range for each reference anchor point in the standard behavior trajectory template to obtain a primary standard dynamic reference range; Obtain the data of the target student performing the specific spatio-temporal behavior multiple times in the state of being recognized as concentrating, extract and generate a personalized behavior trajectory template of the target student; According to the duration of the personalized behavior trajectory template, the time axis of the reference anchor point is scaled proportionally; According to the action amplitude of the personalized behavior trajectory template at the key reference anchor point, adjust the spatial posture parameter fluctuation range of the corresponding reference anchor point, and take the adjusted primary standard dynamic reference range as the final standard dynamic reference range.
4. The method for intelligent analysis of classroom attentiveness based on multi-modal student behavior data according to claim 1, wherein: The step of acquiring multi-modal behavior data of the target student in the real-time teaching process and generating a real-time dynamic model corresponding to the target student based on the multi-modal behavior data comprises: deploying multiple sensors in a teaching environment and synchronously collecting multi-modal behavior data of the target student within a preset time window based on the multiple sensors; extracting a behavior feature vector corresponding to each modality; aligning and fusing the behavior feature vectors from different modalities on a time axis to form a multi-modal time sequence feature sequence; constructing the real-time dynamic model based on the multi-modal time sequence feature sequence.
5. The method for intelligent analysis of classroom attentiveness based on multi-modal student behavior data as claimed in claim 1, wherein: The step of identifying a current classroom scene, matching a standard dynamic model corresponding to the current classroom scene from the standard dynamic model library, and adjusting the standard dynamic model to obtain a target standard dynamic model comprises: identifying a current classroom scene, matching a standard dynamic model corresponding to the current classroom scene from the standard dynamic model library as a preliminary standard dynamic model; determining a target student, acquiring the duration and amplitude of the personalized motion trajectory of the target student, adjusting the dynamic reference range to obtain a personalized dynamic reference range, and taking the standard dynamic model corresponding to the personalized dynamic reference range as the target standard dynamic model.
6. The method for intelligent analysis of in-class attentiveness based on multi-modal student behavior data according to claim 1, wherein: The step of comparing the real-time dynamic model with the target standard dynamic model to determine the degree of conformity between the two comprises: acquiring a scene relative time axis corresponding to the real-time dynamic model and the target standard dynamic model respectively; based on the same time nodes, establishing a synchronization chain on the scene relative time axis; traversing each time node in the scene relative time axis corresponding to the real-time dynamic model, extracting the first multi-dimensional pose parameter of the key point in the real-time dynamic model at the time node; based on the synchronization chain, calling the second multi-dimensional pose parameter of the corresponding key point in the target standard dynamic model and the personalized dynamic reference range of the second multi-dimensional pose parameter at the time node in the scene relative time axis corresponding to the target standard dynamic model; calculating the difference information between the first multi-dimensional pose parameter and the personalized dynamic reference range of the second multi-dimensional pose parameter, and generating the comprehensive conformity between the real-time dynamic model and the target standard dynamic model based on the difference information.
7. The method for intelligent analysis of in-class attentiveness based on multi-modal student behavior data according to claim 1, wherein: The step of calculating the difference information between the first multi-dimensional pose parameter and the personalized dynamic reference range of the second multi-dimensional pose parameter, and generating the comprehensive conformity between the real-time dynamic model and the target standard dynamic model based on the difference information comprises: traversing each time node in the scene relative time axis; for each traversed time node, determining whether the first multi-dimensional pose parameter of the real-time dynamic model falls within the personalized dynamic reference range of the second multi-dimensional pose parameter; if the value of any dimension of the first multi-dimensional pose parameter does not fall within the reference range of the corresponding dimension, it is determined that the pose parameter of the node does not completely fall within the personalized dynamic reference range; counting the number of nodes that are determined to completely fall within the personalized dynamic reference range among all traversed time nodes; The ratio of the number of the nodes to the total number of time nodes in the synchronization chain is calculated, and the ratio is output as a degree of coincidence between the real-time dynamic model and the target standard dynamic model.
8. The method for intelligent analysis of in-class attentiveness based on multi-modal student behavior data as claimed in claim 1, wherein: The step of comparing the degree of coincidence with the preset at least two coincidence threshold values to determine the real-time concentration level of the target student includes: The preset at least two coincidence threshold values are a first coincidence threshold value and a second coincidence threshold value, wherein the first coincidence threshold value is less than the second coincidence threshold value, and the second coincidence threshold value is the minimum coincidence standard for meeting the concentration requirement; When the degree of coincidence is greater than or equal to the second coincidence threshold value, it is determined that the real-time concentration level of the target student is a concentration requirement meeting level; When the degree of coincidence is less than the second coincidence threshold value and greater than or equal to the first coincidence threshold value, it is determined that the real-time concentration level of the target student is a first concentration requirement not meeting level; When the degree of coincidence is less than the first coincidence threshold value, it is determined that the real-time concentration level of the target student is a second concentration requirement not meeting level, wherein the concentration level of the second concentration requirement not meeting level is lower than that of the first concentration requirement not meeting level.