Dynamic feedback method and system for attention in pre-school education classroom driven by somatosensory interaction
By constructing a classroom behavior scenario library and dynamically weighting the haptic data, the problems of cross-scenario adaptability and data processing delay in haptic interaction attention monitoring technology have been solved, realizing the real-time and accurate assessment of preschool classroom attention and supporting personalized teaching strategy adjustments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XINGAN VOCATIONAL & TECH COLLEGE
- Filing Date
- 2025-09-04
- Publication Date
- 2026-05-29
Smart Images

Figure CN121092918B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent education technology, and more specifically, this application relates to a method and system for dynamic feedback of attention in preschool classrooms driven by haptic interaction. Background Technology
[0002] In the field of intelligent education, classroom attention monitoring is a core component of optimizing teaching quality. Its ability to acquire students' attention status in real time supports teachers in dynamically adjusting teaching strategies and interaction modes, thereby significantly improving classroom participation and knowledge transfer efficiency. Especially in preschool education, preschool children (3-6 years old) exhibit significantly different physical characteristics compared to adults: their standard deviation of body movement amplitude is 62% higher than that of adults, and their average attention span is less than eight minutes.
[0003] Current attention monitoring technologies based on haptic interaction, while indirectly assessing attention levels by capturing unstructured haptic features such as limb movement trajectories and micro-expression changes, still face two major technical bottlenecks: First, they have weak cross-scenario adaptability, making it difficult to be compatible with diverse classroom formats such as gamified teaching and group collaboration; second, data processing latency is too high, failing to meet the timeliness requirements of dynamic feedback and restricting the effectiveness of personalized teaching interventions.
[0004] Existing classroom motion sensing data collection systems employ fixed collection strategies, such as uniform frequency and fixed area collection, which cannot adapt to the ever-changing teaching process. This results in both redundant non-key motion sensing data and missing key motion sensing data. At the same time, significant noise interference occurs during the processing of motion sensing feature data, leading to insufficient real-time performance and accuracy of attention assessment. Ultimately, this affects the real-time performance and effectiveness of dynamic feedback on classroom attention. Therefore, a motion-sensory interaction-driven method and system for dynamic feedback on preschool classroom attention is proposed to solve this problem. Summary of the Invention
[0005] To address the aforementioned technical issues, this technical solution provides a method and system for dynamic attention feedback in preschool classrooms driven by motion-sensing interaction, resolving the problems mentioned in the background section.
[0006] To achieve the above objectives, the technical solution of the present invention is as follows:
[0007] In a first aspect, this application provides a method for dynamic attention feedback in preschool classrooms driven by motion-sensing interaction, the method comprising:
[0008] Acquire classroom sample data, perform scenario analysis on it, extract scenario feature data corresponding to each scenario and store it to build a classroom behavior scenario library;
[0009] Collect real-time classroom information from multiple sources and convert it into multimodal text data. After filtering and association to form joint matching data, perform threshold matching with scene feature data in the classroom behavior scene database.
[0010] If the match is successful, a collection control command is generated based on the scene feature data to drive the collection device to perform the collection operation and obtain the original somatosensory data with timestamp alignment. If the match fails, the information collection time is extended and the joint matching data is regenerated. After multiple consecutive matching failures, an error message is pushed to the terminal device. For example, the information collection time can be extended by 10 seconds, and it can be extended three times in a row. After the error message is displayed, manual intervention is required.
[0011] The information entropy values of various somatosensory feature data in the original somatosensory data are calculated, and dynamic weights are generated by combining scene feature data. Valid somatosensory data is then selected through a collaborative filtering mechanism of multi-dimensional scoring and dynamic weights.
[0012] Effective haptic data is input into a pre-built attention assessment model to generate attention assessment data, which is then pushed to the terminal device.
[0013] Secondly, this application provides a motion-sensing interactive-driven dynamic attention feedback system for preschool classrooms, used to implement the motion-sensing interactive-driven dynamic attention feedback method for preschool classrooms described in any of the above claims, including:
[0014] The data acquisition module is used to acquire classroom sample data, perform scenario analysis on it, extract scenario feature data corresponding to each scenario, and store it to build a classroom behavior scenario library;
[0015] The scene matching module is used to collect real-time classroom multi-source information and convert it into multimodal text data. After filtering and association to form joint matching data, it performs threshold matching with scene feature data in the classroom behavior scene library. If the matching is successful, it generates collection control instructions based on scene feature data to drive the collection device to perform collection operations to obtain the original somatosensory data aligned with the timestamp. If the matching fails, it extends the information collection time and regenerates the joint matching data. After multiple consecutive matching failures, it pushes an abnormal prompt to the terminal device.
[0016] The somatosensory data processing module is used to calculate the information entropy values of various somatosensory feature data in the raw somatosensory data, and generate dynamic weights by combining them with scene feature data;
[0017] The attention assessment and push module is used to input effective somatosensory data into a pre-built attention assessment model, generate attention assessment data, and push it to the terminal device.
[0018] Thirdly, this application provides a computer device, the computer device including a memory and a processor, the memory storing code, and the processor being configured to acquire the code and execute the above-described feedback adjustment method based on haptic interaction.
[0019] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described feedback adjustment method based on haptic interaction.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0021] This application analyzes classroom sample data, extracts teaching links and highly correlated somatosensory feature data to construct a scenario library, and combines real-time multi-source information transformation multimodal text data for scenario matching. After successful matching, the collection object, frequency and region are dynamically generated, realizing the precise adaptation of the collection strategy to the teaching scenario, effectively reducing the redundancy of non-key somatosensory data, while avoiding the loss of key somatosensory data information, and improving the targeting and efficiency of data collection.
[0022] This application calculates the information entropy of the original somatosensory data, constructs a dynamic weight model by combining scene features and the covariance between features, and calculates the retention probability by nonlinearly aggregating time correlation score, spatial correlation score and feature stability score, and filters effective data by combining dynamic weights. This significantly reduces noise interference and the computational cost of the attention evaluation model, thereby improving the accuracy and real-time performance of attention evaluation. Attached Figure Description
[0023] The disclosure of this invention is illustrated with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. Wherein:
[0024] Figure 1 The flowchart is a method for dynamic feedback of attention in preschool classrooms driven by somatosensory interaction, as proposed in this invention.
[0025] Figure 2 This is a flowchart of the real-time classroom attention dynamic feedback execution steps in this invention;
[0026] Figure 3 This is a structural block diagram of the motion-sensing interactive-driven dynamic feedback system for preschool classroom attention proposed in this invention. Detailed Implementation
[0027] It is readily understood that, based on the technical solution of this invention, those skilled in the art can propose various interchangeable structural methods and implementations without altering the essential spirit of the invention. Therefore, the following detailed embodiments and accompanying drawings are merely illustrative examples of the technical solution of this invention and should not be considered as the entirety of the invention or as limitations or restrictions on the technical solution of this invention.
[0028] In existing technologies, classroom motion sensing data collection systems typically employ fixed collection strategies, which cannot dynamically adjust data collection parameters according to actual teaching scenarios. This fixed mode can easily lead to the accumulation of redundant data in the ever-changing teaching process, and may also result in the omission of key motion sensing information.
[0029] In preschool education settings, different teaching stages have significantly different requirements for somatosensory data collection: group storytelling requires capturing facial micro-expressions, while craft activities require accurately tracking fine hand movements. However, current standardized collection strategies lead to severe feature mismatch issues: storytelling continuously collects redundant full-body posture data, while craft activities miss key indicators of eye-hand coordination. This scenario-feature mismatch directly reduces the effective data acquisition rate and significantly affects the scenario generalization ability of attention assessment models.
[0030] To address the aforementioned issues, a mechanism is needed that can sense the teaching scenario and adaptively adjust the data collection strategy.
[0031] Reference Figure 1 As shown, a motion-sensing interactive-driven method for dynamic attention feedback in preschool classrooms includes:
[0032] Acquire classroom sample data, perform scenario analysis on it, extract scenario feature data corresponding to each scenario and store it to build a classroom behavior scenario library;
[0033] Collect real-time classroom information from multiple sources and convert it into multimodal text data. After filtering and association to form joint matching data, perform threshold matching with scene feature data in the classroom behavior scene database.
[0034] If the match is successful, a collection control command is generated based on the scene feature data to drive the collection device to perform the collection operation to obtain the original somatosensory data with timestamp alignment. If the match fails, the information collection time is extended and the joint matching data is regenerated. After multiple consecutive matching failures, an abnormal prompt is pushed to the terminal device.
[0035] The information entropy values of various somatosensory feature data in the original somatosensory data are calculated, and dynamic weights are generated by combining scene feature data. Valid somatosensory data is then selected through a collaborative filtering mechanism of multi-dimensional scoring and dynamic weights.
[0036] Effective haptic data is input into a pre-built attention assessment model to generate attention assessment data, which is then pushed to the terminal device.
[0037] Specifically, the system first establishes a feature template library for different scenario types through classroom sample data analysis, including the timeline, spatial area, and key somatosensory indicators of the teaching process. During real-time teaching, the system continuously collects multi-source information and transforms it into structured data, which is then dynamically matched with the scenario library. When a match is successful, the system automatically calls a preset collection strategy to obtain the student's somatosensory data. If multiple matches fail, the system prompts for manual intervention. The collected raw somatosensory data undergoes information entropy analysis to quantify feature uncertainty, and dynamic weights are assigned based on scenario features. Noisy data is filtered out through a multi-dimensional scoring mechanism. Finally, the filtered effective data is input into the attention assessment model, which outputs the real-time attention status and pushes it to the teacher's terminal.
[0038] In one optional implementation, classroom sample data is acquired and subjected to scenario analysis. Scenario feature data corresponding to each scenario is extracted and stored to construct a classroom behavior scenario library, specifically including:
[0039] The classroom sample data includes teaching activity record data, student sensory data, and corresponding attention assessment data; wherein, classroom sample data refers to a historical dataset containing teaching activity records, student sensory data, and corresponding attention assessment results, which can be obtained through classroom videos, sensor logs, and teacher scoring records, and is used to extract the scene types of different teaching segments in teaching activities;
[0040] Based on the teaching activity record data, the teaching steps and corresponding timelines are extracted, and scene type labels and the core area of the scene where students are active are set for each teaching step. The core area of the scene refers to the set of coordinates of the main activity range of students in different teaching steps. Specifically, it can be achieved by dividing the classroom floor plan into areas and collecting coordinates from positioning devices, which is used to limit the spatial correlation range of the somatosensory data.
[0041] By linking teaching sessions with student sensory data using timestamps and spatial coordinates, the correlation between various sensory features and attention assessment data in the student sensory data is analyzed. The correlation is calculated as follows: first, the Pearson correlation coefficient and maximum information coefficient between each sensory feature and attention score are calculated, and sensory features with both indicators exceeding 0.6 are retained. The statistical correlation, feature importance, and scene fit are weighted at a ratio of 6:3:1 to obtain the normalized value of the correlation. The statistical correlation is the absolute value of the Pearson coefficient, the feature importance is the importance score of random forest analysis, and the scene fit is the historical significance of the sensory feature in the classroom sample data, which is automatically adjusted according to the teaching session, such as increasing the weight of head rotation angle in face-to-face teaching sessions.
[0042] For somatosensory feature data with a correlation greater than a preset correlation threshold, secondary verification using the Pearson correlation coefficient is performed to retain somatosensory feature data that is significantly correlated with attention assessment data. Priority and basic weights are then assigned according to the magnitude and proportion of correlation. Secondary verification using the Pearson correlation coefficient refers to conducting a statistical correlation test on the initially screened somatosensory feature data. Specifically, this can be achieved by calculating the ratio of the covariance to the standard deviation of the two sets of data, which is used to eliminate the interference of random correlations.
[0043] For example, a histogram of the correlation degree of all somatosensory feature data under this teaching segment is plotted, and the value of the significant jump point in the distribution curve is selected as the correlation degree threshold.
[0044] Calculate the frequency of change of somatosensory feature type data under different attention assessment data, and determine the threshold of feature change frequency;
[0045] The teaching segment data is associated with the corresponding somatosensory feature type data with a correlation degree greater than the preset correlation degree threshold and stored as scene feature data. The scene type label of the teaching segment is used as the index label to form a classroom behavior scene library.
[0046] Specifically, the timeline of teaching segments is divided based on event nodes in the teaching activity records. For example, teacher-led instruction, group discussions, and experimental operations are marked as independent segments. Each teaching segment is assigned a scenario type label. For example, the core area of a teacher-led instruction scenario is the student's seat, and the core area of an experimental operation scenario is the experimental table. After aligning the somatosensory feature data with the teaching segments via timestamps, its correlation with attention assessment data is analyzed. For example, the correlation between posture tilt angle, blink frequency, and head rotation angle and attention scores is verified using the Pearson correlation coefficient. Features with high correlation are given higher priority; for example, features with a correlation of 0.8 have higher priority than features with a correlation of 0.6. Feature change frequency thresholds are set based on classroom sample data analysis. For example, if the posture tilt angle changes at a frequency ≥3 times / minute when the attention assessment score is <6, the frequency threshold is set to 3; if the blink frequency changes at a frequency ≥5 times / minute when the attention score is <6, the frequency threshold is set to 5. The resulting classroom behavior scenario library uses scenario type labels as an index to store the core area coordinates, feature priorities, and change frequency thresholds for each segment.
[0047] In one optional implementation, real-time classroom multi-source information is collected and converted into multimodal text data. After filtering and association to form joint matching data, threshold matching is performed with scene feature data in a classroom behavior scene database. Specifically, this includes:
[0048] Multiple text data points are extracted from multi-source information in the classroom, and after being filtered, they are linked together to form joint matching data.
[0049] The joint matching data is matched with the index tags in the classroom behavior scenario library for keywords. When the matching degree exceeds the preset matching threshold, the scenario feature data corresponding to the index tag is output.
[0050] To facilitate understanding of the above embodiments, the following description will take a specific application scenario of the above embodiments as an example, including S31-S36:
[0051] S31. Collect teacher's voice data in real time through the speech recognition module, perform frame-by-frame processing on the voice data, extract semantic feature vector data of each frame of voice, and match the semantic feature vector data with a preset scene conversion keyword library. The scene conversion keyword library may contain core instruction words such as start, end, discussion, experiment.
[0052] S32. Calculate the matching score, filter out the speech segments with a matching score ≥ 6 and convert them into the first text data;
[0053] S33. The image recognition module scans the blackboard content line by line, extracts the character data of the text area, and uses the TextRank algorithm to calculate the weight value of each word in the character data. The weight of the topic word is ≥0.7 and the weight of the auxiliary word is <0.6. Words with a weight value ≥0.7 are selected to form the second text data.
[0054] S34. Calculate the temporal correlation score and semantic correlation score between the first text data and the second text data:
[0055] Time correlation score: 1 point is awarded when the time difference between the two data collections is ≤3 minutes; otherwise, the score decreases linearly to 0 points as the time difference increases.
[0056] Semantic similarity score: The semantic similarity between the two is calculated using the BERT model, ranging from 0 to 1.
[0057] S35. When the sum of the temporal correlation score and the semantic correlation score is ≥1.2, the two are merged into joint matching data; otherwise, if the voice data matching score is ≥8, the first text data is directly used as joint matching data.
[0058] S36. Perform keyword matching between the joint matching data and the index labels of the scene feature data in the classroom behavior scene library, and calculate the matching degree (0%~100%). When the matching degree exceeds the dynamic threshold, output the scene feature data corresponding to the index label. If the matching degree does not reach the threshold, discard the current joint matching data and wait for the next round of collection. The initial value of the dynamic matching threshold can be taken in the range of 60%~75%. After every 100 matchings, obtain the median and standard deviation of the matching degree of successful matching for each scene, and use the median minus 0.5 times the standard deviation to cover the current matching threshold.
[0059] Specifically, the voice module collects voice data, which is usually the essential first text data. Other data, such as blackboard writing and text and graphics displayed by the teacher, can be used as second text data. This optimizes trigger sensitivity and improves trigger accuracy. Through the above technical solution, this application effectively solves the problem of scene recognition deviation caused by the heterogeneity of multi-source data. It achieves collaborative verification of cross-modal data through a joint matching mechanism, avoiding misjudgment caused by the abnormality of a single data source.
[0060] In one optional implementation, acquisition control commands are generated based on scene feature data to drive the acquisition device to perform acquisition operations and obtain raw somatosensory data, specifically including:
[0061] Extract teaching segment data and somatosensory feature type data from scene feature data, and extract the corresponding scene core area and feature change frequency threshold;
[0062] Based on the feature change frequency threshold and the core area of the scene, and combined with the priority of the somatosensory feature data, the acquisition object, acquisition frequency and acquisition area are generated. The acquisition frequency is positively correlated with the feature change frequency threshold data.
[0063] The data collection object, collection frequency, and collection area are integrated into collection control data and sent to the collection device. The device is then controlled to collect the students' raw sensory data according to the collection control data and timestamped.
[0064] Compared with existing technologies, traditional fixed acquisition strategies use uniform parameters to cover all teaching scenarios, such as always acquiring data at a rate of 1 frame per second, resulting in data redundancy in non-key areas. In contrast, this solution uses a dynamic parameter adjustment mechanism to automatically optimize the acquisition strategy according to the characteristics of the actual teaching process. For example, it can acquire data on sitting posture tilt angle and blink frequency during teacher face-to-face instruction, and focus on acquiring data on students' line of sight during blackboard explanation, thereby improving the targeting and effectiveness of somatosensory data acquisition.
[0065] Through the above technical solution, this application dynamically matches the characteristics of the teaching process with the equipment acquisition parameters, ensuring the complete acquisition of key somatosensory data while reducing the amount of redundant somatosensory data. This effectively solves the scene adaptability problem caused by fixed acquisition strategies. For example, it automatically reduces the acquisition frequency to save storage resources during the phase of gradual feature changes, and increases the acquisition frequency to improve the completeness of somatosensory data acquisition when the feature change frequency threshold is high, thus providing an accurate data foundation for subsequent attention assessment.
[0066] To facilitate understanding of the above embodiments, the following description will take a specific application scenario of the above embodiments as an example, including S41-S44:
[0067] S41. Retrieve the scene feature data corresponding to teacher face-to-face instruction from the classroom behavior scene database:
[0068] Teaching segment: The timeline is 0-20 minutes, currently at minute 8. The core area of the scene is the student seating area, with coordinates ranging from (X: 4.0m to 4.7m, Y: 5.0m to 5.3m, Z: 1m to 1.5m).
[0069] Somatosensory feature data: sitting tilt angle (correlation 0.8, priority 1, base weight 0.8 / 1.5), blink frequency (correlation 0.7, priority 2, base weight 0.7 / 1.5).
[0070] Characteristic change frequency thresholds: sitting tilt angle ≥ 3 times / minute, blinking frequency ≥ 5 times / minute;
[0071] S42. Determine the acquisition parameters based on the priority, feature change frequency threshold, and core area of the scene feature data:
[0072] Data collection targets are prioritized, with the sitting tilt angle collected first, followed by blink frequency. For example, each 10-second collection cycle is divided into different time periods. During the 0-6 second period, only the sitting tilt angle data is collected. During the 6-10 second period, if the sitting tilt angle data collection is completed, the blink frequency is collected. If it is not completed, the sitting tilt angle data collection continues, and the remaining time is used to collect the blink frequency. The resolution is also adjusted adaptively according to the data collection targets.
[0073] Acquisition frequency: positively correlated with the feature change frequency threshold (the higher the threshold, the higher the acquisition frequency). The acquisition frequency for sitting posture tilt angle is 3Hz (3 times per second), and the acquisition frequency for blink frequency is 1Hz (1 time per second).
[0074] Data collection area: limited to the coordinate range of the core area of the scene, excluding the podium area and corridor;
[0075] S43. Integrate the above parameters into a data acquisition control command and send it to the data acquisition device deployed in the classroom via a wireless communication module;
[0076] S44. The data acquisition device executes the operation according to the control command to obtain raw somatosensory data. All raw somatosensory data are associated with the timestamp (8th minute) of the acquisition control command and stored for subsequent data processing.
[0077] In one optional implementation, the information entropy values of various somatosensory feature data in the original somatosensory data are calculated, and dynamic weights are generated by combining them with scene feature data, specifically including:
[0078] Impulse noise removal and Z-score standardization are performed on the original somatosensory data, and features are extracted to obtain multiple types of somatosensory feature data related to attention evaluation. The information entropy value of each type of somatosensory feature data is calculated using the information entropy algorithm.
[0079] Calculate the standardized covariance matrix among various types of somatosensory feature data, and select the three feature pairs with the largest absolute covariance values;
[0080] Based on the classroom behavior scenario library, obtain the basic weight, historical maximum entropy value and scenario adaptation coefficient corresponding to the current scenario type;
[0081] Construct a dynamic weight model and calculate the dynamic weights. The dynamic weight model is as follows:
[0082] ;
[0083] In the formula, Somatosensory characteristics Dynamic weights, For scene adaptability coefficient, It is a non-linear decay exponent. For feature interaction coefficients, Somatosensory characteristics The basic weights, Somatosensory characteristics The current information entropy value, Somatosensory characteristics The historical maximum entropy value, Somatosensory characteristics and covariance, respectively tactile characteristics and standard deviation The three feature pairs with the largest absolute covariance;
[0084] For example, The value range is [0,1]. The value range is [0.1, 2]. The value range is set to [0, 0.5]. The optimal value for each scenario type is determined through classroom sample data validation. Since preschool children are more sensitive to scenario changes, a larger scenario adaptation coefficient is chosen, such as... =0.8, =0.5, =0.2;
[0085] Compared with existing technologies, traditional methods usually adopt a fixed weight allocation strategy, such as uniformly assigning a static weight of 0.3-0.5 to all somatosensory feature data. This cannot adapt to the dynamic changes in the importance of features under different teaching scenarios. However, this solution introduces a scenario adaptation coefficient and a historical entropy decay mechanism, which can automatically adjust the weight calculation logic according to the real-time teaching process. For example, it can enhance the weight allocation of head turning angle in the blackboard explanation process.
[0086] To facilitate understanding of the above embodiments, the following description will take a specific application scenario of the above embodiments as an example, including S51-S54:
[0087] S51. Extract the somatosensory feature data of preschool children and calculate the information entropy:
[0088] The somatosensory feature data extracted from the raw somatosensory data includes, but is not limited to, the following three types, taking the extraction of a continuous 10-second effective segment as an example:
[0089] Seated tilt angle: [8,10,9,11,8,10,9,12,10,8], in degrees;
[0090] Blinking frequency: [12,11,13,12,14,13,11,12,13,14], in times per minute;
[0091] Head turning angle: [15,16,14,17,15,16,14,17,16,15], in degrees;
[0092] The Shannon entropy formula is: ,in, The probability of an eigenvalue occurring;
[0093] The eigenvalues and probabilities of the sitting posture tilt angle are: 8 (3 / 10), 9 (2 / 10), 10 (3 / 10), 11 (1 / 10), 12 (1 / 10). Substituting these values into the Shannon entropy formula, we get an information entropy of 2.17. Similarly, we can get an information entropy of 1.97 for the blink frequency.
[0094] S52. Calculate the covariance and standard deviation among various types of somatosensory characteristic data:
[0095] Seated tilt angle With head turning angle Covariance: ;
[0096] Seated tilt angle With blinking frequency Covariance: ;
[0097] blink frequency With head turning angle Covariance: ;
[0098] Since only three somatosensory features are listed, the three feature pairs with the largest absolute values are... For the above three;
[0099] The standard deviation of the sitting tilt angle was 1.24, the standard deviation of the blink frequency was 1.10, and the standard deviation of the forward lean was 1.05.
[0100] S53. Obtain the basic weights, historical maximum entropy, and scene adaptation coefficients corresponding to the current scene type:
[0101] Using the teacher face-to-face teaching scenario mentioned earlier, the corresponding parameters are retrieved from the classroom behavior scenario database: sitting posture tilt angle (relevance 0.8, priority 1, basic weight 0.8 / 1.5) and blinking frequency (relevance 0.7, priority 2, basic weight 0.7 / 1.5).
[0102] Calculate the information entropy of all sitting tilt angles and blink frequencies under the teacher face-to-face scene type in the scene library according to the method in S51, and select the maximum value as the historical maximum entropy value. The historical maximum entropy value of sitting tilt angle can be 2.5, and the historical maximum entropy value of blink frequency can be 2.2.
[0103] The scene adaptation coefficient is set to 0.8, the nonlinear decay exponent is set to 0.5, and the feature interaction coefficient is set to 0.2.
[0104] S54. Substitute the above relevant parameters into the dynamic weight model for calculation:
[0105] The dynamic weight of the seated tilt angle is:
[0106] , Zhongyu The relevant feature pairs are and ;
[0107] The dynamic weight of the sitting tilt angle was calculated to be 0.376, and the dynamic weight of the blink frequency was calculated to be 0.28. The dynamic weight of the head turning angle was not calculated here because its base weight is 0.
[0108] The resulting dynamic weights can be directly used for the subsequent screening of effective sensory data;
[0109] In summary, dynamic weighting assigns higher weight to auditory features that are more closely related to the current scene during evaluation, reducing interference from irrelevant features. Covariance features are used to assess the interaction between sensory features, improving the comprehensiveness of feature importance quantification. Dynamic weights directly participate in the retention probability calculation, ensuring that important sensory features may still be retained even when their stability is slightly low. For example, if a feature has a slightly low overall stability score but a high dynamic weight, its retention probability may still meet the standard, thus avoiding the accidental deletion of high-quality features.
[0110] In one optional implementation, effective sensory data is selected through a collaborative filtering mechanism combining multi-dimensional scoring and dynamic weights, specifically including:
[0111] Calculate the temporal correlation score of the somatosensory feature data, wherein the temporal correlation score is the degree of matching between the collection time of the somatosensory feature data and the timeline of the teaching process;
[0112] Calculate the spatial correlation score of the somatosensory feature data, wherein the spatial correlation score is the degree of matching between the spatial coordinates corresponding to the somatosensory feature data and the coordinates of the core area of the scene;
[0113] Calculate the feature stability score of the somatosensory feature data, wherein the feature stability score is the normalized score of the variance of the features in consecutive frames;
[0114] Among them, the temporal correlation score can be calculated using a timestamp comparison algorithm to ensure the relevance of the collected data to the current teaching stage; the spatial correlation score can be calculated using an inverse proportional function of coordinate distance to filter out data interference that deviates from the core area of the scene; and the feature stability score can be obtained by normalizing the variance calculated by a sliding window to eliminate the influence of noisy data on attention assessment.
[0115] Based on temporal correlation score, spatial correlation score, and feature stability score, a nonlinear aggregation method is used to calculate the spatiotemporal stability comprehensive score. The spatiotemporal stability comprehensive score is multiplied by the dynamic weight to obtain the retention probability.
[0116] Somatosensory feature data with a retention probability greater than a preset retention threshold are selected as valid somatosensory data.
[0117] For example, the preset retention threshold is set to 0.3. The rule for setting the value is to collect more than 500 sets of retention data. The retention data includes the original somatosensory features and the attention assessment data (teacher scoring records) corresponding to the retention probability. The accuracy of attention assessment corresponding to different retention probability intervals, such as 0-0.2, 0.2-0.3, 0.3-0.4, and above 0.4, is calculated. The retention threshold is set according to the requirement of cross-validation accuracy ≥ 85%.
[0118] The calculation formula for the nonlinear aggregation method is as follows:
[0119] ;
[0120] In the formula, To preserve probability, These are time-related scores, spatial-related scores, and feature stability scores. It is a non-linear adjustment index. This represents the dynamic weights for the corresponding somatosensory features; the nonlinear aggregation method amplifies the impact of low-scoring dimensions. For example, when the score of a certain dimension is significantly lower than other dimensions, the overall score will decrease exponentially. =1.5, adjusted when min(T,S,F)<0.4. =2.0;
[0121] Compared with existing technologies, traditional methods for filtering somatosensory feature data usually employ fixed thresholds or single-dimensional filtering mechanisms, such as filtering based solely on timestamps or spatial coordinates. This can easily lead to mis-screening or omissions in complex teaching scenarios. In contrast, this application achieves scene-type-aware adaptive somatosensory feature data filtering, effectively reducing data redundancy caused by rigid collection strategies and avoiding the omission of key somatosensory feature data. Furthermore, through a non-linear aggregation mechanism of multi-dimensional scoring, it can accurately identify effective data that meets the spatiotemporal characteristics and stability requirements of the current teaching segment, providing high-quality input for subsequent attention assessment and thereby improving the accuracy and real-time performance of the classroom attention feedback system.
[0122] In summary, the core function of retaining probabilities is to filter out high-quality, highly relevant, and effective data from the original sensory data. This reduces the amount of data input to the attention assessment model, lowers the computational cost of the attention assessment model, and improves the real-time performance of its output assessment results.
[0123] The necessity of dynamic weights and retention probabilities is explained in a unified way: Dynamic weights provide a benchmark for feature importance for retention probabilities, avoiding the retention of irrelevant features simply because of high stability or spatiotemporal matching. Retention probabilities, through spatiotemporal correlation and stability screening, provide data quality assurance for dynamic weights, ensuring that the data of high-weight features are reliable.
[0124] In one optional implementation, effective somatosensory data is input into a pre-built attention assessment model to generate attention assessment data and push it to the terminal device, specifically including:
[0125] Based on the quantified values of various somatosensory features in the effective somatosensory data, combined with the corresponding dynamic weights, the attention concentration index is calculated.
[0126] The pre-built attention assessment model refers to a machine learning model trained on classroom sample data. Specifically, the random forest algorithm can be used to fit student sensory data with corresponding attention assessment data. Cross-validation is used to ensure the model's generalization ability. In the sample data, attention assessment data refers to the quantitative value of the teacher's scoring record, which can be set to a range of 1 to 10 points. Cross-validation accuracy ≥85% means that the average prediction accuracy calculated by the model through the k-fold cross-validation method during training must reach a preset threshold. For example, when using five-fold cross-validation, the accuracy should not be less than 85%. Attention distraction risk label refers to the discrete risk level based on the attention concentration index, which can be divided into three categories: low, medium, and high risk.
[0127] An attention distraction risk label is generated based on the attention concentration index, and the attention concentration index and the risk label are associated as attention assessment data and pushed to the terminal device.
[0128] Through the above technical solution, this application solves the problem of evaluation bias in fixed models in changing teaching scenarios. By combining dynamic weights with machine learning models, it achieves accurate mapping between somatosensory features and attentional states, effectively reducing the data misjudgment rate caused by insufficient scenario adaptation, and ensuring that the attention assessment data obtained by teachers includes both quantitative indices and intuitive risk labels, providing a reliable basis for adjusting teaching strategies.
[0129] See Figure 3 As shown, this solution proposes a motion-sensing interactive-driven dynamic attention feedback system for preschool classrooms, used to implement the aforementioned motion-sensing interactive-driven dynamic attention feedback method for preschool classrooms, including:
[0130] The data acquisition module is used to acquire classroom sample data, perform scenario analysis on it, extract scenario feature data corresponding to each scenario, and store it to build a classroom behavior scenario library;
[0131] The scene matching module is used to collect real-time classroom multi-source information and convert it into multimodal text data. After filtering and association to form joint matching data, it performs threshold matching with scene feature data in the classroom behavior scene library. If the matching is successful, it generates collection control instructions based on scene feature data to drive the collection device to perform collection operations to obtain the original somatosensory data aligned with the timestamp. If the matching fails, it extends the information collection time and regenerates the joint matching data. After multiple consecutive matching failures, it pushes an abnormal prompt to the terminal device.
[0132] The somatosensory data processing module is used to calculate the information entropy values of various somatosensory feature data in the raw somatosensory data, and generate dynamic weights by combining them with scene feature data;
[0133] The attention assessment and push module is used to input effective somatosensory data into a pre-built attention assessment model, generate attention assessment data, and push it to the terminal device.
[0134] In another embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above embodiments.
[0135] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps described above.
[0136] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the steps described above.
[0137] The technical scope of this invention is not limited to the content described above. Those skilled in the art can make various modifications and variations to the above embodiments without departing from the technical concept of this invention, and all such modifications and variations should fall within the protection scope of this invention.
Claims
1. A dynamic feedback method for preschool classroom attention driven by motion-sensing interaction, characterized in that, The method includes: Acquire classroom sample data, perform scenario analysis on it, extract scenario feature data corresponding to each scenario and store it to build a classroom behavior scenario library; Collect real-time classroom information from multiple sources and convert it into multimodal text data. After filtering and association to form joint matching data, perform threshold matching with scene feature data in the classroom behavior scene database. If the match is successful, a collection control command is generated based on the scene feature data to drive the collection device to perform the collection operation to obtain the original somatosensory data with timestamp alignment. If the match fails, the information collection time is extended and the joint matching data is regenerated. After multiple consecutive matching failures, an abnormal prompt is pushed to the terminal device. The information entropy values of various somatosensory feature data in the original somatosensory data are calculated, and dynamic weights are generated by combining them with scene feature data. Valid somatosensory data is then selected through a collaborative filtering mechanism of multi-dimensional scoring and dynamic weights. Among these: Impulse noise removal and Z-score standardization are performed on the original somatosensory data, and features are extracted to obtain multiple types of somatosensory feature data related to attention evaluation. The information entropy value of each type of somatosensory feature data is calculated using the information entropy algorithm. Calculate the standardized covariance matrix among various types of somatosensory feature data, and select the three feature pairs with the largest absolute covariance values; Based on the classroom behavior scenario library, obtain the basic weight, historical maximum entropy value and scenario adaptation coefficient corresponding to the current scenario type; Construct a dynamic weight model and calculate the dynamic weights. The dynamic weight model is as follows: ; In the formula, Somatosensory characteristics Dynamic weights, For scene adaptability coefficient, It is a non-linear decay exponent. For feature interaction coefficients, Somatosensory characteristics The basic weights, Somatosensory characteristics The current information entropy value, Somatosensory characteristics The historical maximum entropy value, Somatosensory characteristics and covariance, respectively tactile characteristics and standard deviation The three feature pairs with the largest absolute covariance; Calculate the temporal correlation score of the somatosensory feature data, wherein the temporal correlation score is the degree of matching between the collection time of the somatosensory feature data and the timeline of the teaching process; Calculate the spatial correlation score of the somatosensory feature data, wherein the spatial correlation score is the degree of matching between the spatial coordinates corresponding to the somatosensory feature data and the coordinates of the core area of the scene; Calculate the feature stability score of the somatosensory feature data, wherein the feature stability score is the normalized score of the variance of the features in consecutive frames; Based on temporal correlation score, spatial correlation score, and feature stability score, a nonlinear aggregation method is used to calculate the spatiotemporal stability comprehensive score. The spatiotemporal stability comprehensive score is multiplied by the dynamic weight to obtain the retention probability. Somatosensory feature data with a retention probability greater than a preset retention threshold are selected as valid somatosensory data. The calculation formula for the nonlinear aggregation method is as follows: ; In the formula, To preserve probability, These are time-related scores, spatial-related scores, and feature stability scores. It is a non-linear adjustment index. For the corresponding dynamic weights of the somatosensory features; Effective haptic data is input into a pre-built attention assessment model to generate attention assessment data, which is then pushed to the terminal device.
2. The method according to claim 1, characterized in that, The process of acquiring classroom sample data, performing scenario analysis on it, extracting scenario feature data corresponding to each scenario, and storing it to construct a classroom behavior scenario library specifically includes: The classroom sample data includes teaching activity record data, student sensory data, and corresponding attention assessment data; Based on the teaching activity record data, the teaching steps and corresponding timelines are extracted, and scene type labels and core scene areas where student activities are located are set for each teaching step; By linking teaching sessions with students' somatosensory data using timestamps and spatial coordinates, the correlation between various somatosensory characteristics and attention assessment data in students' somatosensory data is analyzed. For somatosensory feature data with a correlation greater than a preset correlation threshold, the Pearson correlation coefficient is used for secondary verification. Somatosensory feature data that are significantly related to attention assessment data are retained, and priority and basic weight are assigned according to the correlation magnitude and proportion. Calculate the frequency of change of somatosensory feature type data under different attention assessment data, and determine the threshold of feature change frequency; The teaching process data is associated with and stored as corresponding haptic feature type data with a correlation degree greater than a preset correlation degree threshold. This is recorded as scene feature data, and the scene type label of the teaching process is used as an index label to form a classroom behavior scene library.
3. The method according to claim 2, characterized in that, The process of collecting real-time classroom multi-source information and converting it into multimodal text data, filtering and associating it to form joint matching data, and then performing threshold matching with scene feature data in the classroom behavior scene database specifically includes: Multiple text data points are extracted from multi-source information in the classroom, and after being filtered, they are linked together to form joint matching data. The joint matching data is matched with the index tags in the classroom behavior scenario library for keyword matching. When the matching degree exceeds the preset matching threshold, the scenario feature data corresponding to the index tag is output.
4. The method according to claim 1, characterized in that, The step of generating acquisition control commands based on scene feature data to drive the acquisition device to perform acquisition operations and obtain raw somatosensory data specifically includes: Extract teaching segment data and somatosensory feature type data from scene feature data, and extract the corresponding scene core area and feature change frequency threshold; Based on the feature change frequency threshold and the core area of the scene, and combined with the priority of the somatosensory feature data, the acquisition object, acquisition frequency and acquisition area are generated. The acquisition frequency is positively correlated with the feature change frequency threshold data. The data collection object, collection frequency, and collection area are integrated into collection control data and sent to the collection device. The device is then controlled to collect the students' raw sensory data according to the collection control data and timestamped.
5. The method according to claim 1, characterized in that, The step of inputting effective somatosensory data into a pre-built attention assessment model, generating attention assessment data, and pushing it to the terminal device specifically includes: Based on the quantified values of various somatosensory features in the effective somatosensory data, combined with the corresponding dynamic weights, the attention concentration index is calculated. Attention distraction risk labels are generated based on the attention concentration index, and the attention concentration index and risk labels are associated as attention assessment data and pushed to the terminal device.
6. A motion-sensing interactive-driven dynamic attention feedback system for preschool classrooms, characterized in that: A method for implementing the motion-sensing interactive-driven dynamic attention feedback method in preschool classrooms as described in any one of claims 1-5 includes: The data acquisition module is used to acquire classroom sample data, perform scenario analysis on it, extract scenario feature data corresponding to each scenario, and store it to build a classroom behavior scenario library; The scene matching module is used to collect real-time classroom multi-source information and convert it into multimodal text data. After filtering and association to form joint matching data, it performs threshold matching with scene feature data in the classroom behavior scene library. If the matching is successful, it generates collection control instructions based on scene feature data to drive the collection device to perform collection operations to obtain the original somatosensory data aligned with the timestamp. If the matching fails, it extends the information collection time and regenerates the joint matching data. After multiple consecutive matching failures, it pushes an abnormal prompt to the terminal device. The somatosensory data processing module is used to calculate the information entropy values of various somatosensory feature data in the raw somatosensory data, and generate dynamic weights by combining them with scene feature data; The attention assessment and push module is used to input effective somatosensory data into a pre-built attention assessment model, generate attention assessment data, and push it to the terminal device.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method described in any one of claims 1-5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-5.