Multi-modal evaluation method, system and equipment for class participation degree of autistic children in man-machine collaborative teaching
By constructing an expert rule base and a deep learning model, the classroom participation of children with autism can be assessed in real time, which solves the problems of low accuracy and non-targeted intervention strategies in existing technologies, and achieves efficient classroom participation assessment and visualized intervention.
Patent Information
- Application Number
- CN202511072246.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-18
AI Technical Summary
Existing human-computer collaborative teaching methods for assessing classroom participation in children with autism suffer from problems such as low accuracy, delay, lack of intuitiveness, and poor targeting of intervention strategies.
We construct an expert rule base, acquire multimodal physiological and emotional data, evaluate classroom participation in real time through a deep learning model, and visualize the results using an educational intervention strategy library.
It improved the accuracy and timeliness of identifying classroom participation in children with autism, visualized classroom participation and intervention strategies, and enhanced the pertinence of intervention strategies.
Smart Images

Figure CN120977543A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of educational artificial intelligence technology, and in particular to a method, system and device for multimodal assessment of classroom participation of children with autism in human-computer collaborative teaching. Background Technology
[0002] In the wave of digitalization, the rapid development of new-generation information technologies such as the internet, big data, and artificial intelligence is driving transformation across various fields. In education, traditional teaching models struggle to meet diverse needs, while the powerful data analysis and personalized support capabilities of artificial intelligence are emerging, giving rise to human-machine collaborative teaching and providing new opportunities for educational innovation. Simultaneously, the rapid development of classroom assessment methods for children, driven by modern information technologies such as artificial intelligence, big data, and the Internet of Things, enables real-time collection and in-depth analysis of multimodal classroom data. This lays the foundation for building human-machine collaborative teaching classroom assessment systems for children, helping educational assessment move towards intelligence and precision. However, classroom assessment for children with autism receives relatively little attention regarding the assessment of classroom participation.
[0003] The existing human-computer collaborative teaching methods for assessing classroom participation of children with autism have the following shortcomings, mainly including: low accuracy in identifying the participation of children with autism, delays and asynchrony, unintuitive presentation of classroom participation, and poor targeting of intervention strategies. Summary of the Invention
[0004] The purpose of this application is to provide a multimodal assessment method, system, and device for classroom participation of children with autism in human-computer collaborative teaching, which can improve the accuracy, objectivity, and timeliness of participation identification of children with autism, realize the visualization of classroom participation and intervention strategies, and improve the pertinence of intervention strategies.
[0005] To achieve the above objectives, this application provides the following solution.
[0006] Firstly, this application provides a multimodal assessment method for classroom participation of autistic children in human-computer collaborative teaching. This method includes: constructing an expert rule base; the expert rule base includes: a privacy protection rule base, a special behavioral feature knowledge graph, a teaching scenario classification rule base, a multimodal annotation database, an expert intervention strategy rule base, and an educational intervention strategy base; acquiring video streams and multimodal physiological and emotional data from the children's classroom; the multimodal physiological and emotional data includes: depth image data, high-speed video data, multi-channel speech data, and biosignal data synchronously collected during the children's classroom; performing feature analysis on the multimodal physiological and emotional data to obtain multimodal features; the multimodal features include: visual... The system incorporates multimodal features, physiological features, vocal features, emotional state features, stereotyped behavior features, social behavior features, and directed attention features. Based on video streams, it uses a teaching scenario classification rule base to identify the current teaching scenario, and then uses an expert intervention strategy rule base to assign weights to each modal feature in the current teaching scenario. Based on the multimodal features and their weights, a trained deep learning model is used to evaluate the classroom participation of children with autism in real time. The training data for the trained deep learning model comes from a multimodal annotation database. Based on the multimodal features and classroom participation, intervention strategies are retrieved using an educational intervention strategy library, and the classroom participation and intervention strategies are visualized in real time.
[0007] Secondly, a multimodal assessment system for classroom participation of autistic children in human-computer collaborative teaching is provided. This system includes: a multimodal perception module, a behavioral feature analysis module, a teaching scenario adaptation module, a participation calculation module, an expert knowledge base module, and a human-computer interaction module. The expert knowledge base module is used to construct an expert rule base, which includes: a privacy protection rule base, a special behavioral feature knowledge graph, a teaching scenario classification rule base, a multimodal annotation database, an expert intervention strategy rule base, and an educational intervention strategy base. The multimodal perception module is used to acquire video streams and multimodal physiological and emotional data from the children's classroom. The multimodal physiological and emotional data includes: depth image data, high-speed video data, multi-channel voice data, and biosignal data synchronously collected during the children's classroom. The behavioral feature analysis module is used to perform feature analysis on the multimodal physiological and emotional data. The analysis yields multimodal features, including visual features, physiological features, vocal features, emotional state features, stereotyped behavior features, social behavior features, and directed attention features. The teaching scenario adaptation module identifies the current teaching scenario based on the video stream using a teaching scenario classification rule base, and assigns weights to each modal feature within the current teaching scenario using an expert intervention strategy rule base. The participation calculation module evaluates the classroom participation of autistic children in real time using a trained deep learning model based on the multimodal features and their weights. The training data for the trained deep learning model comes from a multimodal annotation database. The human-computer interaction module retrieves intervention strategies from an educational intervention strategy library based on the multimodal features and classroom participation, and visualizes classroom participation and intervention strategies in real time.
[0008] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described multimodal assessment method for classroom participation of autistic children in human-computer collaborative teaching.
[0009] According to the specific embodiments provided in this application, the following technical effects are disclosed:
[0010] This application constructs an expert rule base; acquires video streams and multimodal physiological and emotional data from children's classrooms, achieving spatiotemporal consistency in the collection of multimodal data. Through cross-modal feature association, it can reduce the assessment misjudgment rate and effectively solve the problem of subjective experience bias. Feature analysis is performed on the multimodal physiological and emotional data to obtain multimodal features; the current teaching scenario is identified, and the weights of each modal feature in the current teaching scenario are assigned using the expert intervention strategy rule base; based on the multimodal features and their weights, a trained deep learning model is used to evaluate the classroom participation of autistic children in real time, effectively improving the timeliness of classroom participation calculation and reducing calculation errors; based on the multimodal features and classroom participation, intervention strategies are retrieved using an educational intervention strategy library, and the classroom participation and intervention strategies are visualized in real time. Therefore, this application improves the accuracy, objectivity, and timeliness of autistic children's participation identification, realizes the visualization of classroom participation and intervention strategies, and improves the targeting of intervention strategies. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating a multimodal assessment method for classroom participation of children with autism in human-computer collaborative teaching, provided as an embodiment of this application.
[0013] Figure 2 This is a schematic diagram of the structural connection of a multimodal assessment system for classroom participation of children with autism in human-computer collaborative teaching, provided as an embodiment of this application.
[0014] Figure 3 This is a schematic diagram of multimodal physiological and emotional data acquisition provided in an embodiment of this application.
[0015] Figure 4 An intervention strategy grid diagram provided for embodiments of this application.
[0016] Figure 5 A heat map provided for an embodiment of this application.
[0017] Figure 6 The attention curve provided for the embodiments of this application.
[0018] Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application.
[0019] Figure labeling: Multimodal perception module-1; Behavioral feature analysis module-2; Teaching scenario adaptation module-3; Participation calculation module-4; Expert knowledge base module-5 and Human-computer interaction module-6. Detailed Implementation
[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] The technical shortcomings of traditional evaluation methods can be categorized into the following aspects.
[0022] 1) High subjectivity and inefficiency: Although current mainstream assessment systems achieve semi-structured digital recording, they still heavily rely on human behavioral observation and scale scoring, resulting in assessment results limited by subjective experience differences. A single assessment takes a long time on average and requires multiple samplings across different scenarios, making it difficult to meet the real-time feedback needs of classroom teaching.
[0023] 2) General classroom systems lack the ability to perceive special behavioral characteristics: General classroom behavior recognition systems are mainly developed for typical children and have blind spots in the recognition of autism-specific behaviors such as stereotyped movements, gaze tracking, and abnormal speech of vocal organs. This results in a limited ability to capture the behavioral characteristics of children with autism, leading to low accuracy in classroom participation evaluation.
[0024] 3) Structural deficiencies in the multimodal assessment system: Current assessment systems mainly use single-modal data acquisition of pure video and pure physiological signals, lacking a spatiotemporal alignment mechanism for multi-source information, resulting in a high misjudgment rate of the causal relationship between physiological arousal and behavioral performance.
[0025] 4) Insufficient performance of stereotyped behavior recognition algorithms: The training data of general action recognition models has a low proportion of special group samples and does not integrate developmental psychology temporal features, which leads to significant limitations in the detection of repetitive stereotyped behaviors such as swaying and patting, resulting in insufficient performance of stereotyped behavior recognition algorithms.
[0026] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0027] Example 1, as Figure 1 As shown in the figure, this embodiment provides a multimodal assessment method for classroom participation of children with autism in human-computer collaborative teaching. The multimodal assessment method for classroom participation of children with autism in human-computer collaborative teaching includes the following steps.
[0028] S1. Construct an expert rule base; the expert rule base includes: a privacy protection rule base, a knowledge graph of special behavioral characteristics, a teaching scenario classification rule base, a multimodal annotation database, an expert intervention strategy rule base, and an educational intervention strategy base.
[0029] In practical applications, the expert knowledge base serves as the core decision-making hub of the system. Through structured integration of multiple intelligences, Vygotsky's zone of proximal development and other special education theories, DSM-5 / ICD-11 clinical diagnostic criteria, and classroom teaching experience, it transforms educational psychology principles into quantifiable algorithm parameters, providing rule support for multimodal behavior analysis, dynamic scenario adaptation, participation calculation, and human-computer interaction modules. Specific knowledge categories and content are shown in Table 1.
[0030] Table 1 Knowledge Categories of Expert Rule Base
[0031]
[0032] S2. Acquire video streams and multimodal physiological and emotional data from children's classrooms; the multimodal physiological and emotional data includes: depth image data, high-speed video data, multi-channel voice data, and biosignal data collected synchronously in children's classrooms.
[0033] Furthermore, step S2 specifically includes the following steps.
[0034] S21. Obtain the video stream of the children's classroom.
[0035] S22. Acquire multimodal physiological and emotional data, and use a privacy protection rule base to perform real-time biometric privacy desensitization processing on the multimodal physiological and emotional data, and update the multimodal physiological and emotional data.
[0036] In practical applications, heterogeneous sensors using vision, biosignals, and acoustics can be used to collect real-time classroom behaviors, physiological responses, and speech characteristics of children with autism, constructing a multi-dimensional data foundation that is synchronized in time and space. Hardware-level synchronization is employed to accurately capture specific behaviors such as social gaze patterns, stereotyped action patterns, and abnormal speech fluctuations, providing integrated perceptual data for subsequent quantitative analysis and supporting dynamic adjustments to teaching strategies and human-machine collaborative intervention decisions.
[0037] S3. Perform feature analysis on multimodal physiological and emotional data to obtain multimodal features; multimodal features include: visual features, physiological features, voice features, emotional state features, stereotyped behavior features, social behavior features, and directed attention features.
[0038] Furthermore, step S3 specifically includes the following steps.
[0039] S31. Based on high-speed video data, the YOLOv5 algorithm is used for target detection and action recognition, and a convolutional neural network is used for facial expression analysis to obtain visual features.
[0040] In practical applications, for a given observation time segment, the YOLOv5 algorithm is used for target detection and action recognition. Based on the captured movement trajectories of the head, hands, eyes, and other body parts of autistic children, key posture and action features are extracted. Furthermore, a convolutional neural network (CNN) is used to analyze micro-expressions, eye movements, and body postures in the visual data to extract behavioral features related to emotions and social behaviors, as shown in the following formula.
[0041] s1 = CNN YOLO (I video ).
[0042] In the formula, I video It is image data from depth camera arrays and high-speed camera units, CNN YOLO This indicates the use of a deep learning model combining CNN and YOLOv5, where s1 represents the extracted visual behavioral features.
[0043] S32. Based on biological signal data, long short-term memory networks are used to model time-series physiological signal data to obtain physiological characteristics.
[0044] In practical applications, for a given observation time segment, based on physiological signal data such as skin conductance, heart rate variability (HRV), and body movement collected by wearable flexible sensors, a Long Short-Term Memory (LSTM) network is used to model the time-series physiological signal data. This captures joint information from the original acoustic signals to obtain behavioral characteristics related to the physiological state. The formula is as follows.
[0045] s2 = LSTM(HRV, SCR, ΔHR).
[0046] HRV, SCR, and ΔHR represent heart rate variability, skin conductance response, and heart rate change per unit time, respectively, and are collected by a flexible wearable device. LSTM is a long short-term memory network, and s2 is a physiological state-related behavioral feature.
[0047] S33. Based on multi-channel speech data, the Wav2Vec 2.0 speech emotion analysis algorithm is used to extract speech features.
[0048] In practical applications, for a given observation time segment, a multi-channel microphone array is used, along with adaptive filters and sound source localization technology, to process the speech signal and extract features such as the child's vocalization, intonation changes, and speech fluency. Combined with the Wav2Vec 2.0 speech emotion analysis algorithm, speech abnormalities in children with autism are identified, including vocal impairments and unclear pronunciation, as shown in the following formula.
[0049] s3=Wav2Vec(I audio ).
[0050] Among them, I audio s3 is the speech signal captured from the microphone array, Wav2Vec is a deep learning model for speech emotion recognition, and s3 is the extracted speech behavior features.
[0051] S34. Based on depth image data, the VGG-Face facial expression recognition network is used to identify emotional state features.
[0052] In practical applications, for a given observation time segment, facial image data acquired by a ring-shaped array of depth cameras within a multimodal data acquisition unit is used. The VGG-Face expression recognition network is then employed to infer the emotional states of autistic children, such as anxiety, pleasure, and silence, in real time, and outputs emotional state feature information. The formula is as follows.
[0053] s4 = VGG Face (I face ).
[0054] Among them, I face It is a facial image captured by a high-speed camera unit, VGG Face It is a facial expression recognition model, and s4 is the emotional state feature information output by the model.
[0055] S35. Based on high-speed video data, stereotypical behavioral features are identified using deep convolutional neural networks and temporal convolutional networks.
[0056] In practical applications, for a given observation time segment, autistic children often exhibit stereotyped behaviors such as rocking, patting, and repetitive movements, which are typically difficult to capture accurately in a regular classroom setting. The behavior feature analysis module utilizes multimodal data, combining deep convolutional neural networks (CNNs) and temporal convolutional networks (TCNs) to identify and classify action patterns, as shown in the following formula.
[0057] s5 = CNN TCN (I video ,ΔSCR).
[0058] Among them, I videoFor video data, ΔSCR represents the temporal change of the skin conductance response, and CNN... TCN It is a deep network that combines image convolution and temporal convolution, where s5 represents the stereotypical behavioral features identified.
[0059] S36. Based on high-speed video data, OpenPose is used to obtain the frequency of eye contact between autistic children and other classroom participants per unit time, and the proportion of dialogue overlap between autistic children and other classroom participants is calculated based on multi-channel voice data to obtain social behavior characteristics.
[0060] In practical applications, for a given observation time segment, based on video data acquired by a depth camera array, OpenPose is used to extract the posture, gaze direction, and facial orientation of autistic children and classroom participants. The frequency of eye contact between autistic children and other classroom participants is calculated per unit time using the eye contact intersection threshold. Combined with audio data acquired by a multi-channel microphone array, the proportion of dialogue overlap time between autistic children and other classroom participants is calculated to obtain social behavior characteristics.
[0061] Furthermore, the formula for calculating social behavioral characteristics is as follows.
[0062]
[0063] In the formula, s6 represents social behavioral characteristics, and N inter T represents the number of times a child with autism makes eye contact with other classroom participants during the observation time segment. overlap The time overlap of speech between autistic children and other classroom participants within the observation time segment is defined as T, where T is the length of the observation time segment, α and β are weighting coefficients (in this embodiment, α and β are 0.7 and 0.3 respectively), and s6 is the extracted social behavior feature (in this embodiment, α and β are 0.7 and 0.3 respectively).
[0064] The method (formula) for calculating social behavior characteristics by weighting the frequency of eye contact and the proportion of dialogue overlap time is proposed for the first time. Compared with a single indicator, this calculation method can better cover and characterize the social behavior of the tested children.
[0065] S37. Based on biosignal data, directional attention features are extracted according to the dual-channel attention theory of cognitive neuroscience.
[0066] In practical applications, for a given observation time segment, based on the Posner model (a two-channel attention theory in cognitive neuroscience), attentional features, namely directional attention, are extracted from multi-source, multi-modal signals acquired by a multimodal perception module. Directional attention measures the ability to maintain focus and is calculated by combining visual attention duration with physiological arousal level.
[0067] Furthermore, the formula for calculating the directional attention feature is as follows.
[0068]
[0069] In the formula, s7 represents the directional attention feature; T on The effective fixation time is represented by T, which is the length of the observation time segment. SCL′ is the first derivative of the skin conductance level, reflecting the strength of transient sympathetic nerve activation.
[0070] By incorporating fixation duration and skin conductance levels into the calculation of directional attention, the results of directional attention level calculation are more accurate.
[0071] In practical applications, multi-source, multimodal data is analyzed, interpreted, and its features are extracted to identify the behavioral characteristics of children with autism. Through deep learning analysis of visual, physiological signals, and sound data, the specific behavioral patterns of children with autism are identified and classified, providing accurate behavioral labels and data support for subsequent participation calculations.
[0072] S4. Based on the video stream, the current teaching scenario is identified using the teaching scenario classification rule base to obtain the current teaching scenario. Then, the weight of each modal feature in the current teaching scenario is assigned using the expert intervention strategy rule base to obtain the weight of each modal feature in the current teaching scenario.
[0073] In practical applications, the core feature vector thresholds set by the expert rule base generate scenario adaptation requirement labels. In this embodiment, the teaching scenario is divided into three scenarios: group teaching, group interaction, and individual tutoring. The core behavioral features and adaptation targets obtained are shown in Table 2.
[0074] Table 2 Comparison of Teaching Scenario Types and Core Behavioral Characteristics
[0075]
[0076] S5. Based on multimodal features and their weights, a trained deep learning model is used to evaluate the classroom participation of children with autism in real time.
[0077] Based on the progressive adaptation principle of Vygotsky's zone of proximal development theory, and combined with the obtained teaching scenario types, a weighted fusion model is used to generate a scenario adaptation weight matrix, as shown in the following formula.
[0078]
[0079] Among them, R i R jThe scenario-feature correlation score is provided by the expert rule base, where n is the number of features, α is the smoothing parameter, and in this embodiment, n = 7, α = 1.2, and ω... i That is, features s in a specified teaching scenario i The feature weights are assigned as shown in Table 3 in this implementation example.
[0080] Table 3 Feature Weight Allocation Table for Teaching Scenarios
[0081] Feature Dimension Collective teaching weight Group interaction weight Individual tutoring weight Motion trajectory (s1) 0.05 0.10 0.05 Physiological arousal (s2) 0.15 0.10 0.25 Voice interaction (s3) 0.05 0.25 0.05 Emotional state (s4) 0.10 0.10 0.35 Stereotyped behavior (s5) 0.25 0.05 0.15 Social behavior (s6) 0.05 0.3 0.05 Directed attention (s7) 0.35 0.10 0.10
[0082] Furthermore, the deep learning model is a deep learning network with a BiLSTM+Multi-HeadAttention spatiotemporal attention mechanism.
[0083] Furthermore, the training process of a deep learning network is as follows.
[0084] 1) Based on multimodal features and their weights, BiLSTM is used to encode feature vectors and construct the hidden state matrix; the specific process is as follows.
[0085] The scene weight ω i The feature sequence at time step t Dynamic weighting is performed, and then BiLSTM is used to encode the dynamically weighted vector at time step t to extract the hidden state matrix h. t The resulting encoding formula is as follows.
[0086]
[0087] Among them, h t The dimension is determined by the hyperparameters of BILSTM; in this embodiment, h t The dimension is 64, which is used to map feature vectors to a high-dimensional space.
[0088] 2) Employ a multi-head attention mechanism to extract the dynamic correlation between the time dimension and the feature dimension in the hidden state matrix, and construct a deep learning model; the specific process is as follows.
[0089] Selecting a time window T, a multi-head attention mechanism is employed to establish the dynamically weighted vector code h from step 3.4.1. t The calculation model of classroom participation over time series captures the dynamic correlation between the time dimension and the feature dimension, as shown in the following formula.
[0090]
[0091] Where P represents the standardized classroom participation score, with a value range of [0,1], T represents the window, and in this embodiment, the time window T = 30 seconds; H = (h t ,ht+1 ,…,h t+T () represents the encoded sequence of a time window starting at time step t; (·) ′ d represents the matrix transpose k for h t The dimension of W; Softmax is a commonly used normalization function in deep learning; σ represents the Sigmoid activation function; W q W k W p b p These are trainable parameters.
[0092] 3) Using historical children's classroom data from the expert rule base as training data, and aiming to minimize the loss function, the deep learning model is trained to obtain a well-trained deep learning model. The specific process is as follows.
[0093] The training data comes from over 200 complete classroom recordings featuring children with autism in an expert knowledge base, along with over 40,000 classroom segments annotated by educational experts. The loss function is L2-regularized root mean square error, as shown in the formula below.
[0094]
[0095] Where N represents the number of training samples, i.e., classroom segments annotated by education experts; y n P represents the actual classroom participation level labeled by experts for sample n. n is the classroom participation rate output by the model; ||W||2 is the model parameter norm; β is the L2 regularization coefficient, which is 0.001 in this embodiment.
[0096] S6. For example Figures 4-6 As shown, intervention strategies are retrieved from the education intervention strategy database based on multimodal characteristics and classroom participation, and the classroom participation and intervention strategies are visualized in real time.
[0097] Furthermore, the classroom participation and intervention strategies are visualized in real time, including the following:
[0098] A dynamic line graph is used to visualize classroom participation. The line graph represents the trend of classroom participation of children with autism over time, and a typical children's classroom participation curve is set as a reference baseline.
[0099] The intervention strategies are visualized using dynamic multidimensional tables.
[0100] Optionally, for system visualization, the following displays are also possible.
[0101] A real-time multidimensional radar chart is used to visualize emotional states such as happiness, sadness, anger, fear, disgust, and surprise. Each axis of the radar chart represents the state score of the corresponding emotional state for children with autism.
[0102] A bar chart is used to visualize the frequency of stereotyped behaviors such as rocking, patting, nodding, twisting fingers, shrugging shoulders, biting hands, and scratching skin. Each bar in the bar chart represents the cumulative frequency of the corresponding stereotyped behavior in the classroom of autistic children.
[0103] A pie chart was used to visualize the orientation attention level of children with autism. The percentage of the pie chart represents the orientation attention score of children with autism in different teaching scenarios in the classroom.
[0104] Heatmaps were used to visualize the distribution of attention, with different colors on the heatmap representing the duration of autistic children’s gaze on different areas in the classroom.
[0105] In practical applications, based on multimodal assessment results and expert knowledge base rules, step-by-step data visualization and lightweight intervention prompts enable two-way collaborative decision-making between educators and the system. This is achieved through transparent presentation that focuses on enhancing classroom participation and providing supplementary teaching strategy suggestions, as detailed below.
[0106] 1) Visualization of key behavioral characteristics: Key behavioral characteristics are selected from the behavioral characteristic analysis module, and appropriate charts are used to visualize behavioral characteristics that reflect the classroom participation of children with autism, dynamically displaying their classroom performance in real time. In this embodiment, real-time multidimensional radar charts, bar charts, and pie charts are used to visualize emotional stability, frequency of stereotyped behaviors, and orientation attention index, respectively.
[0107] 2) Dynamic spatial heat map generation: The heat map is used to mark the focus areas of attention of the blackboard, teachers, and students. The color saturation of the red-yellow-green gradient blocks is positively correlated with the gaze duration of autistic children on the designated area, which is used to dynamically display the attention distribution of autistic children in real time.
[0108] 3) Visualization of real-time classroom participation scores: A dynamic line chart is used to visualize the classroom participation of children with autism. By setting the average real-time participation of the corresponding children in the historical classroom as a reference, teachers can observe the children's real-time classroom participation and compare it with the historical situation.
[0109] 4) Lightweight intervention prompts: Based on multimodal behavioral characteristic data and real-time classroom participation, combined with an expert knowledge base, dynamic multidimensional tables are used to provide educators with real-time lightweight intervention prompts, providing support for educators' intervention decisions when classroom participation of autistic children decreases.
[0110] The technical effects of this application are as follows: In general, this application improves the accuracy, objectivity, and timeliness of identifying the participation of children with autism, realizes the visualization of classroom participation and intervention strategies, and improves the pertinence of intervention strategies. Specific technical effects are as follows.
[0111] 1) This application achieves spatiotemporally consistent acquisition of multimodal data by integrating multi-source heterogeneous sensors such as vision, physiological signals, and speech, combined with hardware-level synchronization and spatial alignment algorithms. Compared to traditional single-modal assessment systems, this solution reduces the assessment misjudgment rate and effectively solves the problem of subjective experience bias through cross-modal feature association. Simultaneously, it integrates multiple deep learning models to accurately identify autism-specific behaviors, improving the accuracy of behavioral feature capture, which is significantly superior to traditional general classroom systems.
[0112] 2) Based on Vygotsky's zone of proximal development theory and the framework of multiple intelligences, a differentiated strategy can be generated in real time through a teaching scenario classification rule base and a dynamic weight allocation model. For three scenarios—group teaching, individual interaction, and individual tutoring—the weights of core behavioral features are dynamically adjusted, which can improve the efficiency of teaching strategy adaptation, increase classroom participation in autistic children, and reduce the probability of emotional outbursts.
[0113] 3) A BiLSTM+Multi-HeadAttention hybrid model is used to dynamically fuse multimodal temporal features with teaching scenario weights. Compared to traditional mean-weighted or static models, this approach uses a bidirectional LSTM to capture long-term dependencies in time series data and adaptively allocates feature weights through a multi-head attention mechanism. This solves the problem of insufficient modeling of temporal correlation and scenario dynamics in traditional classroom participation evaluation methods, effectively improving the timeliness of classroom participation calculation and reducing calculation errors.
[0114] 4) A structured expert knowledge base was constructed, integrating DSM-5 / ICD-11 diagnostic criteria, ABA intervention strategies, and educational psychology theories, transforming 200+ clinical cases and 40,000+ labeled data into quantifiable algorithm parameters. A BiLSTM+Multi-HeadAttention model was used to dynamically fuse expert rules and data features, improving the consistency between classroom participation assessment results and expert scores.
[0115] 5) By using various charts such as multidimensional radar charts, spatiotemporal heat maps, and time-series line graphs, the real-time classroom participation scores and various characteristic indicators of children with autism are visualized, improving the interpretability and timeliness of their classroom behavior. A multidimensional dynamic table combined with an expert rule base provides teachers with lightweight intervention prompts and strategy recommendations based on the classroom participation status of children with autism, improving the system's effectiveness and helping children with autism regain their attention in the classroom in a timely manner.
[0116] Example 2, as Figure 2 As shown, this embodiment also provides a multimodal assessment system for classroom participation of children with autism in human-computer collaborative teaching. The multimodal assessment system for classroom participation of children with autism in human-computer collaborative teaching includes: a multimodal perception module 1, a behavioral feature analysis module 2, a teaching scenario adaptation module 3, a participation calculation module 4, an expert knowledge base module 5, and a human-computer interaction module 6.
[0117] Module 5 of the expert knowledge base is used to build the expert rule base; the expert rule base includes: a privacy protection rule base, a knowledge graph of special behavioral characteristics, a teaching scenario classification rule base, a multimodal annotation database, an expert intervention strategy rule base, and an educational intervention strategy base.
[0118] The multimodal perception module 1 is used to acquire video streams and multimodal physiological and emotional data from children's classrooms. The multimodal physiological and emotional data includes: depth image data, high-speed video data, multi-channel voice data, and biosignal data collected synchronously in children's classrooms.
[0119] The multimodal sensing module 1 includes a multimodal data acquisition unit and a data synchronization and fusion unit.
[0120] 1) such as Figure 3 As shown, the multimodal data acquisition unit consists of a depth camera array arranged in a ring around the face, a high-speed camera unit, a flexible wearable biosignal device, and a multi-channel microphone array.
[0121] Depth camera array captures: subtle temperature changes on the face.
[0122] High-speed camera unit acquires: eye movement trajectory and micro-expression change data.
[0123] Flexible wearable biosignal devices collect: skin conductance response, heart rate, and limb movement characteristics.
[0124] Multi-channel microphone array: voice data.
[0125] 2) The data synchronization and fusion unit adopts the PTPv2 hardware-level time synchronization protocol to ensure the consistency of time sequence of multi-source data and reduce time synchronization error to the microsecond level; it adopts the Kabsch algorithm to establish a three-dimensional coordinate system transformation model and deploys the SIFT-3D feature point matching algorithm to achieve cross-modal data alignment; it adopts YOLOv5 to perform differential privacy processing on video and image data and perform real-time desensitization processing to protect the biometric privacy of children with autism.
[0126] The behavioral feature analysis module 2 is used to perform feature analysis on multimodal physiological and emotional data to obtain multimodal features. The multimodal features include: visual features, physiological features, voice features, emotional state features, stereotyped behavior features, social behavior features, and directed attention features.
[0127] The teaching scenario adaptation module 3 is used to identify the current teaching scenario based on the video stream and the teaching scenario classification rule library to obtain the current teaching scenario. Then, it uses the expert intervention strategy rule library to assign weights to each modal feature in the current teaching scenario to obtain the weights of each modal feature in the current teaching scenario.
[0128] The participation calculation module 4 is used to evaluate the classroom participation of children with autism in real time based on multimodal features and their weights using a trained deep learning model; the training data of the trained deep learning model comes from a multimodal annotation database.
[0129] The human-computer interaction module 6 is used to retrieve intervention strategies from the education intervention strategy library based on multimodal characteristics and classroom participation, and to visualize classroom participation and intervention strategies in real time.
[0130] Example 2: This application also provides a computer device, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 7 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores and processes data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements the methods described above.
[0131] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0132] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0133] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0134] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A multimodal assessment method for classroom participation of children with autism in human-computer collaborative teaching, characterized in that, The multimodal assessment method for classroom participation of children with autism in human-computer collaborative teaching includes: Construct an expert rule base; the expert rule base includes: a privacy protection rule base, a special behavior feature knowledge graph, a teaching scenario classification rule base, a multimodal annotation database, an expert intervention strategy rule base, and an educational intervention strategy base; Acquire video streams and multimodal physiological and emotional data from children's classrooms; the multimodal physiological and emotional data includes: depth image data, high-speed video data, multi-channel voice data, and biosignal data synchronously collected in the children's classrooms; Feature analysis is performed on multimodal physiological and emotional data to obtain multimodal features; the multimodal features include: visual features, physiological features, voice features, emotional state features, stereotyped behavior features, social behavior features, and directed attention features; Based on the video stream, the current teaching scenario is identified using a teaching scenario classification rule base. The current teaching scenario is then obtained, and the weights of each modal feature in the current teaching scenario are assigned using an expert intervention strategy rule base. Based on multimodal features and their weights, a trained deep learning model is used to evaluate the classroom participation of children with autism in real time; the training data of the trained deep learning model comes from a multimodal annotation database. Based on multimodal characteristics and classroom participation, intervention strategies are retrieved using an educational intervention strategy database, and classroom participation and intervention strategies are visualized in real time.
2. The multimodal assessment method for classroom participation of children with autism in human-computer collaborative teaching according to claim 1, characterized in that, Acquiring video streams and multimodal physiological and emotional data from children's classrooms, specifically including: Obtain the video stream of the children's classroom; Acquire multimodal physiological and emotional data, and use a privacy protection rule base to perform real-time biometric privacy desensitization processing on the multimodal physiological and emotional data, and update the multimodal physiological and emotional data.
3. The multimodal assessment method for classroom participation of children with autism in human-computer collaborative teaching according to claim 1, characterized in that, Feature analysis was performed on multimodal physiological and emotional data to obtain multimodal features, specifically including: Based on high-speed video data, the YOLOv5 algorithm is used for target detection and action recognition, and a convolutional neural network is used for facial expression analysis to obtain visual features. Based on biosignal data, a long short-term memory network is used to model time-series physiological signal data to obtain physiological characteristics; Based on multi-channel speech data, the Wav2Vec 2.0 speech emotion analysis algorithm was used to extract speech features; Based on depth image data, the VGG-Face facial expression recognition network is used to identify emotional state features; Based on high-speed video data, stereotypical behavioral features are identified using deep convolutional neural networks and temporal convolutional networks. Based on high-speed video data, OpenPose was used to obtain the frequency of eye contact between autistic children and other classroom participants per unit time, and the proportion of dialogue overlap time between autistic children and other classroom participants was calculated based on multi-channel voice data to obtain social behavior characteristics. Based on biosignal data, directional attention features were extracted according to the dual-channel attention theory of cognitive neuroscience.
4. The multimodal assessment method for classroom participation of children with autism in human-computer collaborative teaching according to claim 3, characterized in that, The formula for calculating the social behavior characteristics is as follows: In the formula, s6 represents social behavioral characteristics, and N inter T represents the number of times a child with autism makes eye contact with other classroom participants during the observation time segment. overlap The time overlap of speech between autistic children and other classroom participants within the observation time segment is defined as T, where T is the length of the observation time segment, and α and β are weighting coefficients.
5. The multimodal assessment method for classroom participation of children with autism in human-computer collaborative teaching according to claim 3, characterized in that, The formula for calculating the directed attention feature is as follows: In the formula, s7 represents the directional attention feature; T on The effective fixation time is T; the observation time segment length is SCL'; and the first derivative of the skin conductance level is SCL'.
6. The multimodal assessment method for classroom participation of children with autism in human-computer collaborative teaching according to claim 1, characterized in that, The deep learning model is a deep learning network with a BiLSTM+Multi-Head Attention spatiotemporal attention mechanism.
7. The multimodal assessment method for classroom participation of children with autism in human-computer collaborative teaching according to claim 6, characterized in that, The training process of the deep learning network is as follows: Based on multimodal features and their weights, BiLSTM is used for feature vector encoding to construct a hidden state matrix; A multi-head attention mechanism is used to extract the dynamic correlation between the time dimension and the feature dimension in the hidden state matrix, and a deep learning model is constructed. Using historical classroom data of children in the expert rule base as training data, and aiming to minimize the loss function, the deep learning model is trained to obtain a well-trained deep learning model.
8. The multimodal assessment method for classroom participation of children with autism in human-computer collaborative teaching according to claim 1, characterized in that, Real-time visualization of classroom participation and intervention strategies, specifically including: A dynamic line graph is used to visualize classroom participation. The line graph represents the trend of classroom participation of children with autism over time, and a typical children's classroom participation curve is set as a reference baseline. The intervention strategies are visualized using dynamic multidimensional tables.
9. A multimodal assessment system for classroom participation of children with autism in human-computer collaborative teaching, characterized in that, The multimodal assessment system for classroom participation of children with autism in human-computer collaborative teaching includes: a multimodal perception module, a behavioral feature analysis module, a teaching scenario adaptation module, a participation calculation module, an expert knowledge base module, and a human-computer interaction module; The expert knowledge base module is used to construct an expert rule base; the expert rule base includes: a privacy protection rule base, a special behavior feature knowledge graph, a teaching scenario classification rule base, a multimodal annotation database, an expert intervention strategy rule base, and an educational intervention strategy base; The multimodal perception module is used to acquire video streams and multimodal physiological and emotional data from the children's classroom; the multimodal physiological and emotional data includes: depth image data, high-speed video data, multi-channel voice data, and biosignal data synchronously collected in the children's classroom; The behavioral feature parsing module is used to perform feature analysis on multimodal physiological and emotional data to obtain multimodal features; the multimodal features include: visual features, physiological features, voice features, emotional state features, stereotyped behavior features, social behavior features, and directed attention features; The teaching scenario adaptation module is used to identify the current teaching scenario based on the video stream using a teaching scenario classification rule library, obtain the current teaching scenario, and use an expert intervention strategy rule library to assign weights to each modal feature in the current teaching scenario, thereby obtaining the weights of each modal feature in the current teaching scenario. The participation calculation module is used to evaluate the classroom participation of children with autism in real time based on multimodal features and their weights using a trained deep learning model; the training data of the trained deep learning model comes from a multimodal annotation database. The human-computer interaction module is used to retrieve intervention strategies from the education intervention strategy library based on multimodal characteristics and classroom participation, and to visualize classroom participation and intervention strategies in real time.
10. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the multimodal assessment method for classroom participation of autistic children in human-computer collaborative teaching as described in any one of claims 1-8.
Citation Information
Patent Citations
Intelligent teaching system for ASD (Autism Spectrum Disorder) children
CN105280044A
Student classroom participation degree analysis system based on classroom videos
CN111046823A
Learning participation degree determination method and device and computer equipment
CN112085392A
Intelligent classroom teaching optimization method and system combining behavior recognition and Internet of Things
CN119831100A
Cited By
AI interaction-based personalized knowledge graph dialogue generation method for old people
CN121658617A
Medical education integrated multi-modal data acquisition and analysis method for autism rehabilitation
CN121662256A
A multimodal data acquisition and analysis method integrating medical and educational approaches for autism rehabilitation
CN121662256B