A method, system, terminal and medium for assessing cognitive expression of children

By introducing immersive scenarios and multimodal visual behavior analysis into the assessment of children with autism, and combining them with intelligent behavior assessment models, the problem of low accuracy of assessment results in existing technologies has been solved, and automated and accurate assessment of children with autism's emotional cognition and expression has been achieved.

CN122266741APending Publication Date: 2026-06-23LISHUI MATERNAL & CHILD HEALTH HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LISHUI MATERNAL & CHILD HEALTH HOSPITAL
Filing Date
2026-01-29
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Current methods for screening, assessing, and intervening in autism rely on manual clinical assessments and lack a unified machine-assisted assessment system, resulting in low accuracy of assessment results and a lack of systematic behavioral and mechanistic analysis.

Method used

By employing multimodal visual behavior analysis technology in immersive scenarios, combined with functional game paradigms and intelligent behavior assessment models, and through viewpoint-independent feature extraction and temporal optimization, we can achieve automated assessment of emotional cognition and expression in children with autism.

Benefits of technology

It improves the accuracy and automation of the assessment of emotional cognition and expression in children with autism, and provides a unified and automated intelligent behavior analysis model to support the comprehensive assessment of emotional development abilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122266741A_ABST
    Figure CN122266741A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, and discloses a child cognitive expression evaluation method, a child cognitive expression evaluation system, a terminal and a medium. The child cognitive expression evaluation method comprises the following steps: acquiring internal information and video information of a child in an immersive scene; analyzing the video information to obtain a plurality of high-level feature vectors; constructing an evaluation model; inputting the internal information and the plurality of high-level feature vectors into the evaluation model; and outputting an evaluation result of the child on cognitive expression. The application can improve the accuracy of the evaluation result of the child cognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to an assessment method, system, terminal, and medium for children's cognitive expression. Background Technology

[0002] Autism Spectrum Disorder (ASD) is clinically characterized by impaired social communication, repetitive behaviors, and restricted interests. Research indicates that evidence-based early intervention can optimize the prognosis and rehabilitation of children with autism, offering significant benefits to families and society, with earlier intervention yielding better results. Early behavioral assessment of autism is crucial for alleviating symptoms and improving the quality of life for children with autism.

[0003] Current methods for autism screening, assessment, and intervention primarily rely on behavioral observation and intervention by clinical physicians. This process depends on experienced clinicians, and the lengthy training periods and limited availability of rehabilitation professionals make it difficult to meet the growing demand for intervention and assessment services. Furthermore, current assessment methods and standards are diverse, lacking a universally recommended single assessment or treatment approach. Additionally, the results of manual assessments are often based on the subjective feelings of professionals and lack specific assessment indicators. These problems hinder the screening, assessment, and intervention process for children with autism.

[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0005] The main purpose of this application is to provide a method, system, terminal and medium for assessing children's cognitive expression, aiming to solve the problem that the existing technology relies on the subjective feelings of professionals to assess autistic patients, lacks specific assessment indicators, and leads to low accuracy of assessment results.

[0006] The first aspect of this application provides a method for assessing children's cognitive expression, the method comprising the following steps: Acquire internal information and video data about children in immersive scenarios; The video information is analyzed to obtain multiple high-level feature vectors; An evaluation model is constructed, and the internal information of the paradigm and multiple high-level feature vectors are input into the evaluation model to output the evaluation results of the child's cognitive expression.

[0007] In one possible implementation, the paradigm internal information includes reaction time, accuracy, similarity, and completion time, and the video information includes time-synchronized video data from multiple viewpoints; The acquisition of internal information and video information of children in immersive scenarios specifically includes: Acquire the reaction time and accuracy of children's interaction in immersive scenarios based on the emotional cognition assessment paradigm, and acquire the similarity and completion time of children's interaction in immersive scenarios based on the emotional expression assessment paradigm; Acquire multiple time-synchronized video data from multiple perspectives of children interacting according to the emotional cognitive assessment paradigm and the emotional expression assessment paradigm.

[0008] In one possible implementation, the plurality of high-level feature vectors include viewpoint attention information, action attention information, and emotion attention information; The analysis of the video information yields multiple high-level feature vectors, specifically including: Feature extraction was performed on multiple video data points synchronized from multiple perspectives to obtain the child's facial expression features, posture features, and gesture features; Cross-modal attention fusion is performed on the facial expression features, posture features, and gesture features to obtain viewpoint attention information, action attention information, and emotion attention information.

[0009] In one possible implementation, the step of extracting features from multiple time-synchronized video data from multiple viewpoints to obtain the child's facial expression features, posture features, and gesture features specifically includes: Based on the video data corresponding to time synchronization from multiple perspectives, multiple high-dimensional visual feature vectors are obtained; The high-dimensional visual feature vectors are fused to obtain a unified visual feature vector. Multimodal feature extraction is performed on the unified visual feature vector to obtain the child's facial expression features, posture features, and gesture features.

[0010] In one possible implementation, the cross-modal attention fusion of the facial expression features, the posture features, and the gesture features to obtain viewpoint attention information, action attention information, and emotion attention information specifically includes: The facial expression features, posture features, and gesture features are time-series optimized to obtain optimized facial expression features, posture features, and gesture features; A cross-modal attention mechanism is used to perform multimodal feature fusion on the optimized facial expression features, posture features, and gesture features to obtain viewpoint attention information, action attention information, and emotion attention information.

[0011] In one possible implementation, the evaluation model includes a first evaluation module and a second evaluation module; The step of inputting the internal information of the paradigm and multiple high-level feature vectors into the evaluation model and outputting the evaluation results of the child's cognitive expression specifically includes: The reaction time, accuracy, similarity, and completion time are input into the first evaluation module to obtain a cognitive score and an expression score. The viewpoint attention information, the action attention information, and the emotion attention information are input into the second evaluation module to obtain a refined score; The child's cognitive expression is assessed based on the cognitive score, the expression score, and the refined score.

[0012] In one possible implementation, the step of inputting the internal information of the paradigm and multiple high-level feature vectors into the evaluation model, and outputting the evaluation results of the child's cognitive expression, further includes: The evaluation results were compared and verified to obtain the verification results; If the verification result does not meet the standard, the evaluation model is optimized by inputting the internal information of the paradigm and the multiple high-level feature vectors into the optimized evaluation model to obtain the optimized evaluation result. The optimized evaluation result is then compared and verified to obtain the optimized verification result, until the optimized verification result meets the standard.

[0013] A second aspect of this application also provides a system for assessing children's cognitive expression, wherein the system is applied to the method for assessing children's cognitive expression described in any of the above-described solutions; the system for assessing children's cognitive expression includes: The information acquisition module is used to acquire internal information and video information of children in immersive scenarios; The data analysis and fusion module is used to analyze the video information and obtain multiple high-level feature vectors; The quantitative assessment module is used to construct an assessment model, inputting the internal information of the paradigm and multiple high-level feature vectors into the assessment model, and outputting the assessment results of the child's cognitive expression.

[0014] A third aspect of this application also provides a terminal, wherein the terminal includes: a memory, a processor, and a child cognitive expression evaluation program stored in the memory and executable on the processor, wherein when the child cognitive expression evaluation program is executed by the processor, it implements the steps of the child cognitive expression evaluation method as described above.

[0015] A fourth aspect of this application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an evaluation program for children's cognitive expression, and when the evaluation program for children's cognitive expression is executed by a processor, it implements the steps of the evaluation method for children's cognitive expression as described above.

[0016] Beneficial effects: This application provides a method, system, terminal and medium for assessing children's cognitive expression. This application analyzes interactive videos of children in immersive scenarios to achieve viewpoint-independent feature extraction and temporal optimization, and performs a hybrid assessment based on feature vectors and internal paradigm information to achieve the purpose of automated assessment of childhood autism and improve the accuracy of assessment results. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A schematic diagram illustrating a preferred embodiment of the assessment system for children's cognitive expression according to this application; Figure 2 A flowchart illustrating a preferred embodiment of the assessment method for children's cognitive expression according to this application; Figure 3 This is a flowchart illustrating the specific implementation steps of the entire execution process in a preferred embodiment of the assessment method for children's cognitive expression in this application. Figure 4 This is a schematic diagram illustrating the extraction of multimodal temporal features from a video stream in a preferred embodiment of the assessment method for children's cognitive expression according to this application; Figure 5 This is a schematic diagram illustrating multimodal feature fusion using a cross-modal attention mechanism, as a preferred embodiment of the assessment method for children's cognitive expression in this application. Figure 6 A structural diagram of a preferred embodiment of the assessment system for children's cognitive expression according to this application; Figure 7 This is a structural diagram of a preferred embodiment of the terminal of this application.

[0019] Explanation of reference numerals in the attached figures: 100. Information Acquisition Module; 200. Data Analysis and Fusion Module; 300. Quantitative Evaluation Module. Detailed Implementation

[0020] To make the objectives, technical solutions, and effects of this application clearer and more explicit, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. The described embodiments are only possible technical implementations of this application and not all possible implementations. Based on the embodiments in this application, those skilled in the art can obtain other embodiments without creative effort, and these embodiments are also within the protection scope of this application.

[0021] In related technologies, to address the two challenges of traditional manual assessment methods—requiring a large number of rehabilitation professionals and the subjectivity of assessment indicators—researchers have introduced machine-assisted assessment and intervention strategies. Currently, many intelligent systems and methods exist for the assessment and intervention of children with autism. However, most interventions primarily use computers or mobile devices, lacking immersion, specificity, and flexibility, and are still difficult to replace manual methods. Furthermore, there are gaps in the assessment of emotional cognition and expression, lacking a systematic machine-based approach to the emotional behavior and mechanisms of children with autism. This application systematically integrates technologies such as immersive environments, functional play paradigms, and multimodal visual-behavioral analysis to provide a complete, closed-loop machine-assisted assessment system specifically targeting the core impairment of emotional cognition and expression abilities in children with autism.

[0022] This application primarily addresses the problems encountered when using machine-assisted assessment to evaluate the emotional cognition and expression abilities of children with autism: First, the lack and inadequacy of machine-assisted assessment systems. Current assessments of emotional cognition and expression in children with autism mostly rely on manual clinical methods, lacking a systematic machine-assisted assessment framework. Existing systems, while few, lack immersion and specificity, indicating significant room for development. Second, poor uniformity and low automation in assessment scenarios. Machine-assisted assessments of autistic emotions lack a unified, automated intelligent behavioral analysis model, making it difficult to conduct intelligent, comprehensive, and unified assessments of children's emotional development abilities. Third, a lack of systematic behavioral and mechanistic analysis. Due to the significant behavioral differences and unclear mechanisms in autism, there is currently a lack of in-depth and systematic research from the perspectives of machine and data analysis into children's emotional behavioral patterns, expressive differences, and subtypes.

[0023] First, the following is combined with the appendix Figure 1 The system architecture of the embodiments of this application will be described.

[0024] This application presents an AI-based assessment system for the emotional cognition and expression of children with autism (an assessment system for children's cognitive expression), integrating machine-assisted systems, computer-assisted systems, and an immersive system. It aims to automate and objectively assess the emotional cognition and expression abilities of children with autism. A large touchscreen provides close-range immersive interaction, while a three-view fisheye camera module is used for somatosensory interaction sensing and video information acquisition. Specifically, this application first constructs an immersive assessment and intervention scenario for autism at the hardware level to enable interactive engagement for children. Then, a functional game paradigm is designed to guide a progressively deeper process from both emotional cognition and expression perspectives. Simultaneously, a data acquisition system is designed to collect children's visual modalities and system data during the assessment process. Subsequently, based on the collected data, an intelligent behavioral assessment and analysis model for the immersive emotional assessment system is designed, combining prior rules from clinical scales to achieve the assessment and analysis of emotional abilities in the autism children's paradigm process. Finally, by integrating the system scenario and analysis model, a methodology for assessing the emotional cognition and expression of children with autism is established, and a controlled experiment is conducted to verify the system's effectiveness.

[0025] Compared to existing intelligent systems for assessing and intervening in autism, this application guides children through the assessment process using an immersive assessment system scenario and a functional game paradigm for assessing emotional cognition and expression. Simultaneously, it designs an intelligent behavioral assessment and analysis model for an immersive emotional assessment system, achieving a comprehensive assessment of the emotional behavior of children with autism through various visual algorithms and system data. Finally, it conducts application verification and efficacy evaluation of this intelligent assessment method.

[0026] To address the issue of low accuracy in assessments that rely on the subjective feelings of professionals and lack specific assessment indicators, this application analyzes interactive videos of children in immersive scenarios to achieve viewpoint-independent feature extraction and temporal optimization. It then performs a hybrid assessment based on feature vectors and internal paradigm information to automate the assessment of childhood autism and improve the accuracy of the assessment results.

[0027] The technical solutions of this application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0028] The preferred embodiment of this application describes a method for assessing children's cognitive expression, such as... Figure 2 As shown, the assessment method for children's cognitive expression includes the following steps: In step S101, the paradigm internal information and video information of the child in the immersive scene are obtained.

[0029] Understandably, see Figure 3This application constructs an immersive system assessment scenario for the emotional cognition and expression of children with autism, establishes an intelligent behavioral assessment and analysis model for the immersive emotional assessment system, and uses children's multi-visual modal information, system data, and clinical rule priors to achieve automated emotional assessment of children with autism.

[0030] In one possible implementation, the internal information of the paradigm includes reaction time, accuracy, similarity, and completion time, and the video information includes time-synchronized video data from multiple perspectives. The process involves acquiring the reaction time and accuracy of children interacting in an immersive scene according to the emotional cognition assessment paradigm, acquiring the similarity and completion time of children interacting in an immersive scene according to the emotional expression assessment paradigm, and acquiring multiple time-synchronized video data from multiple perspectives of children interacting according to both the emotional cognition and emotional expression assessment paradigms.

[0031] Specifically, this application establishes an immersive assessment system scenario and designs a progressive functional game paradigm. See also Figure 1 (a) and Figure 1 In (b), the immersive scene consists of three ultra-short-throw projectors forming a continuous U-shaped image. The projection screen is made of rigid PVC foam board. A three-view fisheye camera module and LiDAR are used for motion-sensing interaction and video information acquisition. An image fusion unit is connected to a desktop computer to control and output the image. This system scene can deploy relevant functional game paradigms and evaluation models, providing a foundation for machine-assisted evaluation methods. The paradigm combines the guidance and evaluation phases. At the guidance level, children's emotional cognitive abilities are established based on expression and scenario design paradigms. At the evaluation level, a progressive process of cognition-expression-cognition is followed. First, children's cognition of others' expressions is evaluated; then, children are guided to imitate and evaluate corresponding expressions; finally, the children's expressions are collected to evaluate their cognition of their own expressions. During the paradigm process, algorithms and system data are further integrated to provide narration, guidance, and feedback on children's actions, thereby stimulating children's attention and enhancing immersion. Specifically, it includes the following two core sub-paradigms: Paradigm 1 is an emotional cognitive functional game designed to assess children's ability to recognize and understand others' facial expressions. At the start of the game, the system greets the child with a gentle voice: "Little friend, shall we play a game of finding facial expressions?" Then, three clear cartoon face photos appear in the center of the screen, representing different basic emotions: "happy," "sad," and "angry." The system prompts with a voice, "Please find the 'happy' expression." The child needs to point or touch the icon they believe to be the correct expression. A lidar or fisheye camera in the scene captures the child's body movements to determine their choice. If the selection is correct, the system immediately provides positive audiovisual feedback, accompanied by a voice message: "You're great!" If the selection is incorrect, the system provides corrective feedback and a voice prompt: "Wrong, let's try again." Paradigm Two is a functional game focused on emotional expression. This paradigm follows the cognitive task and aims to assess children's ability to imitate and make specific facial expressions. At the start of the game, a clear cartoon expression, such as a wide smile, appears in the center of the screen, accompanied by a voice prompt: "Please make a 'happy' expression." The system gives the child a few seconds to prepare and imitate. During this time, a three-view fisheye camera module records a high-definition video stream of the child's face from different angles. The model analyzes the similarity between the child's expression and the target expression in real time, and a progress bar on the screen displays the accuracy of the imitation, providing immediate visual feedback to the child. When the similarity exceeds a preset threshold, the system determines that the imitation is successful. The system provides feedback based on the assessment results: "Wow, your smile is so bright!" or offers targeted guidance: "Let's try again, raise the corners of your mouth a little more, it will look even more like it!" The second paradigm also includes a cognitive reconfirmation step. After the child successfully imitates, the system plays the video of the child's own expression and asks: "Look, this is the expression you made, isn't it a 'happy' expression?" This step is used to assess children's cognition and understanding of their own facial expressions, completing a closed-loop assessment process of "cognition-expression-self-cognition".

[0032] To achieve intelligent and automated evaluation, the data acquisition system is divided into two levels: internal paradigm information acquisition and children's behavioral video acquisition. For internal paradigm information acquisition, the communication transmission between the paradigm and the database is designed. Based on the characteristics of the paradigm and referring to the structured teaching method assessment scale, the children's selection time (reaction time), selection accuracy rate in Paradigm 1, and facial expression similarity and imitation completion time in Paradigm 2 are transmitted to the database for storage. For children's behavioral video acquisition, real-time synchronous frame acquisition of three-view video is performed using a three-view fisheye camera, while simultaneously acquiring internal screen recording information for later behavioral analysis model extraction and modeling.

[0033] It should be noted that data acquisition directly supports the multimodal feature extraction in step S102 (such as visual features of video frames) and the rule evaluation in step S103 (such as accuracy and reaction time), forming a complete data flow of "scenario building - paradigm execution - data acquisition - model evaluation". This provides a solid foundation for subsequent intelligent behavior analysis (step S102) and hybrid evaluation (step S103), ensuring that the evaluation process is both in line with clinical logic and possesses the automation and accuracy of artificial intelligence.

[0034] In step S102, the video information is analyzed to obtain multiple high-level feature vectors.

[0035] In one possible implementation, the multiple high-level feature vectors include viewpoint attention information, action attention information, and emotional attention information. Feature extraction is performed on multiple time-synchronized video data from multiple viewpoints to obtain the child's facial expression features, posture features, and gesture features; cross-modal attention fusion is then performed on the facial expression features, posture features, and gesture features to obtain viewpoint attention information, action attention information, and emotional attention information.

[0036] This application generates high-level feature vectors that support sentiment assessment through multimodal feature extraction, viewpoint invariant learning, temporal optimization, and cross-modal fusion.

[0037] In one possible implementation, multiple high-dimensional visual feature vectors are obtained based on the video data corresponding to time synchronization from multiple viewpoints; feature fusion is performed on the multiple high-dimensional visual feature vectors to obtain a unified visual feature vector; multimodal feature extraction is performed on the unified visual feature vector to obtain the child's facial expression features, posture features, and gesture features.

[0038] In one possible implementation, the facial expression features, posture features, and gesture features are temporally optimized to obtain optimized facial expression features, posture features, and gesture features; a cross-modal attention mechanism is used to perform multimodal feature fusion on the optimized facial expression features, posture features, and gesture features to obtain viewpoint attention information, action attention information, and emotion attention information.

[0039] Specifically, an intelligent behavior assessment and analysis model is built, and based on the system's in-game data and children's video information captured by the camera, the video is precisely cut into short video clips corresponding to different assessment tasks. Figure 4 This demonstrates the entire process from video clips to the final generation of multimodal features. Visual features are extracted from video frames using a pre-trained ResNet-50 residual network. It is assumed that at any given time... From the first Single-frame video images captured by a camera First, the data is preprocessed to perform size normalization, data augmentation, and standardization, and then input into the ResNet-50 backbone network to obtain high-dimensional visual feature vectors. .

[0040] ; in, This indicates the standardized preprocessing procedure performed on the input video frames.

[0041] To eliminate the interference of camera angle on feature extraction and ensure that the model learns the child's inherent behavioral patterns rather than artifacts from a specific perspective, this invention introduces a perspective-invariant feature learning framework based on contrastive learning for the video data collected from three different perspectives.

[0042] In the feature space, feature vectors from different perspectives at the same time are considered positive sample pairs, and all unrelated feature vectors are considered negative samples. By narrowing the distance between positive sample pairs and widening the distance between negative sample pairs, the perspective invariance of the features is ensured, significantly improving the discriminative ability of the features.

[0043] In this application, at time t From the k Feature vectors extracted from each camera viewpoint Positive samples (defined as anchor points) are from the same time point. t From any other perspective j ( j≠k Extracted feature vectors The negative sample set has two sources: one is feature vectors extracted from video frames of other children, and the other is feature vectors extracted from different temporal video frames of the child themselves, defined as... A contrastive learning loss function is used as a constraint on viewpoint consistency. This loss function aims to maximize the mutual information between the anchor point and the positive sample, as shown in the following formula: ; in, To compare the learning loss values, used to constrain viewpoint-invariant feature learning; At the same time t The feature vector extracted from any other perspective j (j≠k) is the positive sample; For the negative sample set; The cosine similarity function is used to calculate the similarity between feature vectors. This is a temperature hyperparameter used to adjust the degree of attention the loss function pays to the samples; here it is taken as... , This represents the total number of samples in the negative sample set.

[0044] After this step, the features from the three perspectives can be fused to obtain a unified visual feature vector. 。 Then, various computer vision algorithms are used to extract frame-level multi-visual modal features of children. Human skeletal pose features and hand pose features are extracted for each frame through human pose estimation and gesture recognition. Head pose and gaze features are extracted through head pose estimation and gaze estimation. Facial features and expression categories are extracted through face detection and expression recognition.

[0045] After extracting features independently from each frame, the resulting raw feature sequences (such as facial landmark coordinates) often contain high-frequency noise and jitter due to lighting changes, rapid object movement, or minor errors in the model itself during video capture. This unstable feature sequence severely interferes with the learning performance of subsequent models. A temporal filtering module employs a Kalman filter to smooth and predict the extracted frame-level features, optimizing their stability and continuity. For head pose, a 12-dimensional state vector is used. This is to comprehensively describe the complete three-dimensional spatial motion state of the head at time t.

[0046] ; in, , , This refers to the three-dimensional translation position of the head. , , For head pitch, yaw, and roll angles; , , This refers to the three-dimensional translation speed of the head. , , The angular velocity component of the head; Next, two mathematical models are established to support the "prediction-correction" process: Motion model: With observation model: ; in, Indicates time The state vector, This represents the state vector from the previous time step. Represents the state transition matrix. This represents a small uncertainty, such as the child suddenly turning their head at a sudden speed. This represents the raw head pose data that the visual algorithm directly reads from the current video frame. This represents the observation matrix. Representative to Inaccuracy estimation. For each frame of the video stream, the system uses a motion model to predict the head pose of the current frame and obtains the actual observation data of the current frame. It dynamically weighs the reliability of the predicted value and the observation value, and intelligently fuses the two to obtain the optimized best pose estimate.

[0047] After temporal optimization of the original features of each modality, these low-level features located in different feature spaces need to be fused and enhanced to extract high-level key information that can directly support sentiment assessment. Multimodal feature fusion is performed using a cross-modal attention mechanism. Through different query and context combinations, three types of key information vectors—viewpoint, action, and sentiment—are generated in a targeted manner as the foundation for establishing the sentiment assessment model. The process is as follows: Figure 5 As shown.

[0048] For the viewpoint, let the optimized head pose and gaze direction features be the query vector. The feature set of all interactive objects in an immersive scene As context, where, This is the feature vector of the first interactive object in the immersive scene. This is the feature vector of the second interactive object in the immersive scene. For the first in immersive scene The feature vectors of interactive objects. First, the query vector and context features are converted into query vectors respectively through linear mapping. Key vector Value vector Then, by calculating the relevance between the query vector and each object in the context, the context information is weighted and summed to obtain the final viewpoint attention information. This vector encodes information about the scene objects that children are most interested in.

[0049] ; in, Let be the dimension of the key vector. The function represents the transformation of an input vector into a probability distribution.

[0050] For motion analysis, meaningful motion patterns need to be identified from a continuous sequence of body poses. A self-attention mechanism is employed to analyze the sequence of body pose features within a time window (e.g., 30 frames). It can also serve as a Query, Key, and Value. It is a moment The body posture feature vector, It is a moment The body posture feature vector, It is a moment The model generates body posture feature vectors. Through self-attention computation, it can associate temporally discontinuous but semantically related postures (such as a hand-raising action followed by a pointing action), thereby understanding the complete action intent. The final action attention information... It is the result of the self-attention mechanism aggregating the information of the entire sequence: ; in, For query vectors, For key vectors, For value vectors, For self-attention; Finally, a comprehensive analysis of emotional information (emotional and attentional information) is conducted. The extraction of information, through the integration of cues from facial expressions, voice, and body language, yields a robust assessment of emotional state. This is based on facial expression features. As a query, action information extracted from other modalities will be used. Harmony and prosodic features As context, after linear mapping, the original facial expression judgment is enhanced or corrected by a cross-modal attention mechanism that utilizes body movements and vocal intonation.

[0051] ; in, For query vectors, For key vectors, It is a value vector.

[0052] The final module will output three high-level key information vectors: These will collectively serve as inputs to the final evaluation model, laying the foundation for accurate evaluation.

[0053] It should be noted that this application transforms the original video data into highly semantic feature vectors through a chain of "multi-view acquisition - feature extraction - viewpoint invariant learning - temporal optimization - cross-modal fusion", ensuring that the evaluation model can not only capture subtle behavioral features, but also resist interference from viewpoint, noise and other factors, thereby providing accurate and interpretable sentiment evaluation basis.

[0054] In step S103, an evaluation model is constructed by inputting the internal information of the paradigm and multiple high-level feature vectors into the evaluation model, and outputting the evaluation results of the child's cognitive expression.

[0055] This application constructs a hybrid evaluation engine that integrates rule-based priors and data-driven approaches. Through a dual path of clinical rule quantification and machine learning modeling, it transforms the high-level features of step S102 into quantitative scores that conform to clinical logic.

[0056] In one possible implementation, the evaluation model includes a first evaluation module and a second evaluation module. The reaction time, accuracy, similarity, and completion time are input into the first evaluation module to obtain a cognitive score and an expression score; the viewpoint attention information, motor attention information, and emotional attention information are input into the second evaluation module to obtain a refined score; and the evaluation result is obtained based on the cognitive score, the expression score, and the refined score.

[0057] Specifically, an assessment model is established to transform these feature vectors into quantifiable competency scores that conform to clinical logic. To balance the accuracy and clinical interpretability of the assessment, this application creatively proposes a hybrid assessment engine based on rules and intelligent models. This engine comprises two parallel modules whose assessment results are ultimately fused to achieve complementary advantages.

[0058] The first module is a rule-based assessment module based on prior clinical scales. It transforms clinical assessment knowledge and logic into a set of automatically executable computer program rules. This module is a collection of "IF-THEN" logical judgment statements. It receives three types of extracted key information vectors as input and performs logical matching and scoring according to the task requirements of the current functional game paradigm. For Paradigm 1 emotional cognition, the scoring incorporates the accuracy of the child's choices. (Value is 1 or 0) and reaction time . The preset reaction time threshold is set to 7 seconds. Score The calculation formula is: ; The indicator function is set to 1 if the condition is true and 0 otherwise. Scoring for Paradigm 1: This formula means 2 points for quick and accurate completion; 1 point for accurate but slow response; and 0 points for error. Paradigm 2, which assesses emotional expression, primarily considers the similarity of the child's imitation of facial expressions. ) and the time required to complete the imitation The scoring logic is as follows: ; in, and These are the high and low thresholds for imitation similarity, respectively, 0.85 and 0.65. This is the completion time threshold, set to 10 seconds. A score of 2 indicates that the child's imitation is both accurate and fast; a score of 1 indicates that the child's imitation is relatively accurate, or accurate but slow; and a score of 0 indicates that the child's imitation is inaccurate.

[0059] Another module is a data-driven intelligent assessment module that leverages machine learning capabilities to discover more subtle and complex behavioral patterns in data that are difficult to describe with explicit rules. This application proposes an advanced hierarchical spatiotemporal attention module. This module aims to model a sequence of key information within a time window as a dynamically evolving flow of children's cognitive-behavioral states, and automatically learn and locate the key behaviors most important for assessment.

[0060] First enter It is a complete information sequence from time 1 to time T. At each time t, the input vector... The three high-level feature vectors at that moment are viewpoint, action, and emotion. It is assembled. After passing through a Transformer encoder layer and using a self-attention mechanism, a new sequence enhanced with contextual information is obtained, which is the entire hidden state sequence. .

[0061] ; in, It is the core encoding module in the Transformer architecture, used to perform context-aware feature enhancement on the input sequence.

[0062] For the One evaluation project uses learnable query vectors to scan the entire hidden state sequence. In the assessment task When, calculate the first Importance score of each moment .

[0063] ; in, Represent the query vector for the task; For a moment The hidden state; , The learnable weights and bias matrices.

[0064] These scores are then normalized to obtain the attention weights at each time step. .

[0065] ; These attention weights are used to sum the hidden states at all time points to obtain a context generalization vector specific to task i. Finally, the context generalization vector specific to each task is... Each result is fed into a separate prediction head to derive the final evaluation score for that item. .

[0066] ; in, In order to target the A custom-designed independent multilayer perceptron for each evaluation task.

[0067] This assessment score It presents a detailed, multi-dimensional profile of children's emotional development. It provides a quantitative score for each specific assessment item, such as the recognition of happy expressions and the expression of sad expressions. The final result is a multi-dimensional assessment that helps clinicians or rehabilitation therapists clearly identify children's specific strengths and weaknesses in emotional abilities, thus providing precise data support for developing personalized intervention and rehabilitation plans.

[0068] To obtain the final score for each child This invention weighted and fused the scores from the rule-based evaluation module and the intelligent evaluation module: ; This is a weighted hyperparameter that can be adjusted according to the actual application scenario: when a stronger clinical logical explanation is required, it can be increased. Value; when pursuing higher evaluation accuracy, it can be reduced. This value allows for flexible adjustment of the final score's emphasis between interpretability and accuracy. This application adopts... The value is 0.7. The model's loss function is: ; in, It is the absolute value of the difference between the score from the intelligent assessment module and the expert clinical score; It is a contrastive learning loss used to ensure the viewpoint independence of features; It is the regularization term of the model, used to prevent overfitting. These are hyperparameters used to balance different loss terms; here they are set to 1.0 and 0.8 respectively.

[0069] It should be noted that this application employs a hybrid engine—a rule module to ensure clinical interpretability and an intelligent module to capture subtle behavioral patterns—to transform the features of step S102 into quantifiable scores. The rule module directly maps to the logic of clinical scales, ensuring that the results meet professional standards; the intelligent module models dynamic behavioral flows through spatiotemporal attention, uncovering complex patterns that are difficult for rules to cover. Finally, a balance between interpretability and accuracy is achieved through weighted fusion, and model performance is jointly optimized through multiple loss functions to ensure that the evaluation results are both scientific and clinically applicable.

[0070] In one possible implementation, the evaluation results are compared and verified to obtain a verification result; if the verification result does not meet the standard, the evaluation model is optimized by inputting the internal information of the paradigm and the multiple high-level feature vectors into the optimized evaluation model to obtain an optimized evaluation result, and the optimized evaluation result is compared and verified to obtain an optimized verification result, until the optimized verification result meets the standard.

[0071] Specifically, a clinical validation experiment was designed to compare the system's automated evaluation results with the "gold standard" human evaluation results from senior clinical experts to assess the system's effectiveness. Two core statistical indicators were used in the comparison and validation of the system's evaluation results with the clinical expert human evaluation results.

[0072] First, the Pearson correlation coefficient: used to measure the score of the system output. With expert clinical scores The degree of linear correlation between them .

[0073] ; in, Indicates the first Quantitative scores for each assessment subject This represents the average score for all evaluated objects. Experts indicated their opinion on the first Clinical scores of each assessment subject This represents the average clinical score given by experts to all subjects being assessed.

[0074] The closer the value is to 1, the more consistent the evaluation results of the system of this invention are with the clinical judgment of experts. Here, it is set... A value greater than 0.8 is considered to indicate a high positive correlation.

[0075] Secondly, the Kappa coefficient Used to measure the consistency of judgments on emotion categories, eliminating occasional consistency.

[0076] ; in, It is the observed consistency rate. It is the random consistency rate. This application designs... A value greater than 0.75 is considered to indicate high consistency.

[0077] If the system's evaluation results show a high positive correlation with the clinical experts' evaluation results in Pearson correlation analysis and a high consistency in Kappa coefficient analysis, it can fully demonstrate that the system and method described in this invention have high effectiveness and clinical reliability. Its evaluation results can accurately and stably reproduce the experts' professional judgment and have value for clinical promotion and application.

[0078] This application uses cross-validation with dual indicators (linear correlation + classification consistency) to ensure that the system not only matches the experts in quantitative scores, but also reaches a professional level in emotional category judgment. Ultimately, it proves that the system has clinical application value and can provide a scientific and efficient automated tool for assessing the emotional abilities of children with autism.

[0079] Next, referring to the accompanying drawings, an assessment system for children's cognitive expression according to an embodiment of this application is described, which is applied to the assessment method for children's cognitive expression in any of the above-described solutions.

[0080] Figure 6 This is a structural diagram of a child cognitive expression assessment system according to an embodiment of this application.

[0081] like Figure 6 As shown, the assessment system for children's cognitive expression includes: an information acquisition module 100, a data analysis and fusion module 200, and a quantitative assessment module 300.

[0082] Specifically, the information acquisition module 100 is used to acquire internal information and video information of children in immersive scenarios; The data analysis and fusion module 200 is used to analyze the video information to obtain multiple high-level feature vectors; The quantitative assessment module 300 is used to construct an assessment model, input the internal information of the paradigm and multiple high-level feature vectors into the assessment model, and output the assessment results of the child's cognitive expression.

[0083] Figure 7 A structural diagram of a terminal provided in an embodiment of this application. The terminal may include: The memory 501, the processor 502, and the computer program stored on the memory 501 and capable of running on the processor 502.

[0084] When the processor 502 executes the program, it implements the method for assessing children's cognitive expression provided in the above embodiments.

[0085] Furthermore, the terminal also includes: Communication interface 503 is used for communication between memory 501 and processor 502.

[0086] The memory 501 is used to store computer programs that can run on the processor 502.

[0087] Memory 501 may include high-speed RAM memory, and may also include non-volatile memory. volatile memory), for example, at least one disk storage.

[0088] If the memory 501, processor 502, and communication interface 503 are implemented independently, then the communication interface 503, memory 501, and processor 502 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EIS) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 7 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0089] Optionally, in a specific implementation, if the memory 501, processor 502, and communication interface 503 are integrated on a single chip, then the memory 501, processor 502, and communication interface 503 can communicate with each other through an internal interface.

[0090] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of this application.

[0091] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for assessing children's cognitive expressions.

[0092] One embodiment of this application provides a computer program product, including a computer program that, when executed by a processor, implements the features described in this application. Figure 2 The corresponding embodiments provide methods for assessing children's cognitive expression.

[0093] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0094] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0095] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0096] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable storage medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable storage medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable storage medium could be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0097] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0098] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0099] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0100] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

[0101] It should be understood that the application of this application is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for assessing children's cognitive expression, characterized in that, The assessment methods for children's cognitive expression include: Acquire internal information and video data about children in immersive scenarios; The video information is analyzed to obtain multiple high-level feature vectors; An evaluation model is constructed, and the internal information of the paradigm and multiple high-level feature vectors are input into the evaluation model to output the evaluation results of the child's cognitive expression.

2. The method for assessing children's cognitive expression according to claim 1, characterized in that, The internal information of the paradigm includes reaction time, accuracy, similarity and completion time, and the video information includes time-synchronized video data from multiple perspectives; The acquisition of internal information and video information of children in immersive scenarios specifically includes: Acquire the reaction time and accuracy of children's interaction in immersive scenarios based on the emotional cognition assessment paradigm, and acquire the similarity and completion time of children's interaction in immersive scenarios based on the emotional expression assessment paradigm; Acquire multiple time-synchronized video data from multiple perspectives of children interacting according to the emotional cognitive assessment paradigm and the emotional expression assessment paradigm.

3. The method for assessing children's cognitive expression according to claim 2, characterized in that, The multiple high-level feature vectors include viewpoint attention information, action attention information, and emotion attention information; The analysis of the video information yields multiple high-level feature vectors, specifically including: Feature extraction was performed on multiple video data points synchronized from multiple perspectives to obtain the child's facial expression features, posture features, and gesture features; Cross-modal attention fusion is performed on the facial expression features, posture features, and gesture features to obtain viewpoint attention information, action attention information, and emotion attention information.

4. The method for assessing children's cognitive expression according to claim 3, characterized in that, The step of extracting features from multiple time-synchronized video data from multiple viewpoints to obtain the child's facial expression features, posture features, and gesture features specifically includes: Based on the video data corresponding to time synchronization from multiple perspectives, multiple high-dimensional visual feature vectors are obtained; The high-dimensional visual feature vectors are fused to obtain a unified visual feature vector. Multimodal feature extraction is performed on the unified visual feature vector to obtain the child's facial expression features, posture features, and gesture features.

5. The method for assessing children's cognitive expression according to claim 4, characterized in that, The cross-modal attention fusion of the facial expression features, posture features, and gesture features to obtain viewpoint attention information, action attention information, and emotion attention information specifically includes: The facial expression features, posture features, and gesture features are time-series optimized to obtain optimized facial expression features, posture features, and gesture features; A cross-modal attention mechanism is used to perform multimodal feature fusion on the optimized facial expression features, posture features, and gesture features to obtain viewpoint attention information, action attention information, and emotion attention information.

6. The method for assessing children's cognitive expression according to claim 5, characterized in that, The evaluation model includes a first evaluation module and a second evaluation module; The step of inputting the internal information of the paradigm and multiple high-level feature vectors into the evaluation model and outputting the evaluation results of the child's cognitive expression specifically includes: The reaction time, accuracy, similarity, and completion time are input into the first evaluation module to obtain a cognitive score and an expression score. The viewpoint attention information, the action attention information, and the emotion attention information are input into the second evaluation module to obtain a refined score; The child's cognitive expression is assessed based on the cognitive score, the expression score, and the refined score.

7. The method for assessing children's cognitive expression according to claim 1, characterized in that, The process of inputting the internal information of the paradigm and multiple high-level feature vectors into the evaluation model, and outputting the evaluation results of the child's cognitive expression, further includes: The evaluation results were compared and verified to obtain the verification results; If the verification result does not meet the standard, the evaluation model is optimized by inputting the internal information of the paradigm and the multiple high-level feature vectors into the optimized evaluation model to obtain the optimized evaluation result. The optimized evaluation result is then compared and verified to obtain the optimized verification result, until the optimized verification result meets the standard.

8. An assessment system for children's cognitive expression, characterized in that, The assessment system for children's cognitive expression is applied to the assessment method for children's cognitive expression according to any one of claims 1-7; the assessment system for children's cognitive expression includes: The information acquisition module is used to acquire internal information and video information of children in immersive scenarios; The data analysis and fusion module is used to analyze the video information and obtain multiple high-level feature vectors; The quantitative assessment module is used to construct an assessment model, inputting the internal information of the paradigm and multiple high-level feature vectors into the assessment model, and outputting the assessment results of the child's cognitive expression.

9. A terminal, characterized in that, The terminal includes: a memory, a processor, and a child cognitive expression evaluation program stored in the memory and executable on the processor, wherein when the child cognitive expression evaluation program is executed by the processor, it implements the steps of the child cognitive expression evaluation method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an evaluation program for children's cognitive expressions, which, when executed by a processor, implements the steps of the evaluation method for children's cognitive expressions as described in any one of claims 1-7.