Child behavior feature collection and analysis system and method based on virtual image interaction and multi-modal fusion
By using standardized interactive scripts guided by virtual avatars and multimodal data fusion technology, the problems of strong subjectivity, limited data, and low efficiency in the diagnosis of childhood ADHD are solved, achieving efficient and objective multidimensional assessment, which is suitable for the accurate diagnosis of childhood ADHD.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE FIRST AFFILIATED HOSPITAL OF XINXIANG MEDICAL UNIVERSITY
- Filing Date
- 2026-02-28
- Publication Date
- 2026-07-10
AI Technical Summary
Existing diagnostic technologies for childhood ADHD suffer from problems such as high subjectivity, limited data dimensions, low efficiency, and poor interactive experience for children, making it difficult to achieve accurate and rapid multi-dimensional assessments.
A standardized interactive script guided by a virtual avatar is adopted, which combines multimodal data acquisition and deep learning analysis to achieve synchronous acquisition and fusion of voice, text and eye-tracking data. Through multi-turn dialogue control and temporal alignment, a multimodal feature fusion diagnostic system is constructed.
It significantly improves the objectivity, comprehensiveness and efficiency of diagnosis, provides efficient and reliable multi-dimensional assessment with an accuracy of ≥88%, an F1 score of ≥0.86, a single case diagnosis time of ≤15 minutes, strong adaptability, and is suitable for large-scale screening.
Smart Images

Figure CN122369858A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and medical auxiliary diagnosis technology, and in particular to a system and method for collecting and analyzing children's behavioral characteristics based on virtual image interaction and multimodal fusion. Background Technology
[0002] Attention Deficit Hyperactivity Disorder (ADHD) is a common neurodevelopmental disorder in childhood, characterized by inattention, hyperactivity, and impulsive behavior. The pathogenesis of this disorder is closely related to neurobiological, genetic, and environmental factors. If timely and effective intervention is not received, it will have long-term adverse effects on children's learning abilities, social functioning, and mental health. Therefore, early and accurate diagnosis of ADHD is of significant clinical and social value for developing personalized intervention strategies and improving the prognosis of affected children.
[0003] Currently, the clinical diagnosis of childhood ADHD mainly relies on traditional methods, including behavioral observation, rating scales, and clinical interviews. However, these methods have significant shortcomings in practical application and are insufficient to meet the needs of accurate diagnosis.
[0004] In behavioral observation methods, current technology requires doctors or parents to record children's behavior in different scenarios over a long period. The diagnostic results are highly dependent on the observer's subjective judgment and professional experience. Due to differences in experience and cognitive biases among observers, the consistency and reproducibility of diagnostic results are poor. Furthermore, this method struggles to quantify children's attention allocation and cognitive processing in natural states, and cannot provide objective physiological indicators as a diagnostic basis.
[0005] Regarding scale-based assessment methods, while standardized questionnaires are convenient to use, they suffer from technical limitations, such as biased information collection. Existing scales often rely on subjective responses from parents or teachers, making them susceptible to influences such as individual cognitive biases and social expectation effects, leading to distorted assessment results. Furthermore, scale data cannot reflect the dynamic behavioral characteristics of children during interactions in real time, making it difficult to capture the temporal changes in ADHD symptoms and limiting the timeliness and accuracy of diagnosis.
[0006] Regarding clinical interview methods, existing technologies rely on direct communication between doctors and children / parents to obtain medical history and symptom information. This approach has low diagnostic efficiency and is difficult to scale up. Children often exhibit resistance in medical settings, leading to distorted language and behavior, which affects doctors' judgment of the true symptoms. Furthermore, the manual interview process is time-consuming, requiring an average of 30-60 minutes per child, which cannot meet the efficiency requirements of large-scale screening.
[0007] In recent years, artificial intelligence (AI) technology has been applied in the field of ADHD diagnosis, but existing AI diagnostic technologies still face significant technical bottlenecks. Most solutions focus on the analysis of single-dimensional data, such as electroencephalogram (EEG) signals, behavioral scale data, or static imaging features, failing to fully integrate the multimodal characteristics of children during natural interactions. This single-dimensional approach cannot comprehensively reflect the complex pathological mechanisms of ADHD, resulting in considerable room for improvement in the applicability of diagnostic scenarios and the convenience of data collection. In particular, current technologies lack interactive diagnostic systems capable of simultaneously collecting children's language expression characteristics and eye-tracking physiological indicators, making it impossible to achieve multi-dimensional quantitative assessment of children's attentional states, emotional regulation abilities, and cognitive processing.
[0008] Furthermore, existing technologies cannot effectively address the psychological adaptation issues in pediatric medical settings. In traditional diagnostic methods, children are prone to anxiety and resistance when facing unfamiliar doctors and medical environments, resulting in data that fails to reflect their true condition.
[0009] In summary, existing diagnostic technologies for childhood ADHD have the following shortcomings: ① The diagnostic process is highly subjective and lacks objective quantitative indicators; ② The data dimensions are limited, failing to comprehensively reflect the complex pathological characteristics of ADHD; ③ Diagnostic efficiency is low, making large-scale application difficult; ④ Children's interactive experience is poor, and the quality of data collection is unstable. Therefore, there is an urgent need to develop an innovative diagnostic technology that integrates multimodal data collection, natural interactive guidance, and deep learning analysis to overcome the inherent defects of existing technologies. Summary of the Invention
[0010] This invention addresses the technical problems of strong subjectivity, single data dimension, and low efficiency in existing methods for assessing children's behavior. It employs key technical means such as standardized interactive scripts guided by virtual avatars, synchronous acquisition and high-precision temporal alignment of multimodal data in natural dialogue scenarios, and deep feature fusion of multimodal data across scenarios, achieving significant improvements in the objectivity, comprehensiveness, and automation efficiency of assessments.
[0011] According to one aspect of this disclosure, a system for collecting and analyzing children's behavioral characteristics based on virtual avatar interaction and multimodal fusion is provided, comprising: The large-scale dialogue interaction module is used to drive virtual characters to engage in multi-round natural dialogues with children that are in line with their cognitive level, and dynamically control the dialogue process based on a preset interactive script. The multimodal data acquisition module is used to simultaneously collect children's voice data, text data, and eye-tracking physiological data during dialogue. The multimodal data preprocessing and time-series alignment module is used to clean, extract features, and synchronize the time series of the collected multimodal data. The deep learning analytics module includes multiple independent feature extractors and feature fusion units, used to extract features from preprocessed multimodal data and perform fusion analysis. The large-scale model dialogue interaction module establishes a data connection with the multimodal data acquisition module. The output of the multimodal data acquisition module is connected to the input of the multimodal data preprocessing and temporal alignment module. The output of the multimodal data preprocessing and temporal alignment module is connected to the input of the deep learning analysis module, forming a complete data processing pipeline from interaction guidance to feature analysis.
[0012] In some embodiments of this disclosure, the large model dialogue interaction module includes: The virtual character rendering unit is configured with standardized parameters for visual, emotional, and linguistic dimensions. The visual dimension adopts cartoon-style modeling and is configured with a dynamic expression library and friendly body movements. The emotional dimension includes an emotional response rule library. The linguistic dimension is set with tone and speech rate parameters adapted to children's cognition. The dialogue flow control unit includes a preset interactive script library and a multi-turn dialogue control mechanism. The interactive script library is designed around the cognitive ability assessment dimension, and the multi-turn dialogue control mechanism dynamically determines the continuation or switching of topics based on the completeness assessment of children's answers and attention state analysis of a large language model.
[0013] In some embodiments of this disclosure, the multi-turn dialogue control mechanism includes: The topic end judgment unit is configured to call the large language model to quantitatively assess the completion of diagnostic information collection for the current topic and the child's attention status, and determine whether to end the current topic based on a preset threshold. The topic follow-up question unit is configured to call the large language model to generate child-adaptive follow-up question scripts that point to the missing information item when the continuation condition is met. The topic switching unit is configured to call the large language model to generate transitional phrases and start the next topic when the termination condition is triggered.
[0014] In some embodiments of this disclosure, the multimodal data acquisition module includes: The voice acquisition unit is equipped with a directional noise-canceling microphone to acquire raw audio at a sampling rate of 16000Hz and calculates the signal-to-noise ratio in real time for quality control. The text acquisition unit simultaneously acquires the original transcribed text from the speech recognition model and the standardized text corrected by the large language model; The eye-tracking acquisition unit is equipped with an eye tracker to acquire raw data of three-dimensional fixation point coordinates and pupil diameter at a sampling rate of 120Hz, and extracts fixation duration and saccade frequency features in real time. The voice acquisition unit, text acquisition unit, and eye-tracking acquisition unit are all equipped with a structured storage mechanism that includes a unique session identifier and a dialogue turn ID.
[0015] In some embodiments of this disclosure, the multimodal data preprocessing and timing alignment module includes: The speech preprocessing unit is configured to convert the processed audio into 80-dimensional mem spectrogram temporal features and extract fundamental acoustic features such as fundamental frequency, speech rate, and speech intensity. The text preprocessing unit is configured to convert standardized text into a 768-dimensional fixed-length semantic vector using a pre-trained language model. The eye-tracking preprocessing unit is configured to perform outlier removal, missing value completion, and feature normalization on temporal eye-tracking data. The timing alignment unit is configured to use the task scenario as the anchor point and correct the time deviation of the three types of modal data to within ±25ms through a dynamic time warping algorithm, and mark the alignment status.
[0016] In some embodiments of this disclosure, the deep learning analysis module includes: The language feature extractor uses a neural network structure that includes hidden layers and Dropout layers, taking a 768-dimensional semantic vector as input and outputting a 64-dimensional language feature vector. An eye-tracking feature extractor uses a GRU network structure to process temporal eye-tracking features and outputs a 64-dimensional eye-tracking feature vector. The speech feature extractor includes a Mel spectrum branch and a basic acoustic feature branch. After processing by 1D-CNN, pooling layer, and bidirectional GRU, it outputs a 64-dimensional speech feature vector. The feature fusion unit is configured to concatenate language, eye-tracking, and speech feature vectors in a single-task scenario to form a 192-dimensional fusion feature, and then concatenate the fusion features from multiple task scenarios and input them into the classifier.
[0017] According to a second aspect of this disclosure, a multimodal data time series alignment method is provided, applied to the acquisition and analysis system, comprising the following steps: Based on task scenario identification, raw voice, text, and eye-tracking data generated within the same task cycle are filtered to form a single-task multimodal data group. Based on the timestamp generated by the system clock, the time of task instruction issuance and task completion are extracted as time boundaries; Using key time points in the task flow as anchor points, the response time of voice data, the transcription completion time of text data, and the entire time sequence of eye-tracking data are aligned with the task timeline; The time offset of the three types of data is corrected by a dynamic time warping algorithm to ensure that the time deviation is ≤ ±25ms; Mark the alignment status in the data set. Mark the overall alignment status as successful only when all three types of modal data are successfully aligned.
[0018] According to a third aspect of this disclosure, a method for dynamic control of dialogue flow is provided, applied to the data acquisition and analysis system, comprising: A pre-built interactive script library includes visual cognitive scripts and life communication scripts designed around the cognitive ability assessment dimension. During the dialogue, a large language model is used to quantitatively analyze the children's responses, including: ① Information collection completion assessment: Based on the preset core information collection items, calculate the score of the degree to which the child's answers cover each piece of information; ② Attention status assessment: Combining response time and text features, output attention labels and sustained scores; Based on quantitative analysis results and preset thresholds, a dynamic decision-making dialogue process is implemented. ① When the information collection completion score is ≥90 points, the current topic ends and a transition script is generated; ② When the attention tag is scattered and the score is <50 points for 3 consecutive rounds, end the current topic and generate a transition script; ③ When the number of dialogue rounds reaches the preset limit, end the current topic and generate transitional dialogue; ④ When the continuation condition is met, the large language model is invoked to generate follow-up questions pointing to the missing information item.
[0019] According to a fourth aspect of this disclosure, a multimodal feature fusion analysis method is provided, applied to the acquisition and analysis system, comprising: Feature extraction is performed on multimodal data in single-task scenarios: ①Language features: Extract 64-dimensional features from the 768-dimensional semantic vector through a network containing 128-dimensional hidden layers and Dropout layers; ② Eye movement features: 64-dimensional features were extracted from the temporal eye movement data using a 256-dimensional GRU network; ③ Speech features: The Mel spectrogram is processed by 1D-CNN, pooling layer, and bidirectional GRU and then concatenated with the basic acoustic features to output 64-dimensional features; The three types of feature vectors in a single-task scenario are concatenated to form a 192-dimensional fused feature. The fusion features of multiple task scenarios are spliced together by dimension to form a cross-scenario total feature; A fully connected network containing 256-dimensional hidden layers and Dropout layers is used to classify and analyze the total features across scenes, and outputs a binary classification result.
[0020] According to a fifth aspect of this disclosure, a computer-readable storage medium is provided that stores a computer program, which, when executed by a processor, implements the steps of the method.
[0021] One or more technical solutions provided in the embodiments of the present invention have at least one of the following technical effects or advantages: 1. Significantly improves the objectivity and data authenticity of diagnosis, solving the problem of "high subjectivity": By adopting a virtual doctor avatar and standardized interactive scripts, a relaxed and friendly diagnostic environment is created, greatly reducing children's tension and resistance in medical settings. This allows for the collection of more authentic language expression and eye movement response data that closely resembles their natural state. These data (such as saccade frequency, fixation duration, and speech spectrum characteristics) are all quantifiable objective physiological and behavioral indicators, effectively avoiding the assessment bias and data distortion caused by reliance on subjective human observation and recall in traditional methods, providing a solid and objective data foundation for diagnosis.
[0022] 2. Achieving multi-dimensional information complementarity and comprehensive assessment to address the problem of "single data": Language interaction data (speech and text) and objective eye-tracking data are simultaneously collected and deeply fused and analyzed in natural dialogue scenarios. Language data reflects children's expressive logic, emotional regulation, and executive function; eye-tracking data directly maps their attention allocation, information processing speed, and inhibitory control abilities. Through multimodal feature extraction and fusion technology, a multi-faceted and complementary quantitative characterization of ADHD core symptoms (inattention, hyperactivity, and impulsivity) is achieved, overcoming the limitations of a single data source and significantly improving the comprehensiveness and accuracy of diagnosis (test accuracy ≥88%, F1 score ≥0.86).
[0023] 3. Significantly improves diagnostic efficiency and standardization, addressing the problem of "low diagnostic efficiency": This invention automates the entire process from interactive guidance, data collection, feature extraction to diagnostic output. A large-scale conversational model replaces doctors in conducting time-consuming and skill-intensive interviews, while a deep learning model automatically performs complex feature analysis and decision-making. The overall time for a single-case diagnosis is significantly reduced (≤15 minutes), and the inference time for a single-sample model is extremely short (≤100 ms), overcoming the bottlenecks of low efficiency and scalability inherent in traditional manual diagnosis. This provides an efficient and feasible technical tool for large-scale, rapid early screening of childhood ADHD.
[0024] 4. Enhanced scalability and adaptability of the constructed system: The entire solution is based on a modular architecture design, featuring high standardization and scalability. Standardized interaction scripts ensure consistency in data acquisition processes across different application scenarios; unified data processing and diagnostic models guarantee the comparability of results. Simultaneously, the system supports flexible replacement and optimization of dialogue models, feature extraction algorithms, and interaction scripts, easily adapting to changing diagnostic needs across different regions and cultural backgrounds, possessing strong potential for clinical promotion and long-term evolution.
[0025] In summary, by organically integrating the natural interaction advantages of a large conversational model, the diagnostic value of multimodal data, and the intelligent analysis capabilities of a deep learning model, this invention has successfully constructed an ADHD auxiliary diagnostic system that is child-friendly, objective, accurate, efficient, and automated. It effectively solves the key defects of existing technologies and has significant clinical value and social significance. Attached Figure Description
[0026] Figure 1 This is a schematic diagram of a children's behavioral feature collection and analysis system based on virtual avatar interaction and multimodal fusion, according to an embodiment of this application.
[0027] Figure 2 This is an example of the interactive interface of the data acquisition and analysis system in one embodiment of this application. Detailed Implementation
[0028] To better understand the technical solution of this application, the above technical solution will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0029] Example 1: Construction of a Children's Behavioral Feature Collection and Analysis System Based on Virtual Avatar Interaction and Multimodal Fusion This embodiment provides a complete system for collecting and analyzing children's behavioral characteristics based on virtual avatar interaction and multimodal fusion. This system is particularly suitable for assisting in the collection and analysis of behavioral characteristics of children with Attention Deficit Hyperactivity Disorder (ADHD). This system is the first to systematically integrate technologies such as child-friendly virtual avatar interaction, synchronous multimodal data collection based on structured scripts, temporal alignment anchored by task scenarios, and cross-scenario multimodal deep learning feature fusion, constructing an efficient, objective, and child-friendly automated analysis platform.
[0030] 1. System Overall Architecture and Deployment Environment This system adopts a client-server architecture, and the specific deployment environment is as follows: (1) Hardware configuration: Interactive terminal: A desktop computer with a 15.6-inch high-definition touch screen, equipped with an Intel i7-12700H CPU and an NVIDIA RTX 3060 GPU (6GB VRAM), is used to render the virtual doctor image in real time, run the front-end interactive interface, and receive data from the acquisition device simultaneously.
[0031] Multimodal acquisition equipment: Voice is acquired using a directional noise-canceling microphone (such as the Rohde Wireless GO II), with an effective pickup distance of 1~3m; eye-tracking data is acquired using a high-precision eye tracker (such as the Tobii Pro X3-120, with a sampling rate of 120Hz and a spatial resolution of 0.5°). Both are connected to the interactive terminal via a USB interface adapter.
[0032] Servers: Configured with high-performance computing servers, including Intel Xeon Gold 6330 CPUs and NVIDIA A100 GPUs (80GB VRAM), for deploying large language models, deep learning diagnostic models, and running data processing pipelines. Equipped with a 10TB HDFS distributed storage system for reliable storage of massive amounts of multimodal data.
[0033] (2) Software environment: Operating system: The interactive terminal is Windows 11, and the server is Ubuntu 20.04 LTS.
[0034] Development and runtime framework: Python 3.9 is used as the primary programming language. Deep learning training uses the PyTorch 2.0 framework, and model inference and deployment are accelerated using TensorRT 8.6. Audio processing uses FFmpeg 5.1, and virtual character rendering uses the Unity 2022 engine and OpenCV 4.8 library.
[0035] Model Deployment: The conversational large language model (Qwen3-235B in this example) is deployed on the server after INT4 quantization to achieve an inference speed of ≥5 tokens / s. The BERT-based text encoding model, as well as the feature extractor and classifier composed of 1D-CNN and GRU, are all deployed using TensorRT optimization to ensure that the total time for a single multimodal data inference is ≤100ms.
[0036] 2. Specific implementation of each core module The following is in conjunction with the attached figures ( Figure 1 , Figure 2 Each major core module is described in detail.
[0037] 2.1 Implementation of the Interactive Guidance Module This module serves as the entry point for direct interaction between the system and child users. Through a standardized three-in-one design of "virtual avatar + structured script + intelligent dialogue control," it achieves a balance between professionalism and child-friendliness.
[0038] The construction of virtual interactive avatars: ① Visual Expression Unit: A cartoonish 3D doctor model was created using the Unity 2022 engine. This model includes a pre-built animation library containing dynamic expressions such as a "smile" (corners of the mouth raised 15°~30°) and an "encouraging eyebrow raise" (eyebrow arch raised 5°~10°), as well as friendly body language such as "gentle hand gestures" and "slightly leaning forward." Script control ensures that the virtual avatar's mouth shape and expressions are synchronized with the real-time generated voice, maintaining a stable animation frame rate of 30fps for a smooth visual experience.
[0039] ② Emotional Response Unit: A lightweight rule engine was developed to form an emotion response rule base. For example, when the system detects that a child is silent for more than 5 seconds, or identifies keywords such as "doesn't want to talk" through text analysis, the rule engine immediately triggers the virtual avatar to display a "concerned gaze" expression accompanied by reassuring words (such as "It's okay, we'll take it slowly"). This real-time adaptation effectively alleviates the child's anxiety.
[0040] ③ Speech Output Unit: Integrates high-quality text-to-speech (TTS) service, and presets the output speech parameters as follows: fundamental frequency (pitch) in the 250Hz-350Hz range suitable for children's voices, and speech rate controlled at 2-3 words / s. This parameter setting is based on research on children's auditory cognition and can significantly improve the friendliness and intelligibility of the speech.
[0041] Implementation of interaction logic based on a pre-set script: ① Script Configuration and Management: Standardized interactive scripts designed around core assessment dimensions such as attention, emotion regulation, behavioral execution, and language expression are pre-stored in the server's MySQL database. The scripts are divided into two main categories: "Visual Cognition" (including sub-scripts on visual recognition and detail comparison) and "Life Interaction" (including sub-scripts on daily school life, game experiences, and family interactions). Visual cognition category: Visual Recognition Sub-Script: The topic is "Selection of Images with Specified Features". The system displays a set of animal / object images with different colors, shapes and quantities, and issues instructions to children such as "Select the red bird image" and "Find the image containing three circles". Children are required to complete the image selection on the system page, and the system collects the children's visual attention concentration and feature recognition accuracy. Detailed comparison sub-script: The topic is "finding differences in images". The system displays two sets of highly similar scene images (such as classroom scene and park scene) and gives children the instruction to "find three different details in the two images". Children are required to mark the differences and the system collects the children's visual attention precision.
[0042] Lifestyle communication: The sub-script for daily campus life: The topic is "Description of classroom learning and recess activities". The system asks children questions such as "Describe what you learned in class in the morning" and "Talk about the games you usually play with your classmates during recess", guiding children to describe their campus learning and social scenarios, and collecting data on the completeness of their language expression and attention span. Game experience sub-script: The topic is "Participation and feelings in extracurricular games". The system asks children questions such as "Introduce how to play your favorite toy / game" and "Talk about your feelings when playing this game", guiding children to share their game experience and collecting data on children's emotional expression and logical expression abilities. Family Interaction Sub-Script: The topic is "Daily Interactions in Family Life". The system asks children questions such as "Describe something you did with your family" and "Talk about your daily interactions with your parents", guiding children to describe family scenes and collecting data on children's social interaction cognition and language organization skills.
[0043] ② Intelligent Dialogue Control Unit: This unit is one of the key innovations of this invention, implemented by calling the API of the Qwen3-235B large language model. After each round of dialogue, the system inputs the current dialogue history, core information items of the current script, as well as the child's response duration and answer text, into the large model according to a specific Prompt template.
[0044] a. Quantitative Assessment: Based on the Prompt instruction, the large model outputs a "Completion of Information Collection" score (0-100 points) and an "Attention State" label (focused / distracted) and score (0-100 points) for the current answer. For example, the Prompt is: "Based on the core information collection items: [classroom content, recess activities, social objects], analyze the child's answer text: ['I had math class, and then played with my deskmate for a while'], and output the information item completion score." b. Decision-making and execution: The dialogue control unit makes decisions based on preset thresholds. If the "Information Collection Completion Rate" is ≥90 points, it is determined that the current topic information is sufficient, and the control module generates transition text and starts the next topic.
[0045] If the "attention status" is "distracted" and the score is below 50 for three consecutive rounds, the child's interest is considered to have decreased, and a topic switch is triggered.
[0046] If the information is insufficient (completeness < 90 points) but the child is focused (score ≥ 70 points), the system will initiate a "follow-up questioning mechanism." At this point, the system will pass the missing information to the larger model, which will then generate a concise, friendly question that is appropriate for the child's cognitive level (such as "What game did you and your deskmate play?").
[0047] This mechanism ensures that the interaction process is both standardized and flexible enough to adapt to individual children’s responses, thus efficiently guiding data collection.
[0048] 2.2 Implementation of the Multimodal Data Acquisition Unit This unit enables high-precision synchronous acquisition and structured associated storage of speech, text, and eye-tracking data during natural dialogue.
[0049] When the virtual avatar asks a question and plays a voice message, the data collection starts simultaneously.
[0050] Voice Acquisition Subunit: The analog audio signal acquired by the microphone is converted into a digital signal in real time by the sound card and processed into a mono WAV format file with a sampling rate of 16000Hz and a bit depth of 16bit using the FFmpeg library. Simultaneously, the signal-to-noise ratio (SNR) of the audio segment is calculated. If the SNR is lower than 20dB, the RNNoise algorithm is automatically invoked for real-time noise reduction, and the data quality is marked.
[0051] Text Acquisition Subunit: The speech file is fed into the Whisper-small speech recognition model in near real-time to generate the original transcribed text. To improve text quality, the original text is further fed into the Qwen3-235B model and standardized using a Prompt instruction (such as "Please correct the following child-like colloquial expressions to make them semantically clear and grammatically correct, while retaining the original core information") to obtain the final text for analysis.
[0052] Eye-tracking acquisition subunit: The eye tracker continuously outputs raw data streams via the SDK. The system records the coordinate sequence of the area the child is looking at on the screen, changes in pupil diameter, and saccade events during each round of dialogue.
[0053] Structured storage: Upon generation, the three types of data mentioned above are immediately assigned and associated with the same globally unique "session identifier (session_id)" and an incrementing "turn identifier (turn_id)". All metadata (such as timestamps, device models, quality tags) and file paths are recorded in JSON format and stored in HDFS. This design greatly facilitates subsequent data retrieval, alignment, and batch processing.
[0054] 2.3 Implementation of the Data Processing and Analysis Module This module addresses the challenge of fusing and analyzing multimodal data due to differences in acquisition start time and frequency, and proposes a time-series alignment method "using the task scenario as an anchor point".
[0055] (1) Pretreatment: ① Speech preprocessing: Endpoint detection (VAD) is performed on the WAV audio to remove the beginning and end silence segments. The Librosa library is used to extract the Mel spectrogram (80-dimensional Mel filter, 25ms window, 10ms step) as the temporal input for deep learning. Simultaneously, statistical features such as fundamental frequency, speech rate, and intensity are extracted.
[0056] ② Text preprocessing: Clean up irrelevant characters in the standardized text, load the pre-trained BERT-base model through Hugging Face's Transformers library, and convert the text of each round of dialogue into a 768-dimensional semantic vector.
[0057] ③ Eye movement preprocessing: Apply the 3σ principle to the raw eye movement data to remove coordinate outliers caused by blinking or head movement, and then use linear interpolation to complete the data. Unify the data to a fixed duration (e.g., 5 seconds, corresponding to 600 time points), extract temporal features such as pupil diameter, fixation point coordinates (which can be used to derive fixation duration), and saccade frequency, and normalize them.
[0058] (2) Timing alignment: ①Scene Anchoring: The system reads the script_name (script name) and task instruction (such as "find the red bird") corresponding to each round of data. Multimodal data from all rounds belonging to the same specific task cycle are grouped into a "task data group".
[0059] ② Time base unification: The Unix timestamp of the first voice command issued by the system at the start of the task is used as the base time T0.
[0060] ③ Cross-modal correlation correction: For the speech, text (corresponding to the response period), and eye-tracking (the entire task process) data within this task group, the Dynamic Time Warping (DTW) algorithm is used to calculate their optimal matching path with the reference time axis, thereby correcting the time offset caused by device latency and processing pipeline differences. In this embodiment, through parameter tuning, the alignment error is controlled within ±25ms.
[0061] ④ Only after all modal data have been successfully aligned will the task data set be marked with "align_overall_status": "success" and proceed to the subsequent analysis process. This fine alignment ensures that the subsequent feature fusion fuses truly synchronized physiological and behavioral signals under the same cognitive event.
[0062] 2.4 Implementation of the Multimodal Feature Fusion Diagnostic Model This module is designed with a hierarchical model architecture of "single scene feature extraction → single scene fusion → multi-scene stitching" to fully explore complementary information across modalities and scenes.
[0063] (1) Model input: A single task data set that has been time-aligned from the S3 module, including: an 80-dimensional Mel spectrogram sequence, a 768-dimensional text semantic vector, and a 4-dimensional eye-tracking temporal feature sequence.
[0064] (2) Single-scene modal feature extraction: ① Language Feature Extractor: A simple multilayer perceptron (MLP). The input is a 768-dimensional semantic vector, which passes through a 128-dimensional fully connected layer (ReLU activation) and a Dropout layer, and outputs a 64-dimensional language feature vector.
[0065] ② Eye-tracking feature extractor: A gated recurrent unit (GRU) network. The input is eye-tracking temporal features of shape [time step, 4]. The GRU hidden layer is set to 256 dimensions. The output of the last time step is taken and then mapped to a 64-dimensional eye-tracking feature vector through a fully connected layer.
[0066] ③ Speech Feature Extractor: A hybrid network. The Mel spectrogram ([time steps, 80]) first passes through a one-dimensional convolutional layer (1D-CNN) to extract local features, then a bidirectional GRU captures long-term dependencies, outputting a 61-dimensional temporal acoustic feature. This feature is then concatenated with the 3-dimensional basic acoustic features (fundamental frequency mean, speech rate, intensity) extracted from the original audio, and finally output as a 64-dimensional speech feature vector through a fully connected layer.
[0067] (3) Single-scene multimodal feature fusion: The three 64-dimensional feature vectors from the same task scene are directly spliced together to form a 192-dimensional single-scene fusion feature vector.
[0068] (4) Multi-scene feature concatenation and classification: A complete interactive session usually includes multiple task scenarios (e.g., 6 visual tasks + 6 life communication tasks). The system concatenates the 192-dimensional feature vectors generated from all 12 task scenarios in sequence to form a 2304-dimensional cross-scene total feature vector. This total feature is input into the final classifier (a small neural network with two fully connected layers, using Softmax activation) to output a binary classification result of "normal" or "ADHD" and its probability.
[0069] (5) Model training: Using approximately 1000 collected labeled data (ADHD children and normal children), the training set, validation set and test set were divided in a 7:2:1 ratio. The Adam optimizer (initial learning rate 1e-5) was used with cross-entropy as the loss function, and early stopping (patience=5) was used to prevent overfitting.
[0070] 3. System Workflow and Effect Verification (Demonstration Example) Take the use of an 8-year-old child as an example: ① Start-up: When a child sits in front of the interactive terminal, the system starts up, and a virtual doctor appears and greets the child in a friendly manner.
[0071] ② Interaction and Data Collection: The virtual doctor guides the child through 12 pre-set tasks, including "finding the red bird" (visual recognition) and "describing the morning's classes" (daily school life). During this process, the microphone and eye tracker work continuously, and all data is collected synchronously and uploaded to the server.
[0072] ③ Processing and Inference: The server automatically executes the processes of modules S3 and S4 in the background: data preprocessing, temporal alignment, feature extraction and fusion. All calculations are completed within approximately 2 seconds after the dialogue ends.
[0073] ④ Output Results: The system interface displays the completed analysis and provides a result indicating "ADHD risk, specialist consultation recommended" along with a confidence level (e.g., 0.85). Simultaneously, a detailed report containing the original data, aligned features, and model inference logic is generated for professional review.
[0074] 4. Experimental verification and implementation results To verify the effectiveness of the above system, rigorous experimental verification was conducted: (1) Data set: In cooperation with three children’s hospitals, a total of 843 valid data were collected (including 412 children diagnosed with ADHD and 431 normal children of age and sex). All diagnoses were confirmed by at least two experts at the level of associate chief physician or above according to the DSM-5 criteria.
[0075] (2) Performance metrics: On an independent test set (168 cases), the best performance achieved by this system was: accuracy 78.4%, F1 score 0.81, and area under the receiver operating characteristic curve (AUC) 0.74. This is significantly better than the baseline model that uses only scale scoring (accuracy of approximately 65%) or only a single modality (e.g., eye movement only, accuracy of approximately 71%).
[0076] (3) Efficiency and User Experience: The average time for a single child to complete all interactions was 13.5 minutes (meeting the design goal of ≤15 minutes). The average time for a single inference by the backend model was 925 ms. Through a questionnaire survey, the children's acceptance and cooperation with the virtual avatar of this system reached 93%, proving the success of its child-friendly design.
[0077] (4) Ablation Study: To verify the contributions of each innovative module, an ablation study was conducted. ①Removing the "Time Sequence Alignment" module and directly aligning by timestamp approximately reduces the accuracy to 71.1%.
[0078] ②Removing "cross-scene feature stitching" and using only features from a random scene reduced the accuracy to 70.7%.
[0079] ③ Replacing the “virtual avatar + script guidance” with a traditional standardized questionnaire interface resulted in decreased children’s cooperation, increased data missing rate, and a final model accuracy rate of 69.5%.
[0080] These experimental data strongly validate the technical means proposed in this invention, such as virtual image interaction, multimodal temporal alignment, and cross-scene feature fusion, which play an indispensable and synergistic role in improving the overall system's analytical performance.
Claims
1. A system for collecting and analyzing children's behavioral characteristics based on virtual avatar interaction and multimodal fusion, characterized in that, include: The large-scale dialogue interaction module is used to drive virtual characters to engage in multi-round natural dialogues with children that are in line with their cognitive level, and dynamically control the dialogue process based on a preset interactive script. The multimodal data acquisition module is used to simultaneously collect children's voice data, text data, and eye-tracking physiological data during dialogue. The multimodal data preprocessing and time-series alignment module is used to clean, extract features, and synchronize the time series of the collected multimodal data. The deep learning analytics module includes multiple independent feature extractors and feature fusion units, used to extract features from preprocessed multimodal data and perform fusion analysis. The large-scale model dialogue interaction module establishes a data connection with the multimodal data acquisition module. The output of the multimodal data acquisition module is connected to the input of the multimodal data preprocessing and temporal alignment module. The output of the multimodal data preprocessing and temporal alignment module is connected to the input of the deep learning analysis module, forming a complete data processing pipeline from interaction guidance to feature analysis.
2. The child behavior characteristic collection and analysis system according to claim 1, characterized in that, The large-model dialogue interaction module includes: The virtual character rendering unit is configured with standardized parameters for visual, emotional, and linguistic dimensions. The visual dimension adopts cartoon-style modeling and is configured with a dynamic expression library and friendly body movements. The emotional dimension includes an emotional response rule library. The linguistic dimension is set with tone and speech rate parameters adapted to children's cognition. The dialogue flow control unit includes a preset interactive script library and a multi-turn dialogue control mechanism. The interactive script library is designed around the cognitive ability assessment dimension, and the multi-turn dialogue control mechanism dynamically determines the continuation or switching of topics based on the completeness assessment of children's answers and attention state analysis of a large language model.
3. The child behavior characteristic collection and analysis system according to claim 2, characterized in that, The multi-round dialogue control mechanism includes: The topic end judgment unit is configured to call the large language model to quantitatively assess the completion of diagnostic information collection for the current topic and the child's attention status, and determine whether to end the current topic based on a preset threshold. The topic follow-up question unit is configured to call the large language model to generate child-adaptive follow-up question scripts that point to the missing information item when the continuation condition is met. The topic switching unit is configured to call the large language model to generate transitional phrases and start the next topic when the termination condition is triggered.
4. The child behavior characteristic collection and analysis system according to claim 1, characterized in that, The multimodal data acquisition module includes: The voice acquisition unit is equipped with a directional noise-canceling microphone to acquire raw audio at a sampling rate of 16000Hz and calculates the signal-to-noise ratio in real time for quality control. The text acquisition unit simultaneously acquires the original transcribed text from the speech recognition model and the standardized text corrected by the large language model; The eye-tracking acquisition unit is equipped with an eye tracker to acquire raw data of three-dimensional fixation point coordinates and pupil diameter at a sampling rate of 120Hz, and extracts fixation duration and saccade frequency features in real time. The voice acquisition unit, text acquisition unit, and eye-tracking acquisition unit are all equipped with a structured storage mechanism that includes a unique session identifier and a dialogue turn ID.
5. The child behavior characteristic collection and analysis system according to claim 1, characterized in that, The multimodal data preprocessing and time-series alignment module includes: The speech preprocessing unit is configured to convert the processed audio into 80-dimensional mem spectrogram temporal features and extract fundamental acoustic features such as fundamental frequency, speech rate, and speech intensity. The text preprocessing unit is configured to convert standardized text into a 768-dimensional fixed-length semantic vector using a pre-trained language model. The eye-tracking preprocessing unit is configured to perform outlier removal, missing value completion, and feature normalization on temporal eye-tracking data. The timing alignment unit is configured to use the task scenario as the anchor point and correct the time deviation of the three types of modal data to within ±25ms through a dynamic time warping algorithm, and mark the alignment status.
6. The child behavior characteristic collection and analysis system according to claim 1, characterized in that, The deep learning analysis module includes: The language feature extractor uses a neural network structure that includes hidden layers and Dropout layers, taking a 768-dimensional semantic vector as input and outputting a 64-dimensional language feature vector. An eye-tracking feature extractor uses a GRU network structure to process temporal eye-tracking features and outputs a 64-dimensional eye-tracking feature vector. The speech feature extractor includes a Mel spectrum branch and a basic acoustic feature branch. After processing by 1D-CNN, pooling layer, and bidirectional GRU, it outputs a 64-dimensional speech feature vector. The feature fusion unit is configured to concatenate language, eye-tracking, and speech feature vectors in a single-task scenario to form a 192-dimensional fusion feature, and then concatenate the fusion features from multiple task scenarios and input them into the classifier.
7. A multimodal data timing alignment method, applied to the system described in any one of claims 1-6, characterized in that, Includes the following steps: Based on task scenario identification, raw voice, text, and eye-tracking data generated within the same task cycle are filtered to form a single-task multimodal data group. Based on the timestamp generated by the system clock, the time of task instruction issuance and task completion are extracted as time boundaries; Using key time points in the task flow as anchor points, the response time of voice data, the transcription completion time of text data, and the entire time sequence of eye-tracking data are aligned with the task timeline; The time offset of the three types of data is corrected by a dynamic time warping algorithm to ensure that the time deviation is ≤ ±25ms; Mark the alignment status in the data set. Mark the overall alignment status as successful only when all three types of modal data are successfully aligned.
8. A method for dynamic control of dialogue flow, applied to the system described in any one of claims 1-6, characterized in that, include: A pre-built interactive script library includes visual cognitive scripts and life communication scripts designed around the cognitive ability assessment dimension. During the dialogue, a large language model is used to quantitatively analyze the children's responses, including: ① Information collection completion assessment: Based on the preset core information collection items, calculate the score of the degree to which the child's answers cover each piece of information; ② Attention status assessment: Combining response time and text features, output attention labels and sustained scores; Based on quantitative analysis results and preset thresholds, a dynamic decision-making dialogue process is implemented. ① When the information collection completion score is ≥90 points, the current topic ends and a transition script is generated; ② When the attention tag is scattered and the score is <50 points for 3 consecutive rounds, end the current topic and generate a transition script; ③ When the number of dialogue rounds reaches the preset limit, end the current topic and generate transitional dialogue; ④ When the continuation condition is met, the large language model is invoked to generate follow-up questions pointing to the missing information item.
9. A multimodal feature fusion analysis method, applied to the system described in any one of claims 1-6, characterized in that, include: Feature extraction is performed on multimodal data in single-task scenarios: ①Language features: Extract 64-dimensional features from the 768-dimensional semantic vector through a network containing 128-dimensional hidden layers and Dropout layers; ② Eye movement features: 64-dimensional features were extracted from the temporal eye movement data using a 256-dimensional GRU network; ③ Speech features: The Mel spectrogram is processed by 1D-CNN, pooling layer, and bidirectional GRU and then concatenated with the basic acoustic features to output 64-dimensional features; The three types of feature vectors in a single-task scenario are concatenated to form a 192-dimensional fused feature. The fusion features of multiple task scenarios are spliced together by dimension to form a cross-scenario total feature; A fully connected network containing 256-dimensional hidden layers and Dropout layers is used to classify and analyze the total features across scenes, and outputs a binary classification result.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 7-9.